Synthetic Industry

Inspectable example · updated 2026-10-11

Example stabilisation ledger for five flaky tests, with how to read the run counts

An inspectable synthetic ledger of five tests showing baseline failures and rates, classified causes, fixes, a not-reproduced test and an approved open quarantine, with how the required run count is worked out and what a clean run count can and cannot show.

An example, not a customer case study. Scope and evidence limitations are described below.

This ledger is synthetic

Every test name, count and cause below is invented to show the shape of the evidence a buyer should receive. It is not a record of work for any client and it measures nothing real. A real ledger replaces each row with actual runs from your own suite, with the runner, the environment and the source revision recorded.

If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.

SYNTHETIC EXAMPLE: all names and numbers are invented.
Suite: "example-shop", one runner, one CI route
Baseline: 200 full-suite runs in shuffled order
Clean runs required per fixed test: 3 divided by its baseline failure rate, rounded up, never fewer than 50. No retry setting.
Other conditions: each fixed test also passed 50 of 50 runs alone and 50 of 50 under the pipeline's 4 parallel workers.

test                              baseline  rate   cause class        change made                        required  final    status
--------------------------------  --------  -----  -----------------  ---------------------------------  --------  -------  -----------
test_discount_applies_to_basket   9/200     4.5%   shared state       fresh list per test (fixture)      67        67/67    fixed
test_report_for_last_day_of_month 6/200     3.0%   clock              date passed in, fixed in the test  100       100/100  fixed
test_email_queue_drains           4/200     2.0%   fixed wait         wait for queue-empty condition     150       150/150  fixed
test_parallel_upload_size         0/200     0.0%   not reproduced     none                               n/a       n/a      not counted
test_legacy_import_order          3/200     1.5%   real race in code  quarantined with approval, owner   n/a       excluded quarantined
                                                                      and date set

Break check on a throwaway branch: all three fixed tests failed when the behaviour they guard was deliberately broken.

How to read each row

The baseline column shows how often the test failed before anything changed, and the rate column turns that into a percentage. A test that never failed in the baseline, like the fourth row, was not reproduced, so it cannot be called fixed and is not counted. The cause class is a conclusion with evidence behind it, reached by changing one factor at a time. The required column is worked out from the baseline rate, and the final column shows the test passed that many consecutive runs under the same conditions as the baseline. The last row shows an honest outcome: the cause is a real race in the code, which is beyond a test fix, so the test is quarantined openly, with the buyer's approval, an owner and a date, instead of being hidden or retried until green.

  • Every row names a cause class and a change, or says why it has none.
  • No assertion was removed or loosened, no wait was made longer as the fix and no retry setting was added.
  • The quarantined test appears in the suite summary as excluded, so the exclusion is visible.

What a clean run count can and cannot show

A run of clean results lowers the chance that a flake remains, and it can never prove that none does. The arithmetic is simple. If a test still failed at its baseline rate of 2 percent, one run in 50, the chance of 150 clean runs in a row would be about 5 percent, which is why 150 is the number required for that row. The same sum gives about 5 percent for 100 clean runs at a 3 percent rate and for 67 clean runs at 4.5 percent. But if the rate had only halved, to 1 percent, the chance of 150 clean runs in a row would still be about 22 percent. So a clean count is strong evidence that the old rate is gone and weak evidence against a rarer flake that the baseline never saw. The ledger says so, and states the baseline rate next to the final count, so a reader can judge how much comfort the number gives.

  • Ask for the baseline failure rate, not just the final count.
  • For a rare flake, expect a longer run or a different kind of evidence, such as a code explanation of why the cause is gone.

What would make a row not count

A row would not count if the final runs used different conditions from the baseline, if a test was skipped rather than run, if a retry setting turned a failure into a pass, if the fix was a longer wait with no explanation of the cause, or if the test did not fail when the behaviour it guards was deliberately broken. A buyer who receives a ledger like this should check those five points before accepting it. The one-off job for up to five named tests and the larger suite project both use this shape, and both state that a clean run count is evidence with limits.

Sources and limits