This ledger is synthetic
Every test name, count and cause below is invented to show the shape of the evidence a buyer should receive. It is not a record of work for any client and it measures nothing real. A real ledger replaces each row with actual runs from your own suite, with the runner, the environment and the source revision recorded.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
SYNTHETIC EXAMPLE: all names and numbers are invented.
Suite: "example-shop", one runner, one CI route
Baseline: 200 full-suite runs in shuffled order
Clean runs required per fixed test: 3 divided by its baseline failure rate, rounded up, never fewer than 50. No retry setting.
Other conditions: each fixed test also passed 50 of 50 runs alone and 50 of 50 under the pipeline's 4 parallel workers.
test baseline rate cause class change made required final status
-------------------------------- -------- ----- ----------------- --------------------------------- -------- ------- -----------
test_discount_applies_to_basket 9/200 4.5% shared state fresh list per test (fixture) 67 67/67 fixed
test_report_for_last_day_of_month 6/200 3.0% clock date passed in, fixed in the test 100 100/100 fixed
test_email_queue_drains 4/200 2.0% fixed wait wait for queue-empty condition 150 150/150 fixed
test_parallel_upload_size 0/200 0.0% not reproduced none n/a n/a not counted
test_legacy_import_order 3/200 1.5% real race in code quarantined with approval, owner n/a excluded quarantined
and date set
Break check on a throwaway branch: all three fixed tests failed when the behaviour they guard was deliberately broken.How to read each row
The baseline column shows how often the test failed before anything changed, and the rate column turns that into a percentage. A test that never failed in the baseline, like the fourth row, was not reproduced, so it cannot be called fixed and is not counted. The cause class is a conclusion with evidence behind it, reached by changing one factor at a time. The required column is worked out from the baseline rate, and the final column shows the test passed that many consecutive runs under the same conditions as the baseline. The last row shows an honest outcome: the cause is a real race in the code, which is beyond a test fix, so the test is quarantined openly, with the buyer's approval, an owner and a date, instead of being hidden or retried until green.
- Every row names a cause class and a change, or says why it has none.
- No assertion was removed or loosened, no wait was made longer as the fix and no retry setting was added.
- The quarantined test appears in the suite summary as excluded, so the exclusion is visible.
What a clean run count can and cannot show
A run of clean results lowers the chance that a flake remains, and it can never prove that none does. The arithmetic is simple. If a test still failed at its baseline rate of 2 percent, one run in 50, the chance of 150 clean runs in a row would be about 5 percent, which is why 150 is the number required for that row. The same sum gives about 5 percent for 100 clean runs at a 3 percent rate and for 67 clean runs at 4.5 percent. But if the rate had only halved, to 1 percent, the chance of 150 clean runs in a row would still be about 22 percent. So a clean count is strong evidence that the old rate is gone and weak evidence against a rarer flake that the baseline never saw. The ledger says so, and states the baseline rate next to the final count, so a reader can judge how much comfort the number gives.
- Ask for the baseline failure rate, not just the final count.
- For a rare flake, expect a longer run or a different kind of evidence, such as a code explanation of why the cause is gone.
What would make a row not count
A row would not count if the final runs used different conditions from the baseline, if a test was skipped rather than run, if a retry setting turned a failure into a pass, if the fix was a longer wait with no explanation of the cause, or if the test did not fail when the behaviour it guards was deliberately broken. A buyer who receives a ledger like this should check those five points before accepting it. The one-off job for up to five named tests and the larger suite project both use this shape, and both state that a clean run count is evidence with limits.
Sources and limits
- Martin Fowler: Eradicating Non-Determinism in Tests Checked 2026-10-11.
- Isolation, asynchronous waits, remote services, time and resource leaks are named causes of non-deterministic tests, and quarantined tests should be fixed quickly and limited in number or age.
- pytest: how to use skip and xfail Checked 2026-10-11.
- skip and xfail are separate markers, and by default neither XFAIL nor XPASS fails the suite.