An invented situation
Synthetic example: every name, number and result below is invented for illustration. It is not a client's work, a measurement or a delivery by us.
A test called invoice_total_rounds_half_up is said to fail about once in ten pipeline runs on the same commit. The team re-runs it until green. For this example, assume the work is done on a branch with test data only, and that each count below comes from the invented run logs listed in the table.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
condition | runs | failures | observed rate
test alone, fixed order | 50 | 3 | 6%
test in full suite, shuffled | 50 | 11 | 22%
test in full suite, parallel | 50 | 8 | 16%What the pattern points to
The failure rate climbs when the order is shuffled and when the suite runs in parallel, and is lowest when the test runs alone. That pattern suggests the test depends on state left behind by another test, rather than on its own logic. In this invented case, assume that one other test changes a tax-rate setting and does not reset it, so the rounding test sometimes runs with the wrong rate. A shuffled run that failed can be replayed with the same seed, which turns an intermittent failure into a repeatable one.
- Alone: lowest failure rate.
- Shuffled: highest, so order matters.
- Parallel: raised, so shared state matters.
After an invented fix, and how much a clean streak shows
Assume the fix gives the rounding test its own fixed tax rate and resets the shared setting after the test that changed it. The test then runs 100 times shuffled and under parallelism with no failures and no retry wrapper. If the old 22 per cent failure rate were still present, the chance of 100 clean shuffled runs would be about 0.78 multiplied by itself 100 times, around one in sixty billion. If the true rate were only the 6 per cent seen when running alone, the chance of 100 clean runs would still be about 0.2 per cent. Both figures assume independent runs, so they are a guide, not a proof.
The paid stabilisation job sets the number of clean runs for each test from its own baseline: three divided by the observed failure rate, rounded up, and never fewer than 50, in the conditions where the test failed, plus at least 50 in each other agreed condition. Here the baseline rates are 6, 22 and 16 per cent, so the formula gives 50, 14 and 19 runs and the floor of 50 applies to every condition; the 100 runs above are more than that rule needs.
To check that a fix did not weaken the test, a reviewer breaks the behaviour the test covers on purpose on a throwaway branch and confirms that the test fails there. The paid job requires this check for every test returned as fixed.
- Clean streak: reduces the chance a rare failure remains, but cannot show it is gone.
- Broken-behaviour run: shows the assertion was not weakened; the paid job requires it for each test returned as fixed.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
after fix | runs | failures
test in full suite, shuffled | 100 | 0
test in full suite, parallel | 100 | 0
throwaway branch, behaviour broken on purpose | 1 | fails, as it shouldWhat this example proves, and the next step
Synthetic example: every name, number and result below is invented for illustration. It is not a client's work, a measurement or a delivery by us. The local arithmetic above was checked only as arithmetic; no pipeline, test or fix was run for it, and it is not evidence that any failure rate can be removed.
For named flaky tests of your own, the stabilisation job covers up to five in one repository and is priced per test, from £395 for one test (an untested price), after a quote based on the test names and a redacted failure message for each. The total for several tests is the sum of their quotes. Each test is paid for after its repeat runs pass and you sign off its row; a test that is not reproduced or not stabilised is not charged for. The stabilisation ledger example shows several tests at once; this page follows one test's cause. Send the test names and failure messages with secrets removed. Do not send code or credentials in a first message.
Sources and limits
- GitHub: re-running workflows and jobs Checked 2026-10-11.
- A re-run reuses the same commit SHA and reference as the original run.
- Jest command line options Checked 2026-10-11.
- Jest can shuffle test order using a displayed seed so a failing order can be replayed.