A re-run shows nothing
When a test fails and passes on the same code, the code is not the only input. A re-run uses the same commit and reference as the original, so a green re-run only tells you that this time the other inputs lined up: order, timing, shared state, the clock or a random value. That is why the habit of re-running until green hides a problem instead of fixing it, and why the useful question is not whether the test passed but how often it fails.
If you have not already decided who owns a flaky test, the ownership guide covers that question. This guide is about the evidence that a fix worked.
Turn a failure rate into a number of runs
Suppose a test fails once in every ten runs. If nothing has been fixed, the chance of it passing five times in a row is 0.9 multiplied by itself five times, about 59 per cent, so five clean runs prove very little. Twenty clean runs happen by luck about 12 per cent of the time, and thirty about 4 per cent. The lower the failure rate, the more runs you need before a clean streak means something. The table below is arithmetic, not data from a source: it assumes each run is independent, which real runs are not always, so treat it as a lower bound on how much evidence you need.
As a rule of thumb, to be about 95 per cent sure that a test that still fails with probability p would have shown a failure, you need about three divided by p clean runs. At one failure in a hundred, that is about three hundred runs. The worked stabilisation ledger shows how a baseline rate and a final run count are written side by side so that a reader can judge how much comfort the count gives.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
failure rate before fix | clean runs | chance of that streak if nothing was fixed
10% (1 in 10) | 5 | 59%
10% (1 in 10) | 20 | 12%
10% (1 in 10) | 30 | 4.2%
5% (1 in 20) | 20 | 36%
5% (1 in 20) | 50 | 7.7%
5% (1 in 20) | 100 | 0.6%
1% (1 in 100) | 100 | 37%
1% (1 in 100) | 300 | 4.9%Where the dependence hides
Look for what the test does not control. Order: another test leaves a record, a file or a global value behind. Many runners can shuffle test order with a seed that is printed, so a failing order can be replayed; Jest documents this with a randomize option. Time: a test compares against the current time, a day boundary or a fixed wait. Randomness: a value generated without a fixed seed. Resources: tests running in parallel share a database, a port or a temporary folder. Network: a call to a service that is sometimes slow. Each of these can be removed from the test rather than hidden.
- Order and shared state
- Clock and fixed waits
- Unseeded randomness
- Shared database, port or folder under parallel runs
- Outside network calls
A safe way to investigate on a branch
Use test data only. First run the single test alone many times and count failures: this gives the baseline rate. Then run it shuffled with a seed and keep any seed that fails. Then run it under the same parallelism the pipeline uses. The pattern of which runs fail points at the cause: failing only when shuffled suggests order, only in parallel suggests shared resources, only near midnight suggests the clock. Record the counts, because they decide how many clean runs you need afterwards.
Do not add a retry wrapper, loosen an assertion or skip the test to make the pipeline green. Those are decisions about what you want the suite to check, and they belong to you, not to the fix. To check that a fix did not quietly weaken the test, break the behaviour the test covers on purpose on a throwaway branch and confirm that the test now fails.
How the paid outcome is accepted
The stabilisation job covers up to five named tests in one repository. For each test returned as fixed, it is accepted when the baseline shows at least one failure under the agreed conditions and the same test then passes a number of consecutive runs worked out from that baseline (three divided by the observed failure rate, rounded up, and never fewer than fifty) under the same conditions, and at least fifty consecutive runs in each other agreed condition; when a deliberately broken version of the behaviour the test guards makes it fail; when a test that could not be made to fail is reported as not reproduced rather than counted as fixed; when anything quarantined was approved by you in writing and is listed openly with a reason, an owner and a review date; and when no assertion was removed or loosened, no fixed wait was lengthened and no retry was added. It is priced per test, from £395 for one test (an untested price), after a quote based on the test names and the runner; the total for several tests is the sum of their quotes, and you pay for each test after you sign off its row.
Sources and limits
- GitHub: re-running workflows and jobs Checked 2026-10-11.
- A re-run reuses the same commit SHA and reference as the original run, and a workflow run can be re-run at most 50 times.
- Jest command line options Checked 2026-10-11.
- Jest can shuffle the order of tests within a file using a seed that is displayed, so a failing order can be replayed.
- Existing single-bug repair Checked 2026-10-11.