1. Is the failure real and repeatable?
Start by separating a test that fails every time from one that fails sometimes. A test that always fails with the same message is reporting something repeatable, and the repair is a bug fix with a regression test, or a job repair if the whole CI job fails the same way. Only a test that gives different results for the same code is unreliable in the sense used here. Collect failed-then-passed pairs from your CI history, with the test name and the first error line, and see whether a few names recur. If the same few tests recur, name them. If a different test fails each time, the cause is likely the environment, not five bad tests.
- Same failure every run: repair the cause rather than diagnosing flakiness.
- Same few tests, sometimes: continue to step 2.
- A different test each time: look at the runner, the shared database and the machine.
2. Does it depend on what ran before?
Run the test alone, then with its file, then with the whole suite, and if possible in a different order. If the result changes with the order, an earlier test is leaving state behind: a wide fixture, a module-level variable, a shared file or a database row. The shared-state guide shows how to prove it and how to isolate the test. Order dependence is also what makes a test pass alone and fail in the suite, which is the most common first sign. Fix the isolation. Do not fix the order.
- Test passes alone and fails in the suite: use the shared-state guide.
- Browser tests collide when run together: use the accounts and parallel-workers guide.
3. Does it depend on time, randomness or the machine?
If the order is not the cause, look at what the test reads from outside: the date and time, a random value, an environment variable, the locale, the order of items from a dictionary or database. Failures that cluster around month ends, midnight or daylight-saving changes point to the clock. Control that input in the test and replay the failing conditions deliberately. If the failure then appears every time, you have found the cause. If it does not, the cause is somewhere else.
- Failures cluster at certain dates or on one machine: use the clock, randomness and environment guide.
4. Does it depend on waiting?
In browser and asynchronous tests, a test often acts or asserts before the system is ready, and a fixed pause fails on a slow day. Read the timeout as a question about which condition was not met, use the waiting the tool already does and replace each pause with the visible result it stood in for. If the failures happen only in CI, collect a trace and compare the differences that matter before changing the test. Lengthening timeouts is rarely the repair.
- Timeout on a slow page: use the guide on waiting for a condition.
- Fails only in CI: use the guide on collecting a trace first.
5. What do you do with a test you cannot fix yet?
Some tests cannot be fixed quickly, for example because the cause is a real race in the product. Mark the test openly, with a reason, an owner and a date to revisit, and keep the list of what protection is lost. Retries and permanent skips only rename the problem. The quarantine guide compares skip, xfail and a separate suite, and the stabilisation ledger example shows how to record each outcome.
- Set a limit on how many tests may be quarantined and for how long.
- Ask who owns the cleanup, and read the ledger example for the shape of the record.
Paid routes, by size
Up to five named flaky tests: a fixed-scope stabilisation, from a published test price of £395 for one test and quoted per test, accepted by baseline failures shown and a run of clean results sized from the baseline failure rate, without a retry setting. An unknown number across a suite: a project from £2,400, quoted from what your CI history shows, with a price cap, before we have any access to your code. The measurement of the suite happens only after you agree. A job that fails the same way every time is a different fix. Prices are untested hypotheses, nothing is charged before sign-off on a fixed job, we have not delivered these jobs for a client, and a clean run count lowers but cannot remove the chance that a flake remains. Send test names and redacted failure messages, never code or credentials.
Sources and limits
- Martin Fowler: Eradicating Non-Determinism in Tests Checked 2026-10-11.
- Isolation, asynchronous waits, remote services, time and resource leaks are named causes, and quarantine is a stop-gap that needs fixing quickly.
- Playwright: test retries Checked 2026-10-11.
- Tests that fail and then pass on retry are reported as flaky, and retries are off by default.
- pytest: how to use skip and xfail Checked 2026-10-11.
- By default neither XFAIL nor XPASS fails the suite.