Job test-stabilise-named-flaky-tests · revised 11 October 2026
Stabilise flaky tests, up to five, priced per test, each one fixed or openly quarantined
Up to five named tests that fail at random, priced per test from £395 for one test. Each is reproduced, its cause found, then fixed or, with your approval, quarantined. You get a pull request.
You might be seeing
- A named test fails in some runs and passes when the same commit is run again
- A test passes when run alone but fails in the full suite, or the other way round
No passwords, keys, card details or admin invites needed to start.
What usually happened
A small, nameable set of tests gives different results for the same code. The cause sits in the test or in the code it exercises: shared state, order dependence, a clock, a fixed wait, an unreplaced external service or a real race. A rerun hides it without removing it. Unlike a job that fails the same way every time, no single log shows the cause, so it has to be provoked and measured before it can be fixed.
Who it’s for: An engineering lead whose team reruns the build by habit because a few named tests fail at random, and who wants each failure traced to a cause.
Usually starts when: The same few tests turn builds red and then pass on a rerun, and people have started merging after a rerun without reading the failure.
The result: Each named test either passes a run count worked out from its baseline failure rate, in the conditions it failed in and in shuffled and parallel runs, after a fix whose cause is explained and with a check that it still fails when the behaviour it guards is deliberately broken, or is quarantined in the open, with your written approval, with a reason, an owner and a date to revisit. You receive one pull request and the before and after run counts.
Check whether this job fits
Check these against your CI history without sharing code or credentials. The result is a fit check for this fixed job, not permission to run your tests.
Checks you can run yourself
Find failed-then-passed pairs
In your CI history, look for a commit whose run failed and whose rerun passed with no change. Write down the failing test name and the first error line in each pair, with secrets and customer details removed.
Look for: The same few test names across several pairs. Those are the candidates to name in an enquiry. A different test each time points to the environment rather than five bad tests.
What you get
- A pull request containing the fixes and any approved quarantine markers
- A stabilisation ledger: for each test, the baseline failure count and rate, the classified cause, the change made, the required and achieved run counts in each condition, and the result of the broken-behaviour check
- A quarantine list showing each quarantined test's reason, owner and review date, and the protection lost while it is excluded
- Instructions to repeat the same run-count check yourself
Included
- Up to five named tests in one repository, one test suite and runner, and one CI-equivalent environment agreed in writing
- Reproduce each failure by running the test repeatedly: alone, within its file, in the full suite and, where the runner supports them, in shuffled order with the seed recorded and in parallel with the worker count your pipeline uses, recording failures per run
- Classify each cause with its evidence: shared state, order dependence, time or randomness, fixed waits, an unreplaced external service, resource contention or a real race in the code
- Fix the cause in the test, or in the code only where we show the cause is a defect; where a fix is not justified in scope, propose quarantine and apply it only to the tests you approve in writing, each with a visible marker, a reason, an owner and a review date
- For each fixed test, break the behaviour it guards on a throwaway branch and show that the test then fails, so the assertion is shown to be still doing its job
Not included
- A whole CI job that fails the same way each time, such as a bad command or a missing dependency; that is the GitHub Actions job repair
- Tests that fail every time because they correctly report a defect; that is a bug-fix job
- More than five tests, or the whole suite; a suite-wide effort is the larger stabilisation project
- Speeding up slow tests or raising coverage
- Retry settings or plugins that rerun a failing test until it passes: a rerun is not a fix and is not added as one
- Deleting a test, quietly skipping it, or changing an assertion so that it checks less; quarantine happens only in the open and only with your written approval
- Intermittent failures that happen only in production
How we know it’s done
Agreed with you before work starts. Each check produces evidence you keep.
For each test returned as fixed, the baseline record shows at least one failure and an observed failure rate under the agreed conditions, and the same test then passes a required number of consecutive runs, with no retry setting, in the conditions in which it failed. The required number is 3 divided by the observed baseline failure rate, rounded up, and never fewer than 50: a test that still failed at its baseline rate would pass that many runs in a row only about 5 percent of the time. The test also passes at least 50 consecutive runs in each other agreed condition the runner supports: alone, in shuffled order with the seed recorded, and under the parallel workers your pipeline uses.
Evidence: The ledger rows with baseline failure counts and rates, the required and achieved run counts per condition, runner and source revisions, plus the run logs or summaries.
For each test returned as fixed, a deliberately broken version of the behaviour it guards, on a throwaway branch, makes the test fail. If that behaviour cannot be broken without editing the test itself, the ledger says why and the row counts only if your maintainer accepts that in writing.
Evidence: A failing run on a throwaway branch for each fixed test, linked from the ledger.
Each test that was never made to fail in the baseline is returned as not reproduced and is not counted as fixed.
Evidence: The baseline run counts for that test and the stated reason it is not counted.
Every quarantined test was approved by you in writing and carries a visible marker with a reason, an owner and a review date, and is reported as skipped or excluded in the suite summary rather than silently absent.
Evidence: Your written approvals, the quarantine list and a suite run summary that names each excluded test.
No assertion was removed or loosened, no fixed wait was lengthened as the fix, and no retry-until-green setting was added; your authorised maintainer accepts the pull request.
Evidence: The independently reviewed diff, the complete changed-file list and your written sign-off.
Sign-off. You read the ledger and the quarantine list, check the run counts against your own CI history, sign off in writing test by test and merge the pull request. Payment for each test follows its sign-off.
If it fails. Payment is per test. If a test cannot be reproduced or is not stabilised, it is not counted and nothing is due for it within this fixed scope; we explain what we found and what a different scope would need. Quarantine counts, and is payable at the quoted price for that test, only when you approved it in writing and it is complete and visible. Wider work needs a new written agreement.
When it fits, and when we stop
It fits when
- You can name each test and say roughly how often it fails, or on which runs it failed
- The tests can run repeatedly on a CI route or equivalent machine without deploying anything or using production secrets, within your CI allowance
- The code and tests can be shared through an authorised company-controlled route after agreement, and a named person on your side can review and merge the pull request
We stop and tell you if
- A test cannot be made to fail within the agreed number of baseline runs, so there is no failure to explain: we report the runs and propose instrumentation, or stop
- The cause lies outside the repository or the agreed environment, such as a shared test database, a provider outage or runner hardware we cannot reach
- Running the tests needs production credentials, or a run would send real email, charge a card or change live data
- The cause is a real defect that needs a larger change than the agreed scope: we explain it and agree a separate bug-fix job
- The run count needed for a test (three divided by its baseline failure rate) is more than the CI allowance or time you have approved: we report the baseline rate, the count needed and what the approved allowance can show, and we do not describe a shorter run as proof
What could go wrong
Before merge, closing the pull request leaves your default branch unchanged. After merge, your maintainer can revert the commit; reverting restores the earlier behaviour, including the flakiness. Quarantine markers are plain, searchable annotations that can be removed.
Scroll the table sideways to read it all.
| Risk | How we handle it |
|---|---|
| A test passes the repeated runs only because its cause is rare, and a clean count is mistaken for proof that the problem is gone. | The required run count is set from the baseline failure rate, so that a test still failing at that rate would pass it only about 5 percent of the time. The ledger states the baseline count and rate and the number of clean runs, and says plainly that a clean count lowers the odds of a remaining flake but does not prove it is gone, least of all a flake rarer than the baseline. |
| A test is made to pass by weakening what it checks. | The reviewer compares each assertion before and after, and a deliberately broken version of the behaviour must still make each fixed test fail. |
| A quarantined test hides a real defect and is forgotten. | Each quarantine is visibly marked with a reason, an owner and a review date, appears in the suite summary and is listed in the handover with what protection is lost meanwhile. |
| Repeated runs consume CI allowance or touch shared resources. | The run count, the environment and the CI usage are agreed in advance, and nothing runs against production services or with production secrets. |
An independent reviewer checks that no assertion was removed or loosened, that no wait was lengthened as the fix, that each fixed test failed in its broken-behaviour check, and that every quarantine was approved by you, is visible and is dated. Your authorised maintainer reviews and merges under your existing rules; we do not bypass that approval.
Need to keep it working?
If flakiness turns out to be spread across the suite, we quote the larger stabilisation project separately.
Ongoing work is separately scoped and quoted: no monitoring, response-time guarantee or automatic subscription is included in this job.
Explore an ongoing engineering lane, or mention the responsibility you need in your enquiry.
What you can check
This is a new service. We have not delivered this job for a client yet.
Other ways to get this done
- Martin Fowler's article on non-deterministic tests lists the usual causes (lack of isolation, asynchronous behaviour, remote services, time, resource leaks) and a quarantine approach your own team can follow. martinfowler.com
- pytest's documentation explains how skip and xfail markers work, which is what a clearly labelled quarantine needs. docs.pytest.org
Questions
Why not just add retries?
A retry reruns a failing test until it passes, which hides the cause and can mask a real defect. This job finds the cause, then fixes it or quarantines the test in the open. Retries are not added as the repair.
Can you promise the test will never fail again?
No. A long run of clean executions lowers the chance that the flake remains, and the ledger states the counts, but it cannot prove the cause is gone. We say what we found and how sure that evidence leaves us.
How many repeat runs is enough?
It depends on how often the test failed in the baseline. We take the observed failure rate and require three divided by that rate in clean runs, rounded up and never fewer than 50. A test that failed in 2 of 100 baseline runs, for example, would need 150 clean runs. If it still failed at its old rate it would pass that many runs in a row only about 5 percent of the time. That shows the old rate is gone; it cannot show a rarer flake is gone.
What does quarantine mean here?
The test is marked visibly, excluded or reported as skipped, and listed with a reason, an owner and a date to revisit. We quarantine a test only if you approve it in writing. It is a temporary, honest exclusion, not a deletion or a silent skip. A test quarantined this way counts as delivered for that test and is payable at its quoted price once you sign off its row.
Is this the same as fixing a failing CI job?
No. A CI job that fails the same way every time has one repeatable cause in its log. This job is for individual tests that give different results for the same code.
Send an enquiry
Send us
- The names of up to five tests, the test runner and language, and where the tests run (hosted CI or a self-managed runner)
- How often each test fails, as best you know, and a redacted failure message for each
- Whether each test touches a database, the network, the clock or a browser, in a line each
- Do not send repository code, credentials, customer data or an access invitation in the first enquiry
Later, once you agree
- The agreed source revision and branch route through an authorised company-controlled repository or code-export route
- Your written approval to run the named tests repeatedly, the CI minutes or machine time that allows, and, after the baseline, the run count for each test worked out from its baseline failure rate as set out in the acceptance checks
- Your written decision, test by test, on whether a test we cannot fix in scope may be quarantined
- Synthetic or test-only fixtures, and the named person who accepts the result
You keep the repository, the CI account, the runners and every secret. We work on a branch through a company-controlled identity, never a personal login, and run only the agreed tests on an agreed route. You approve CI usage, review the pull request and merge it. We do not change repository settings or deploy.
Email fallback: open your mail app
If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “test-stabilise-named-flaky-tests” as the subject.