Synthetic Industry

Project test-stabilise-whole-test-suite-project · revised 11 October 2026

Project

Make a flaky test suite reliable enough that a red build means something

We measure which tests in one suite are unreliable, then fix or openly quarantine each until the whole suite passes an agreed number of full runs in a row. You accept the finished suite.

This asks for a proposal by email. Nothing is charged, and nothing starts, until you have agreed the scope, the price and the terms in writing.

The result you are buying

In a suite with many intermittent failures, nobody knows which tests are unreliable, how often they fail or why. Fixing them one at a time from memory never ends, and the team learns to ignore failures, including real ones. The project measures the whole suite first, then repairs it in batches until the suite itself is dependable.

Who it’s for: A founder or engineering lead whose team no longer trusts a red build because an unknown number of tests fail at random.

Usually starts when: Reruns are routine, merges happen over red builds, or a release slipped because nobody could tell a real failure from a flaky one.

The result: The whole suite passes an agreed number of consecutive full runs on the agreed CI route, every test that proved unreliable is fixed or openly quarantined with a reason and an owner, and a ledger shows each test's failure counts before and after. You accept the finished suite.

How the work fits together

The project is complete when every test the baseline showed to be unreliable has been fixed or openly quarantined, the whole suite has passed the agreed number of consecutive full runs, and you have accepted the finished suite.

  1. Fix one failing GitHub Actions job and show a green run Job First

    One failing job in one workflow runs green on your pull request branch, on the same trigger, with the before and after logs attached. You review and merge the change.

    If a job fails the same way every time for a non-flaky reason, such as a broken command, it is repaired first so the measurement is meaningful.

  2. Stabilise flaky tests, up to five, priced per test, each one fixed or openly quarantined Job Included

    Up to five named tests that fail at random, priced per test from £395 for one test. Each is reproduced, its cause found, then fixed or, with your approval, quarantined. You get a pull request.

    One batch of up to five tests at a time, until every test the baseline showed to be unreliable has been dealt with

    Each batch meets the same acceptance as the one-off job.

  3. Pin one module's current behaviour with a regression test suite Job Optional

    A test suite around one named module records what it does today, so a refactor or upgrade shows exactly which behaviour changed. You get the tests, a seeded-change check and a gaps note.

    On request, for a module guarded only by quarantined tests

    Quarantine removes protection; a regression suite can replace what is lost.

  4. Keep your GitHub Actions CI green, month after month Standing service Optional

    We watch the workflows you name. When one fails on your default branch we investigate it and open a fix pull request, or tell you what only you can change.

    After the project, a monthly service can keep the named workflows passing.

How an engagement works

The price cap covers the agreed suite and up to the number of unreliable tests the quote assumed. If the baseline finds more, or tests turn out to be defects, or flakiness appears later in new tests, that is quoted separately and nothing is done beyond the cap without your written agreement.

How it starts

  1. You tell us the suite, the CI route and what a red build costs you today, and describe or export what your CI history shows without sharing code: how long a full run takes, how many runs it holds, and runs of the same commit where a test failed and then passed.

  2. From that description alone, without access to your code or CI, we propose the scope, the number of unreliable tests the quote assumes, the number of runs for acceptance and a price cap.

  3. You agree the scope, price cap and terms in writing. Nothing starts, and we have no access to your repository or CI, before then.

  4. After you agree, you share the repository and approve the baseline runs; we run the suite many times on a copy to measure it and check the assumption in the quote. If the baseline finds more unreliable tests than assumed, we stop and show you the evidence, and the price changes only if you agree a new written price.

  5. We work through the unreliable tests in batches, sending a pull request and the evidence for each batch.

  6. We run the final consecutive full runs and send the ledger, the quarantine list and the summary.

Who decides what

You approve the repeated runs and the CI usage, decide on each quarantine and merge each pull request. We do not deploy or change repository settings.

Handover

Each batch arrives as a pull request. Quarantine markers, the ledger and the measurement method stay in your repository or are handed over as files you keep.

Sharing your product safely. Describe the suite and the CI route in words and send no code, secrets or customer data. After you agree the project, share the repository through an authorised company-controlled route and approve the measurement runs.

What is included, and what is not

  • A pull request for each batch of fixes, with its evidence
  • A flake ledger covering every test that failed at least once in the baseline: its cause, the change and its final repeated-run result
  • A quarantine list with each quarantined test's reason, owner and review date, and the protection lost meanwhile
  • The measurement method as a script or a scheduled job you can rerun, and a final summary comparing the suite's failure rate before and after

Included

  • One repository, one test suite and one CI route, agreed in writing
  • Measure first, once you have agreed the scope and price cap and given access: run the whole suite many times, an agreed number, and record failures per test so the unreliable set is found by evidence
  • Repair in batches by cause: fix the cause in each test, or in the code where we show a defect, or quarantine the test openly with a reason, an owner and a review date
  • Repeat the full-suite run until it passes the agreed number of consecutive times, and leave a repeatable way to keep measuring

Not included

  • Repairing real defects that tests correctly report; each is a separate bug-fix job
  • Speeding up the suite, raising coverage or rewriting it in another framework
  • More than one suite or repository, or buying more CI capacity or runners
  • Retry settings that rerun a failing test until it passes: reruns are not a fix and are not added as one
  • Production changes, deployments or changes to repository settings

How we know it’s done

Agreed with you before work starts. Each check produces evidence you keep.

  1. The whole suite passes the agreed number of consecutive full runs on the agreed CI route with no retry setting enabled, the same configuration throughout and never fewer than 20 runs.

    Evidence: The CI run links or logs for the consecutive runs, with the source revision, runner and configuration.

  2. The flake ledger lists every test that failed at least once in the baseline with its cause, the change made and its final result, and no such test is missing or silently skipped.

    Evidence: The ledger compared with the baseline run records.

  3. Every quarantined test has a visible marker, a reason, an owner and a review date, and the suite summary names each excluded test.

    Evidence: The quarantine list and a suite summary from the final runs.

  4. No assertion was removed or loosened, no fixed wait was lengthened as a fix and no retry-until-green setting was added.

    Evidence: The independently reviewed diffs for every batch and the complete changed-file list.

Sign-off. You read the ledger and the quarantine list, check the final runs against your CI history, and accept the finished suite in writing.

If it fails. A test we cannot stabilise is named with the reason and what it would take, and is quarantined only with your approval; the price is adjusted to match what was done. Nothing is billed as delivered that you have not accepted.

When it fits, and when we stop

It fits when

  • The whole suite can run repeatedly on a CI route or equivalent machine without deploying anything or using production secrets, within your CI allowance
  • Each test can be run with test data only; tests that need live services are named and excluded or replaced in the proposal
  • A named person on your side can approve repeated runs, review pull requests and accept the finished suite

We stop and tell you if

  • A full run takes so long, or costs so much CI allowance, that the agreed number of runs cannot be reached: we propose a smaller scope
  • Most failures come from the shared environment, such as a database or runner we cannot reach, rather than from tests
  • The baseline finds far more unreliable tests than the quote assumed: we stop, show you the baseline evidence and propose a new scope and price; nothing beyond the agreed cap is done or charged unless you agree it in writing
  • Running the suite needs production credentials, or would send real email, charge cards or change live data

What could go wrong

Every batch is a separate pull request that your team merges, so any one can be reverted on its own. If the project stops, merged batches stay and the rest is handed back with notes, including the measurement so far.

Scroll the table sideways to read it all.

RiskHow we handle it
A clean run of the agreed length is read as proof that no flakiness remains.The summary states the baseline failure rate and the number of clean runs, and says that a clean count lowers the odds of a remaining flake but cannot prove there is none.
Quarantine removes protection and the excluded tests are never revisited.Every quarantine has a visible marker, a reason, an owner and a review date, and the list names the protection lost; you decide on each one.
Repeated full runs use more CI allowance than expected.The number of runs and the CI usage are agreed before the baseline starts, and the project scope is reduced rather than exceeding them.
The quote was based on a description of your CI history, and the baseline shows the suite is worse than assumed.The quote states the number of unreliable tests it assumes and a price cap. If the baseline finds more, we stop and show you the evidence, and nothing beyond the cap is done or charged unless you agree a new written price.

Each batch is reviewed separately from the work that produced it, and you accept the finished suite. No human supervisor is included unless your proposal names one. At launch the work is largely automated, and we say so.

Stays with a person

  • You approve repeated runs and CI usage
  • You decide on each quarantine, and merge each pull request

Access we would need

  • Read access to a repository or fork you control, and your approval to run the suite repeatedly on an agreed CI route; no production access

Questions

How is this different from stabilising five named tests?

The one-off job starts from tests you can already name. This project starts by measuring the whole suite, so the unreliable set is found by evidence, and it ends only when the whole suite passes the agreed number of full runs.

How can you quote before measuring the suite?

We cannot measure before you agree, because that needs access to your code and CI runs that only you can approve. So we quote from what you can describe or export from your CI history without sharing code, state how many unreliable tests the quote assumes and set a price cap. The baseline then checks that assumption. If it finds more, we stop and show you the evidence, and the price changes only if you agree a new written price.

How is this different from the faster, steadier CI project?

That project brings a repository's CI to an agreed duration and a reliable run history and can include several causes besides flakiness. This one is about the reliability of the tests themselves: it measures every test in one suite and ends when the whole suite passes an agreed number of full runs in a row.

Can you promise the suite will never flake again?

No. New tests can be flaky, and a clean run count lowers but does not remove the chance that something remains. The handover includes the measurement method so you can keep checking.

What if a flaky test turns out to be a real bug?

We name it with the evidence and take it off this list. Fixing it with a regression test is a separate job.

Send an enquiry

Send us

  • The repository link or hosting platform, the test runner and language, and the CI service that runs the suite
  • Roughly how long a full run takes and how often a build goes red for reasons nobody can explain
  • What your CI history shows without sharing code: how many runs it holds, and a few pairs of runs on the same commit where a test failed and then passed, with test names and redacted failure lines only
  • Whether any tests call live services, and who can approve repeated runs and the CI usage they cause
  • Do not send repository code, credentials, customer data or an access invitation in the first enquiry

Later, once you agree

  • The repository through an authorised company-controlled code-export route, with a branch route for the pull requests, given only after you have agreed the scope and price cap in writing
  • Your written approval for the baseline and acceptance runs, the CI minutes or machine time they use and the CI route agreed
  • Synthetic or test-only fixtures, and the named person who accepts the finished suite

You own the repository, the CI account, the runners and every secret. We work on branches through a company-controlled identity, never a personal login, and run only the agreed suite on an agreed route. You approve CI usage, review and merge each pull request, and decide on each quarantine. We do not deploy or change repository settings.

A public HTTPS link only, without login details, query strings or fragments. No code or logs.

Sending emails your enquiry and contact address to our team through our mail provider (Resend). It is not kept in a website database. Do not send passwords, keys, recovery links, confidential code or customer records. Your contact email is unverified; nothing is ordered, charged or reserved. Privacy notice.

Email fallback: open your mail app

If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “test-stabilise-whole-test-suite-project” as the subject.