Name the symptom before the solution
Teams say the same thing for three different problems: we do not trust our tests. In the first the tests exist but give different results for the same code, so a red build means nothing. In the second the tests are thin or absent in the area you want to change, so a green build means little. In the third the tests are fine but nobody can see what happens in production, so you learn about failures from customers. Each needs a different, small piece of work. Write down which describes your last bad week, with one concrete example, before you look at any offer.
- Red builds that pass on a rerun: unreliable tests.
- Changes made nervously in one area: missing safety net.
- Failures found by customers: missing visibility.
- A flow that breaks and the build does not notice: missing end-to-end check.
Match the symptom to a bounded job
Martin Fowler's note on non-deterministic tests says they make failures untrustworthy and spread that distrust to healthy tests, so measure and fix them. A named handful of unreliable tests is a small job; a suite where nobody knows how many tests are unreliable is a measure-first project. For code nobody wants to touch, characterisation tests record what a module does now so the next change is visible. Fowler also says coverage is of little use as a number for how good the tests are, and that two better signs are that bugs rarely escape into production and developers are rarely afraid to change code. For blind spots in production, error tracking and structured logs make failures visible.
- Up to five named flaky tests: a fixed-scope stabilisation, from £395 for one test, quoted per test.
- An unknown number of flaky tests across a suite: a project, from £2,400, quoted from your CI history with a price cap before any measurement.
- One module, no tests: a regression suite, a published test price of £595.
- Three flows nothing watches: a browser smoke suite, a published test price of £995.
- No view of production errors: error tracking on one app, a published test price of £595.
What to ask a provider, including us
Ask how the work will be shown to have worked. For tests, the answer should be a test you could repeat: a failure shown before the change and a long run of clean results after it, with no assertion weakened and no retry setting added to hide failures. For a safety net, the answer should be that deliberate changes make tests fail, not a coverage percentage. For visibility, the answer should be a deliberate test error you can follow from the alert to the cause. Ask what is out of scope, who holds the accounts and keys, and what the provider will not promise. A provider who cannot say how acceptance works has not defined the job.
- We hold no keys and no accounts; you hold them and merge every change.
- We have not delivered these jobs for a client, and every price is an untested hypothesis.
A first conversation
Describe the symptom, give one example, say how long it has been going on and say who decides whether to buy. Name the runner or tool, the CI service and whether a non-production copy exists. Do not send code, keys, logs or customer data. We confirm fit, scope and the access route in writing before any work, nothing is charged before you sign off on a fixed job, and a project is quoted from what your CI history shows, with a price cap, before we have any access to your code; the measurement of the suite happens only after you agree. Quality work does not remove the need for your own judgement about priorities; it makes the facts behind them easier to see.
- The first enquiry needs the symptom, the stack in general terms and a decision-maker.
- Further scope, such as monthly upkeep, is agreed separately.
Sources and limits
- Martin Fowler: Test Coverage Checked 2026-10-11.
- Coverage is of little use as a numeric statement of how good the tests are; two better signs of enough testing are that bugs rarely escape into production and that developers are rarely afraid to change code.
- Martin Fowler: Eradicating Non-Determinism in Tests Checked 2026-10-11.
- Non-deterministic tests make failures untrustworthy and spread distrust to healthy tests in the same suite.
- Michael Feathers: Characterization Testing Checked 2026-10-11.
- Characterization tests document actual behaviour and act as a safety net for refactoring.