Require comparable evidence, not endless retries
Keep passing and failing attempts with their revision, event, job, matrix configuration, runner and first useful error. A rerun keeps the original SHA/ref, but dependencies, external services or timing may still differ; equal revision does not prove an identical environment. Compare existing logs before authorising another run with potential deployment or external effects. This is an investigation plan, not an executed flaky-test diagnosis.
- Count reruns as attempts, not successful repairs.
- Do not claim a failure rate from a selectively retained sample.
- Retain the original failed attempt when only a subset of jobs reran.
Assign three separate decisions
Name an investigation owner who gathers and tests the cause, a maintainer or release-risk owner who approves any temporary gate change, and an application or test owner who accepts the actual correction. These may be the same authorised person, but the decisions remain separate. An analysis tool can assist the first responsibility; it does not acquire the authority to waive a required check.
- Investigation: describe the suspected condition and evidence needed to disconfirm it.
- Temporary quarantine: document lost coverage, an explicit review date or removal condition, and compensating checks; do not present it as a passing repair.
- Correction: reproduce the defective behaviour on a safe route and retain the regression and relevant checks.
Do not sell a flake as the £149 repeatable-job repair
The existing £149 one-job offer expressly excludes intermittent tests and real application bugs. The £295/month standing CI offer can investigate repeated failures and propose a fix or customer-approved quarantine within its agreed scope; it excludes application repairs, merging, runner/secret/billing changes and guaranteed response times. Ask how investigation and unresolved causes are treated before agreeing the monthly allowance. If the cause is product behaviour, scope a product repair separately rather than weakening CI.
- No emergency or round-the-clock coverage follows from either offer.
- The customer's authorised maintainer controls merges and gate changes.
- A quarantine is risk containment, not evidence the original assertion passes.
Choose the smallest useful owner
Use your existing engineer and free run-log documentation when they can finish the investigation. Consider a tool only with acceptable access, review effort and evidence. Consider recurring CI ownership when the missing work is repeated workflow investigation, not merely notifications. The proposed US$10,000/month lane is a separate fit discussion only for an existing product with a continuing bounded feature/bug/maintenance backlog; it is not a flaky-test upsell or on-call service. It has one active item and at most four agreed monthly items, not guaranteed completions. Initial enquiries need redacted context, responsibility and budget questions, never code or keys. Prices are untested and written scope, safe access and terms precede work.
Sources and limits
- GitHub rerunning workflows and jobs Checked 2026-10-11.
- A rerun uses the original SHA/ref and original triggering actor's privileges. A passing rerun does not establish a cause or repair.
- GitHub run logs and attempt archives Checked 2026-10-11.
- Partial-rerun archives contain only rerun jobs; earlier attempts may be needed for the full record.
- One-off CI repair excludes intermittent tests Checked 2026-10-11.
- Existing recurring CI investigation and customer approval boundaries Checked 2026-10-11.
- Existing proposed broader engineering lane Checked 2026-10-11.