Job ci-job-killed-out-of-memory-or-disk · revised 11 October 2026
Stop a CI job being killed for running out of memory or disk space
The named job completes on the same runner size with a measured margin left, instead of being killed, with the cause named and no runner upgrade needed unless we tell you and you approve it.
You might be seeing
- The log ends with Killed, an out-of-memory message or no space left on device
- The job passes on a developer machine with more memory, or passes when fewer tests run in parallel
No passwords, keys, card details or admin invites needed to start.
What usually happened
The job asks for more memory or disk than its runner has. Memory peaks come from too many parallel workers, a bundler holding everything at once or a test that loads a large fixture; disk fills from build outputs, caches, container images and dependency downloads that are never cleaned. The runner kills the process, which looks like a random failure.
Who it’s for: A small team whose build or test job dies part-way through with a killed process, an out-of-memory message or no space left on device, and whose only idea is to pay for a bigger runner.
Usually starts when: A job that used to pass now ends with the process killed, an exit code of 137, or a message that the device has no space left, usually after the project or its test data grew.
The result: The named job completes on the same runner size on three consecutive runs, with its peak memory and disk use measured and a margin left that was agreed before work starts. The change reduces what the job needs rather than moving it to a larger runner.
Check whether this job fits
Use the end of the failing log. The wording separates a memory kill, a full disk and a normal failure.
Checks you can run yourself
Read the last lines and the runner label
Open a failed run and copy the last twenty lines of the failing step and the runner label shown in the run details. Do not re-run the job just for this check.
Look for: Words such as Killed, out of memory or no space left on device, and the runner type. Remove secrets and customer details first.
What you get
- A pull request with the change and a note naming the step and the resource that ran out
- A table of peak memory and free disk by step before and after, from the run logs
- Links to the baseline failure and three passing runs
- Steps to revert the change, and a note on whether a larger runner would be the better fix
Included
- One job in one pipeline in one repository, on the runner size it uses now
- Measure peak memory and free disk across the job's steps on a branch run, and identify the step that reaches the limit
- Reduce demand with the smallest change: fewer parallel workers, a split step, cleaning large intermediate files, a smaller install or a narrower cache
- Run the job three times in a row and record the measured margins
Not included
- Buying or configuring a larger runner or changing the billing plan
- Memory leaks inside the application's own production code
- Jobs that fail with a normal error rather than being killed: on GitHub Actions that is the existing single-job repair; on GitLab CI it is not covered by a fixed-price outcome, so ask for a quote
- Self-hosted runner hardware, swap or disk administration
- Guaranteed headroom for future growth of the project
How we know it’s done
Agreed with you before work starts. Each check produces evidence you keep.
The named job completes without being killed on three consecutive runs on the same runner size, with the same steps and tests executed.
Evidence: Three run links with the full step list for each and the runner label.
Peak memory and free disk at the limiting step are recorded before and after, and the remaining margin is at least the amount agreed before work starts.
Evidence: The before and after table taken from the run logs.
No test, check or step was removed or skipped, and the job's duration stays within the bound agreed before work starts.
Evidence: The reviewed diff and the durations of the three runs.
Your authorised maintainer accepts the evidence and merges the pull request.
Evidence: Your written sign-off and the merged pull request record.
Sign-off. You inspect the three runs, the measurements and the diff, sign off in writing and merge the pull request. Payment follows sign-off.
If it fails. If the job is still killed, or the agreed margin cannot be met on this runner size, you do not pay for this fixed scope. We hand over the measurements and say whether a larger runner is the honest answer.
When it fits, and when we stop
It fits when
- You can name the job and the message or exit status it ends with
- The job can be run on a branch with the same trigger without deploying
- The runner type is known, or can be read from the run details
- Your maintainer can review and merge the change
We stop and tell you if
- The job genuinely needs more memory than the largest runner you will pay for: we say so with the measurements and stop
- The kill comes from a quota or billing limit rather than the runner's resources
- The only way to reproduce it is on a self-hosted machine we cannot reach
What could go wrong
Before merge, closing the pull request leaves your default branch unchanged. After merge, your maintainer can revert the commit and the job returns to its earlier resource use.
Scroll the table sideways to read it all.
| Risk | How we handle it |
|---|---|
| The job passes because tests were skipped or workers were cut so far that it is now too slow. | Acceptance requires the same tests to run and the duration to stay within a stated bound agreed before work starts. |
| Measurement steps or logs leak environment values. | Measurements record only totals, and the reviewer checks the logs for values before they are shared. |
| Three passes at today's size do not cover next month's growth. | We state the measured margin and its limits; growth beyond it may need a larger runner, which is your decision. |
An independent reviewer checks that the passing runs executed the full set of steps and tests, that the margin was measured rather than assumed and that no check was removed to save memory. Your maintainer reviews and merges.
How we deliver
We arrange the work and independent review, then show you the result against the agreed checks. You keep authority over your systems.
- Agree the job, the runner size, the margin target and the branch-run permission in writing
- Read the failing runs and note the step, the message and the resource named
- Add temporary measurement of memory and free disk to a branch run and record the baseline
- Reduce demand at the step that reaches the limit, without skipping tests or lowering what is checked
- Run the job three times in a row and record the margins, then remove the temporary measurement unless you ask to keep it
- Have an independent reviewer check the runs and the diff, then hand over the pull request and revert notes
This is a one-off job, not emergency cover or a subscription. We confirm eligibility, the total price, a start window and a delivery date before you accept. Work starts only after agreed inputs, secure access and necessary permissions are in place. Platform, runner and supplier charges are excluded unless the written quote includes them. No charge or booking is created by an enquiry.
Need to keep it working?
If the project keeps growing toward the limit, a monthly build-time budget can keep watching how long the pipeline takes; it does not measure memory or disk.
Ongoing work is separately scoped and quoted: no monitoring, response-time guarantee or automatic subscription is included in this job.
Explore an ongoing engineering lane, or mention the responsibility you need in your enquiry.
What you can check
This is a new service. We have not delivered this job for a client yet.
Other ways to get this done
- A larger hosted runner may be the right answer if the job genuinely needs the resources. GitHub lists the memory and storage of its standard runners, so you can compare them with the job's needs. docs.github.com
- If the job runs containers, Docker explains how the kernel's out-of-memory handling kills processes and how a container's memory limit works. docs.docker.com
Questions
Will you just move the job to a bigger runner?
No, that is your decision and your cost. We reduce what the job needs and tell you if the measurements say a bigger runner is the honest fix.
Does an exit code of 137 always mean memory?
No. It means the process was killed by a signal. Memory is one cause, so we read the measurements before naming one.
What about containers that fill the disk?
Image and layer growth can be the cause. If the job builds images, cleanup in that step is in scope; a wider image rebuild is a separate job.
How is this different from the £149 single-job repair?
The single-job repair (GitHub Actions only) fixes a job that fails with a readable error from a command. This job is for a job whose process is killed or whose disk fills, with no command error to read: we measure peak memory and disk, agree a margin in advance and accept three runs that keep to it. If your log shows a normal error from a command, use the single-job repair; on GitLab CI that case is not covered by a fixed-price outcome, so ask for a quote.
Send an enquiry
Send us
- The job name, runner label and the last lines of its log before it was killed, with secrets removed
- Whether the job passes with fewer parallel workers or fewer tests, if you have tried
- Roughly when it started failing and what grew around then
- Do not send credentials, source code or an access invitation in the first enquiry
Later, once you agree
- The pipeline file and the commands the job runs, through an authorised company-controlled route
- Read access to the named run logs
- Written approval for repeated branch runs, and the person who reviews and merges
You own the repository, the CI account and the runners. We work from authorised files and redacted logs on a branch through a company-controlled identity. Repeated runs need your written approval, and you decide whether to buy a larger runner.
Email fallback: open your mail app
If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “ci-job-killed-out-of-memory-or-disk” as the subject.