Synthetic Industry

Troubleshooting guide · updated 2026-10-11

A CI job runs for hours and is then cancelled: find what is waiting

Why a hung job shows no error, the usual waiting processes, a safe way to locate the one in your job, and what a lasting fix looks like.

Why a hang shows no error

Most CI problems fail: a command exits with an error and the log shows why. A hang is different. Some process in the job never exits, so there is nothing to fail and the only thing that ends the job is the platform's time limit. On GitHub Actions a job runs for up to 360 minutes by default, so a hung job can hold a runner for six hours. On GitLab CI, a job that passes its configured timeout plus 15 minutes is dropped and the reason is recorded; GitLab also lists a drop for a running job with no updates for 30 minutes, but its documentation does not say what counts as an update, so do not assume that a job that has gone quiet will be ended early. Either way the visible result is a cancelled job and a quiet log, which is easy to mistake for flakiness.

The usual things a job waits for

Five causes cover most cases. An interactive prompt: an installer or CLI asks a yes or no question on a terminal that does not exist in CI, and waits for an answer that never comes. A tool left in watch mode: a test runner or build tool that, by default or by a stray flag, keeps running after the work is done to watch for changes. A background process: a server started for tests that nothing stops, or a worker that keeps the step open. An unclosed handle: a database or network connection that keeps the process alive after the last test. And an unbounded wait: a loop that polls a service which has already failed and retries forever.

Be careful to separate these from a job that is waiting for a runner. That is a queue, not a hang, and the job has not started. Check that the job has a start time before you look for a waiting process.

  • Interactive prompt waiting for input
  • Tool in watch mode
  • Background server or worker that is never stopped
  • Open connection keeping the process alive
  • Unbounded wait for a service that is down

A safe way to find the one in your job

Read the cancelled log first and find the last line printed before the cancellation message. Note the time between that line and the cancellation, and check whether two hung runs stop at the same place. If they do, the step after that line, or the process that line started, is the suspect.

To confirm without waiting hours, use a branch and give the suspect step its own short time limit. GitHub lets a step carry a limit, so the step fails after a few minutes with its output intact instead of occupying the job for hours. Do not raise the job's timeout to make the problem disappear: that only moves the failure later. Check every step in the job for deploy or publish commands before any repeated run, so a test run cannot change anything real.

What a lasting fix looks like

The fix is specific to the cause: a non-interactive flag or an explicit yes, a single-run mode instead of watch mode, an explicit shutdown of the background process at the end of the step, a closed connection, and a bounded wait that fails with a readable message when its condition is not met. A short step-level limit is worth keeping as a backstop, because it turns a six-hour hang into a clear failure within minutes.

What does not fix it: retrying until the job passes, a longer timeout, or deleting the step. A pass that depends on luck still wastes runner time on the runs that hang.

How the paid outcome is accepted

The fixed job covers one job in one pipeline. It is accepted when that job completes on its own, within a duration agreed before work starts, on five consecutive runs of the same trigger, with its timeout unchanged and every previously running step still running. Five clean runs are evidence rather than proof, and the handover says so. It is £245 for one job, untested, and is paid after you sign off. On GitHub Actions, a job that fails quickly with an error belongs to the existing single-job repair instead. On GitLab CI, a job that starts and then fails is not covered by a fixed-price outcome; ask for a quote.

Sources and limits

  • GitHub workflow syntax: time limits Checked 2026-10-11.
    • A job's default time limit is 360 minutes and a step can have its own limit, with a maximum of 360 minutes.
  • GitLab job troubleshooting Checked 2026-10-11.
    • GitLab drops a pending job after one hour if no runner matches and after 24 hours if one does, a running job after 30 minutes with no updates, and a running job once its configured timeout plus 15 minutes has passed, and records a reason for each. The page does not define what counts as an update.
  • Existing GitHub Actions single-job repair Checked 2026-10-11.