Synthetic Industry

Collection · updated 2026-10-11

A background job or schedule misbehaves: did it run, did it run twice, did it finish?

An ordered process map for scheduled jobs and queue workers: whether the run happened, why it did not, whether it repeated, where failed messages went, and who is told, each linked to a guide and a fixed-scope job.

1. Did it run at all?

Establish when the job last produced its result and whether the scheduler or worker was alive then. For a scheduled job, check the schedule line, the account, the environment and where its output went. If it did not run, or ran and did nothing, start with the cron guide. If it ran, go to step 3.

  • Cron environment, mail, percent signs and day fields: the cron guide.
  • A host that was off at the scheduled minute does not run the job later.

2. Who would have been told?

If the answer is nobody, that is the finding. A heartbeat check that expects a signal on each run and alerts a named person when one is late closes the gap, and it works even when the failing host cannot send anything. Test it with a deliberately failed and a deliberately skipped run.

  • The heartbeat guide covers schedule, grace time and signals.
  • One job on a test host, fixed and watched: the cron job.

3. Did it run twice?

Pick one affected operation and list every effect for its reference. Seconds apart suggests a retry or a visibility window shorter than the run; hours apart a re-enqueue; deploy times a restart. Queue delivery is at least once, so the job itself must be safe to repeat. A lock is not enough; a database record keyed by the business operation is.

  • The idempotency guide explains the record and the crash case.
  • One job type with replay tests: the idempotency job.

4. Where did the failed messages go?

Check the queue timers against the real run time, the retry count and the dead-letter queue. A dead-letter queue with no alarm is a slower way to lose work. Do not redrive messages until the job is safe to repeat.

  • The queue guide covers visibility timeouts, retry counts and dead letters.
  • Keeping watch each month: the queue and schedule watching service.

5. What did the repeats already do?

Repeated side effects may already sit in your tables as duplicate rows, each with related records. Merge them only by an agreed rule, on a restored copy, repointing references first and reconciling the counts. Where several jobs and schedules in one service need the same treatment, the background-jobs project lists them as one agreed set with a closing check using deliberate failures. No change should reach production without a restore point you have tested.

  • Duplicates already in a table: the duplicate-rows guide and the merge job.
  • Several jobs at once: the reliable background jobs project.

Sources and limits