Synthetic Industry

Troubleshooting guide · updated 2026-10-11

Alert when a scheduled job does not run: heartbeat checks, grace time and their blind spots

How a heartbeat monitor inverts alerting, how to choose a schedule and grace time, what start and fail signals add, and what the check cannot see.

Inverting the alert

Most monitoring asks a question: is the site up? Is the disk full? A scheduled job needs the opposite. When it has stopped, there is nothing to ask and no error to catch; the evidence is an absence. A heartbeat monitor, sometimes called a dead-man's switch, handles this by waiting for the job to report in. The job requests a unique URL each time it runs, and the monitor raises the alarm when an expected request does not arrive.

Healthchecks.io documents the pattern clearly, and it is one vendor's description of a general idea that other tools implement. The key property is that the alert does not depend on the failing system being able to send anything. If the host is down, the disk full, or cron itself stopped, the monitor notices because the signal never came.

  • Use a monitor outside the host that runs the job.
  • Keep the monitor account under your control.

Schedule, grace time and the late state

A check needs an expected schedule, either a fixed period or the same cron-style expression the job uses, and a grace time. After the expected moment passes without a signal the check is late, and when the grace time also runs out it is down and alerts go out. Healthchecks gives the example of an hourly job with a five-minute grace: last ping at 12:00, late at 13:00, down at 13:05. For a schedule expressed in cron syntax, lateness begins at the next scheduled run time.

Choose the grace time from the job's real run-time variation, not a hopeful guess. Too short and a slow night raises a false alarm; too long and a real failure waits to be noticed. If people start ignoring alerts, the grace time or the schedule is wrong, and that is itself a finding.

  • Record each job's typical and worst run time over a month.
  • Add the monitor for the schedule the job should keep, not the one it currently keeps.

Start, success and failure signals

A bare request means success. Adding a start signal lets the monitor measure how long the job takes and treat a job that starts but never finishes as a failure, because the grace time then also limits the run. A failure signal lets the job report an explicit error immediately, rather than waiting for the grace time to pass. In practice the wrapper script does three things: signal start, run the job, then signal success only if it exited cleanly or failure if not.

The common mistake is to send the success signal unconditionally at the end of the script, so a failed job still reports success. Test both paths on purpose: run the job so that it fails, and run it so that it is skipped, and watch the alerts arrive.

  • Signal success only after a clean exit.
  • Keep the monitoring URL out of the repository and supply it through the environment.

What a heartbeat cannot see

A heartbeat tells you that a run happened and whether it exited cleanly. It does not tell you that the run did the right thing. A job that exits successfully after processing nothing will report success. Add a check inside the job for what matters, such as the number of records processed, and fail on zero where zero is wrong.

The monitor URL is also a credential of a kind: anyone with it can send fake signals, so treat it as a secret and know how to rotate it. The cron outcome configures a heartbeat in an account you own, wires start, success and failure signals, and shows a deliberately failed and a deliberately skipped run each reaching a named person. It does not cover the job's own business logic beyond what stops it running.

  • Alert on suspicious success: zero rows, empty output.
  • Rotate the check URL if it leaks.

Sources and limits

  • Healthchecks.io documentation Checked 2026-10-11.
    • The monitored job requests a URL each time it runs and the service raises an alarm when an expected check-in does not arrive; checks have a schedule and a grace time, and become late and then down.
    • A /start signal enables run-time measurement and a /fail signal reports failure explicitly; the check URL identifiers are effectively secrets because anyone holding one can send fake signals.
  • Debian crontab(5) manual page (cron 3.0pl1) Checked 2026-10-11.
    • Output from a cron job is mailed to the crontab owner unless MAILTO says otherwise.