Synthetic Industry

Job cron-job-silent-failure-monitored-and-alerting · revised 11 October 2026

Find why a scheduled job stopped, fix it, and get an alert if it ever misses a run

One cron job that silently stopped is diagnosed and fixed, overlapping runs are guarded, and a missed-run alert is wired up and shown to reach a named person.

You might be seeing

  • The job's output is missing or stale and no error was ever reported
  • The command works when run by hand but not from the schedule
  • Mail from the scheduler goes to a mailbox nobody reads

No passwords, keys, card details or admin invites needed to start.

What usually happened

Cron runs a command in a minimal environment, sends its output as mail to a mailbox that may not exist, treats an unescaped percent sign in the command as a line break, and combines day-of-month and day-of-week fields in a way people often misread. A previous run that has not finished can overlap the next, and a host that was off at the scheduled minute simply skips the run. All of this fails quietly, because nothing checks that the job ran.

Who it’s for: A founder or engineering lead whose nightly or hourly job (a report, a sync, a clean-up) stopped and nobody noticed until its output was missing.

Usually starts when: Someone finds that a job has not produced its result for days or weeks, or a job is suspected of running but doing nothing, and there is no alert that would have said so.

The result: The named job runs on its schedule on a test host, overlapping runs are prevented, and a failed or missed run raises an alert to a named person within the grace time you set. We show it with one deliberately failed run and one deliberately skipped run.

Check whether this job fits

These questions point at the usual causes. Nothing here needs access to your server.

Does the command work when run by hand?
Where does the job's output or error mail go?
What runs the job?

Answer the questions to see whether this job fits.

Nothing is sent anywhere until you choose to email us.

Send an enquiry about this outcome

Checks you can run yourself

  1. Check the schedule line for the common traps

    Open the crontab of the account that runs the job, with the command read-only. Do not edit it yet.

    crontab -l

    Look for: An unescaped percent sign in the command, a relative path or a bare program name, no mail setting, and both day-of-month and day-of-week fields filled in. Send the shape of the line with secrets removed.

What you get

  • A diagnosis note naming the cause with the evidence that showed it
  • The corrected crontab entry and wrapper script as a pull request or patch
  • The heartbeat monitor settings (schedule, grace time and alert route) applied in the account you own, with the test alerts shown
  • A short runbook: what each alert means, how to re-run the job safely and how to pause the monitor

Included

  • One scheduled job on one Linux host that uses crontab
  • Find why it is not running or not working: environment, schedule fields, quoting, permissions, mail, overlap or a missing run
  • Fix the cause and add overlap protection so a second run does not start while the first is working
  • Add start, success and failure signals to an external heartbeat monitor that you create in an account you own, with an alert to a named person

Not included

  • Bugs in the job's own business logic, beyond what stops it running
  • systemd timers, Kubernetes CronJobs and cloud schedulers; each is its own scope
  • The monitoring service subscription, which is yours
  • More than one job
  • Changes to your production host except through your own review and deployment

How we know it’s done

Agreed with you before work starts. Each check produces evidence you keep.

  1. On the test host, the job runs from the schedule under a minimal environment and produces the expected output on the agreed schedule.

    Evidence: Log of at least two scheduled runs with the output produced.

  2. Starting a second run while the first is working is refused, and the refusal is logged.

    Evidence: Log of an overlapping attempt.

  3. A deliberately failed run raises an alert to the named person, and a deliberately skipped run raises a missed-run alert after the grace time.

    Evidence: Screenshots or messages from both alerts with timestamps.

  4. The runbook lets someone not involved in the work re-run the job and pause the monitor.

    Evidence: The independent reviewer's rerun note.

Sign-off. You receive the two test alerts yourself or confirm that the named person did, read the runbook, and sign off in writing. Payment follows sign-off.

If it fails. If the alerts do not reach the named person in the agreed test, you do not pay for this fixed scope and keep the findings. If the cause is outside cron, we explain what to fix and stop.

When it fits, and when we stop

It fits when

  • The job is run by cron on a Linux host and its script is available to us in a repository or as a pasted file
  • You can name the schedule it should follow and what it should produce
  • A test host or container that mimics the production user and environment can be prepared
  • You create the heartbeat monitor and the alert route in an account you own, and either invite a user limited to that one monitor's project or apply the settings from our written steps

We stop and tell you if

  • The job cannot be run on a test host without live side effects
  • The cause is a full disk, expired credential or outage elsewhere, which is reported to you rather than fixed here
  • The schedule belongs to another scheduler we are not given
  • No person can be named to receive alerts

What could go wrong

The change is a corrected schedule entry and a wrapper script. Restoring the original entry you saved returns the earlier behaviour, and the monitor can be paused or deleted in your own account. We change nothing on your production host.

Scroll the table sideways to read it all.

RiskHow we handle it
The monitor stays quiet because the signal is sent even when the job fails.Success is signalled only after the job exits cleanly, failure is signalled on any other exit, and both are tested.
A stale lock stops every later run.The lock is released when the process ends, and the runbook says how to check for a lock held by a live process.
A monitoring URL leaks and anyone can send fake signals.The URL is kept out of the repository, supplied through the environment, and the runbook says how to rotate it.

A second reviewer checks the fix against the evidence, confirms the lock is released when the job fails, and confirms that the test alerts really reached the named person. Your engineer deploys the change.

How we deliver

We arrange the work and independent review, then show you the result against the agreed checks. You keep authority over your systems.

  • Agree the job, the schedule, the expected output and the alert recipient in writing
  • Reproduce the failure on a test host under the same user and a cron-like minimal environment
  • Identify the cause from the schedule line, the environment, mail and logs, and record the evidence
  • Fix it, add a lock so overlapping runs are refused, and add start, success and failure signals
  • With the monitor you created, set the schedule and a grace time, then run one failing and one skipped run on the test host to prove the alerts
  • Independent review of the diff and the alert evidence, then hand over the runbook

This is a one-off job, not emergency cover or a subscription. We confirm eligibility, the total price, a start window and a delivery date before you accept. Work starts only after agreed inputs, secure access and any permissions are in place. Hosting and database charges are excluded unless the written quote includes them. No charge or booking is created by an enquiry.

Need to keep it working?

Discuss a monthly service that watches your scheduled jobs and queues if you have more than the one.

Ongoing work is separately scoped and quoted: no monitoring, response-time guarantee or automatic subscription is included in this job.

Explore an ongoing engineering lane, or mention the responsibility you need in your enquiry.

What you can check

This is a new service. We have not delivered this job for a client yet.

Other ways to get this done

  • The crontab manual page explains the environment cron sets, the MAILTO setting, how percent signs are treated and how the day fields combine. Your engineer can check each against your entry. manpages.debian.org
  • Healthchecks.io documents the heartbeat pattern: the job pings a URL and an alert fires when a ping is late. You can use it directly or self-host a similar monitor. healthchecks.io

Questions

Can the alert come from the server itself?

An alert sent by the thing that failed may not be sent. A heartbeat monitor outside the host notices when nothing arrives, which is why we use one.

Which monitoring service do I need?

Any heartbeat monitor you own. The runbook is written so you can move to another.

What if the job needs a bigger fix than the schedule?

If the job's own logic is broken, that is a separate bug-fix scope. We tell you what we found and stop.

Send an enquiry

Send us

  • The schedule line as written, with the command shortened to its shape and any secrets removed
  • What the job should produce, and when it last did
  • The host operating system and who owns the account that runs it
  • Do not send credentials, keys or the full script in the first enquiry

Later, once you agree

  • The wrapper script and any files it calls, through an authorised company-controlled route
  • A test host or container, and the production user's environment settings with secrets replaced
  • The heartbeat monitor you created in an account you own, with either a user limited to its project invited for us or the settings applied by you from our written steps
  • The name and contact of the person who receives alerts, and the grace time you want

You own the host, the code, the monitoring account and every credential. We work on a test host or container that you prepare, from a copy of the schedule line, the wrapper script and the production user's environment settings with secrets replaced, and we hand work back as a pull request or patch. We never ask for production passwords or a login to your production host, and we do not connect to it. You deploy the change yourself, after saving the original schedule entry.

A public HTTPS link only, without login details, query strings or fragments. No code or logs.

Sending emails your enquiry and contact address to our team through our mail provider (Resend). It is not kept in a website database. Do not send passwords, keys, recovery links, confidential code or customer records. Your contact email is unverified; nothing is ordered, charged or reserved. Privacy notice.

Email fallback: open your mail app

If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “cron-job-silent-failure-monitored-and-alerting” as the subject.