Standing service queue-keep-scheduled-jobs-and-queues-watched · revised 11 October 2026
Standing service
Keep your scheduled jobs and job queues watched, month after month
Named cron jobs and queues get heartbeat checks and queue alarms; a missed run, backlog or dead-lettered message reaches a named person, and up to three alerts a month are investigated.
This starts a conversation by email. Nothing is charged, and nothing is monitored, until we have agreed scope and terms with you in writing.
The responsibility you hand over
Scheduled jobs and queue workers fail without any user-facing error: a job simply does not run, a backlog grows behind a stuck worker, or messages end up in a dead-letter queue that nobody reads. The work they do is missing for days before someone notices. Alerts that come from the failing system itself are often never sent, so the check has to come from outside and has to name a person.
Who it’s for: A founder or engineering lead whose product depends on nightly jobs and queue workers that nobody checks until output goes missing.
Usually starts when: A job stopped or a queue backed up without anyone knowing, or the team has several jobs and queues and no list of who is told when one fails.
The result: The scheduled jobs and queues you name are covered by outside heartbeat checks and agreed queue alarms. A missed or failed run, or a backlog past its threshold, reaches a named person; on Amazon SQS and Sidekiq so does an oldest message past its threshold or a message in the dead-letter queue (SQS) or the dead set (Sidekiq), while Celery queues are watched on backlog length only, unless you name a broker-side dead-letter queue as a queue of its own. Up to three separate alerts a month are investigated and explained, one fix pull request a month is prepared, and each month you see what happened.
What stays true, and what we do about it
No response-time guarantee is published for this new service. A target is agreed in writing before it starts, set to what a service at this stage can keep.
Hours are agreed in writing before the service starts. At launch the service is not staffed round the clock, so we do not offer round-the-clock cover.
What must remain true
- Every named job has a check that alerts a named person if a run is missed or fails
- Every named queue has an alarm on its agreed backlog threshold that alerts the named person; each Amazon SQS and Sidekiq queue also has alarms on oldest-item age and on the dead-letter queue (SQS) or dead set (Sidekiq), while a Celery queue has backlog length only, unless you name a broker-side dead-letter queue as a queue of its own
- Each alert is investigated and ends in an explanation, with one fix pull request a month where a fix is possible
What we watch
- Heartbeat check events for each named job
- Queue size, and for Amazon SQS and Sidekiq oldest-item age and dead-letter counts, read by the alarms or the check script, without message contents
- Failed-job lists with identifiers and error text, redacted
When something happens
Scroll the table sideways to read it all.
| When | What we do |
|---|---|
| A named job misses a run or reports a failure. | We read the redacted log and schedule, reproduce on a test host where we can, and open a pull request for one fix a month or explain what only you can change. Priced as: Cron job fixed and watched |
| A job or queue is reported to have repeated its side effect. | We reproduce the repeat on staging and explain the cause. Making the job safe to repeat is the separate job linked here, bought at its own price. Priced as: Make one job idempotent |
| A queue's backlog passes its threshold or, on Amazon SQS and Sidekiq, its oldest message passes its threshold or a message reaches the dead-letter queue or dead set, and the alarm alerts the named person. | We examine the failure pattern in the redacted list, and report the cause and the options. |
| A month ends. | We send a short written summary: alerts seen, causes and what is still open. |
We do on our own
- Read heartbeat events, queue readings and redacted failure lists
- Reproduce failures on a test host or staging queue
- Open one pull request a month that changes job code, schedules or monitoring settings
We ask you first
- Redriving or deleting messages, or re-running a job that has side effects
- Changing production settings, secrets or schedules
- Adding jobs or queues beyond the agreed list
- More than three separate alerts to investigate, or more than one fix, in a month
We escalate to you when
- A dead-letter queue is filling with messages that look like lost customer work
- The cause is outside the application, such as an expired credential or a provider outage
- More than three separate alerts arrive in a month, which is more than the monthly fee covers
How you know it held. Each month the summary lists every alert seen, its cause and how it ended, and the register shows each job and queue with its last heartbeat, its latest queue reading and its threshold status, so you can compare it with your own records.
How we keep it true
This service is never finished. Each month's summary shows whether the jobs and queues stayed healthy, and it continues until you end it.
Find why a scheduled job stopped, fix it, and get an alert if it ever misses a run Job Each time it fires
One cron job that silently stopped is diagnosed and fixed, overlapping runs are guarded, and a missed-run alert is wired up and shown to reach a named person.
Up to one fix pull request a month, as agreed
A fix for a job that already has a heartbeat check is smaller than the one-off job, which also builds the check and the overlap guard. A second fix in the same month is bought as that job at its own price.
Stop one background job doing its work twice, with an idempotency check and replay test Job Optional
One background job that sometimes repeats its side effect gets a database-enforced idempotency record with a claim state and a lease, and a test that replays it through failures.
Bought separately when a repeated side effect is reported; not part of the monthly fee
Making a job safe to repeat is the job you can buy on its own. This service reproduces the repeat and explains it.
What is included, and what is not
- A register of the jobs and queues covered, with schedule, owner and thresholds
- Written settings for the heartbeat checks and queue alarms, and the check script for Sidekiq or Celery queues, with the checks and alarms created in accounts you own
- An investigation note for each of up to three separate alerts a month, and a fix pull request for one of them each month
- A monthly summary: alerts seen, causes and how each ended
Included
- Up to ten named scheduled jobs and three named queues in one application
- Heartbeat checks for each job with a schedule, a grace time and an alert route
- A backlog threshold for each queue and, for Amazon SQS and Sidekiq queues, oldest-item age and dead-letter (SQS) or dead-set (Sidekiq) thresholds, each carried by a real alarm that sends to the named person: on Amazon SQS, CloudWatch alarms in your AWS account on the queue's visible-message count and oldest-message age and on the dead-letter queue's visible-message count; on Sidekiq, a small check script on a host you control that reads queue size, queue latency and the dead set through Sidekiq's API; on Celery, the same kind of script reading the queue length from the broker, which gives backlog length only, unless you name a broker-side dead-letter queue as a queue of its own. The script reports a failure to a heartbeat check when a threshold is passed
- Investigation of up to three separate alerts a month, with a written explanation of each; alerts with one cause count as one
- Up to one fix pull request a month, for a cause an investigation finds, and a monthly written summary
Not included
- Round-the-clock paging or any response-time guarantee
- Merging code or changing production settings
- The monitoring service subscription and the cloud account charges, which are yours
- Rebuilding the job system or moving to another queue product
- Business-logic changes to the jobs beyond what an alert reveals
- More than one fix pull request, or more than three separate investigated alerts, in a month; those are quoted or bought as the one-off jobs at their own prices
How we know it’s done
Agreed with you before work starts. Each check produces evidence you keep.
For each named job, a deliberately failed run and a deliberately skipped run each raise an alert to the named person when the checks are set up.
Evidence: Alert messages with timestamps for each job.
For each named queue, lowering its backlog threshold below the current reading (or filling a test queue set up the same way) raises an alert to the named person. For an Amazon SQS queue, a test message placed in a test dead-letter queue raises the dead-letter alert; for a Sidekiq queue, a test job made to fail until it lands in the dead set of a test instance raises the dead-set alert; for a Celery queue the backlog alert is the whole test, unless you named a broker-side dead-letter queue, which is then tested as a queue of its own. The real thresholds are restored afterwards.
Evidence: Alert messages with timestamps for each queue, and the register showing the restored thresholds.
For each named queue, the register shows the agreed thresholds and the latest reading against them.
Evidence: The register, which you can compare with your own queue console.
Each month's summary lists every alert seen on the named jobs and queues and how each ended.
Evidence: The written summary, which you can compare with your own alert history.
Sign-off. You read each monthly summary and review each pull request. A fix counts as delivered when you accept it.
If it fails. If we cannot fix a failure, we say so, explain what we found and what it would take, and it does not count against the monthly allowance. If the service is not working for you, you can end it at the end of any month.
When it fits, and when we stop
It fits when
- Jobs run by cron, and queues on Sidekiq, Celery or Amazon SQS, in one application
- You create the heartbeat checks and the queue alarms in accounts you own from our written settings, or you give us a user or role limited to creating them
- Queue size, age and dead-letter counts can be read without message contents: from CloudWatch for SQS, from Sidekiq's API for Sidekiq, or from the broker's queue length for Celery, where only the length is watched
- A person on your side receives alerts and reviews pull requests
We stop and tell you if
- Alerts have no named recipient
- Most alerts have causes outside the application that you will not or cannot fix
- Message contents would have to be shared to investigate
- Alerts regularly exceed the agreed monthly number, so we agree a different scope
What could go wrong
Every fix is a pull request that your team merges, so you can revert it as you would any other change. Ending the service leaves your jobs and queues as they are; checks and thresholds live in accounts you own.
Scroll the table sideways to read it all.
| Risk | How we handle it |
|---|---|
| Alerts arrive so often that people ignore them. | Thresholds are agreed with you and tuned each month; repeated noisy alerts are reported as a finding. |
| A redrive of dead-lettered messages repeats side effects. | We never redrive without your approval, and we check the job is safe to repeat first. |
| Queue metrics expose message contents. | We ask for counts and ages only; any content that arrives is deleted and reported. |
Each pull request is checked by a reviewer separate from the work that produced it. No human supervisor is included unless your agreement names one. At launch the work is largely automated, and we say so.
Stays with a person
- You review and merge every pull request
- You approve any redrive, deletion or schedule change
Access we would need
- Heartbeat checks and queue alarms created in accounts you own, by you or through a user or role limited to them
- Queue readings without message contents
- A fork or branch for pull requests
Questions
How is this different from one cron fix?
A one-off fix repairs one job once. This watches the jobs and queues you name every month, so a missed run or a growing backlog reaches a person and does not wait for someone to notice the output is missing. One fix a month is included; more are the one-off job at its own price.
How do the queue alerts work?
On Amazon SQS they are CloudWatch alarms in your account on the queue and its dead-letter queue. On Sidekiq and Celery a small script on a host you control reads the queue and reports a failure to a heartbeat check when a threshold is passed. In both cases you create them from our written settings or give us a user or role limited to that, and we prove each one with a test. Celery queues are watched by length only.
Do you read our messages?
No. We ask for counts, ages and redacted failure lists. Message contents stay in your systems.
What if our queue is not on your list?
Say which one in your enquiry. A queue we do not cover is declined, not guessed at.
Send an enquiry
Send us
- The list of jobs with their schedules, and the queue names and system
- Which jobs and queues matter most, and who should be told
- How often things have failed in the past, as best you know
Later, once you agree
- The heartbeat checks and queue alarms created in accounts you own, either by you from our written settings or by us through a user or role limited to creating them
- A host you control where the queue check script can run, for Sidekiq or Celery queues
- A fork or branch we can open pull requests from
Your jobs, queues, hosts and secrets stay yours. The checks and alarms live in accounts you own, created by you from our written settings or by us through a user or role you limit to creating them. We read queue size, age and counts with message contents excluded, and we change code only on a branch or fork, through pull requests your team reviews and merges.
Email fallback: open your mail app
If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “queue-keep-scheduled-jobs-and-queues-watched” as the subject.