The queue does not know your job is still working
A queue cannot see inside your worker. When a worker receives a message the queue starts a timer; if the message is not deleted before the timer runs out, the queue assumes the worker died and offers the message to someone else. In Amazon SQS that timer is the visibility timeout, 30 seconds by default. A job that legitimately takes two minutes will therefore reappear twice before it finishes, and run in parallel with itself.
This is the root of many repeat-work complaints. The job is not failing; it is slow, and the queue's patience is shorter than the job. The cure is to align the timer with the real run time, not to add locks.
- Measure the real run time of the slowest legitimate job, including a bad day.
- Compare it with the visibility timeout or ack timeout on the queue.
Set the timeout and extend it with a heartbeat
SQS lets you change the timeout per queue or per message with ChangeMessageVisibility, and recommends a heartbeat that keeps extending it while processing continues, so a crashed worker still releases the message soon but a slow healthy one does not. There is a ceiling: the timeout cannot exceed 12 hours from first receipt, and extending does not reset that. Setting it far too high has its own cost, because a message whose worker really died waits that long to reappear.
Other systems have their own versions of the same trade. Celery acknowledges before running by default, so a task that has started is not run again; acks_late moves the acknowledgement to after the task, so a crash mid-task can repeat it, which is why Celery asks for idempotent tasks if you use it. Sidekiq notes that even a finished job can run again if Redis goes down before completion is recorded.
- Pick the failure you prefer: lost work or repeated work, and then design for it.
- Monitor the age of the oldest message as well as the queue depth.
Where failed messages go: dead-letter queues
A message that fails every time should not circulate forever. SQS dead-letter queues take messages that were received more than maxReceiveCount times without being deleted. Set the count high enough to allow genuine retries, since a value of 1 would send a message away after a single failure. The dead-letter queue must live in the same account and Region as its source, and AWS documents alarms that fire when messages arrive in one.
A dead-letter queue nobody reads is just a slower way to lose work. Give it an alarm that reaches a named person, and decide in advance who looks at its contents and what they do. Mind retention: for standard queues the expiry clock runs from the original enqueue time, so a message can arrive in the dead-letter queue already partly aged, and AWS recommends a longer retention period on the dead-letter queue than on the source.
- Alarm on any message entering a dead-letter queue.
- Write down what happens to a dead-lettered message: inspect, fix, redrive or discard.
Redrive carefully; what the watching service covers
Moving messages back from a dead-letter queue to the source is called a redrive. Do it only after the cause is fixed, and only for jobs that are safe to repeat; otherwise a redrive of a thousand messages is a thousand repeated side effects. If your jobs are not yet idempotent, make that the first step.
The monthly queue-watching service sets thresholds for depth, oldest-item age and dead-letter counts, routes alerts to a named person, and investigates each one, using counts and redacted failure lists rather than message contents. It never redrives or deletes without your approval. It covers Sidekiq, Celery and SQS consumers, not other brokers.
- Keep a register of queues with owner, thresholds and alert route.
- Review the dead-letter queue weekly even when no alarm has fired.
Sources and limits
- Amazon SQS: Visibility timeout Checked 2026-10-11.
- A received message stays in the queue but is invisible to other consumers for the visibility timeout, 30 seconds by default; if it is not deleted in time it becomes visible again and can be received again.
- ChangeMessageVisibility extends or shortens the timeout; the maximum is 12 hours from first receipt and extending does not reset it; a heartbeat to extend it is recommended for varying run times.
- Amazon SQS: Dead-letter queues Checked 2026-10-11.
- A redrive policy sets maxReceiveCount, the number of receives before a message moves to the dead-letter queue, which must be in the same account and Region and can be monitored with alarms.
- For standard queues expiry is based on the original enqueue time, so a dead-letter queue retention period should be longer than the source queue's.
- Celery: Tasks Checked 2026-10-11.
- By default a worker acknowledges before running a task; acks_late acknowledges after, at the price that a task may run again after a crash.
- Sidekiq wiki: Best practices Checked 2026-10-11.
- Even a finished job can run again if Redis goes down before completion is acknowledged.