Delivery is at least once, so repeats are normal
Queue systems prefer to deliver a message twice than to lose it. Sidekiq's documentation says plainly that it runs jobs at least once, not exactly once, and tells you to write jobs that are idempotent. Celery's documentation warns that if you acknowledge late, a task may run again after a worker crashes in the middle. Amazon SQS standard queues can deliver a copy of a message again and ask applications to tolerate it. A worker restart during a deploy, a timeout shorter than the real run time or a slow acknowledgement can all produce a repeat.
So the question is not how to stop the queue repeating. It is what happens when the job sees the same message again. If its work is writing a row, sending an email or calling an outside service, each repeat is a second row, a second email or a second call.
- Find the repeat's timing: seconds apart suggests a retry, hours apart a re-enqueue, deploy times a restart.
- Check whether two different producers enqueue the same operation.
Why a lock is not enough
The first fix people try is a lock around the job: take a lock named after the operation, do the work, release it. That narrows the window but does not close it. A worker can crash while holding the lock, so it expires and a later delivery proceeds, possibly after the first run had already done its work. A lock tells you the job is running now; it says nothing about whether it has already finished.
What you need is a durable record of what has been done, keyed by something that identifies the business operation, not the job. A job id changes on retry and re-enqueue. An order number, an event reference or a payment id does not. Both Sidekiq and Celery advise passing identifiers rather than objects, and re-reading current state inside the job.
- Choose the operation key with the business owner, then confirm it in writing.
- Make the key part of what the job is given, not something it computes from time.
An idempotency record with a state and a lease, enforced by the database
Create a small table with a unique constraint on the job type and the operation key, a status (claimed or completed) and the time of the claim. The job first tries to insert its key as claimed. In PostgreSQL, INSERT ... ON CONFLICT DO NOTHING either inserts or skips, atomically, even under concurrency, and RETURNING tells you whether a row was actually inserted, so the job knows whether it is the first. If it is, it does the work and then marks the record completed.
If the insert is skipped, the job reads the record. Completed means the work is done, so it stops. Claimed and recent means another run is working on it, so it stops too. Claimed and older than the lease, a time you set a little beyond the longest normal run, means a worker may have died holding the claim: without this third case the operation would be skipped for ever. For that case you apply a policy you wrote down in advance.
The ordering decides what a crash leaves behind. Where the effect is a write in the same database, do the claim, the effect and the completion in one transaction and the problem disappears. Where it is an outside call, a worker that dies after the claim may or may not have made the call. Use a reference the other system accepts, or look the result up, and when you cannot, choose deliberately: hold the operation for a person, accept a rare duplicate, or accept a rare miss. None of the three makes an email through a provider with no lookup exactly-once.
- Same database: one transaction for the effect, the record and the completion.
- Outside system: pass your operation key as its reference where it allows one, and look the result up when a claim has gone stale.
- Uncertain case: choose hold, rare duplicate or rare miss, deliberately, and test it.
What the paid job proves, and what it does not
The idempotency outcome adds this record, with its state and lease, to one job type on Sidekiq, Celery or an SQS consumer, and proves it with replay tests: the same message delivered three times in a row, delivered to two workers at once, a simulated worker death after the claim and another after the effect, and two distinct operations to show it does not skip real work. Tests are run on the old code too, so you see them fail first.
It does not make the queue exactly-once, clean up duplicates already created, or make an outside system safe when it offers no reference and no lookup. For that last case the handover states the policy you chose and what can be reconciled afterwards. Nothing runs against your production data.
- Ask for a query that finds past repeats, then decide separately what to do about them.
- Set a retention period for the record table and confirm it.
Sources and limits
- Sidekiq wiki: Best practices Checked 2026-10-11.
- Sidekiq promises at-least-once rather than exactly-once execution, so jobs should be idempotent and transactional, even a job which has completed can be run again if Redis goes down before the acknowledgement, and arguments should be simple identifiers rather than objects.
- Celery: Tasks Checked 2026-10-11.
- With acks_late a task may be executed more than once if a worker crashes mid-execution, so such tasks should be idempotent.
- A task queued inside a database transaction can start before the commit; delay_on_commit (Celery 5.4) sends it after the commit.
- Amazon SQS: At-least-once delivery Checked 2026-10-11.
- Standard queues can deliver a message copy again, and applications should be idempotent.
- PostgreSQL 18: INSERT, ON CONFLICT Checked 2026-10-11.
- ON CONFLICT DO NOTHING skips a conflicting row and DO UPDATE guarantees an atomic insert-or-update outcome under concurrency; RETURNING returns only rows actually inserted or updated.