Synthetic Industry

Job queue-duplicate-job-runs-made-idempotent · revised 11 October 2026

Stop one background job doing its work twice, with an idempotency check and replay test

One background job that sometimes repeats its side effect gets a database-enforced idempotency record with a claim state and a lease, and a test that replays it through failures.

You might be seeing

  • The same email, record or API call happens twice for one business event
  • The repeats cluster around deploys, worker restarts or slow periods
  • Adding a lock reduced the repeats but did not remove them

No passwords, keys, card details or admin invites needed to start.

What usually happened

Common queue systems deliver at least once: a message can be delivered again after a worker crash, a timeout or a slow acknowledgement. If the job's work is not safe to repeat, every redelivery becomes a duplicate side effect. A lock only narrows the window. What is missing is a record of what has been done, keyed by a stable identifier of the business operation and enforced by the database, with a state that says whether the work was only claimed or finished, and a rule for the case where a worker died holding a claim and nobody can tell whether the effect happened.

Who it’s for: An engineering lead whose queue-driven job occasionally repeats an email, record or call, and who cannot say how often or why.

Usually starts when: Two rows or two calls appear for one event, a customer receives two notifications, or a downstream system reports a repeat, and retries or a deploy restart is suspected.

The result: For the named job, the side effect happens once for the same business operation however many times the message is delivered, where the effect is a row written in your own database or a call to an outside service that accepts a reference or can be looked up. For any other effect, such as an email from a provider with no reference or lookup, the job follows a written, tested policy for the case where it cannot tell whether the effect happened; that policy limits the harm but does not make the effect once-only. A replay test covers duplicate delivery, concurrent delivery and a crash at each step.

Check whether this job fits

These questions find out whether a repeat can be recognised and recorded. No code or data is needed.

Is there a stable reference for one business operation, such as an order or event number?
What does the job do that must not repeat?
Which queue system do you use?

Answer the questions to see whether this job fits.

Nothing is sent anywhere until you choose to email us.

Send an enquiry about this outcome

Checks you can run yourself

  1. Look for the repeat in your own data

    Pick one affected operation and list every row, log line or message sent for its reference. Do this on a read-only copy or in your logging tool.

    Look for: Whether the repeats come close together (a retry), far apart (a re-enqueue) or at deploy times. Send the timing pattern, not the data.

What you get

  • A pull request with the idempotency record (key, state, claimed-at and lease), its migration and the changed job
  • The replay test suite and its results on the old and new code
  • A one-page note on the operation key, what counts as a repeat, the lease length and the written policy for a stale claim
  • A query that lists past repeats in your data so you can decide separately what to do about them

Included

  • One job type on one queue system (Sidekiq, Celery or an Amazon SQS consumer) in one application
  • Define the stable identifier of the business operation and add an idempotency record with a unique constraint on job type and operation key, a state (claimed or completed), the time it was claimed and a lease length
  • Make the job claim first, act, then mark the record completed; for an effect that is a row in your own database, the claim, the effect and the completion share one transaction
  • A written policy for a stale claim (older than the lease, so a worker may have died holding it): look the effect up by its reference, hold the operation for a person, or accept a rare duplicate or a rare miss, as you choose
  • Replay tests for duplicate delivery, concurrent delivery, a worker that dies after the claim and one that dies after the effect

Not included

  • Cleaning up duplicates that already exist; that is a separate reconciliation job
  • An exactly-once guarantee for an outside effect that offers no reference or lookup, such as an email through most mail providers; the note states the policy and what can be reconciled
  • Creating invoices in Xero or QuickBooks Online, or receiving signed webhooks: those have their own jobs
  • Changing the queue infrastructure, broker or hosting
  • More than one job type
  • Running anything against production data

How we know it’s done

Agreed with you before work starts. Each check produces evidence you keep.

  1. Delivering the same message three times, one after another, produces the side effect once and leaves one idempotency record in the completed state.

    Evidence: Test output showing the effect count and the record count and state.

  2. Delivering the same message concurrently to two workers produces the side effect once.

    Evidence: Concurrent test output with the effect count over at least 100 runs.

  3. With a simulated worker death after the claim and before the effect, and another after the effect and before completion, a redelivery after the lease follows the written policy: for a row in your database or an outside call that accepts a reference or can be looked up, the effect happens exactly once; for any other effect, the count of effects is what the policy states in advance (a held operation, or at most one extra or one missing effect).

    Evidence: Test output for both crash points with the effect counts the policy permits written beforehand, and the policy text.

  4. A claim younger than the lease is not taken over by a second worker, and a claim older than the lease is handled by the policy rather than skipped for ever.

    Evidence: Test output with the claim age, the lease and the result for each case.

  5. Two different business operations with similar data are both processed once.

    Evidence: Test output showing two effects for two operations.

Sign-off. You run the replay tests, read the policy for the uncertain case, then sign off in writing and merge. Payment follows sign-off.

If it fails. If the replay tests show the effect happening more than once for the same operation where this job promises once, or a stale claim handled against its written policy, you do not pay for this fixed scope and keep the findings. If the outside system cannot be made safe from our side and no policy is acceptable to you, we say so and stop.

When it fits, and when we stop

It fits when

  • The job runs on Sidekiq, Celery or an Amazon SQS consumer and its code is in a repository we are given access to
  • A stable identifier for the business operation exists or can be defined, such as an order or event reference
  • The application has a relational database where an idempotency record with a unique constraint can be added
  • The job can run on staging with a test queue and synthetic data

We stop and tell you if

  • No stable identifier can be defined for the operation
  • The side effect is in an outside system with no reference, no lookup and no way to detect a repeat, and you cannot choose a policy for the uncertain case (hold for a person, accept a rare duplicate or accept a rare miss)
  • The job's duplicates come from two different enqueuers that disagree on what an operation is, which needs a business decision first
  • The job cannot be run on staging without live external calls

What could go wrong

The change is a pull request with a migration for the idempotency table. Reverting the commit restores the old job; the new table can be left in place or dropped by a separate migration. No production data is touched by us.

Scroll the table sideways to read it all.

RiskHow we handle it
A worker dies holding a claim, so a later delivery cannot tell whether the effect happened and either repeats it or never sends it.The record carries a state and a claimed-at time. A claim older than the lease is handled by the written policy: lookup by reference, hold for a person, or accept a rare duplicate or miss. A replay test covers each crash point.
The operation key is wrong and distinct operations are skipped as repeats.The tests include two distinct operations with similar data, and you confirm the key definition in writing.
The idempotency table grows without bound.The note proposes a retention period for you to confirm; cleanup is not applied by us.

A second reviewer checks the order of claim, effect and completion and the stale-claim policy for each crash point, confirms the unique constraint is what enforces the guard rather than an in-memory check, and reads the replay results. Your engineer reviews and merges.

How we deliver

We arrange the work and independent review, then show you the result against the agreed checks. You keep authority over your systems.

  • Agree the job, the operation reference, what counts as a repeat, the lease length and the outside-call behaviour in writing
  • Reproduce a repeat on staging: duplicate delivery, then concurrent delivery, then a worker that dies after the claim and one that dies after the effect
  • Add the idempotency record with a unique constraint, a claimed or completed state and a lease, and make the job claim, act and complete in that order; keep the claim, effect and completion in one transaction where the effect is a row in your database
  • Write down and implement the policy for a stale claim: look the effect up by reference where the other system allows it, otherwise hold for a person or accept a rare duplicate or miss, as you chose
  • Re-run every replay case on the new code, and keep the failing run on the old code for comparison
  • Independent review of the diff and test runs, then hand over the pull request and the note

This is a one-off job, not emergency cover or a subscription. We confirm eligibility, the total price, a start window and a delivery date before you accept. Work starts only after agreed inputs, secure access and any permissions are in place. Hosting and database charges are excluded unless the written quote includes them. No charge or booking is created by an enquiry.

Need to keep it working?

Discuss a project covering several job types, or a monthly review of queue health and failed jobs.

Ongoing work is separately scoped and quoted: no monitoring, response-time guarantee or automatic subscription is included in this job.

Explore an ongoing engineering lane, or mention the responsibility you need in your enquiry.

What you can check

This is a new service. We have not delivered this job for a client yet.

Other ways to get this done

  • Sidekiq's best-practice page says jobs run at least once, not exactly once, and should be idempotent and transactional. github.com
  • Celery documents that acknowledging late means a task may run again if a worker crashes mid-task, so such tasks should be idempotent. docs.celeryq.dev
  • Amazon SQS documents at-least-once delivery for standard queues and advises idempotent consumers. docs.aws.amazon.com

Questions

Would a lock around the job be enough?

A lock narrows the window but a lock holder can crash and a lock can expire. The unique record in the database is what enforces once.

Can you clean up the duplicates we already have?

This job gives you a query that finds past repeats. Cleaning them up is a separate reconciliation job that depends on what each repeat caused.

Does this make the queue exactly-once?

No queue guarantee changes. The job is made safe to repeat. A row in your own database, or a call to a service that takes a reference, is done once. For an effect that cannot be checked afterwards, such as most emails, we cannot promise once; we write down and test what the job does when it cannot tell, and you choose between holding for a person, a rare duplicate and a rare miss.

What if the job creates invoices in Xero or QuickBooks, or receives webhooks?

Those have their own fixed-scope jobs with their own tests. This one is for a queue job in your application.

Send an enquiry

Send us

  • The queue system and version, the job name and what its side effect is, in plain words
  • How the repeats were noticed and roughly how often
  • What identifies one business operation (for example an order number), without sending any values
  • Do not send credentials, repository access or customer data in the first enquiry

Later, once you agree

  • The job code and queue configuration through an authorised company-controlled repository route
  • A staging queue and database with synthetic data
  • Details of the outside call the job makes, and whether it accepts a reference or can be queried
  • The longest normal run time of the job, to set the lease, and your choice of policy for a stale claim
  • The person who reviews and merges the pull request

You own the database, the code and every credential. We work on a copy you prepare, such as a staging database restored from a backup with personal data removed or replaced, and we hand work back as a pull request or a script. We never ask for production passwords, and we do not connect to your production database. You apply any change to production yourself, after a restore point that you have tested.

A public HTTPS link only, without login details, query strings or fragments. No code or logs.

Sending emails your enquiry and contact address to our team through our mail provider (Resend). It is not kept in a website database. Do not send passwords, keys, recovery links, confidential code or customer records. Your contact email is unverified; nothing is ordered, charged or reserved. Privacy notice.

Email fallback: open your mail app

If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “queue-duplicate-job-runs-made-idempotent” as the subject.