The invented job and key
Everything below is invented for illustration. It is not a measurement of any real system and not a customer delivery. A job sends one confirmation email for an order, through a mail provider that has no reference field and no way to look a message up. The operation key is the invented order reference ORD-1001. A table of handled operations has a unique constraint on (job type, operation key), a status and the time the claim was made.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
job type: send_confirmation
operation key: ORD-1001
record table: handled_operations UNIQUE (job_type, operation_key)
status: claimed | completed claimed_at
lease: 10 minutes (longer than the job's longest normal run, 90 seconds)
policy for a stale claim (claimed, older than the lease): hold the order for a person
run: INSERT ... ON CONFLICT DO NOTHING RETURNING id -> a returned row: this run owns the work
send the email
UPDATE status = completed
skip: no row returned -> read the existing record:
completed -> stop, the work is done
claimed, younger than lease -> stop, another run is working on it
claimed, older than lease -> a worker may have died: apply the policyThe attempts and what must happen
Each row is one replay test. The first column is how the message arrives; the rest is what must be true afterwards.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
attempt emails record afterwards expected
1 first delivery 1 completed sent once
2 same message again after success 1 completed skipped
3 same message to two workers at once 1 completed one claims, one skips
4 different order ORD-1002, similar data 2 two, both completed both processed
5 worker dies after the claim, before the email 0 claimed held for a person after the lease
6 worker dies after the email, before completed 1 claimed held for a person after the lease
7 redelivery of case 5 or 6 inside the lease same claimed (young) skipped; no second worker takes overThe uncertain case
Cases 5 and 6 leave the same record behind, a claim older than the lease, because the database cannot see whether the email left. A queue does not remove that uncertainty. The job follows a policy written down in advance, and the test for the chosen policy is the proof that it is followed. Where the effect is a row in the same database, the claim, the row and the status change share one transaction, and cases 5 and 6 cannot happen. Where the other system takes a reference or can be looked up, the job passes the operation key and checks it. The policy choices for this email job, with the emails each sends in cases 5 and 6:
- Hold for a person: case 5 sends 0 and case 6 sends 1. The order waits for a person, so nothing is sent twice and nothing is lost silently.
- Accept a rare duplicate (send again after the lease): case 5 sends 1 and case 6 sends 2.
- Accept a rare miss (skip for good): case 5 sends 0, so the customer is never emailed, and case 6 sends 1.
- Look up by reference, only where the provider offers one: case 5 sends 1 and case 6 sends 1. This provider offers none, so it is not available here.
Limits, and the priced enquiry
This matrix is a specification, not evidence of any delivery. A queue does not become exactly-once. A row in your own database, or a call to a service that takes a reference, is done once; an email through a provider with no lookup is handled by the policy you choose. The fixed-scope idempotency job builds and runs a matrix like this for one job type on Sidekiq, Celery or an SQS consumer, starting from £395 as an untested proposal and paid after you sign off. It does not clean up duplicates that already exist.
Sources and limits
- PostgreSQL 18: INSERT, ON CONFLICT Checked 2026-10-11.
- ON CONFLICT DO NOTHING skips a conflicting row, RETURNING only returns inserted or updated rows, and the outcome is atomic under concurrency.
- Celery: Tasks Checked 2026-10-11.
- With acks_late a task may be executed more than once if a worker crashes in the middle of execution, so such tasks should be idempotent.