The invented policy
This policy is authored for the example, not recommended for any provider. First wait ceiling 0.5 seconds, doubling after each failure, capped at 8 seconds. The wait after the nth failure is the ceiling for n multiplied by a random draw between 0 and 1, supplied here so the arithmetic can be checked. If the response carries Retry-After, wait exactly that long instead. At most 5 attempts and 20 seconds of waiting in total. Only reads and writes that carry an operation identifier are retried.
- Ceiling after failure n = min(8, 0.5 x 2^(n-1)) seconds.
- Draws are invented: they stand in for a random generator.
- Attempt numbers count requests, so 5 attempts means 4 waits.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
TRACE A GET /items
attempt 1 -> 429, Retry-After: 2 wait 2.00 (header wins) elapsed wait 2.00
attempt 2 -> 503, no header ceiling 1.0, draw 0.40 -> 0.40 elapsed wait 2.40
attempt 3 -> 503, no header ceiling 2.0, draw 0.75 -> 1.50 elapsed wait 3.90
attempt 4 -> 200 result: success, 4 attempts, 3.90 s waiting
TRACE B GET /items, provider down
attempt 1 -> 503 ceiling 0.5, draw 0.90 -> 0.45 total 0.45
attempt 2 -> 503 ceiling 1.0, draw 0.20 -> 0.20 total 0.65
attempt 3 -> 503 ceiling 2.0, draw 0.60 -> 1.20 total 1.85
attempt 4 -> 503 ceiling 4.0, draw 0.50 -> 2.00 total 3.85
attempt 5 -> 503 limit reached result: failed, 5 attempts, 3.85 s waitingA date-form Retry-After and a bad request
Trace C: a 429 response whose Date header reads 10:00:00 carries Retry-After set to the date 10:00:07. The client waits seven seconds from the response's own date, then retries and succeeds. Trace D: a 400 response is not retried at all; the client returns the error with one attempt. Both are specifications to test, with the header formats documented by MDN, not observations of any provider.
- If a date is in the past, retry after the shortest wait your policy allows.
- If the wait would exceed the 20-second total, stop and flag instead of sleeping.
- 401 and 404 follow the 400 rule unless the agreement says otherwise.
A create that times out
Trace E: the client records operation op-0001 as pending, sends a create with that identifier as the idempotency or event ID, and the request times out. The policy treats the outcome as unknown and retries with the same identifier. The fake provider, which records every create it receives, should end with exactly one record for op-0001. If the provider offers no identifier, the policy does not retry: the operation is marked unknown, a person is told, and a lookup is run before anything is sent again.
- Pass: fake provider holds 1 record for op-0001 after 2 requests.
- Fail: 2 records.
- No identifier: 1 request sent, status unknown, 0 automatic retries.
Fifty callers at once, and what was exercised
Trace F: fifty simulated callers all receive 429 with no header at the same instant. The first-failure ceiling is 0.5 seconds, so with evenly spread draws, ten callers retry in each 100-millisecond bucket. A test passes if at least four of the five buckets are occupied and no bucket holds more than twenty-five; real random draws will be uneven. The arithmetic here was worked by hand and re-added; no client, provider or clock was run. The one-client retry outcome is priced from a fixed £295 as an untested test price, with payment after the agreed checks pass and you sign off.
Sources and limits
- AWS Builders' Library: timeouts, retries and backoff with jitter Checked 2026-10-11.
- Backoff needs a cap, jitter and a limit on the number of retries; APIs with side effects are not safe to retry unless they provide idempotency.
- MDN: Retry-After header Checked 2026-10-11.
- Retry-After is a number of seconds or an HTTP date.