Synthetic Industry

Troubleshooting guide · updated 2026-10-11

Rate limited or timing out: wait as the provider asks, then back off with spread

Read Retry-After in both of its forms, add capped exponential backoff with randomness, bound the attempts and total time, and test the policy with a scripted fake provider.

Why blind retrying makes it worse

An engineer at AWS describes retries as selfish: the client spends the server's capacity to improve its own odds. When a dependency is overloaded, retries add load exactly when it can least take it, and when layers each retry, the multiplication is steep: the article's example of three retries at each of five layers lets one failing call become 243 calls at the bottom. So a good policy retries rarely, spreads its retries, and gives up.

  • Retry in one layer, not at every layer.
  • Give every remote call a timeout.
  • Cap the number of attempts and the total time.

Do what the provider says first

A 429 response means too many requests in a given time and may carry a Retry-After header. That header is either a number of seconds or an HTTP date, and 503 responses use it too. Slack's rate-limit page tells callers to wait the stated seconds; Pipedrive rejects with 429 once its daily token budget is used and the wait may be many hours, so the right behaviour there is to park the work rather than hold a worker. When no header is present, fall back on your own backoff.

  • Convert a date to seconds from the response's own date or your clock.
  • Never wait less than the stated time.
  • If the wait would exceed the total budget, stop and flag the job instead of sleeping.

Back off with jitter and a cap

Exponential backoff doubles the wait after each failure up to a ceiling. If every caller computes the same waits they will all return at the same moment, which is why the article calls for jitter, a random spread over the wait. A simple form is to pick a random wait between zero and the current ceiling. Add a limit on the number of retries and, ideally, a local budget so a sustained failure stops generating retries instead of sustaining them.

  • Choose the first wait, the multiplier, the ceiling and the attempt count in writing.
  • Log each attempt's number, wait and result without bodies or credentials.
  • Use fixed seeds or injected randomness so tests are repeatable.

What not to retry

Validation errors will fail again, expired credentials will fail again, and a missing resource usually stays missing. Retrying them wastes budget and can trigger blocks. Google's calendar error guide, for instance, recommends backoff for rate-limit and server errors but treats a 400 as permanent. Writes need separate treatment because a timed-out write may have succeeded; the companion guide covers that.

  • Retry: 429, 503 and connection errors, with limits.
  • Do not retry: 400, 401, 403 for permission, 404 unless you know why.
  • Surface the final failure with its attempt count.

Test with a fake provider

Script a fake provider that returns a chosen sequence of responses and records the time of each request. Test: two 429s with a two-second wait then success; endless 503s ending at the attempt limit; a 400 that is not retried; fifty callers failing together whose retries are spread across the window. A worked trace of such a schedule is in the example linked below.

How the paid outcome is accepted

The retry hardening outcome for one named client is accepted when those scripted sequences pass in your own test suite. The fixed £295 test price is untested and payment follows your sign-off. Raising a provider quota and redesigning your queue are outside it.

Sources and limits

  • AWS Builders' Library: timeouts, retries and backoff with jitter Checked 2026-10-11.
    • Retries amplify load: three retries at each of five layers can multiply load on the database 243 times.
    • Exponential backoff needs a cap and jitter; a cap alone leaves callers at the same rate, so the number of retries should also be limited, for example locally with a token bucket.
    • A timeout should be set on every remote call.
  • MDN: Retry-After header Checked 2026-10-11.
    • Retry-After is either an HTTP date or a non-negative number of seconds, and is used with 503, 429 and redirects.
  • RFC 6585: 429 Too Many Requests Checked 2026-10-11.
    • 429 signals too many requests in a given time and may carry Retry-After; the response must not be cached.
  • Pipedrive API rate limiting Checked 2026-10-11.
    • A daily token budget resets at midnight server time and requests are rejected with 429 once it is exhausted; heavy traffic after 429s with an API token can be blocked with a 403.
  • Slack Web API rate limits Checked 2026-10-11.
    • A 429 response carries Retry-After in seconds.