What Should an Automation Do When an API Throttles It?
What Should an Automation Do When an API Throttles It?
Wait the exact amount of time the API told it to wait, then retry with randomised backoff, then give up loudly rather than quietly. Most broken automations we inherit fail on the third part. They retry forever, or they stop and tell nobody.
Throttling is not an error in the normal sense. The service is working perfectly. It is telling you, politely and in a documented way, that you are asking too fast. The correct response is to slow down, not to try harder.
This sounds obvious. It is also the single most common reason a working automation starts losing records three months after it was built.
What Does an API Actually Send When You Go Too Fast?
An HTTP 429 status code, and usually a header telling you when to come back. That header is Retry-After, defined in Section 10.2.3 of RFC 9110, the HTTP Semantics specification published in June 2022.
Slack's Web API documentation is a clean example. When you exceed a limit it responds with "HTTP 429 Too Many Requests" and includes a Retry-After header carrying "the number of seconds until you can retry," with the documentation's own example showing Retry-After: 30.
If a service hands you that number, honour it. Do not compute your own delay, do not halve it because you are in a hurry, and do not ignore it because your retry library has its own opinion. The service knows its own state better than your code does.
How Tight Are Real Rate Limits?
Tighter than most people building an automation assume. Two published examples make the point.
Airtable's API documentation sets a limit of "5 requests per second per base," alongside a broader ceiling of "50 requests per second for all traffic using personal access tokens from a given user or service account." Exceed it and you get a 429, and the documentation states you must wait "30 seconds before subsequent requests will succeed."
Slack's Web API is organised into tiers instead: Tier 1 at "1+ per minute," Tier 2 at "20+ per minute," Tier 3 at "50+ per minute" and Tier 4 at "100+ per minute," with some methods on their own special tier.
Five requests per second sounds generous until you loop over two thousand rows. One request per minute is a very different design constraint from one hundred. You cannot write one retry policy and assume it fits every service your automation touches.
Why Is a Fixed Retry Delay the Wrong Answer?
Because every one of your failed requests will retry at the same moment. If forty calls get throttled and all forty sleep for exactly five seconds, you have created a synchronised wave that hits the recovering service together. That is the thundering herd problem, and it can keep a service throttling you indefinitely.
The fix is to add randomness, which is a much bigger lever than it looks. The canonical reference here is Marc Brooker's post on the AWS Architecture Blog, published on 4 March 2015, which compared four strategies by simulation.
It found that plain exponential backoff performed worst, producing "substantially more work" and longer completion times than any jittered variant. Full jitter used less work than equal jitter while slightly outperforming it. Brooker's conclusion was that jittered backoff should be "a standard approach for remote clients."
Which Backoff Formula Should You Use?
Full jitter, unless you have a specific reason not to. Here are the four strategies Brooker compared, in his own notation.
| Strategy | Sleep calculation | When we reach for it |
|---|---|---|
| Exponential backoff | sleep = min(cap, base * 2 ^ attempt) | Never, for remote calls. It synchronises retries. |
| Full jitter | sleep = random(0, min(cap, base * 2 ^ attempt)) | The default. Least work, good completion times. |
| Equal jitter | sleep = min(cap, base * 2 ^ attempt) / 2 + random(0, min(cap, base * 2 ^ attempt) / 2) | When you need a guaranteed minimum wait. |
| Decorrelated jitter | sleep = min(cap, random(0, 3 * last_sleep)) | Comparable to full jitter, slightly more work. |
Pick a sensible cap so a long outage does not produce an hour long sleep, and pick a sensible base so your first retry is not instant. Then leave the formula alone.
How Many Times Should You Retry?
A small, fixed number, with a hard ceiling on total elapsed time. Five attempts over a couple of minutes is a reasonable starting point for most business automations.
Unlimited retries feel safe and are not. They turn a temporary problem into a permanent one, they burn your rate limit budget on a service that is already struggling, and on a metered API they burn money. We have inherited automations that kept retrying a hopeless call for hours, because nobody set a ceiling.
The important design decision is what happens after the last attempt. Not whether the automation stops, but where the failed item goes.
Where Should a Failed Item Actually Go?
Into a queue you can see, with enough context to replay it later. A dead letter queue, a table of failures, even a spreadsheet: the mechanism matters far less than the fact that the item still exists and someone knows it does.
Each failed record needs the payload, the timestamp, the error, and how many attempts it survived. With those four things you can replay the batch after the outage. Without them you have lost data and you will find out when a customer asks.
This is the part that separates an automation from a script. A script that fails has failed. An automation that fails has a remainder to process, and the remainder is the whole job. Our piece on stopping an automation from failing silently goes deeper into the alerting side of this, which is the other half of the same problem.
What Stops a Retry From Creating Duplicates?
An idempotency key. When you retry a request that creates something, you have no way of knowing whether the original request succeeded and only the response was lost. Without a key, your retry creates a second record.
The pattern is simple. Generate a stable identifier for the operation, from the source record rather than from the clock, and send it with every attempt. Well designed APIs will recognise a repeat and return the original result instead of doing the work twice. Stripe, among others, has supported this for years.
If the API you are calling has no idempotency support, you have to build the check yourself: look for the record before you create it. That is slower and it is still cheaper than explaining duplicate invoices.
Should You Avoid Throttling Instead of Handling It?
Yes, where you can, and it is usually easier than it sounds. Most throttling we see is self-inflicted by a design that asks for far more than it needs.
Three changes do most of the work. Use batch endpoints where the API offers them, so one call carries fifty records instead of fifty calls carrying one. Cache anything that does not change between runs. And stop polling for things the service will tell you about, which is the argument we made in webhooks versus polling.
A deliberate pause helps too. If an API allows five requests per second, sending four is free insurance. Running flat out against a documented ceiling means every small hiccup becomes a 429.
What Does a Sound Retry Policy Look Like Written Down?
Short enough to fit in a paragraph, and written down somewhere the next person will find it. Ours reads roughly like this, per integration.
Honour Retry-After whenever the service sends it. Otherwise use full jitter with a base of one second and a cap of thirty. Stop after five attempts or two minutes, whichever comes first. Send every exhausted item to the failure table with its payload and error. Alert on the first failure of a run, not on every one. Carry an idempotency key on anything that writes.
Write that once, apply it everywhere, and revisit it only when a specific integration proves it needs something different. Most teams do not need a sophisticated policy; they need a consistent one that nobody has to reinvent under pressure. Our broader notes on handling automation failures cover how this fits into a larger stack.
If you have an automation that has been quietly dropping records and you would rather not spend a week working out which ones, we are happy to take a look. You can reach us at phoenix.studio.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.