Rate Limits and Retries: Backoff That Does Not Make It Worse
Token buckets, jitter, retry budgets and circuit breakers for LLM APIs — how to stay under the limit instead of discovering it, and why naive retries amplify outages.
Retry logic is the most commonly copy-pasted code in any codebase that talks to an LLM, and the most commonly wrong. The default version — catch an exception, sleep, try again — turns a two-minute upstream wobble into a self-inflicted incident, because every client in your fleet retries at the same moment.
The fix is not more retries. It is pacing that keeps you under the limit, backoff that spreads load rather than synchronising it, and a budget that stops retrying when retrying is clearly not working.
Know which limit you are hitting
LLM providers meter on several axes at once. OpenAI enforces requests per minute, requests per day, tokens per minute and tokens per day, plus image and audio limits on the relevant models, and whichever you exhaust first produces the 429.
That distinction changes the remedy entirely. Hitting a request limit means you are sending too many calls, so batching or queueing helps. Hitting a token limit means your calls are too large, so trimming context or capping output helps and sending fewer requests barely moves the needle.
The response headers tell you which, on every request, not just failures:
x-ratelimit-limit-requests: 500
x-ratelimit-remaining-requests: 499
x-ratelimit-reset-requests: 120ms
x-ratelimit-limit-tokens: 30000
x-ratelimit-remaining-tokens: 29368
x-ratelimit-reset-tokens: 1.264s
Export those two remaining counters as gauges. Watching them approach zero is how you learn about a limit before your users do, and it distinguishes "we are near capacity" from "one bad client is spraying requests".
Pace on the client, do not discover the limit
Backoff is a recovery mechanism. If it is running constantly, you are using 429s as a flow control signal, which is expensive and slow.
A token bucket on your side is a better primary control. Refill at slightly under your allowance, take a permit before each call, and let requests queue rather than fail:
import time, threading
class TokenBucket:
def __init__(self, rate_per_sec, capacity):
self.rate = rate_per_sec
self.capacity = capacity
self.tokens = capacity
self.updated = time.monotonic()
self.lock = threading.Lock()
def take(self, amount=1):
while True:
with self.lock:
now = time.monotonic()
self.tokens = min(
self.capacity,
self.tokens + (now - self.updated) * self.rate,
)
self.updated = now
if self.tokens >= amount:
self.tokens -= amount
return
deficit = (amount - self.tokens) / self.rate
time.sleep(deficit)
For token-per-minute limits, take permits equal to your estimated prompt size rather than one per request. An approximate count from a tokeniser is fine; being roughly right is enormously better than not modelling it at all.
Set the refill rate around eighty percent of your documented allowance. The headroom absorbs estimation error and the fact that other processes share the same quota.
Backoff needs jitter, not just exponents
Exponential backoff without randomisation synchronises your clients. Everyone fails at t=0, everyone retries at t=1, everyone fails again together. The retry storm looks exactly like the original outage and lasts longer.
Full jitter — sleep a random duration between zero and the current ceiling — spreads the retries across the window:
import random, time
def backoff_sleep(attempt, base=0.5, cap=30.0):
ceiling = min(cap, base * (2 ** attempt))
time.sleep(random.uniform(0, ceiling))
When the response carries Retry-After, honour it and skip your own calculation. It is the minimum number of seconds the provider wants you to wait, and it is better information than your exponent. Some deployments, including Azure OpenAI, send retry-after-ms in milliseconds instead, so read both.
Retry only what is retryable
The retryable set is small: connection errors, request timeouts, 429 rate limits, and 5xx server errors including 503 and Anthropic 529 overloads.
Everything else is a bug in your request. A 400 will be a 400 forever. A 401 will not authenticate itself. A 429 carrying insufficient_quota is a billing state, not congestion, and retrying it burns latency and produces a worse error message than the honest one.
Cap total retry time, not just attempt count. Five attempts with a thirty-second ceiling can hold a request open for well over a minute, and if that request is behind a user-facing page load you have chosen the worst possible failure — slow, then failed.
Beware stacked retry layers
This one causes real outages. The OpenAI Python SDK retries twice by default with exponential backoff and a ten-minute request timeout. If you wrap it in your own three-attempt retry decorator, you now have up to nine upstream calls per logical request, and a framework or gateway in between can multiply it again.
Pick one layer to own retries and disable the others:
client = OpenAI(max_retries=0, timeout=30.0)
Then implement retries where you can instrument them. If you cannot count retries in your metrics, you cannot tell an upstream incident from an amplification loop you built.
Retry budgets and circuit breakers
Per-request retry limits still allow unbounded aggregate retries when everything fails at once. A retry budget fixes that: allow retries to be at most some fraction of total requests — ten percent is a common starting point — and when the budget is exhausted, fail fast.
A circuit breaker does the same job at coarser grain. Track the recent failure rate for a provider; once it crosses a threshold, stop calling for a cooldown period and reject immediately. After the cooldown, let a single probe request through and only reopen on success.
Both mechanisms exist to protect a struggling upstream from you. That matters more with inference providers than with ordinary services, because capacity is genuinely scarce and a retry storm makes recovery slower for everyone sharing the pool.
Concurrency is the limit you forgot to set
Requests per minute is not the same constraint as requests in flight. A hundred concurrent streaming completions, each open for forty seconds, is a very different load profile from a hundred requests spread across a minute — and it is usually the one that exhausts connection pools and file descriptors on your side first.
Put an explicit semaphore around outbound calls. Pick the number by measuring where your own latency starts degrading, then leave it fixed. An unbounded worker pool that scales with incoming traffic is a load generator pointed at your provider.
What to do when you are simply over budget
If you are hitting limits constantly at steady state, retry tuning is not the answer — capacity is. The options are a quota increase, spreading load across providers with a fallback route, moving batchable work to an offline queue where latency does not matter, or reducing tokens per request.
Also worth checking: how much of your traffic is retries of requests nobody is waiting for any more. Cancelling upstream work when a client disconnects frees real capacity, and it is usually a smaller change than anything else on that list.
A checklist
- Export remaining-requests and remaining-tokens as metrics.
- Pace with a token bucket at about eighty percent of your allowance.
- Retry only connection errors, timeouts, 429 rate limits and 5xx.
- Use full jitter; honour
Retry-Afterwhen present. - Own retries in exactly one layer; disable SDK retries.
- Cap total retry duration, not only attempt count.
- Add a retry budget or circuit breaker before you need one.
- Bound concurrency explicitly.
Common questions
How many retries should I allow?
Two or three for background work, often zero or one for anything a user is waiting on. Cap total retry duration as well, so a slow failure does not become a very slow failure.
Why is jitter necessary if I already use exponential backoff?
Without randomisation every client retries at the same instant, so the load spikes repeat in lockstep. Jitter spreads those retries across the window and lets the upstream recover.
Do I need my own retry logic if the SDK already retries?
You need exactly one layer of it. Stacked retries multiply, so disable the SDK behaviour and handle it where you can measure it, or keep the SDK behaviour and add none of your own.