Rate limits will find you: backpressure for LLM APIs
429s aren't an edge case at scale, they're a Tuesday. Here's how to design queueing, concurrency caps, and retries so a rate limit degrades gracefully instead of cascading.
Most teams discover their rate limit strategy the day it fails: a marketing campaign
drives traffic, a batch job kicks off at the same time as peak usage, and suddenly every
request is getting 429s. The naive fix — retry immediately — turns a momentary limit
into an outage, because every failed request just adds another retry on top of an
already-saturated queue. Rate limits aren’t a bug in the provider’s API. They’re a
constraint you have to design for, the same way you design for disk space or database
connections.
The failure mode: retry storms
A retry storm looks like this:
- Traffic exceeds your requests-per-minute (RPM) or tokens-per-minute (TPM) limit.
- The provider returns
429. - Your code retries immediately, adding to the same second’s request volume.
- The retry also gets rate-limited, so it retries again.
- Load compounds until every client is retrying, and the limit is nowhere near clearing.
This is the same shape as a database connection pool getting exhausted, and it needs the same fix: don’t let failures generate more load than the requests that caused them.
Three layers that actually work
1. A concurrency cap in front of the provider, not just a retry policy
Before you touch retries, cap how many in-flight requests you allow to a given model/provider at once. A semaphore sized to roughly 70-80% of your known RPM/TPM limit keeps you under the ceiling most of the time, so you’re handling rate limits as an exception, not a steady-state condition.
const limit = pLimit(20) // max 20 concurrent calls to this model
const results = await Promise.all(
requests.map(req => limit(() => callModel(req)))
)
This alone eliminates most rate-limit errors, because you stop sending more concurrent work than the provider will accept.
2. Exponential backoff with jitter, and respect Retry-After
When you do get a 429, back off — and randomize the wait so a batch of clients doesn’t
resync and retry in lockstep:
async function callWithBackoff(fn, attempt = 0) {
try {
return await fn()
} catch (err) {
if (err.status !== 429 || attempt >= 5) throw err
const retryAfter = err.headers?.['retry-after']
const base = retryAfter ? Number(retryAfter) * 1000 : 500 * 2 ** attempt
const delay = base + Math.random() * base * 0.25
await sleep(delay)
return callWithBackoff(fn, attempt + 1)
}
}
Most providers (OpenAI, Anthropic) return a retry-after header or equivalent — use it
instead of guessing. A capped attempt count matters too: infinite retries just move the
outage later and make it invisible until a queue backs up.
3. A queue with priority, not a flat FIFO
Not all requests are equal. An interactive chat response blocking a user is not the same as a nightly embeddings backfill. When you’re near the limit, a flat queue lets the backfill starve the user-facing request. Two queues — interactive and batch — with the interactive queue always drained first, keeps degradation proportional to what users actually notice.
flowchart LR
A[Interactive request] --> C{Concurrency gate}
B[Batch/background request] --> C
C -->|priority: interactive first| D[Provider API]
D -->|429| E[Backoff + jitter]
E --> C
D -->|200| F[Response]
What to actually monitor
Rate limiting is invisible until it isn’t, so track it before you need to:
- 429 rate as a percentage of total calls, per model and per provider — a rising trend is your early warning, not the eventual outage.
- Queue depth and wait time for both interactive and batch lanes — this tells you whether users are feeling it yet.
- Retry count distribution — if most successful calls take 2+ retries, your concurrency cap is set too high for current traffic.
The part teams skip: capacity planning per provider, not per app
If you call multiple models or providers, your rate limits are per-provider-account, not per-feature. A new feature that adds calls to the same underlying account eats into the same TPM budget as everything else already running. Track headroom at the account level and treat “add a new LLM-backed feature” the same way you’d treat “add a new heavy database query” — check the budget before you ship, not after the first traffic spike.
Rate limits are a resource constraint like any other. Teams that treat them as an occasional error to retry around get outages. Teams that build a concurrency gate, real backoff, and priority queueing get a system that slows down gracefully under load and comes back on its own — which is the actual bar for production reliability.
Want something like this built for your team?
Get a quote →