A limit the caller cannot see is just an outage with extra steps — publish the budget, and well-built clients will pace themselves.
Picture a bucket of coins. The service drips coins in at a steady rate, and the bucket has a maximum depth. Every request spends one coin: if there are coins you are served immediately, so a caller who has been quiet can fire a short burst; when the bucket is empty, you wait.
The other half — the half people forget — is telling the caller what is in the bucket. How deep it is, how many coins are left, and how long until the next one. Given those numbers a client can pace itself and never get rejected. Given nothing but a bare 429, it guesses, and guessing means retrying.
And when you do have to say no, say when to come back — with a little randomness, so that everyone you rejected does not come back at the same instant.
Both fleets have ten workers and the same fifty jobs, and each gets its own with the same settings. The cooperative fleet reads the rate-limit headers and paces itself; the naive fleet retries every 100 ms and ignores Retry-After. They share one gateway with 40 req/s of capacity. Watch who finishes first — and what each one costs.
Nothing has been sent yet. Press play.
elapsed 0.0 s · gateway capacity 40 req/s
| measure | cooperative | naive |
|---|---|---|
| jobs done | 0/50 | 0/50 |
| requests sent | 0 | 0 |
| rejected 429 | 0 | 0 |
| dropped at the edge | 0 | 0 |
| tokens in bucket | 8.0/8 | 8.0/8 |
| finished at | — | — |
last response the cooperative client parsed
waiting for the first response…
bucket depth C = 8 tokens # the burst you forgive
refill rate r = 4 tokens/second # the rate you actually sell
tokens <- min(C, tokens + r * dt) # continuous refill
serve if tokens >= 1, then tokens <- tokens - 1
from full, a caller may fire 8 at once, then settles to 4/s.
in the first 10 s it can get at most C + r*10 = 8 + 40 = 48 requests.
empty bucket, tokens = 0.25:
time to one token = (1 - 0.25) / 4 = 0.19 s
Retry-After: 1 # delay-seconds is a whole number (RFC 9110),
# so round up. never down.
HTTP/1.1 200 OK
RateLimit-Limit: 8
RateLimit-Remaining: 3
RateLimit-Reset: 2
RateLimit-Policy: 8;w=2 # 8 per 2 s -> 4/s sustained
HTTP/1.1 429 Too Many Requests
RateLimit-Remaining: 0
Retry-After: 1 # the emergency brake, not the speedometer
no jitter sleep = Retry-After 300 clients return in the same 20 ms
full jitter sleep = random(0, backoff) spread, but ignores the server's floor
best of both sleep = RA + random(0, RA) honours the floor AND spreads
backoff = min(cap, base * 2^attempt) base 0.5 s, cap 30 s
retry budget: retries <= 10% of live requests, per client, per minute
(a budget bounds the herd even when every retry looks reasonable)
| mechanism | fits when | trade-off |
|---|---|---|
| token bucket | bursty callers you want to forgive — UIs, batch jobs, webhooks replayed after a blip | needs per-key state; a full bucket can still hit a fragile downstream all at once |
| fixed window counter | cheap, coarse quotas: "10,000 per day" on a plan | up to 2× the limit across a window boundary; everyone resets on the hour |
| sliding window log | precise fairness, small key space, audit-grade counting | memory and CPU grow with request volume |
| concurrency cap | protecting a slow dependency: a database, a third-party API | bounds pressure, not total work — a caller can still consume your whole day |
| queue + | async work you own end to end, where is negotiable | the queue must be bounded and shed, or it just relocates the outage |
You own a partner export API. The published quota is 10 req/s sustained with a burst of 50. A partner's nightly job pulls 12,000 records, one request each, and fans out to 40 workers that retry immediately on failure. At 03:10 your gateway is at 200 req/s from that one API key, and other tenants start timing out.
The arithmetic settles the argument. The floor for that job is 50 + 11,950 / 10 ≈ 1,195 s, just under twenty minutes — and no amount of retrying changes it, because the bucket, not the client, sets the pace. The 190 extra requests per second buy exactly zero additional records; they buy a noisy-neighbour incident, and eventually a temporary block that makes the job take longer than twenty minutes.
The fix is symmetrical. On your side: return RateLimit-Policy: 10;w=1 and remaining on every response, 429 with Retry-After: 1, and shed the heaviest key at the edge so the other tenants are unaffected. On theirs: one shared client-side limiter across all 40 workers seeded from the remaining header, Retry-After + random(0, Retry-After) on rejection, and a retry budget capped at 10% of live traffic. Same twenty minutes, roughly 12,050 requests instead of 200,000, and nobody gets paged.
1. An incident throttles 300 clients at the same instant; each is told Retry-After: 2. What do you add to keep the queue draining?
2. Bucket depth 20, refill 5 tokens/s. A client that has been idle for a minute fires 30 requests in one instant. What happens?