Rate limits a good client can cooperate with

A limit the caller cannot see is just an outage with extra steps — publish the budget, and well-built clients will pace themselves.

The idea

Picture a bucket of coins. The service drips coins in at a steady rate, and the bucket has a maximum depth. Every request spends one coin: if there are coins you are served immediately, so a caller who has been quiet can fire a short burst; when the bucket is empty, you wait.

The other half — the half people forget — is telling the caller what is in the bucket. How deep it is, how many coins are left, and how long until the next one. Given those numbers a client can pace itself and never get rejected. Given nothing but a bare 429, it guesses, and guessing means retrying.

And when you do have to say no, say when to come back — with a little randomness, so that everyone you rejected does not come back at the same instant.

The race: two fleets, one quota each

Both fleets have ten workers and the same fifty jobs, and each gets its own with the same settings. The cooperative fleet reads the rate-limit headers and paces itself; the naive fleet retries every 100 ms and ignores Retry-After. They share one gateway with 40 req/s of capacity. Watch who finishes first — and what each one costs.

served (200) rejected (429) dropped at the edge

Nothing has been sent yet. Press play.

elapsed 0.0 s · gateway capacity 40 req/s

measurecooperativenaive
jobs done0/500/50
requests sent00
rejected 42900
dropped at the edge00
tokens in bucket8.0/88.0/8
finished at——

last response the cooperative client parsed

waiting for the first response…

How it works

  1. Pick the sustained rate first. r is what you can serve all day without hurting anything downstream. Then pick the depth C — how much of a burst you are willing to forgive from a caller who has been quiet.
  2. Refill continuously, charge one token per request. No window boundaries, no cliff edges.
  3. Publish the budget on every response, not only on rejections. A client that only learns the limit by being rejected has to discover it by overshooting it.
  4. When you reject, say when to come back. 429 plus Retry-After, in whole seconds, rounded up.
  5. A good client mirrors the counter. It keeps its own bucket seeded from RateLimit-Remaining and spends at the published rate, so 429 becomes an exception rather than a control loop.
  6. every wait. Randomness is what turns a synchronised wave into a queue that drains.
  7. Make rejection cheap and attributable. Shed the heaviest caller at the edge so polite callers do not pay for a noisy one.
bucket depth   C = 8 tokens          # the burst you forgive
refill rate    r = 4 tokens/second   # the rate you actually sell

tokens <- min(C, tokens + r * dt)         # continuous refill
serve if tokens >= 1, then tokens <- tokens - 1

from full, a caller may fire 8 at once, then settles to 4/s.
in the first 10 s it can get at most  C + r*10 = 8 + 40 = 48 requests.

empty bucket, tokens = 0.25:
  time to one token = (1 - 0.25) / 4 = 0.19 s
  Retry-After: 1     # delay-seconds is a whole number (RFC 9110),
                     # so round up. never down.

The headers that let a client pace itself

HTTP/1.1 200 OK
RateLimit-Limit: 8
RateLimit-Remaining: 3
RateLimit-Reset: 2
RateLimit-Policy: 8;w=2        # 8 per 2 s  ->  4/s sustained

HTTP/1.1 429 Too Many Requests
RateLimit-Remaining: 0
Retry-After: 1                 # the emergency brake, not the speedometer

Why jitter, and how much

no jitter      sleep = Retry-After            300 clients return in the same 20 ms
full jitter    sleep = random(0, backoff)     spread, but ignores the server's floor
best of both   sleep = RA + random(0, RA)     honours the floor AND spreads

backoff = min(cap, base * 2^attempt)     base 0.5 s, cap 30 s
retry budget: retries <= 10% of live requests, per client, per minute
              (a budget bounds the herd even when every retry looks reasonable)

When to use it

mechanismfits whentrade-off
token bucketbursty callers you want to forgive — UIs, batch jobs, webhooks replayed after a blipneeds per-key state; a full bucket can still hit a fragile downstream all at once
fixed window countercheap, coarse quotas: "10,000 per day" on a planup to 2× the limit across a window boundary; everyone resets on the hour
sliding window logprecise fairness, small key space, audit-grade countingmemory and CPU grow with request volume
concurrency capprotecting a slow dependency: a database, a third-party APIbounds pressure, not total work — a caller can still consume your whole day
queue + async work you own end to end, where is negotiablethe queue must be bounded and shed, or it just relocates the outage

Watch out for

Worked example

You own a partner export API. The published quota is 10 req/s sustained with a burst of 50. A partner's nightly job pulls 12,000 records, one request each, and fans out to 40 workers that retry immediately on failure. At 03:10 your gateway is at 200 req/s from that one API key, and other tenants start timing out.

The arithmetic settles the argument. The floor for that job is 50 + 11,950 / 10 ≈ 1,195 s, just under twenty minutes — and no amount of retrying changes it, because the bucket, not the client, sets the pace. The 190 extra requests per second buy exactly zero additional records; they buy a noisy-neighbour incident, and eventually a temporary block that makes the job take longer than twenty minutes.

The fix is symmetrical. On your side: return RateLimit-Policy: 10;w=1 and remaining on every response, 429 with Retry-After: 1, and shed the heaviest key at the edge so the other tenants are unaffected. On theirs: one shared client-side limiter across all 40 workers seeded from the remaining header, Retry-After + random(0, Retry-After) on rejection, and a retry budget capped at 10% of live traffic. Same twenty minutes, roughly 12,050 requests instead of 200,000, and nobody gets paged.

Check yourself

1. An incident throttles 300 clients at the same instant; each is told Retry-After: 2. What do you add to keep the queue draining?

2. Bucket depth 20, refill 5 tokens/s. A client that has been idle for a minute fires 30 requests in one instant. What happens?