Serving a model inside a latency budget

Serving design is one question — precompute or compute now — followed by arithmetic you can do out loud.

The idea

A model in a notebook has no budget. A model behind an endpoint has two: a number of milliseconds the page will wait, and a number of cents each answer is allowed to cost. Almost every serving design is assembled from a small set of moves that trade one of those against the other.

The first decision is whether the answer can exist before the request does. That depends on three things: how many distinct inputs there are, how fresh the answer must be, and whether the features are even known ahead of time. A churn score over a bounded user list precomputes happily. A ranking that depends on what this person clicked eleven seconds ago cannot.

Everything after that is subtraction. Write the budget down, subtract the network and the feature fetch, and see what is left for the model. If the model does not fit, you move something.

Build a serving path

Toggle stages to assemble a path, then raise the traffic dial until something breaks. The budget is fixed at 200 ms .

p99 latency
—
cost / 1k requests
—
served by fallback
—
200 ms budget

presets:

How it works

Six moves, in order. The first two decide the shape; the rest decide whether it survives.

  1. Count the keys. Bounded and enumerable ahead of time — 4M users, one score a day — precomputes. Unbounded — query × user × context × time — cannot be a table, so it is computed on demand.
  2. Ask how fresh. If minutes-old is fine, precompute and cache. If the answer must reflect this session's last event, the features do not exist until the request does, and precomputing is off the table regardless of key count.
  3. Do the budget arithmetic out loud. Subtract everything that is not the model, and see what remains.
  4. If it does not fit, pick a move. A cheap candidate generator in front of an expensive scorer; a smaller distilled model; batching; or pushing the expensive step off the request path entirely and caching its output.
  5. Price it per 1,000 requests, then multiply by traffic. Cost failures arrive earlier and more quietly than latency failures.
  6. Decide how it degrades. Headroom for the spike, a bounded queue with an explicit shed policy, and a defined answer for the cache miss and for the model being down.

The budget, worked

page budget                          200 ms  (p99, not the mean)
  network + TLS                     -  20 ms
  auth + feature fetch              -  25 ms
  serialisation + render            -   8 ms
  ----------------------------------------
  left for scoring                    147 ms

single full model, p99               300 ms   ->  does not fit
----------------------------------------
retrieve 500 candidates (ANN), p99    25 ms
rerank top 50 (distilled), p99        70 ms
  ----------------------------------------
  two-stage, p99                       95 ms  ->  fits, 52 ms spare

The bill, worked

5,000,000 requests / day  ->  5,000 units of 1k

full model        $6.000 / 1k  ->  5,000 x 6.000  = $30,000 / day
two-stage         $1.350 / 1k  ->  5,000 x 1.350  =  $6,750 / day
+ cache at 90%    $0.145 / 1k  ->  5,000 x 0.145  =    $725 / day
                  ^ 0.010 lookup + 0.10 x 1.350 on the misses

The two-stage split bought the latency. The cache bought the money. They are different moves and they fail in different ways — the split loses recall, the cache loses freshness.

When to use it

request-time scoring
fits when the features only exist at request time, or the key space is effectively unbounded — search ranking, fraud on the live transaction, session-aware recommendations.
trade-off you pay the model on every single request, and the model's p99 becomes your p99. There is nowhere to hide.
precomputed and cached
fits when the key set is bounded and staleness of minutes-to-hours is acceptable — profile scores, churn risk, per-item embeddings, per-customer segment.
trade-off freshness, plus you owe a defined answer for the cold key. "Compute it inline on a miss" quietly reintroduces the full latency for the unlucky 1%.
two-stage: retrieve, then rerank
fits when the candidate space is large and the good scorer is expensive — feeds, search, ads, any top-k over millions of items.
trade-off recall lost by the cheap generator can never be recovered by the reranker. The first stage sets your ceiling; measure recall@k of stage one separately.
asynchronous scoring with a defined cache-miss answer
fits when the answer is allowed to lag the request — document enrichment, batch fraud review, content moderation with a hold state, embedding backfill.
trade-off you must say, in the design, what the request returns while the score is pending. "Nothing yet" is a product decision, not an implementation detail.

Watch out for

Worked example

The interview asks you to serve a personalised product feed: 4M items, 5M requests a day with a 4× evening peak, a 200 ms page budget, and a ranker whose good version runs 300 ms p99.

Start with the keys. The ranking depends on this session's browsing, so the features do not exist before the request — precomputing the ranking is out. But the item side does precompute: encode all 4M items offline into a vector index, nightly, and never pay for that on the request path.

Now the arithmetic. 200 minus network, features and serialisation leaves about 147 ms. The good ranker at 300 ms does not fit, so split it: an ANN retrieve of 500 candidates at 25 ms, then a distilled reranker over the top 50 at 70 ms. Ninety-five milliseconds, fifty to spare. Put a content-keyed cache in front for the non-personalised slice of the response and the bill drops from roughly $6,750 a day to under $1,000 — until the evening peak widens the working set and the hit rate falls, at which point cost rises before latency does.

Finally, degradation. Size the reranker fleet for peak plus headroom, bound its queue, and shed at admission rather than letting requests age in a queue. What the shed traffic gets is a cached popularity-and-category list — worse ranking, same page, no error. Say that number out loud in the interview: at 2× the provisioned peak, roughly a third of traffic sees the fallback ranking, and p99 stays inside budget because we shed at the door.

Check yourself

1. A fraud score must reflect the transaction happening right now; the key is card × merchant × amount × timestamp. Precompute or compute on demand?

2. Your p99 is 250 ms against a 200 ms budget. The cache hit rate is 85%. You raise it to 95%. What happens?