Serving design is one question — precompute or compute now — followed by arithmetic you can do out loud.
The idea
A model in a notebook has no budget. A model behind an endpoint has two: a number of milliseconds the page will wait, and a number of cents each answer is allowed to cost. Almost every serving design is assembled from a small set of moves that trade one of those against the other.
The first decision is whether the answer can exist before the request does. That depends on three things: how many distinct inputs there are, how fresh the answer must be, and whether the features are even known ahead of time. A churn score over a bounded user list precomputes happily. A ranking that depends on what this person clicked eleven seconds ago cannot.
Everything after that is subtraction. Write the budget down, subtract the network and the feature fetch, and see what is left for the model. If the model does not fit, you move something.
Build a serving path
Toggle stages to assemble a path, then raise the traffic dial until something breaks. The budget is fixed at 200 ms .
p99 latency
—
cost / 1k requests
—
served by fallback
—
200 ms budget
presets:
How it works
Six moves, in order. The first two decide the shape; the rest decide whether it survives.
Count the keys. Bounded and enumerable ahead of time — 4M users, one score a day — precomputes. Unbounded — query × user × context × time — cannot be a table, so it is computed on demand.
Ask how fresh. If minutes-old is fine, precompute and cache. If the answer must reflect this session's last event, the features do not exist until the request does, and precomputing is off the table regardless of key count.
Do the budget arithmetic out loud. Subtract everything that is not the model, and see what remains.
If it does not fit, pick a move. A cheap candidate generator in front of an expensive scorer; a smaller distilled model; batching; or pushing the expensive step off the request path entirely and caching its output.
Price it per 1,000 requests, then multiply by traffic. Cost failures arrive earlier and more quietly than latency failures.
Decide how it degrades. Headroom for the spike, a bounded queue with an explicit shed policy, and a defined answer for the cache miss and for the model being down.
The budget, worked
page budget 200 ms (p99, not the mean)
network + TLS - 20 ms
auth + feature fetch - 25 ms
serialisation + render - 8 ms
----------------------------------------
left for scoring 147 ms
single full model, p99 300 ms -> does not fit
----------------------------------------
retrieve 500 candidates (ANN), p99 25 ms
rerank top 50 (distilled), p99 70 ms
----------------------------------------
two-stage, p99 95 ms -> fits, 52 ms spare
The bill, worked
5,000,000 requests / day -> 5,000 units of 1k
full model $6.000 / 1k -> 5,000 x 6.000 = $30,000 / day
two-stage $1.350 / 1k -> 5,000 x 1.350 = $6,750 / day
+ cache at 90% $0.145 / 1k -> 5,000 x 0.145 = $725 / day
^ 0.010 lookup + 0.10 x 1.350 on the misses
The two-stage split bought the latency. The cache bought the money. They are different moves and they fail in different ways — the split loses recall, the cache loses freshness.
When to use it
request-time scoring
fits when the features only exist at request time, or the key space is effectively unbounded — search ranking, fraud on the live transaction, session-aware recommendations.
trade-off you pay the model on every single request, and the model's p99 becomes your p99. There is nowhere to hide.
precomputed and cached
fits when the key set is bounded and staleness of minutes-to-hours is acceptable — profile scores, churn risk, per-item embeddings, per-customer segment.
trade-off freshness, plus you owe a defined answer for the cold key. "Compute it inline on a miss" quietly reintroduces the full latency for the unlucky 1%.
two-stage: retrieve, then rerank
fits when the candidate space is large and the good scorer is expensive — feeds, search, ads, any top-k over millions of items.
trade-off recall lost by the cheap generator can never be recovered by the reranker. The first stage sets your ceiling; measure recall@k of stage one separately.
asynchronous scoring with a defined cache-miss answer
fits when the answer is allowed to lag the request — document enrichment, batch fraud review, content moderation with a hold state, embedding backfill.
trade-off you must say, in the design, what the request returns while the score is pending. "Nothing yet" is a product decision, not an implementation detail.
Watch out for
Budgeting the mean and shipping the tail. A 60 ms mean with a 300 ms p99 blows a 200 ms page for one user in a hundred — and that user retries. Related trap: a cache does not move your p99 until the hit rate clears 99%. Below that, the 99th percentile still sits on the miss path. The cache is a cost lever first, a latency lever second.
A cache key nobody can hit twice. Keying on user id when the answer depends only on the content gives every user their own private miss. Key by content hash or feature vector and one computation serves everyone who asks the same question.
Precomputing a product instead of a list. A score per user over 4M users is 4M rows. A score per query × user × device × hour is not a table you can hold — it is a cross-product that grows faster than your storage plan. Count the keys before promising a nightly job.
the expensive modality on the request path. Image and video encoding runs roughly an order of magnitude dearer than text for the same request. Encode once offline, store the vector, and let the request path do cheap arithmetic on it. The same instinct kills "one small fine-tuned model per customer" — that fails on money and GPU memory long before latency; use one shared base with per-customer adapters or embeddings.
No headroom, no shed policy, no answer when the model is down. Provision at the mean and the spike queues; the queue causes timeouts; timeouts cause retries; retries are the spike again. Size for peak with headroom, bound the queue, shed explicitly and fast, and make the shed path return a simpler model or a cached heuristic rather than a 500. Then say out loud what degrades first.
Discovering the privacy constraint after the architecture. On-device or in-region inference is not a deployment detail — it caps model size, rules out the hosted API, and changes which of the shapes above is even available. Ask where the data is allowed to be computed before you pick the model.
Worked example
The interview asks you to serve a personalised product feed: 4M items, 5M requests a day with a 4× evening peak, a 200 ms page budget, and a ranker whose good version runs 300 ms p99.
Start with the keys. The ranking depends on this session's browsing, so the features do not exist before the request — precomputing the ranking is out. But the item side does precompute: encode all 4M items offline into a vector index, nightly, and never pay for that on the request path.
Now the arithmetic. 200 minus network, features and serialisation leaves about 147 ms. The good ranker at 300 ms does not fit, so split it: an ANN retrieve of 500 candidates at 25 ms, then a distilled reranker over the top 50 at 70 ms. Ninety-five milliseconds, fifty to spare. Put a content-keyed cache in front for the non-personalised slice of the response and the bill drops from roughly $6,750 a day to under $1,000 — until the evening peak widens the working set and the hit rate falls, at which point cost rises before latency does.
Finally, degradation. Size the reranker fleet for peak plus headroom, bound its queue, and shed at admission rather than letting requests age in a queue. What the shed traffic gets is a cached popularity-and-category list — worse ranking, same page, no error. Say that number out loud in the interview: at 2× the provisioned peak, roughly a third of traffic sees the fallback ranking, and p99 stays inside budget because we shed at the door.
Check yourself
1. A fraud score must reflect the transaction happening right now; the key is card × merchant × amount × timestamp. Precompute or compute on demand?
2. Your p99 is 250 ms against a 200 ms budget. The cache hit rate is 85%. You raise it to 95%. What happens?