Keeping a model working after launch

A model that passed evaluation is only half a system — the other half is the machinery that notices when it stops being right, and answers well while you fix it.

The idea

Evaluation asks was this model good on data I already had? Operating asks is it good right now, on traffic I have not labelled yet? Those are different questions, and the second one usually cannot be answered directly: the labels arrive days or months late, and sometimes they arrive bent by the model's own decisions.

So you watch the things that move before the metric does — the inputs, the shape of the feed, the mix of decisions, and any cheap outcome correlated with quality. And because detection alone serves nobody, every path needs a fallback: a simpler answer is worth more than an error, and far more than a confident wrong one.

The production monitor

Ninety simulated days. A drift is injected on a day you cannot see. Choose what to alarm on, how late your labels arrive, what the serving path does when something breaks, and what triggers a retrain — then run it and read the scorecard.

monitored signal missing rate / proxy outcome delayed accuracy true served quality (revealed at the end)
alarms you trust (each needs 2 consecutive breach days)
day 0
alarms fired 0
accuracy known through —
served quality to date —
Press play. Days 0–19 are the reference window every alarm is measured against; nothing can fire before day 20.

How it works

1 · fix a reference window, then score every signal as a distance from it

Pick a stretch of traffic you believe was healthy. For each monitored signal, store its mean and spread. Every day afterwards, express the live value as a distance in reference standard deviations, and require two consecutive breach days so one noisy afternoon does not page anyone.

reference window        days 0-19
  feature mean          mu = 50.1    sigma = 1.4
  delayed accuracy      mu = 0.894   sigma = 0.008

day 44  feature = 62.4  ->  z = |62.4 - 50.1| / 1.4  = 8.8   breach (k = 2.6)
day 45  feature = 66.9  ->  z = 12.0                         breach -> alarm fires day 45

label lag L = 14
  accuracy for day 44 is only visible on day 58
  first two consecutive accuracy breaches land on data days 46, 47
  -> accuracy alarm fires on day 47 + 14 = day 61

lead time = 61 - 45 = 16 days of decisions you could act on

For categorical or binned features the same idea is often written as population stability index, PSI = Σ (p_live − p_ref)·ln(p_live / p_ref), with a common rule of thumb of ~0.1 for "look" and ~0.25 for "act". The statistic matters less than the discipline: one reference, one threshold, one owner, one action.

2 · rank signals by how early they move and how tightly they bind to the decision

Input-side signals move first but can be blind. Outcome-side signals are the truth but arrive late. A proxy — a cheap downstream outcome correlated with quality, like an early repayment signal, a review-agreement rate, or a completion rate — sits in the middle and is usually the thing that saves you.

3 · give the serving path an answer for three failures

A timeout, a missing feature, and a model that is confidently wrong all need a defined response. In almost every case the right answer is a degraded answer, not an error: a cached heuristic, a rules table, a population prior, a smaller model. Choose it with the same arithmetic you use for anything else.

a day where 40% of rows lose the feature, 1.5% of requests time out

  0.585 of traffic  model on complete rows   quality 0.895  -> 0.524

fallback = error             0.524 + 0                          = 0.524
fallback = cached heuristic  0.524 + 0.415 * 0.78               = 0.847
fallback = previous model    0.524 + 0.400 * 0.55 + 0.015*0.86  = 0.757
                                     ^ reads the same broken feature

The previous model version is an excellent fallback for a bad retrain and a poor one for a broken feed, because it shares the dependency. A fallback must fail independently of the thing it is covering.

4 · make retraining a comparison, not a calendar entry

And note the arithmetic of : a retrain cannot land before label lag + build + validate. With a 21-day label lag, retraining is never the incident response — the fallback and the rollback are.

5 · close the online/offline gap before you trust either number

When live quality sits below offline quality, the usual three causes are training/serving skew (different code paths computing the "same" feature, different defaults, different freshness), (a feature or a label computed with information not available at decision time), and an offline metric nobody tied to the live decision — an AUC that improved while the fixed operating threshold got worse. Log the exact served feature vector, replay it offline, and diff.

When to use it

signalcatchesblind tocan move
feature distributioncovariate shift; new segments; some changesconcept drift; anything computed only over non-null rowssame day
missing / null / broken feeds, upstream contract changes, stale joinsquality changes with perfectly healthy inputssame day
prediction mixanything that changes what the model decides; feedback loopsdrift that leaves the decision rate intact; masked by an error fallbacksame day
proxy outcomealmost anything that hurts served decisionscause — a campaign or a moves it toohours to days
delayed accuracythe thing you actually care abouteverything inside the label lag; selection-biased labelslabel lag

The trade-off is always the same shape: earlier signals are cheaper and more frequent but less specific, so each one has to be paid for in false alarms. Price that cost explicitly — an alarm with no defined action is a tax on your on-call rota.

Watch out for

Worked example

You are asked to design the operating half of a lending approval model. Defaults are observed at roughly 60 days, so accuracy is two months behind reality, and you only ever see repayment behaviour for applicants you approved — labels that are both late and selected by the model.

The monitoring stack is layered by latency. Same-day: null and schema checks on every feature at the serving boundary, feature distributions computed over all requests including defaults, and approval rate by segment. Days: a proxy — first-payment-missed rate on the newest , plus manual-review agreement on a sampled slice. Months: the delayed default rate, which confirms rather than detects, and which you compute on a small randomly-approved so the labels are not conditioned on your own policy.

The serving path answers a timeout with a cached scorecard and a missing bureau attribute with the same scorecard, never with an error; low-confidence cases route to manual review rather than to a guess. When approval rate falls four points in a week with stable inputs, the evidence points at the model's relationship to reality, not at the feed — so you compare a retrain on recent data against a threshold recalibration, and you ship the recalibration first because it lands in a day and the retrain cannot land for two months. That comparison, said out loud in an interview, is the answer they are listening for.

Check yourself

Chargebacks arrive 30 days late. Feature distributions and prediction mix are both flat, but you suspect quality has fallen. What is the strongest early signal?

A feature service starts returning nulls for 30% of requests. Your fallback is "previous model version". What actually happens to served quality?