A training pipeline has exactly one job — produce data that looks like what the model will meet in production — and almost every way it fails is silent.
A model can only learn from what you hand it. If the rows you hand it contain a whisper of the future, the model will happily learn the whisper, score beautifully offline, and then meet a world where the whisper is missing.
So the pipeline is not plumbing. It is a claim: this row is what the model would have seen, at the moment it would have had to decide. Every stage below is a way of keeping that claim honest — the split, the watermark, the label, the feature definition, the snapshot.
Start with leakage, because it is the failure that makes an offline number excellent and a live model useless.
Each dot stands for about 10 events. Dots sit at their event time; the ones that fall through the axis were still in flight when their day's build closed.
Eight stages, each with a check that would catch it going wrong.
max(feature_source_time) < label_moment for every feature, on every build.max(train_time) <= min(test_time), and zero overlap of entity keys if the label is per entity.end(D) + W. Waiting longer buys completeness and costs freshness — but the number must be in config, not an accident of when cron fired.contact_recency is computed by one piece of code, used both to backfill training rows point-in-time and to serve online. Two implementations means training/serving skew: the same input, two different numbers, and a live drop no offline test reproduces.features.latest are not comparable to each other. Freeze an immutable snapshot with a manifest: source ranges, watermark, code commit, guideline version, row count, content hash. Last week's model is explainable only if last week's dataset can be rebuilt byte-for-byte.label moment t_L = 2024-03-11 14:02 (agent files the claim category)
feature "contact_recency"
computed over [t_L - 30d, t_L) -> legal, servable
computed over [t_L - 30d, t_L + 7d) -> leaks 7 days of the future
offline soars, serving sees null
watermark (daily build closes at end-of-day + W), this lag profile:
W = 0h completeness ~60% data delay 1.0 d
W = 24h completeness ~86% data delay 2.0 d
W = 48h completeness ~93% data delay 3.0 d
W = 96h completeness ~99% data delay 5.0 d
rule: take W where the curve flattens, put it in config,
alert when any day closes below the expected completeness.
is the test window big enough to trust the number?
n = 7 days x 1,200 labelled rows = 8,400
se = sqrt(0.78 x 0.22 / 8400) = 0.0045
95% band = +/- 1.96 x 0.0045 = +/- 0.9 pts
a 0.4 pt "win" inside that band is not a win.
label quality on a 34-class space, 5% double-labelled
observed agreement po = 0.86
chance agreement pe = 0.62 (from the marginal class priors)
kappa = (po - pe) / (1 - pe) = 0.24 / 0.38 = 0.63
read it as: agreement is well above chance but 14% of pairs
disagree -> those confusion pairs go to adjudication, and the
guideline gets a new version (v4) that names the boundary.
Every row keeps guideline_version, so the shift is visible later.
| Situation | Default choice | Trade-off to state out loud |
|---|---|---|
| Time-ordered behaviour: fraud, demand, ranking | Time split, test window strictly after train | Less data to train on; the test window must be long enough to hold a weekly cycle |
| Labels arrive late: conversions, chargebacks, adjudicated review | Watermark set where the completeness curve flattens | Every extra hour of waiting is an extra hour of staleness in the deployed model |
| The target belongs to an entity: churn, lifetime value | Group split by entity, on top of the time split | Fewer independent test units, so wider confidence intervals |
| Human labels on a wide label space | Double-label a sample, measure agreement, adjudicate disputes | Labelling cost roughly doubles on the sampled slice |
| The same feature at train and serve | One definition in a feature store, backfilled point-in-time | The store becomes a hard dependency; a change is now a model change |
| Dozens of experiments a day on shared data | Immutable snapshot plus a manifest hash | Storage cost, and the discipline of never editing in place |
| Messy inputs: scans, audio, user-typed text | A measurable admission gate with a published threshold | Real documents get rejected; that slice needs its own monitoring and route |
days_to_close, resolution_notes_length, a nightly-recomputed aggregate that includes the label day. Offline goes near-perfect; at serving the column is null or , and the model collapses onto a value it never really had. The tell is a suspiciously large single-feature importance. The check is the point-in-time assertion, run in CI on a sample of rows.W a config value, key the snapshot on it, and make reruns either byte-identical or a new snapshot id.Interview scenario: design the training pipeline for document at an insurer — scanned claim documents, 34 categories, the label being the category an operations agent files after reading the document.
The label moment is the agent's filing timestamp, so every feature is computed from the intake OCR pass only. Weeks 1–8 train, weeks 9–10 test, giving two full weekly cycles including the Monday post-weekend spike; the split is also grouped by claim id, because one claim can carry six pages that would otherwise straddle it. The watermark is 48 hours, chosen because about 12% of documents are re-scanned and re-arrive the next working day — at 48h a day closes at roughly 93% complete, and the build alerts if any day closes below 90%. Admission gate: mean per-page OCR confidence at or above 0.72, with rejected pages routed to manual handling and reviewed monthly so the gate itself is monitored. Five percent of documents are double-labelled; kappa is 0.63, and the disputed pairs — two near-identical liability categories — go to adjudication and produce guideline v4. Every snapshot carries a manifest: date ranges, W, the gate threshold, the guideline version, the code commit, a content hash.
Then the first model scores 0.91 offline and 0.74 live. The cause is a feature called pages_final, computed from the document record after the agent finished — because re-scans update the record, that column encodes the agent's own handling. Recomputing it as of the intake timestamp drops offline to 0.76 and lifts live to 0.75. Two weeks later a customer exercises a deletion request: the rows are tombstoned, the three snapshots containing them are re-materialised with new ids, and the two models trained on them are flagged for retraining on the clean snapshot rather than quietly left in production.
Offline AUC jumps from 0.78 to 0.94 the day you add days_since_last_contact, recomputed nightly from the CRM. What do you check first?
30% of your conversion labels land more than 24 hours after the click. The daily build currently takes whatever exists at 02:00. What is the right move?