Training data pipelines: freshness, labels and leakage

A training pipeline has exactly one job — produce data that looks like what the model will meet in production — and almost every way it fails is silent.

The idea

A model can only learn from what you hand it. If the rows you hand it contain a whisper of the future, the model will happily learn the whisper, score beautifully offline, and then meet a world where the whisper is missing.

So the pipeline is not plumbing. It is a claim: this row is what the model would have seen, at the moment it would have had to decide. Every stage below is a way of keeping that claim honest — the split, the watermark, the label, the feature definition, the snapshot.

Start with leakage, because it is the failure that makes an offline number excellent and a live model useless.

One build, four decisions

train row test row arrived after its day closed
label moment feature window earlier later

how the split is drawn
the feature contact_recency is computed…
offline accuracy (what the report says)
—
live accuracy (what production gets)
—
0%gap: 0.0 pts100%

sweep—
events seen—
admitted—
late, dropped—
day 28 completeness—
data delay—
test users also in train—
test rows—

Each dot stands for about 10 events. Dots sit at their event time; the ones that fall through the axis were still in flight when their day's build closed.

How it works

Eight stages, each with a check that would catch it going wrong.

  1. Write down the label moment. For every training row there is a timestamp at which the label became knowable — the agent's decision, the chargeback, the click. A feature is legal only if every input to it existed strictly before that timestamp. This is a point-in-time join, not a table join. Check: assert max(feature_source_time) < label_moment for every feature, on every build.
  2. Split on time by default. Train on the earlier window, test on the later one, because production always evaluates on the future. A random split lets the model see Tuesday when it is scored on Monday, and lets the same entity sit on both sides. Check: max(train_time) <= min(test_time), and zero overlap of entity keys if the label is per entity.
  3. Size the windows by cycles, not by percentages. A test window shorter than seven days cannot contain a weekly cycle, so weekday effects land entirely in train or entirely in test. If the product has , one test window will not measure it — use several rolling folds instead.
  4. Choose a watermark as a stated number. Event time is when it happened; processing time is when it landed. The build for day D closes at end(D) + W. Waiting longer buys completeness and costs freshness — but the number must be in config, not an accident of when cron fired.
  5. Treat labels as a system. Version the guideline; stamp each label with the guideline version and the annotator; double-label a sample and measure agreement; route disputes to adjudication rather than to a majority vote of two. And know which labels only exist because the product already said yes.
  6. One definition, two paths. The feature store's real job is that contact_recency is computed by one piece of code, used both to backfill training rows point-in-time and to serve online. Two implementations means training/serving skew: the same input, two different numbers, and a live drop no offline test reproduces.
  7. Snapshot, don't share a mutable table. Thirty experiments a day against features.latest are not comparable to each other. Freeze an immutable snapshot with a manifest: source ranges, watermark, code commit, guideline version, row count, content hash. Last week's model is explainable only if last week's dataset can be rebuilt byte-for-byte.
  8. Admission gates and deletion. Messy inputs need a measurable threshold at the door (mean OCR confidence, minimum page count) so admission is a number, not a mood — and the rejected slice is monitored, because it is a population. A deletion request must reach the snapshots too: the rows, re-materialise the affected snapshot ids, list every model trained on them, and have a stated retrain-or-retire path.
label moment   t_L = 2024-03-11 14:02   (agent files the claim category)

feature "contact_recency"
  computed over [t_L - 30d, t_L)        -> legal, servable
  computed over [t_L - 30d, t_L + 7d)   -> leaks 7 days of the future
                                           offline soars, serving sees null

watermark (daily build closes at end-of-day + W), this lag profile:
  W =  0h   completeness ~60%   data delay 1.0 d
  W = 24h   completeness ~86%   data delay 2.0 d
  W = 48h   completeness ~93%   data delay 3.0 d
  W = 96h   completeness ~99%   data delay 5.0 d
  rule: take W where the curve flattens, put it in config,
        alert when any day closes below the expected completeness.

is the test window big enough to trust the number?
  n  = 7 days x 1,200 labelled rows = 8,400
  se = sqrt(0.78 x 0.22 / 8400)     = 0.0045
  95% band = +/- 1.96 x 0.0045      = +/- 0.9 pts
  a 0.4 pt "win" inside that band is not a win.
label quality on a 34-class space, 5% double-labelled
  observed agreement  po = 0.86
  chance agreement    pe = 0.62      (from the marginal class priors)
  kappa = (po - pe) / (1 - pe) = 0.24 / 0.38 = 0.63

  read it as: agreement is well above chance but 14% of pairs
  disagree -> those confusion pairs go to adjudication, and the
  guideline gets a new version (v4) that names the boundary.
  Every row keeps guideline_version, so the shift is visible later.

When to use it

SituationDefault choiceTrade-off to state out loud
Time-ordered behaviour: fraud, demand, rankingTime split, test window strictly after trainLess data to train on; the test window must be long enough to hold a weekly cycle
Labels arrive late: conversions, chargebacks, adjudicated reviewWatermark set where the completeness curve flattensEvery extra hour of waiting is an extra hour of staleness in the deployed model
The target belongs to an entity: churn, lifetime valueGroup split by entity, on top of the time splitFewer independent test units, so wider confidence intervals
Human labels on a wide label spaceDouble-label a sample, measure agreement, adjudicate disputesLabelling cost roughly doubles on the sampled slice
The same feature at train and serveOne definition in a feature store, backfilled point-in-timeThe store becomes a hard dependency; a change is now a model change
Dozens of experiments a day on shared dataImmutable snapshot plus a manifest hashStorage cost, and the discipline of never editing in place
Messy inputs: scans, audio, user-typed textA measurable admission gate with a published thresholdReal documents get rejected; that slice needs its own monitoring and route

Watch out for

Worked example

Interview scenario: design the training pipeline for document at an insurer — scanned claim documents, 34 categories, the label being the category an operations agent files after reading the document.

The label moment is the agent's filing timestamp, so every feature is computed from the intake OCR pass only. Weeks 1–8 train, weeks 9–10 test, giving two full weekly cycles including the Monday post-weekend spike; the split is also grouped by claim id, because one claim can carry six pages that would otherwise straddle it. The watermark is 48 hours, chosen because about 12% of documents are re-scanned and re-arrive the next working day — at 48h a day closes at roughly 93% complete, and the build alerts if any day closes below 90%. Admission gate: mean per-page OCR confidence at or above 0.72, with rejected pages routed to manual handling and reviewed monthly so the gate itself is monitored. Five percent of documents are double-labelled; kappa is 0.63, and the disputed pairs — two near-identical liability categories — go to adjudication and produce guideline v4. Every snapshot carries a manifest: date ranges, W, the gate threshold, the guideline version, the code commit, a content hash.

Then the first model scores 0.91 offline and 0.74 live. The cause is a feature called pages_final, computed from the document record after the agent finished — because re-scans update the record, that column encodes the agent's own handling. Recomputing it as of the intake timestamp drops offline to 0.76 and lifts live to 0.75. Two weeks later a customer exercises a deletion request: the rows are tombstoned, the three snapshots containing them are re-materialised with new ids, and the two models trained on them are flagged for retraining on the clean snapshot rather than quietly left in production.

Check yourself

Offline AUC jumps from 0.78 to 0.94 the day you add days_since_last_contact, recomputed nightly from the CRM. What do you check first?

30% of your conversion labels land more than 24 hours after the click. The daily build currently takes whatever exists at 02:00. What is the right move?