A clean test can still lie to you — three ways — and each one has a boring, reliable antidote.
You can randomize perfectly and still reach a false conclusion, because three habits quietly manufacture “wins” out of pure noise. Peeking — stopping the moment it turns significant — inflates several times over. Novelty — a week-one bump that fades — fools you if you decide early. Multiple comparisons — checking twenty metrics — all but guarantees a lucky green somewhere.
The antidotes are dull on purpose: pre-register the sample size, the primary metric, and the decision before you look; run an A/A test to feel how much your own pipeline lies. Play with each lab below and watch the lie form.
Lab 1 · peeking
Each square is one experiment with no real difference. A true finding is impossible; every “win” here is a false positive.
Lab 2 · novelty
A shiny new feature draws a burst of clicks that fades as it stops being new. Its durable effect here is zero — slide the day you decide.
Lab 3 · multiple comparisons
None of these metrics truly moved. Yet under a 5% threshold, each has a 1-in-20 chance of a “significant” p-value by luck. Reroll and count the green.
Peeking. Each test at the 5% level has a 5% false-positive rate at one fixed look. But the test statistic wanders as data arrives, so the more often you look, the more chances it has to stray past the line at least once:
looks (peeks) P(a false "win" at some point)
1 (correct) ~5%
5 ~14%
20 ~25%
40 (daily, 6 wks) ~30% ← this lab
fix: fixed sample size, or a sequential design (alpha-spending / group-sequential)
that budgets the 5% across all the looks you intend to take.
Novelty. The measured lift is durable-effect plus a fading novelty bump. Decide during the fade and you bank the bump:
lift(day) = durable + novelty × e^(-(day-1)/6) durable = 0, novelty = 8 pp
day 1 -> +8.0 pp "ship it!" (almost all novelty)
day 9 -> +2.1 pp "still a win"
day 21 -> +0.3 pp "no real effect" (the truth)
fix: run past the decay (2-3+ weeks), judge the plateau, not the spike.
Multiple comparisons. Under the null, p-values are uniform, so each metric independently trips 5% of the time:
P(at least one false win) = 1 - (1 - 0.05)^20 = 1 - 0.95^20 = 64%
expected false winners = 20 × 0.05 = 1.0 per run
fix (Bonferroni): test each at 0.05 / 20 = 0.0025
-> 1 - (1 - 0.0025)^20 = ~4.9% (family-wise rate back to ~5%)
| Use it when… / trade-off | |
|---|---|
| Fixed sample size, decide once | Default for a confirmatory test. Trade-off: you must estimate the needed size up front and wait for it. |
| Sequential / group-sequential | When you genuinely need to stop early. It lets you peek — but only because it spends the 5% across pre-planned looks. |
| Run past novelty | Any user-visible change. Trade-off: slower reads; hold your nerve through the week-one spike. |
| Correction (Bonferroni / FDR) | When you must scan many metrics. Trade-off: a stricter bar costs power — better to name one primary metric. |
An interviewer says: “We shipped a redesign. Day 3, conversion is up 6% and significant — do we roll out?” The disciplined answer names all three traps at once. Was day 3 the pre-planned readout, or a peek? If you’ve been checking daily, that significance is inflated — hold to the pre-set sample size. Is 6% durable or novelty? A redesign is precisely the change that spikes then fades; wait two to three weeks and read the plateau. And is conversion the one primary metric you committed to, or the best of many you scanned? If it’s a lucky secondary, correct for the family or don’t count it. A strong candidate closes with the cheap insurance: pre-register next time, and keep an A/A test running so you always know your own noise floor.
Check yourself
Your dashboard refreshes hourly and you’ll ship “as soon as it’s significant.” What’s wrong with that plan?
A new feature shows +9% in week one, sliding to +1% by week three. Best read?