Experiment pitfalls: peeking, novelty, multiple comparisons

A clean test can still lie to you — three ways — and each one has a boring, reliable antidote.

The idea

You can randomize perfectly and still reach a false conclusion, because three habits quietly manufacture “wins” out of pure noise. Peeking — stopping the moment it turns significant — inflates several times over. Novelty — a week-one bump that fades — fools you if you decide early. Multiple comparisons — checking twenty metrics — all but guarantees a lucky green somewhere.

The antidotes are dull on purpose: pre-register the sample size, the primary metric, and the decision before you look; run an A/A test to feel how much your own pipeline lies. Play with each lab below and watch the lie form.

Three mini-labs

Lab 1 · peeking

100 A/A tests — where both arms are identical

Each square is one experiment with no real difference. A true finding is impossible; every “win” here is a false positive.

running, no call over the line right now called a false “win”
Day / look
0 of 40
False “wins” called
0 of 100
You accepted
5%
Fixed-horizon rules: you’ll decide once, on day 40. Press play. Then flip stop at first significance to peek daily and watch the false wins pile up.

Lab 2 · novelty

The same feature, read on different days

A shiny new feature draws a burst of clicks that fades as it stops being new. Its durable effect here is zero — slide the day you decide.

durable effect (truth) ≈ 0 ship gate +2pp week 1 week 3 day 1 day 7 day 14 day 21 day 28
day 3
Decision day
3
Measured lift
+5.7 pp
Reads as
a win

Lab 3 · multiple comparisons

Twenty secondary metrics, all pure noise

None of these metrics truly moved. Yet under a 5% threshold, each has a 1-in-20 chance of a “significant” p-value by luck. Reroll and count the green.

Significant this run
—
Threshold
p < 0.05
P(≥1 false win)
64%
Reroll a few times. With twenty independent shots at p < 0.05, a false “winner” shows up in about two out of three runs — then turn on the correction.

How it works

Peeking. Each test at the 5% level has a 5% false-positive rate at one fixed look. But the test statistic wanders as data arrives, so the more often you look, the more chances it has to stray past the line at least once:

looks (peeks)      P(a false "win" at some point)
   1 (correct)          ~5%
   5                    ~14%
  20                    ~25%
  40  (daily, 6 wks)    ~30%      ← this lab

fix: fixed sample size, or a sequential design (alpha-spending / group-sequential)
     that budgets the 5% across all the looks you intend to take.

Novelty. The measured lift is durable-effect plus a fading novelty bump. Decide during the fade and you bank the bump:

lift(day) = durable + novelty × e^(-(day-1)/6)     durable = 0,  novelty = 8 pp
  day  1   ->  +8.0 pp   "ship it!"      (almost all novelty)
  day  9   ->  +2.1 pp   "still a win"
  day 21   ->  +0.3 pp   "no real effect" (the truth)

fix: run past the decay (2-3+ weeks), judge the plateau, not the spike.

Multiple comparisons. Under the null, p-values are uniform, so each metric independently trips 5% of the time:

P(at least one false win) = 1 - (1 - 0.05)^20 = 1 - 0.95^20 = 64%
expected false winners     = 20 × 0.05 = 1.0 per run

fix (Bonferroni): test each at 0.05 / 20 = 0.0025
     ->  1 - (1 - 0.0025)^20 = ~4.9%   (family-wise rate back to ~5%)

When to use it

Use it when… / trade-off
Fixed sample size, decide onceDefault for a confirmatory test. Trade-off: you must estimate the needed size up front and wait for it.
Sequential / group-sequentialWhen you genuinely need to stop early. It lets you peek — but only because it spends the 5% across pre-planned looks.
Run past noveltyAny user-visible change. Trade-off: slower reads; hold your nerve through the week-one spike.
Correction (Bonferroni / FDR)When you must scan many metrics. Trade-off: a stricter bar costs power — better to name one primary metric.

Watch out for

Worked example

An interviewer says: “We shipped a redesign. Day 3, conversion is up 6% and significant — do we roll out?” The disciplined answer names all three traps at once. Was day 3 the pre-planned readout, or a peek? If you’ve been checking daily, that significance is inflated — hold to the pre-set sample size. Is 6% durable or novelty? A redesign is precisely the change that spikes then fades; wait two to three weeks and read the plateau. And is conversion the one primary metric you committed to, or the best of many you scanned? If it’s a lucky secondary, correct for the family or don’t count it. A strong candidate closes with the cheap insurance: pre-register next time, and keep an A/A test running so you always know your own noise floor.

Check yourself

Your dashboard refreshes hourly and you’ll ship “as soon as it’s significant.” What’s wrong with that plan?

A new feature shows +9% in week one, sliding to +1% by week three. Best read?