When you can’t A/B test: quasi-experiments and holdouts

“We launched and revenue rose 30%” is a story, not a result — until you subtract what would have happened anyway.

The idea

Sometimes you can’t randomize: the feature ships to a whole market, a whole region, everyone at once. So you can’t compare treated vs untreated users directly. The trick is to find a comparable series that didn’t get the change and use its movement as your — what your market would have done on its own.

Difference-in-differences subtracts that control’s before/after change from yours, so shared shocks — a holiday, the economy, a trend — cancel out. And over a year of small wins, a long-term holdout is the honest scoreboard: it tells you whether a dozen claimed lifts actually compounded, or just got counted twice.

diff-in-differences · two markets, one launch

your market (got the feature) comparable control (didn’t)
month 6
naive before / after
+32.5
your market only
diff-in-differences
+8.0
after subtracting control
true effect
+8.0
what the feature did
Press walk it through, or drag the launch month. Watch how much of the “win” the naive read gives to the feature.

the four-quarter holdout · do the small wins add up?

Every quarter the team ships a few features, each measured as a small win. Keep 5% of users in a holdout — they get none of them — and after four quarters compare. The claimed wins pile up; the holdout tells the truth.

sum of claimed lifts holdout says (actual)
claimed so far
—
holdout truth
—
over-claim
—
counted but not real
Press advance quarters to add each quarter’s shipped features and watch the gap open up.

How it works

Difference-in-differences is two subtractions. First, each series’ own before/after change. Then, one minus the other — which cancels anything both markets felt. With the launch at month 6:

your market:  after avg 157.5  − before avg 125.0  = +32.5   (feature + trend + holiday)
control:      after avg 129.5  − before avg 105.0  = +24.5   (trend + holiday, no feature)
--------------------------------------------------------------
effect     =  32.5 − 24.5                       = +8.0    (what the feature actually did)

The naive read (+32.5) hands the feature credit for the holiday and the trend. The control rose +24.5 with no feature at all — subtract it and you’re left with the real +8. The whole method rests on one assumption: parallel trends — absent the feature, your market would have moved like the control. Always check that the two lines tracked each other before the launch.

The holdout is the same instinct across time. Instead of trusting the sum of a year’s claimed wins, hold 5% of users out of everything and measure the real gap at the end:

sum of claimed quarterly lifts (Q1..Q4)  =  +32%
holdout: treated 95% vs held-out 5%      =  +11%
--------------------------------------------------
over-claim (overlap, seasonality, decay) =  21 points

When to use it

MethodUse when…Its Achilles’ heel
Randomized A/BYou can split users randomly.Not always possible — some changes ship to everyone.
Diff-in-differencesA comparable market/region didn’t get the change.Fails if the control wasn’t on a parallel trend.
Long-term holdoutMany changes accumulate and you need true cumulative lift.Costs a slice of users; must be protected from spillover.

Watch out for

Worked example

An interviewer says: “We rolled a loyalty program out to our whole US business. Ninety-day revenue rose 18%. Did it work?” The disciplined answer refuses the naive read. You ask for a comparable market that didn’t get loyalty — say Canada — and check the two tracked each other for the prior year. If Canadian revenue rose 14% over the same window (a category-wide tailwind), the diff-in-differences estimate is about +4%, not 18%. Then you widen the lens: this team has shipped nine “+3%” wins this year, which “should” be +27%. A 5% holdout that’s seen none of them shows the real annual lift is +7%. Naming both the counterfactual and the holdout is what a senior interviewer is listening for.

Check yourself

Revenue in your launch market rose 30% in the two months after the feature. What do you most need before crediting the feature?

Eight features each measured +4% this year; leadership expected ~+35% total. The 5% holdout shows +12%. Best explanation?