A/B test design: hypothesis, randomization, primary metric

Turn “make the homepage better” into a fair coin toss that can actually tell you what worked.

The idea

An A/B test isn’t “ship it and watch the numbers.” It’s a small experiment with three parts written down before you look at data: a falsifiable hypothesis (change X moves metric Y by Z for population P), random assignment so the two groups differ only by the change, and one primary metric plus a couple of you commit to in advance.

Randomization is the quiet hero. It’s the only thing that lets you say the change caused the difference, instead of “the two groups were just different people.”

1 · write a falsifiable hypothesis

We believe adding will lift by +8 pp for .

Primary metric committed up front: signup completion rate. Everything else is a guardrail or secondary — you won’t let a good secondary number talk you into shipping a losing primary.

2 · assign users, then measure the lift

high-intent user (base convert 40%) low-intent user (base convert 10%)
Assign by:
+8 pp
Control group — base intent
—
Treatment group — base intent
—
Measured lift (treatment − control)
—
True effect (what really happened)
+8 pp
Choose an assignment method and press run experiment. Watch what happens to the two groups’ make-up.

How it works

The whole game is making the two groups exchangeable — identical on average in every way except the change. Then any difference in the primary metric is caused by the change. Here is the arithmetic the simulator runs:

# 60 visitors. Early sign-ups skew high-intent (they'd convert anyway).
# base convert: high-intent 40%, low-intent 10%.  True effect: +8 pp.

Split by signup date  (a confounder rides along)
  control  = first 30 (mostly high-intent) -> base 34%
  treatment= last  30 (mostly low-intent)  -> base 16%, +8 = 24%
  measured lift = 24% - 34% = -10 pp   # sign flipped! looks harmful.

Random coin flip  (groups balance)
  control   base ~= 25%
  treatment base ~= 25%, +8 = 33%
  measured lift ~= 33% - 25% = +8 pp   # matches the truth.

Same visitors, same real effect. The only thing that changed is how you split them — and that decided whether your number was true or a lie.

When to use it

Reach for an A/B test when…The trade-off
You can randomly assign users and wait for enough of them.Needs traffic and patience — small samples give noisy, unstable lifts.
The metric responds within a testable window (a click, a signup).Slow outcomes (annual churn) are hard; use proxies carefully.
The change is isolable to one variable you can toggle.Bundled changes tell you something moved, not what.

Watch out for

Worked example

An interviewer asks: “The team wants to make the pricing page ‘better.’ Design the test.” A strong answer names all three parts. Hypothesis: adding annual-plan will lift the paid-conversion rate by ~3 pp for logged-out visitors. Primary metric: paid-conversion rate; guardrails: refund rate and page load. Assignment: randomize each visitor 50/50, bucketed by a hashed user id so the split is stable and independent of when or where they arrived. Then estimate the sample size needed to detect a 3 pp lift, run to that size, and read the primary metric once. If someone suggests “just show it to the East region and compare” — that’s the signup-date trap in disguise, and you’d flag the confounder.

Check yourself

Halfway through, the treatment’s primary metric isn’t moving — but a secondary metric (time on page) is up nicely. What’s the disciplined move?

A teammate proposes rolling the new checkout out to everyone who signed up this week, and comparing them to last week’s users. Your read?