Statistical power and sample size

Power is the odds your test notices a real effect. Small effects, rare events, and thin traffic all steal it — so you plan the sample before you run.

The idea

Every test has two ways to fail: cry wolf when nothing happened (a ), or miss a real change (a false negative). Power is one minus that second risk — the chance you’ll detect an effect that’s truly there. Aiming for 80% power means: if the lift you care about is real, you’ll catch it about four times in five.

Four dials trade off against each other: how big an effect you want to catch, how common the event already is, how much traffic you get, and how long you run. Move one and the others must give. A quick formula turns “is two weeks enough?” into a number instead of a shrug.

Power calculator

The two bells are what you’d measure if there’s no real effect (left) versus if the lift is real (right). Power is the shaded slice of the right bell that clears the significance line.

Baseline rate (how often it happens now) 12.0%
Relative lift to detect 5.0%
Daily traffic (both groups) 20,000
Test duration 14 days
Absolute effect (δ)
—
Sample per group so far
—
Needed per group for 80%
—
Days to reach 80% power
—
Power at this duration—
Drag any slider. Watch the right bell slide toward or away from the significance line — that gap is your power.

How it works

For comparing two rates at the usual settings (95% confidence, 80% power), the per-group sample size is a one-liner — the “rule of 16”:

          16 · p(1 - p)
   n  ≈  ----------------      per group
                δ²

  p = baseline rate      δ = absolute effect you want to catch
  (the 16 bakes in 95% confidence + 80% power)

Worked: baseline p = 12%, detect a 5% relative lift
   δ = 0.12 × 0.05 = 0.006   (0.6 percentage points)
   n ≈ 16 × 0.12 × 0.88 / 0.006²
     = 1.6896 / 0.000036
     ≈ 46,933 per group  ->  ~93,900 total
   at 20,000 / day that is ~5 days.  Two weeks is plenty.

Shrink the lift 10x (0.5% relative -> δ = 0.0006):
   δ² is 100x smaller, so n is 100x bigger
     ≈ 4.69 million per group  ->  ~4 months of traffic.

The lesson lives in that δ²: halving the effect you want to detect quadruples the sample. Tiny lifts and rare events (small p) are expensive to prove, and no amount of cleverness dodges the arithmetic — only more traffic or a bigger effect.

When to use it

Question in the roomWhat to compute
“Is two weeks enough?”Needed n vs. traffic×days. If you can’t reach n, you can’t conclude — extend, widen the effect, or don’t ship the test.
“Is this 4% drop significant?”Was the test even powered to see 4%? An underpowered “flat” result means didn’t detect, not no effect.
“Can we detect a 0.5% lift?”Plug it in. Often the honest answer is “not in any reasonable timeframe” — better to say so early.

Watch out for

Worked example

An interviewer says: “Checkout converts at 2.3%. Product wants to detect a 5% relative lift from a new button. We get about 50,000 sessions a week. Two-week test — go or no-go?” Work it: δ = 0.023 × 0.05 ≈ 0.00115. Per group n ≈ 16 × 0.023 × 0.977 / 0.00115² ≈ 272,000, so ~544,000 sessions total. At 50,000/week that’s about 11 weeks — two weeks gives only ~23% power. The strong answer isn’t “run it and see”; it’s “two weeks can’t see this. We either aim for a bigger effect, pool more traffic, or accept it needs ~11 weeks.”

Check yourself

A two-week test on a small feature comes back “no significant difference.” The team wants to conclude the feature does nothing. Best response?

You want to detect a smaller lift than planned — half the size. Roughly what happens to the sample you need?