Interference and choosing the randomization unit

When treated users can touch control users, splitting by user quietly poisons the control group — and your effect estimate goes with it.

The idea

A clean A/B test assumes one user’s treatment doesn’t change another user’s outcome. That assumption — sometimes called SUTVA, no interference — breaks whenever people are connected: invites, marketplaces, shared drivers, teammates on the same account.

When it breaks, a treated user’s effect leaks across the boundary into control, and the naive treatment-minus-control gap is biased. The fix is to randomize a bigger, self-contained unit — a cluster, a city, a time slice — so the leak stays inside one arm. You buy an honest estimate, and you pay for it in power.

Treatment leaks along the edges

treated control (clean) control (contaminated)
Randomize by:
0.60
Bias in the naive estimate
—
Effective sample size
—
Measured effect (true = +8.0)
—
Contaminated controls
—
Press play spread. With user-level randomization, treated and control sit side by side inside every cluster — watch the effect leak across the edges.

How it works

The bias and the cure are two sides of the same number — the strength of interference, s. Here the population is 48 users in 6 connected clusters of 8, and the true per-user effect is +8:

Randomize by USER  (edges cross the boundary)
  control units sit next to treated ones, so control is lifted
  by  s x effect x (share of treated neighbors ~ 0.5)
  measured effect = 8 - 4s        at s=0.6 -> +5.6  (30% too low)
  effective N = 48                 (full power, but the number is a lie)

Randomize by CLUSTER  (no edge crosses the boundary)
  measured effect = 8              unbiased
  but cluster members move together: ICC ~ s
  design effect = 1 + (m-1)*s = 1 + 7s
  effective N = 48 / (1 + 7s)      at s=0.6 -> ~9 of 48
  same power now needs sqrt(1+7s) x the users

At s = 0 there’s no interference, and user-level is perfectly fine — full power, no bias. Every unit of interference you add makes user-level more biased and makes clustering more expensive. That tension is the whole decision.

When to use it

Randomization unitUse it when…The price
UserUsers act independently — no invites, no shared markets, no teammates.Any real interference biases the estimate, usually toward zero.
Cluster (team / company)Interference is contained inside a group but not across groups.Power drops with the intra-cluster correlation — you effectively have far fewer units.
Geo (city / region)Effects spill locally — marketplaces, drivers, delivery supply.Few, uneven regions; long runtimes; confounds with local events.
Time-sliced (switchback)Effects are short-lived and the whole market must be one state at a time.Carryover between slices; needs careful slice length.

Watch out for

Worked example

An interviewer says: “We’re testing a new driver-incentive on a rideshare marketplace. How do you randomize?” A strong answer names the interference first: a treated driver who takes more trips leaves fewer for control drivers in the same city, so user-level randomization makes the incentive look better than it is — the control is depressed by the very treatment you’re measuring. Randomize by city (or switchback by time) so supply and demand balance within an arm. Then say the honest part: you’ll have far fewer independent units, so you power the test on the number of cities and accept a larger minimum detectable effect — a correct, slightly blurry answer beats a sharp, wrong one.

Check yourself

Your user-level test of a “refer a friend” feature shows a small, disappointing lift. A teammate says “let’s just run it longer for more power.” Your read?

You switch from user- to city-level randomization. What must you expect to change?