When treated users can touch control users, splitting by user quietly poisons the control group — and your effect estimate goes with it.
A clean A/B test assumes one user’s treatment doesn’t change another user’s outcome. That assumption — sometimes called SUTVA, no interference — breaks whenever people are connected: invites, marketplaces, shared drivers, teammates on the same account.
When it breaks, a treated user’s effect leaks across the boundary into control, and the naive treatment-minus-control gap is biased. The fix is to randomize a bigger, self-contained unit — a cluster, a city, a time slice — so the leak stays inside one arm. You buy an honest estimate, and you pay for it in power.
Treatment leaks along the edges
The bias and the cure are two sides of the same number — the strength of interference, s. Here the population is 48 users in 6 connected clusters of 8, and the true per-user effect is +8:
Randomize by USER (edges cross the boundary)
control units sit next to treated ones, so control is lifted
by s x effect x (share of treated neighbors ~ 0.5)
measured effect = 8 - 4s at s=0.6 -> +5.6 (30% too low)
effective N = 48 (full power, but the number is a lie)
Randomize by CLUSTER (no edge crosses the boundary)
measured effect = 8 unbiased
but cluster members move together: ICC ~ s
design effect = 1 + (m-1)*s = 1 + 7s
effective N = 48 / (1 + 7s) at s=0.6 -> ~9 of 48
same power now needs sqrt(1+7s) x the users
At s = 0 there’s no interference, and user-level is perfectly fine — full power, no bias. Every unit of interference you add makes user-level more biased and makes clustering more expensive. That tension is the whole decision.
| Randomization unit | Use it when… | The price |
|---|---|---|
| User | Users act independently — no invites, no shared markets, no teammates. | Any real interference biases the estimate, usually toward zero. |
| Cluster (team / company) | Interference is contained inside a group but not across groups. | Power drops with the intra-cluster correlation — you effectively have far fewer units. |
| Geo (city / region) | Effects spill locally — marketplaces, drivers, delivery supply. | Few, uneven regions; long runtimes; confounds with local events. |
| Time-sliced (switchback) | Effects are short-lived and the whole market must be one state at a time. | Carryover between slices; needs careful slice length. |
An interviewer says: “We’re testing a new driver-incentive on a rideshare marketplace. How do you randomize?” A strong answer names the interference first: a treated driver who takes more trips leaves fewer for control drivers in the same city, so user-level randomization makes the incentive look better than it is — the control is depressed by the very treatment you’re measuring. Randomize by city (or switchback by time) so supply and demand balance within an arm. Then say the honest part: you’ll have far fewer independent units, so you power the test on the number of cities and accept a larger minimum detectable effect — a correct, slightly blurry answer beats a sharp, wrong one.
Check yourself
Your user-level test of a “refer a friend” feature shows a small, disappointing lift. A teammate says “let’s just run it longer for more power.” Your read?
You switch from user- to city-level randomization. What must you expect to change?