Product judgment when the data disagrees
When the study and the experiment point opposite ways, they are usually answering two different questions — find the question before you pick a side.
The idea
A research readout and an A/B result almost never contradict each other on the same fact. They differ on who was measured, over what window, and against which number. Change any one of those three and the same evidence can say ship or say hold.
So the useful move is not "which source do I trust?" It is: write down each signal's coverage, re-read both at the same framing, and notice which single framing choice is really carrying your recommendation. Then say out loud what result would change your mind — before you know which side it helps.
Two streams, one decision
How it works
- Split the measurement question from the product question. "Did completion go up?" is a measurement question and the test settles it. "Should we ship this?" is a product question and no single number settles it.
- Write each signal's coverage triple: population × horizon × metric. The study above covers new to category × 28 days × help-seeking. The test covers all traffic × 7 days × completion. They overlap on nothing.
- Re-read both at one shared framing. Some cells will be empty — say so plainly. "The study is silent about 80% of traffic" is a finding, not a gap to paper over.
- Find the pivot. Change one framing choice at a time and see which one flips the call. That choice is your real argument; everything else is decoration.
- Pre-commit the falsifier. Write "I would change my mind if ___" before the next readout exists.
- Decide, and name the limitation in the same sentence as the decision. Often the answer is a segmented rollout plus a , not a global yes or no.
The arithmetic that hides the disagreement
Week-4 retention, 28-day window
------------------------------------------------------------
experienced users 80% of traffic +1.5% (95% CI +-1.4)
new to category 20% of traffic -6.5% (95% CI +-3.3)
blended = 0.80 x (+1.5%) + 0.20 x (-6.5%)
= +1.20% - 1.30%
= -0.10% reported as "flat, no effect"
blended CI = -0.1% +- 1.3% -> [-1.4%, +1.2%]
Why the subgroup interval is so much wider:
1 / sqrt(0.20) = 2.24 only one fifth of traffic is new
1 / sqrt(0.62) = 1.27 only 62% have 28-day follow-up
x 1.15 the "new user" flag is imperfect
-> 1.0% x 2.24 x 1.27 x 1.15 = 3.3%
So: "no significant effect" on the blend is perfectly
consistent with one fifth of users losing 6.5% of
week-4 retention. The blend did not disprove the study.
It averaged it away.
When to use it
| situation | why this fits |
|---|---|
| Qualitative research and an experiment point opposite ways on the same feature | Almost always a coverage difference, not a factual contradiction. Aligning the triple usually dissolves it. |
| A launch is being waved through on "no significant effect" | Check for dilution before accepting flat. Flat on the blend is not flat for everyone. |
| Two teams quote different metrics for the same change | Makes the metric choice visible as a choice, rather than as the truth. |
| A short readout is being used to settle a long-cycle behaviour | Forces the horizon question: does the behaviour you care about even fit inside the window? |
| The trade-off | This framework does not create evidence. If the honest output is "nobody measured the slice that matters", the deliverable is a and a research plan, not a decision. And re-slicing after the fact invites , so the slice has to be pre-committed and then replicated. |
| When to skip it | When every framing agrees, or when the change is cheap, reversible and monitorable. Then ship and watch — deliberation costs more than the mistake. |
Watch out for
- Reading dilution as safety. "-0.1%, not significant" on a blend with an 80/20 mix is compatible with real harm to the minority. Always ask what effect size the blend could be hiding for the smallest group you care about.
- Fishing for the subgroup that agrees with you. Sweep twenty slices at the 5% level and roughly one will look real by chance. The defensible version is narrow: test the one slice the qualitative work named in advance, and require it to replicate in the next cycle before you act on the magnitude.
- Assuming bigger n wins. Sample size buys precision within a population. It buys nothing about a population you never sampled. Fourteen participants cannot size an effect, but they can tell you which group is at risk and why — and that is the thing the experiment could not tell you.
- Letting the release calendar pick the horizon. A 7-day readout on a monthly behaviour measures novelty and the cost of learning, not the steady state. Pick the window from the behaviour's natural cycle, then admit when you do not have it yet.
- Naming your falsifier after you have picked a side. If the criterion appears only once you know which way the number went, it is a rationalisation. Write it down first, in numbers, and let it bind you.
Worked example
An interviewer says: "Research says users hate the new setup. The A/B test says completion is up 11%. What do you do?"
Start by separating the questions. The measurement question — did completion move? — is settled: +10.8% (95% CI +10.3 to +11.3) on all traffic at 7 days. Nothing in the study contradicts that. Then name the coverage: the study watched 14 new-to-category users for 28 days and counted help-seeking; the test watched 480,000 mostly-experienced users for 7 days and counted completions. They share no population, no window and no metric, so "disagreement" is the wrong word.
Now re-read both at one framing. On the blend at 28 days, week-4 retention is -0.1% (CI -1.4 to +1.2) — flat. Slice to the population the study actually studied and the same experiment reads -6.5% (CI -9.8 to -3.2) on retention and +18% (CI +11.4 to +24.6) on support contacts. The study was right about direction and mechanism; the test supplies the size. The pivot is population, and the blend was averaging one fifth of users away.
The recommendation is therefore not a global yes or no: ship to experienced traffic, hold the new-to-category path, and run a 28-day holdout with week-4 retention and support contacts as pre-registered guardrails on that slice. Then commit to the falsifier out loud: "I will ship it to new users if the retention loss replicates above -2% and contacts per 100 stay within +3%; if it comes back below -5% again, we rebuild the comprehension step instead." That last sentence is what separates judgment from preference.
Check yourself
Blended week-4 retention comes back at -0.1% with a 95% CI of -1.4% to +1.2%. What is the fair read?
The 14-person study and the 480,000-user test point opposite ways. What is the best first move?