Product judgment when the data disagrees

When the study and the experiment point opposite ways, they are usually answering two different questions — find the question before you pick a side.

The idea

A research readout and an A/B result almost never contradict each other on the same fact. They differ on who was measured, over what window, and against which number. Change any one of those three and the same evidence can say ship or say hold.

So the useful move is not "which source do I trust?" It is: write down each signal's coverage, re-read both at the same framing, and notice which single framing choice is really carrying your recommendation. Then say out loud what result would change your mind — before you know which side it helps.

Two streams, one decision

The decision: replace the five-step setup wizard with one screen of smart defaults. Two evidence streams landed on your desk this week and they point opposite ways. Move the framing below and watch each stream re-read itself.
-30 -15 0 +15 +30 study test not covered not covered
interval clears zero, better interval clears zero, worse interval crosses zero x-axis: relative change vs control (%)
research study
14 new-to-category users · 28-day diary · observed help-seeking
—
out of coverage
not covered
A/B test
480,000 users, all traffic · 7-day readout · setup completion
—
out of coverage
not covered
population
study looked only at new to category · the test ran on all traffic
time horizon
study ran 28 days · the test's headline was read at 7 days
success metric
study counted help-seeking (a support proxy) · the test's primary metric was completion
ship it
framing—
traffic covered100%
test users read480,000
tourstep 1 of 5

How it works

  1. Split the measurement question from the product question. "Did completion go up?" is a measurement question and the test settles it. "Should we ship this?" is a product question and no single number settles it.
  2. Write each signal's coverage triple: population × horizon × metric. The study above covers new to category × 28 days × help-seeking. The test covers all traffic × 7 days × completion. They overlap on nothing.
  3. Re-read both at one shared framing. Some cells will be empty — say so plainly. "The study is silent about 80% of traffic" is a finding, not a gap to paper over.
  4. Find the pivot. Change one framing choice at a time and see which one flips the call. That choice is your real argument; everything else is decoration.
  5. Pre-commit the falsifier. Write "I would change my mind if ___" before the next readout exists.
  6. Decide, and name the limitation in the same sentence as the decision. Often the answer is a segmented rollout plus a , not a global yes or no.

The arithmetic that hides the disagreement

Week-4 retention, 28-day window
------------------------------------------------------------
experienced users    80% of traffic     +1.5%   (95% CI +-1.4)
new to category      20% of traffic     -6.5%   (95% CI +-3.3)

blended = 0.80 x (+1.5%) + 0.20 x (-6.5%)
        = +1.20%  -  1.30%
        = -0.10%                  reported as "flat, no effect"

blended CI = -0.1% +- 1.3%  ->  [-1.4%, +1.2%]

Why the subgroup interval is so much wider:
  1 / sqrt(0.20) = 2.24   only one fifth of traffic is new
  1 / sqrt(0.62) = 1.27   only 62% have 28-day follow-up
  x 1.15                  the "new user" flag is imperfect
  -> 1.0% x 2.24 x 1.27 x 1.15 = 3.3%

So: "no significant effect" on the blend is perfectly
consistent with one fifth of users losing 6.5% of
week-4 retention. The blend did not disprove the study.
It averaged it away.

When to use it

situationwhy this fits
Qualitative research and an experiment point opposite ways on the same featureAlmost always a coverage difference, not a factual contradiction. Aligning the triple usually dissolves it.
A launch is being waved through on "no significant effect"Check for dilution before accepting flat. Flat on the blend is not flat for everyone.
Two teams quote different metrics for the same changeMakes the metric choice visible as a choice, rather than as the truth.
A short readout is being used to settle a long-cycle behaviourForces the horizon question: does the behaviour you care about even fit inside the window?
The trade-offThis framework does not create evidence. If the honest output is "nobody measured the slice that matters", the deliverable is a and a research plan, not a decision. And re-slicing after the fact invites , so the slice has to be pre-committed and then replicated.
When to skip itWhen every framing agrees, or when the change is cheap, reversible and monitorable. Then ship and watch — deliberation costs more than the mistake.

Watch out for

Worked example

An interviewer says: "Research says users hate the new setup. The A/B test says completion is up 11%. What do you do?"

Start by separating the questions. The measurement question — did completion move? — is settled: +10.8% (95% CI +10.3 to +11.3) on all traffic at 7 days. Nothing in the study contradicts that. Then name the coverage: the study watched 14 new-to-category users for 28 days and counted help-seeking; the test watched 480,000 mostly-experienced users for 7 days and counted completions. They share no population, no window and no metric, so "disagreement" is the wrong word.

Now re-read both at one framing. On the blend at 28 days, week-4 retention is -0.1% (CI -1.4 to +1.2) — flat. Slice to the population the study actually studied and the same experiment reads -6.5% (CI -9.8 to -3.2) on retention and +18% (CI +11.4 to +24.6) on support contacts. The study was right about direction and mechanism; the test supplies the size. The pivot is population, and the blend was averaging one fifth of users away.

The recommendation is therefore not a global yes or no: ship to experienced traffic, hold the new-to-category path, and run a 28-day holdout with week-4 retention and support contacts as pre-registered guardrails on that slice. Then commit to the falsifier out loud: "I will ship it to new users if the retention loss replicates above -2% and contacts per 100 stay within +3%; if it comes back below -5% again, we rebuild the comprehension step instead." That last sentence is what separates judgment from preference.

Check yourself

Blended week-4 retention comes back at -0.1% with a 95% CI of -1.4% to +1.2%. What is the fair read?

The 14-person study and the 480,000-user test point opposite ways. What is the best first move?