A pristine study of somewhere else can tell you less about here than a messy study of home.
The evidence hierarchy ranks designs — randomized trial over quasi-experiment over observational over qualitative. Useful, but it answers only one question: , how believable the estimate is for the setting it came from. It says nothing about transferability — whether that estimate travels to your population, prices, and institutions.
So a distant randomized trial can deserve less weight than a nearby weaker study. The job isn’t to grab the top of the hierarchy; it’s to weigh each study by precision × how much its setting looks like yours — and to keep effect size and statistical significance separate from that.
The evidence scale — predict a local job-training effect
You must estimate how much a job-training program would raise local employment (percentage points). Put studies on the scale, then set how much each foreign setting matches yours.
Grade each study on two axes, then pool. Weight is precision discounted by how far the setting is from yours:
Two axes per study
internal validity -> precision of ITS estimate (1 / variance)
transferability -> how much its setting matches yours (0..1)
Weight for predicting the LOCAL effect
weight = transferability / variance
(a tight foreign estimate with poor context match gets little weight)
Pool
predicted = sum(weight * effect) / sum(weight)
band width grows with 1/sum(weight) + a structural-uncertainty term
qualitative evidence shrinks that structural term by grounding
whether the mechanism operates here
Worked, with the defaults on screen
foreign RCT effect +12 pp, variance 1.5, transfer 0.43
local obs. effect +5 pp, variance 4.0, transfer 0.90
weights 12: 0.29 5: 0.23 -> pooled ~ 8.9 pp
drop "institutions" match and the local study out-weighs the RCT
Watch the separation the tool makes concrete. Each study’s own whisker (its 95% CI) can exclude zero — “significant” where it was run — while the pooled local band still includes zero. Significance in another setting is not evidence of a local effect.
| Weigh evidence this way when… | The limit |
|---|---|
| You’re importing a result from elsewhere to predict a local effect. | Transferability is a judgment, not a measured number — be explicit about it. |
| Studies disagree, or the strongest design comes from the least similar place. | Weights depend on variance and match estimates you may only know roughly. |
| A qualitative or mechanism study can tell you whether the causal story even holds here. | Qualitative work grounds transfer; it rarely sizes the effect on its own. |
An interviewer says: “A famous randomized trial abroad found a job-training program raised employment 12 points. Should we expect that here?” A senior answer refuses the shortcut. It credits the trial’s internal validity, then asks the transfer questions: is our unemployed population similar, is our labor market as slack, do we have the delivery institutions the trial relied on? If institutions differ sharply, the pristine estimate gets discounted — and a rougher local before-and-after study, plus interviews confirming the mechanism, may carry more weight for predicting our effect. The honest output is a range, wider than any single study’s, centered below 12 — and a note that the band still comfortably clears zero, so the direction is safe even if the magnitude isn’t.
Check yourself
A large foreign RCT (very tight CI) and a small local observational study point to different effects. Institutions differ a lot between the two settings. For predicting the local effect, the sound move is…
Each study’s own 95% CI excludes zero, but your pooled local band includes zero. What does that mean?