The evidence base

How closely do emulations match their trials?

Xera OpenScience maintains a living systematic review of every published target trial emulation that reports a randomized benchmark, paired with the trial it emulates. The full database is browsable, filterable and downloadable there; the headline findings are below.

Studies
129
20192026
Comparisons
474
279 in the pooled models
Target trials
94
named randomized benchmarks
Interval coverage
53.4%
149 of 279 comparisons
What it shows

No systematic bias on average. Wide disagreement case by case.

Pooled ratio · HR
1.02
95% CrI 0.96 to 1.07

The emulated hazard ratio divided by the trial hazard ratio. An interval straddling 1 means no detectable average bias in either direction.

Interval coverage
53.4%
149 of 279

How often the emulation’s 95% interval contains the trial estimate. A well-calibrated method would sit near 95%, so agreement on average is not agreement in any single case.

Pooled estimates are Bayesian multilevel models fit in Stan. These figures are read live from the OpenScience API, so this page cannot drift away from the review.