How closely do emulations match their trials?
Xera OpenScience maintains a living systematic review of every published target trial emulation that reports a randomized benchmark, paired with the trial it emulates. The full database is browsable, filterable and downloadable there; the headline findings are below.
No systematic bias on average. Wide disagreement case by case.
The emulated hazard ratio divided by the trial hazard ratio. An interval straddling 1 means no detectable average bias in either direction.
How often the emulation’s 95% interval contains the trial estimate. A well-calibrated method would sit near 95%, so agreement on average is not agreement in any single case.
Pooled estimates are Bayesian multilevel models fit in Stan. These figures are read live from the OpenScience API, so this page cannot drift away from the review.
Browse the studies
Every emulation in the review, filterable by clinical area, data type, method and disclosure.
Compare against the trials
One row per comparison, with the trial estimate and the emulated estimate side by side on a shared scale.
Read the analysis
Pooled models, heterogeneity, coverage, small-study effects and the sensitivity analyses behind them.