How well do emulations reproduce their trials?
279 comparisons from 69 studies against 94 named randomized trials. Estimates below come from Bayesian multilevel models with comparisons nested in studies; this page reports them, it does not recompute them.
Pooled agreement by effect measure
| Measure | Comparisons | Studies | Pooled ratio (95% CrI) | Prediction interval | Between-study SD |
|---|---|---|---|---|---|
| HR Hazard ratio | 228 | 58 | 1.02 0.96 to 1.07 | 0.68 to 1.52 | τ = 0.12 0.09 to 0.15 |
| RR Risk ratio | 34 | 10 | 1.17 0.88 to 1.58 | 0.39 to 3.54 | τ = 0.39 0.24 to 0.58 |
| MD Mean difference | 15 | 7 | Not pooled Raw mean differences use heterogeneous outcome units; coverage and standardized discrepancy are reported instead. | ||
| RD Risk difference | 2 | 2 | Not pooled Fewer than 3 independent TTE reports; a random-effects heterogeneity model would be prior-driven. | ||
Interval coverage
A well-calibrated emulation would sit near 95%. Overall coverage is 53.4%; allowing agreement in either direction raises it to 65.9%.
Concordance
CCC measures agreement between paired estimates on a −1 to 1 scale, penalizing both poor correlation and systematic offset. A pooled ratio near 1 with a middling CCC is exactly the pattern here: no average bias, weak agreement case by case.
Agreement by funding
Exploratory. Distance in pooled standard errors between the emulation and trial estimate; higher means further apart. The study counts are small, so read this as a description of the corpus, not a causal claim about funders.