Pooled results

How well do emulations reproduce their trials?

279 comparisons from 69 studies against 94 named randomized trials. Estimates below come from Bayesian multilevel models with comparisons nested in studies; this page reports them, it does not recompute them.

Pooled ratio · HR
1.02
95% CrI 0.96–1.07
Pooled ratio · RR
1.17
95% CrI 0.88–1.58
Interval coverage
53.4%
149 of 279
Concordance · HR
0.42
Lin's CCC

Pooled agreement by effect measure

ratio of emulation to trial estimate
MeasureComparisonsStudiesPooled ratio (95% CrI)Prediction intervalBetween-study SD
HR
Hazard ratio
228581.02
0.96 to 1.07
0.68 to 1.52τ = 0.12
0.09 to 0.15
RR
Risk ratio
34101.17
0.88 to 1.58
0.39 to 3.54τ = 0.39
0.24 to 0.58
MD
Mean difference
157Not pooled
Raw mean differences use heterogeneous outcome units; coverage and standardized discrepancy are reported instead.
RD
Risk difference
22Not pooled
Fewer than 3 independent TTE reports; a random-effects heterogeneity model would be prior-driven.

Interval coverage

emulation CI contains the trial estimate
HR · 130/22857.0%
95% CI 49.2% to 65.6%
MD · 10/1566.7%
95% CI 27.3% to 94.7%
RD · 0/20.0%
RR · 9/3426.5%
95% CI 17.9% to 38.5%

A well-calibrated emulation would sit near 95%. Overall coverage is 53.4%; allowing agreement in either direction raises it to 65.9%.

Concordance

Lin's CCC
HR · 228 comparisons, 58 studies0.42
95% CI 0.25 to 0.60
RD · 2 comparisons, 2 studies
sparse/exploratory
RR · 34 comparisons, 10 studies0.19
95% CI -0.17 to 0.49
MD · 15 comparisons, 7 studies
not estimable across heterogeneous units

CCC measures agreement between paired estimates on a −1 to 1 scale, penalizing both poor correlation and systematic offset. A pooled ratio near 1 with a middling CCC is exactly the pattern here: no average bias, weak agreement case by case.

Agreement by funding

mean standardized discrepancy
None · 5 studies0.79
Not declared · 5 studies2.07
Other/Mixed · 7 studies0.79
Private/commercial · 15 studies1.53
Public · 37 studies1.23

Exploratory. Distance in pooled standard errors between the emulation and trial estimate; higher means further apart. The study counts are small, so read this as a description of the corpus, not a causal claim about funders.