Methods
XeraTTE is the interface to a living systematic review. Everything here describes that review; the site adds presentation, not analysis.
What a target trial emulation is
A target trial emulation analyzes observational data as though it were a randomized trial. You first write down the trial you would have run: who is eligible, which treatment strategies are compared, how assignment happens, when follow-up starts, what outcome is measured, and how the data will be analyzed. Then you emulate each component in the observational data.
The discipline is the point. Writing the protocol first exposes design errors that otherwise hide in a regression: immortal time bias from starting follow-up before treatment is assigned, prevalent user bias from enrolling people already on treatment, and selection on post-baseline variables. It does not make confounding disappear, which is why benchmarking emulations against real trials is worth doing at all.
The benchmark question
When an emulation and the trial it emulates report the same effect, the two estimates can be compared directly. Across the corpus that comparison is summarized three ways, because they answer different questions:
Pooled ratio of estimates
The emulation estimate divided by the trial estimate, pooled across comparisons. Answers: on average, is there systematic bias in one direction? Ratio measures are pooled on the log scale; the models are Bayesian multilevel models with comparisons nested in studies, fit in Stan.
Interval coverage
How often the emulation’s 95% confidence interval contains the trial’s point estimate. Answers: would trusting a single emulation have led you right? A calibrated method would sit near 95%; the observed figure is 53.4%.
Concordance correlation
Lin’s CCC on the paired estimates, penalizing both weak correlation and systematic offset. Answers: do the two methods rank and size effects the same way?
The randomized-benchmark gate
A comparison is only as good as the trial number it is measured against. Emulation papers sometimes quote a benchmark that does not appear in the trial’s primary report, or quote it on a different scale or follow-up horizon. Every extracted comparison was therefore checked against the trial’s own publication before entering the pooled models.
Where the emulation’s quoted benchmark conflicts with the trial’s primary report, the primary report wins, and the substitution is recorded in the correction ledger. Where the benchmark could not be verified at all, the comparison is withheld and shown with that reason attached rather than silently dropped.
What the gate decided
| Disposition | Comparisons |
|---|---|
| Retained | 279 |
| Withheld as unverifiable | 89 |
| Incompatible benchmark | 7 |
375 otherwise eligible rows were reviewed at this gate. A further set of comparisons never reached it, having been excluded earlier for incompatible effect scales, a missing estimate, or a reference trial that had not reported before the emulation was published. Browse the withheld comparisons.
Search and screening
Study flow
- PubMed records
- 741
- OpenAlex records
- 5,807
- Unified corpus
- 5,966
- Screened on title and abstract
- 3,308
- Included at title and abstract
- 140
- Full texts sought
- 77
- Full texts assessed
- 51
- Eligible after full text
- 45
- Studies in the analysis
- 69
- Comparisons in the analysis
- 279
- Named target trials
- 94
Search recall against a held-out set of known includes: 98.7% (75 of 76).
Known limitations
Outcome direction is not harmonized. Comparisons are pooled on the ratio of emulation to trial estimate, which is direction-free. Pooling the effects themselves would need a clinical direction for every outcome, and that field has not been extracted.
Two measures are descriptive only. MD (Raw mean differences use heterogeneous outcome units; coverage and standardized discrepancy are reported instead.) RD (Fewer than 3 independent TTE reports; a random-effects heterogeneity model would be prior-driven.)
Transparency is a disclosure count. The 0-5 indicator counts which of COI, funding, protocol, data and code are present. It does not establish that shared data are reusable or that a protocol was followed.
Publication-bias tests are weak here. Egger regression on a few dozen reports has little power, and several emulations reuse the same target trial, which the test does not account for.