How this was built

Methods

XeraTTE is the interface to a living systematic review. Everything here describes that review; the site adds presentation, not analysis.

What a target trial emulation is

A target trial emulation analyzes observational data as though it were a randomized trial. You first write down the trial you would have run: who is eligible, which treatment strategies are compared, how assignment happens, when follow-up starts, what outcome is measured, and how the data will be analyzed. Then you emulate each component in the observational data.

The discipline is the point. Writing the protocol first exposes design errors that otherwise hide in a regression: immortal time bias from starting follow-up before treatment is assigned, prevalent user bias from enrolling people already on treatment, and selection on post-baseline variables. It does not make confounding disappear, which is why benchmarking emulations against real trials is worth doing at all.

The benchmark question

When an emulation and the trial it emulates report the same effect, the two estimates can be compared directly. Across the corpus that comparison is summarized three ways, because they answer different questions:

Pooled ratio of estimates

The emulation estimate divided by the trial estimate, pooled across comparisons. Answers: on average, is there systematic bias in one direction? Ratio measures are pooled on the log scale; the models are Bayesian multilevel models with comparisons nested in studies, fit in Stan.

Interval coverage

How often the emulation’s 95% confidence interval contains the trial’s point estimate. Answers: would trusting a single emulation have led you right? A calibrated method would sit near 95%; the observed figure is 53.4%.

Concordance correlation

Lin’s CCC on the paired estimates, penalizing both weak correlation and systematic offset. Answers: do the two methods rank and size effects the same way?

The randomized-benchmark gate

A comparison is only as good as the trial number it is measured against. Emulation papers sometimes quote a benchmark that does not appear in the trial’s primary report, or quote it on a different scale or follow-up horizon. Every extracted comparison was therefore checked against the trial’s own publication before entering the pooled models.

Where the emulation’s quoted benchmark conflicts with the trial’s primary report, the primary report wins, and the substitution is recorded in the correction ledger. Where the benchmark could not be verified at all, the comparison is withheld and shown with that reason attached rather than silently dropped.

What the gate decided

DispositionComparisons
Retained279
Withheld as unverifiable89
Incompatible benchmark7

375 otherwise eligible rows were reviewed at this gate. A further set of comparisons never reached it, having been excluded earlier for incompatible effect scales, a missing estimate, or a reference trial that had not reported before the emulation was published. Browse the withheld comparisons.

Search and screening

Study flow

search to 2026-08-09
PubMed records
741
OpenAlex records
5,807
Unified corpus
5,966
Screened on title and abstract
3,308
Included at title and abstract
140
Full texts sought
77
Full texts assessed
51
Eligible after full text
45
Studies in the analysis
69
Comparisons in the analysis
279
Named target trials
94

Search recall against a held-out set of known includes: 98.7% (75 of 76).

Known limitations

Outcome direction is not harmonized. Comparisons are pooled on the ratio of emulation to trial estimate, which is direction-free. Pooling the effects themselves would need a clinical direction for every outcome, and that field has not been extracted.

Two measures are descriptive only. MD (Raw mean differences use heterogeneous outcome units; coverage and standardized discrepancy are reported instead.) RD (Fewer than 3 independent TTE reports; a random-effects heterogeneity model would be prior-driven.)

Transparency is a disclosure count. The 0-5 indicator counts which of COI, funding, protocol, data and code are present. It does not establish that shared data are reusable or that a protocol was followed.

Publication-bias tests are weak here. Egger regression on a few dozen reports has little power, and several emulations reuse the same target trial, which the test does not account for.