Skip to content

The Trial You Never Ran

Target trial emulation, and why the most famous failure of observational research wasn’t really a failure of the data.

For most of the 2000s, hormone replacement therapy was the story we told to explain why observational studies can’t be trusted. Cohort after cohort had suggested that women taking estrogen and progestin had lower coronary heart disease risk. Then the Women’s Health Initiative, a large randomized trial, found the opposite: risk went up, not down [3]. The gap between the two hardened into a kind of morality tale: proof that confounding will fool you, and that only randomization reveals the truth.

Then came the twist. In 2008, Miguel Hernán and colleagues went back to one of those observational datasets, the Nurses’ Health Study, and re-analyzed it, not with a fancier confounding adjustment but by making it answer the exact question the trial had asked, in the same way the trial had asked it. They defined eligibility, aligned the start of follow-up with the moment therapy began, and compared initiators to non-initiators as an intention-to-treat contrast [3]. Once risk was allowed to vary by time since starting therapy, the observational hazard ratios stopped disagreeing with the trial [3]. The discrepancy had been largely an artifact of how the data were analyzed, not evidence that the data were rotten.

That reframing is the whole idea behind target trial emulation.

The core idea

The framework, formalized by Hernán and Robins [1], inverts the usual order of operations. Instead of asking “what associations live in this database?”, you first write the protocol of the randomized trial you would run if you could: eligibility criteria, the treatment strategies being compared, how and when treatment is assigned, the start of follow-up (“time zero”), the outcome, the follow-up period, the causal contrast you actually care about, and the analysis [5,6]. Only then do you emulate each of those elements with the observational data you have.

This sounds almost too simple to matter. It isn’t a rebranding of “adjust for more confounders.” Its power comes from a different observation: many of the most damaging errors in observational comparative effectiveness research are not confounding at all [2]. They are design errors that a competent trialist would never make, and that a database analyst makes precisely because there is no protocol forcing the discipline of a trial.

Three alignments that prevent self-inflicted injuries

The framework forces you to align three moments in time, all at “time zero”:

  • when eligibility is assessed,
  • when the treatment strategy is assigned,
  • when follow-up begins.

When these drift apart, you manufacture bias. The most notorious is immortal time bias, where a stretch of time during which the outcome could not yet have occurred gets misassigned to the treated group, making a treatment look protective for reasons that are pure bookkeeping [2,4]. Prevalent-user bias, selection on post-baseline events, and reverse causation live in the same neighborhood. None of them is fixed by a better propensity score. They are fixed upstream, by designing the study like a trial before touching the analysis. Hernán and colleagues aptly called these “self-inflicted injuries”: errors we impose on ourselves, then blame on the data [2].

Does it actually work?

Healthy skepticism is warranted, since emulation still runs on non-randomized data, and no design trick randomizes what you didn’t measure. The RCT-DUPLICATE program (Wang, Schneeweiss, and colleagues) offers the most systematic test to date: 32 completed trials were emulated using insurance-claims databases, with the design specified before analysts could tune anything to match the known trial results. Across the trials that could be faithfully emulated, agreement was strong: an overall correlation around 0.82 between the database estimates and the trials [7].

Just as instructive were the misses. Most traced back to measurable differences in design or data: a trial enrolled patients the database couldn’t identify, or measured an outcome the claims data captured differently [7]. That is exactly the kind of discrepancy the target-trial lens surfaces rather than hides. When your emulation and a trial disagree, the framework gives you a structured place to look for why.

What it does not do

Here is the part worth being honest about, because overselling it is how the method gets a bad name. Target trial emulation is a discipline for asking a causal question well. It is not a cure for unmeasured confounding. If a strong prognostic factor is simply absent from your data, no amount of careful emulation will randomize it away. A recent appraisal by Hernán and colleagues makes this explicit: the framework’s contribution is narrower, and more useful, than “observational studies are now as good as trials” [8]. It removes the self-inflicted errors so that what remains is the genuinely hard problem, confounding, which you then have to confront openly, with subject-matter knowledge, sensitivity analyses, and negative controls. The value is in cleanly separating the errors you had no excuse for from the ones you have to argue about.

It is now mainstream enough to have a reporting standard

The clearest signal that a method has matured is when it acquires a reporting checklist. As of September 2025, target trial emulation has one: the TARGET Statement in JAMA, the target-trial analogue of CONSORT for randomized trials and STROBE for observational studies [9]. When a technique gets an EQUATOR-style guideline, it has moved from clever idea to expected practice.

For anyone generating or reviewing real-world evidence (in a regulatory submission, an HTA dossier, or a journal manuscript), this matters practically. The vocabulary of the target trial (specify the estimand, align time zero, state the emulation choices explicitly) is fast becoming the language reviewers expect to see [9]. Studies that don’t speak it increasingly look like they haven’t asked the question carefully.

The real lesson

The lasting lesson of the hormone therapy saga was never “trust randomization, distrust cohorts.” It is subtler and more demanding: a causal question deserves a specified comparison before it deserves a model. Write down the trial you wish you could run. Then measure honestly how close your data can get, and be precise about where it can’t.

That discipline costs a little humility up front. It buys a great deal of credibility later.


What’s your experience? Have you seen a target-trial framing change a study’s conclusion, or expose a bias that a model alone would have missed?


Manuel Pfister

References

  1. Hernán MA, Robins JM. Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available. Am J Epidemiol. 2016;183(8):758-764. doi:10.1093/aje/kwv254
  2. Hernán MA, Sauer BC, Hernández-Díaz S, Platt R, Shrier I. Specifying a target trial prevents immortal time bias and other self-inflicted injuries in observational analyses. J Clin Epidemiol. 2016;79:70-75. doi:10.1016/j.jclinepi.2016.04.014
  3. Hernán MA, Alonso A, Logan R, Grodstein F, Michels KB, Willett WC, Manson JE, Robins JM. Observational studies analyzed like randomized experiments: an application to postmenopausal hormone therapy and coronary heart disease. Epidemiology. 2008;19(6):766-779. doi:10.1097/EDE.0b013e3181875e61
  4. García-Albéniz X, Hsu J, Hernán MA. The value of explicitly emulating a target trial when using real world evidence: an application to colorectal cancer screening. Eur J Epidemiol. 2017;32(6):495-500. doi:10.1007/s10654-017-0287-2
  5. Hernán MA, Wang W, Leaf DE. Target Trial Emulation: A Framework for Causal Inference From Observational Data. JAMA. 2022;328(24):2446-2447. doi:10.1001/jama.2022.21383
  6. Matthews AA, Danaei G, Islam N, Kurth T. Target trial emulation: applying principles of randomised trials to observational studies. BMJ. 2022;378:e071108. doi:10.1136/bmj-2022-071108
  7. Wang SV, Schneeweiss S; RCT-DUPLICATE Initiative. Emulation of Randomized Clinical Trials With Nonrandomized Database Analyses: Results of 32 Clinical Trials. JAMA. 2023;329(16):1376-1385. doi:10.1001/jama.2023.4221
  8. Hernán MA, Dahabreh IJ, Dickerman BA, Swanson SA. The target trial framework for causal inference from observational data: why and when is it helpful? Ann Intern Med. 2025;178(3):402-407. doi:10.7326/ANNALS-24-01871
  9. Cashin AG, Hansford HJ, Hernán MA, et al. Transparent Reporting of Observational Studies Emulating a Target Trial: The TARGET Statement. JAMA. 2025;334(12):1084-1093. doi:10.1001/jama.2025.13350