Pith. sign in

REVIEW 3 major objections 5 minor 79 references

The efficiency gain from adding real-world data to a trial is real but narrow: with A-TMLE it crosses break-even near one residual SD of bias, erodes as the trial grows, and survives a block-jackknife interval in only one of six real fusion

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 08:55 UTC pith:AEMNAXC4

load-bearing objection A genuinely useful, unusually candid empirical-methods paper that maps when A-TMLE fusion actually pays; the block-jackknife guardrail is the most valuable and also the most fragile piece. the 3 major comments →

arxiv 2607.02787 v2 pith:AEMNAXC4 submitted 2026-07-02 stat.ME

When Does Trial-Real-World Data Fusion Improve Precision? Model Auditing and Selection-Aware Inference for Adaptive-TMLE

classification stat.ME
keywords adaptive targeted maximum likelihood estimationtrial–real-world data fusionefficiency gainselection-aware inferenceblock jackknifebias magnitudehighly adaptive lassoreal-world evidence
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Using adaptive targeted maximum likelihood estimation (A-TMLE) as a worked example, this paper asks when fusing a randomized trial with real-world data actually narrows the confidence interval for the treatment effect, and how to state that gain honestly. It finds that the gain is governed mainly by the magnitude of the real-world bias, not by the complexity of the bias function: the variance of the estimator rises quadratically with bias magnitude, so the efficiency ratio crosses parity near a bias of about one residual standard deviation and falls badly at larger bias. The gain also shrinks as the trial sample size grows, so it is finite-sample rather than a super-efficiency result. The paper then treats the gain itself as a data-adaptive estimand and shows that, among ten candidate standard errors, only a block jackknife that re-selects the working model gives near- or above-nominal coverage; the naive standard error undercovers. In six real-data fusions, that conservative interval keeps the trial-only analysis primary in five of six — the toolkit acts as a guardrail, not a booster.

Core claim

The paper's central claim is that the finite-sample efficiency of A-TMLE for RCT-plus-real-world fusion is modest and reference-relative. Writing the ATE as a pooled projection minus a learned bias-correction term, the estimator's influence curve is D = D_A − D_S, and the efficiency gain is R = var(D_rct)/var(D_atmle) relative to a matched, correctly-specified trial-only estimator. An exact population-oracle identity shows var(D_A) = a + b m² under a restricted working model: the bias magnitude m enters quadratically with no linear term and no shape dependence at that order, explaining why magnitude dominates complexity in the simulation map. The gain starts near 1.15 at zero bias, crosses o

What carries the argument

Central object is the A-TMLE decomposition of the trial ATE into a pooled-projection estimand and a bias projection built from the learned enrollment-effect surface τ_S(W,A); the working model is a relaxed highly-adaptive-lasso (HAL) basis selected by cross-validation and then targeted. The efficiency gain R is the ratio of influence-curve variances, var(D_rct)/var(D_atmle), with D = D_A − D_S. The argument is carried by Proposition 1, an exact population-oracle identity var(D_A) = a + b m² showing bias magnitude enters quadratically as the leading variance driver, and by a delete-a-fold block jackknife that re-selects the working model on each leave-fold-out subsample, giving conservative n

Load-bearing premise

The headline efficiency map assumes trial enrollment is completely random and the within-trial effect is homogeneous, so the matched trial-only estimator is correctly specified by construction and the comparison is a pure variance contrast; if enrollment depends on covariates or treatment effects vary, the break-even point and the magnitude-dominance ordering could shift.

What would settle it

Re-run the main 15-cell grid with W-dependent trial enrollment and a heterogeneous within-trial effect (for example, CATE = 1.5 + 0.8W1 − 0.5W2) at n_rct = n_ext = 250. If the efficiency gain then crosses parity at m < 0.5, or if functional complexity rather than magnitude orders the cells, or if the block-jackknife interval's coverage drops below nominal in the low-bias cells, the paper's headline map and selection-aware verdict are contradicted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Practitioners should report the efficiency gain with a block-jackknife interval and claim an efficiency improvement only when its lower bound exceeds one; the naive influence-function standard error is unsafe.
  • The break-even near one residual SD of real-world bias gives a concrete stopping rule: if the external cohort is expected to be biased at or beyond that scale, fusion is unlikely to buy precision at moderate trial sizes.
  • The asymptotic oracle super-efficiency guarantee of A-TMLE does not translate to a finite-sample advantage; the gain erodes as the trial grows and can be below one against an efficient trial-only reference.
  • The learned bias model should be reported as a stress diagnostic (effective basis count, variance attribution, targeting drift), not interpreted as a measure of confounding or a proxy for the gain.
  • When trial enrollment depends on covariates and the bias surface is rough and large, A-TMLE's own ATE interval can undercover with main-terms nuisances; flexible nuisance fits restore coverage, so the envelope matters.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the quadratic magnitude-dominance identity extends beyond restricted oracle models, similar break-even maps should appear for other debiased estimators that borrow through a learned correction; that is a testable transfer, not something the paper establishes.
  • The one real fusion whose interval clears parity rests on a four-basis correction for a prognostically distant external arm, suggesting the method's value may be concentrated precisely in the large, strongly biased external cohorts the public examples do not contain.
  • A practical extension would be to convert the block-jackknife width ratio into a pre-study power or sample-size tool: the conservative interval implies that detecting a gain near 1.15 requires either very large trials or unusually small bias, which could inform whether to collect real-world data at all.
  • The report card's targeting drift, not the basis count, flagged the one coverage-failure corner; an analyst facing a single fusion with no comparison panel cannot yet use drift as a calibrated detector, so a calibration study on real data would be a natural next step.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper uses adaptive targeted maximum likelihood estimation (A-TMLE) as a worked example of adaptive trial-plus-real-world-data fusion and develops three tools: (1) a report card that audits the learned bias model via surface recovery, influence-curve variance attribution, and targeting drift; (2) a simulation-based efficiency map of the finite-sample variance gain of A-TMLE relative to a matched RCT-only AIPW, with the gain driven mainly by bias magnitude rather than functional complexity, crossing parity near a bias of about one residual SD and eroding as the trial grows; and (3) selection-aware inference for the efficiency gain, culminating in a delete-a-fold block jackknife as the only one of ten candidate standard errors with consistently near- or above-nominal coverage. The tools are applied to six real-data fusions from three openly available trials; in five of six the block-jackknife interval includes one, so the paper recommends keeping the RCT-only estimate primary, with only one marginal exception.

Significance. If correct, this is a valuable and timely contribution. The paper is unusually candid about its scope: the headline map is restricted to constant enrollment probability and constant CATE, the block-jackknife result is explicitly empirical rather than a theorem, and the real-data illustrations are partly based on constructed external arms. The simulation work is extensive (15,000 A-TMLE fits with zero failures; 40 fixed-truth cells for the standard-error head-to-head) and the analysis pipeline is reproducible from public code. The identification of a conservative selection-aware interval for the efficiency gain, if independently validated, would be an important guardrail for practice and a useful caution against naive influence-function standard errors after model selection. The paper is well within the scope of the journal and deserves serious consideration.

major comments (3)
  1. [Section 4.6, Table 8, Figure 4] The block jackknife is recommended after examining its coverage on the same 15-cell grid and the robustness slices. Because it was selected as the best of ten candidates using those same results, the reported coverage (0.984–0.998) is a selected maximum and may be optimistic; the additional robustness cells do not break the selection loop because they were evidently examined before the recommendation was finalized. This matters because the Section 3.2 decision rule and the Section 5 primary-analysis verdicts rest entirely on this empirical calibration, and the paper explicitly provides no theorem for this non-smooth statistic. I ask for an independent validation strategy — e.g., a pre-specified holdout set of DGP configurations not examined during method selection, or a separate selection/validation split of the simulation grid — or, failing that, a clear re-labeling of the guardrail as
  2. [Abstract; Sections 2, 4.3, 4.5] The abstract states the efficiency-map findings ('driven mainly by the magnitude of the real-world bias... crosses break-even near a moderate bias and erodes as the trial grows') without the scope caveat that the main grid assumes constant trial-enrollment probability and a constant within-trial effect (Section 2 scope caveat; Section 4.1). Under these restrictions the matched GLM reference is correctly specified on the trial arm, so the comparison is a pure variance contrast. The paper itself notes that under a heterogeneous effect the matched-GLM gain sits at parity even at zero bias, and under selective enrollment A-TMLE's own ATE interval degrades at the rough large-bias corner (Section 4.5). The abstract and the title's promise of a general 'when' map should carry the same caveat; otherwise readers may take a design-specific map as a general characterization.
  3. [Section 4.4, Proposition 1] The population identity var(D_A)=a+b m^2 is derived under a forced intercept-only oracle working model (Φ≡1), which is not the estimator deployed in the simulations or real data. The authors are careful to call this an oracle restriction, but the abstract's phrase 'a dominance an exact population-oracle variance identity explains' could oversell the explanation: the identity does not cover the HAL-selected working model, and the finite-sample shape ordering is actually reversed. The paper addresses this, but the abstract and the 'takeaway' would be clearer if the identity were described as a diagnostic analogue under a restricted oracle model rather than 'the' explanation of the empirical magnitude-dominance finding.
minor comments (5)
  1. [Abstract] Please add a sentence stating that the efficiency map and the parity crossing are established under constant trial enrollment and a homogeneous within-trial effect, with secondary robustness checks for W-dependent enrollment and heterogeneous CATE.
  2. [Table 8] The column 'Cover., own mean' is explicitly circular; consider adding '(circular)' to the column header or footnote so that the diagnostic purpose is immediately clear and not misread as a valid coverage estimate.
  3. [Section 4.4] The sentence about the population shape ordering being the reverse of the empirical finite-sample ordering is easy to misread. Please spell out in one or two sentences that Proposition 1 concerns the oracle intercept-only influence curve, while the deployed HAL-selected estimator is affected by finite-sample basis selection and nuisance estimation error.
  4. [Section 5.3, Table 12] The basis-count contrast between PSID (one basis) and CPS (seven bases) is a nice illustration, but the text already notes the count is not monotone in complexity and may reflect power; consider adding a one-sentence reference to Section 4.3 near the table so the reader does not over-interpret the count.
  5. [Throughout] There are occasional tense shifts between 'we show' and 'we do not establish' that make it harder to track which claims are the paper's contributions versus caveats. A short 'evidence class' table (as in Table 13) for Section 4.6 would help.

Circularity Check

0 steps flagged

No significant circularity: the efficiency map, Proposition 1, and the selection-aware SE recommendation are independent of their inputs, and the only in-sample selection concern is not a definitional reduction.

full rationale

Walking the derivation chain: (i) The efficiency map (Sec. 4.3, Table 4) is a Monte Carlo variance-ratio benchmark against a matched, correctly-specified GLM AIPW; the gain R is defined independently as var(Drct)/var(Datmle), and the 'magnitude not complexity' claim is supported by simulation plus Proposition 1 (Sec. 4.4), whose constants a and b are computed from DGP objects and stated assumptions, not fitted to the gain: 'the population constant a=ccσ2=4.18 reproduces the empirical m=0 value of var(DA) to three digits ... per-shape b={1.49,1.38,1.16} computed, not fit.' No quantity in the map is defined in terms of the gain or vice versa. (ii) The block-jackknife recommendation (Sec. 4.6) is explicitly an empirical calibration result, not a theorem: the paper states 'This is an empirical calibration result for A-TMLE, not a general theorem' and 'we do not claim it is consistent—the statistic is non-smooth.' That it was selected by comparing ten candidates on the studied grid is a winner-selection/external-validity limitation, not a circular reduction; the coverage numbers are measurements, not consequences of the selection rule. (iii) The real-data verdicts use open datasets and are reported as single-dataset illustrations, with the limitation 'the real-data sections are best read as faithful end-to-end demonstrations of the toolkit rather than as evidence on the magnitude of genuine confounding.' (iv) The paper's citations to van der Laan et al. [2026] define the estimator under study and supply the claim being tested; the paper also diverges from that claim, so it is not load-bearing self-citation, and the present author is not an author of the cited A-TMLE paper. No step reduces by equations to its own inputs; therefore score 0.

Axiom & Free-Parameter Ledger

3 free parameters · 8 axioms · 0 invented entities

The central claims rest on the A-TMLE identification assumptions it inherits (trial-enrollment positivity, external overlap, standard rates/regularity), plus two disclosed design restrictions: constant enrollment probability and homogeneous CATE in the main grid, and the Proposition 1 forced intercept-only working model. The paper explicitly labels the latter as an assumed restriction and the scope caveat as making the headline conclusions regime-specific. No fitting of constants to the target gain is load-bearing: the Proposition 1 constants (a=4.18, b≈1.34) are computed from DGP objects, not fit to R. No new entities are postulated; the report-card diagnostics are measurements of existing fitted objects.

free parameters (3)
  • var(D_A) corroboration regression coefficients = intercept 3.25, m² slope 1.40, basis-count slope 0.094 (R²=0.989)
    Fit to the same 15 simulation cells that define the efficiency map; the paper labels the standardized magnitude-vs-complexity effect (0.91 vs 0.11) 'indicative... rather than a clean variance decomposition' because the basis count d̄ is endogenous and collinear with m² (Section 4.4).
  • HAL working-model tuning (knots, degree, penalty) = 5 knots; degree 3 for the working model, degree 2 for the relaxed-HAL reference; penalty multiplicity nλ=1 default
    Hand-chosen defaults inherited from the atmle package; Section 4.5 Panel C shows neither knot enlargement (5→10→20) nor penalty relaxation (nλ=1,3,5) repairs the selective-enrollment coverage collapse, so tuning is not the driver, but it shapes the report-card basis counts and the map.
  • selective-enrollment stress design = central 90% of Π(W) in [0.10,0.90]; 1.4% of covariates outside [0.05,0.95]
    Design stress parameter for the safe-operating-envelope analysis (Section 4.5); the paper argues near-positivity alone does not explain the wiggly-surface collapse because comparison surfaces under the identical Π(W) keep 0.94–0.96 coverage.
axioms (8)
  • domain assumption Trial-enrollment positivity: 0 < Π(W) = P(S=1|W) < 1 P_W-a.e.
    A-TMLE's stated price for fusing external data (Section 2); it places the external covariate support inside the trial's support and identifies the bias projection Ψ#.
  • domain assumption External arm-specific overlap P(A=1|S=0,W) ∈ (0,1) when both external arms are used
    Section 2, Eq. (4) discussion; required for the two-term bias projection used in the ACTG175 and WASH fusions.
  • standard math Consistency/no interference (Y=Y(A)), RCT randomization, trial treatment positivity
    Section 2 identification of the trial-population ATE from within-trial conditionals; standard causal assumptions, inherited from the A-TMLE framework.
  • standard math Donsker/empirical-process and n^{-1/4} rate conditions for asymptotic linearity
    Section 2: asymptotic linearity of the deployed estimator requires both second-order terms (working-model approximation error, TMLE remainder) to be o_P(n^{-1/2}); regularity conditions inherited from van der Laan et al. 2026.
  • ad hoc to paper Forced intercept-only working model (Φ≡1) in Proposition 1
    Section 4.4: 'an assumed restriction, since a homogeneous within-trial effect does not by itself make the pooled projection τA(W) constant'; the magnitude-dominance identity var(D_A)=a+bm² is exact only under this restriction, not for the deployed HAL-selected estimator.
  • ad hoc to paper Constant trial-enrollment probability in the main simulation grid
    Section 4.1 and the Section 2 scope caveat: trial membership is assigned deterministically, so Π(W) is constant and enrollment positivity holds trivially; the paper states the headline report-card and efficiency-map conclusions are strictly about this constant-positivity regime.
  • domain assumption Homogeneous within-trial effect (constant CATE = 1.5) in the main design
    Section 4.1: makes the efficiency comparison a pure variance contrast and the matched-GLM reference correctly specified on the trial arm by construction; the heterogeneous-effect relaxation runs only at n_rct ∈ {250,400}.
  • standard math Conditional mean-zero outcome noise E[U_Y | W, A, S] = 0
    Section 4.4, Proposition 1 proof: kills the linear-in-m term and the bias–noise cross term in var(D_A); the paper notes this uses only conditional mean-zero, not S⊥(A,W).

pith-pipeline@v1.3.0-alltime-deepseek · 36349 in / 22526 out tokens · 209867 ms · 2026-08-02T08:55:58.220218+00:00 · methodology

0 comments
read the original abstract

Augmenting a randomized controlled trial (RCT) with real-world data (RWD) promises greater efficiency, but how much a given fusion delivers, and how to attach honest uncertainty to that gain, are rarely characterized. Using adaptive targeted maximum likelihood estimation (A-TMLE) as a worked example of an estimator that learns a working model and then debiases it, we develop three reproducible tools for reliable evidence from combined trial and real-world data. First, a report card that makes the data-adaptively learned bias model auditable: on simulated data it measures how well the model recovers the true enrollment-effect surface and attributes the estimator's variance to its structural parts. Second, a map of when fusion helps versus hurts, benchmarked against a matched trial-only estimator; the efficiency gain is driven mainly by the magnitude of the real-world bias rather than its functional complexity (a dominance an exact population-oracle variance identity explains), it crosses break-even near a moderate bias and erodes as the trial grows, so the advantage is finite-sample, not super-efficiency. Third, selection-aware inference for the gain, treated as a data-adaptive estimand: the naive standard error undercovers, and among ten candidate standard errors only a block jackknife achieved consistently near- or above-nominal coverage, though conservatively. Across six fusions of three openly available trials (a biomedical HIV trial, a public-health trial, and a job-training trial), only one interval clears one, and only marginally; in the rest, fusion has not earned an efficiency claim over the RCT alone. On real data the toolkit therefore functions mainly as a guardrail: the learned-model dimension is a stress diagnostic, not a proxy for ground truth, and the block-jackknife interval decides whether fusion or the RCT-only analysis should be primary.

Figures

Figures reproduced from arXiv: 2607.02787 by M. Ehsan Karim.

Figure 1
Figure 1. Figure 1: Schematic. A-TMLE decomposes the trial-population ATE into a pooled projection Ψe (via the working model τA) minus a bias projection Ψ# (via the learned τS); the efficient influence curve is D = DA − DS. Contribution (1) audits τS (the report card); contribution (2) maps the gain R; contribution (3) builds a calibrated interval for R via the block jackknife. Alt text: Flow diagram. Trial-plus-real-world da… view at source ↗
Figure 1
Figure 1. Figure 1: Schematic. A-TMLE decomposes the trial-population ATE into a pooled projection Ψe (via the working model τA) minus a bias projection Ψ# (via the learned τS); the efficient influence curve is D = DA − DS. Contribution (1) audits τS (the report card); contribution (2) maps the gain R; contribution (3) builds a selection-aware block-jackknife interval for R. Alt text: Flow diagram. Trial-plus-real-world data … view at source ↗
Figure 2
Figure 2. Figure 2: Recovery surface. The learned bias model τbS(W, A) (solid, from cross-validated relaxed￾HAL) against the truth τS,0(W, A) = −B(W, A) (dashed), sliced at W2 = W3 = 0, by treatment arm (A = 0/A = 1, columns) and scenario (rows). The report card reproduces the structure of the enrollment-effect surface, including the arm-specific W1-dependence. Alt text: Grid of line plots. In each panel the learned bias-mode… view at source ↗
Figure 2
Figure 2. Figure 2: Recovery surface. The learned bias model τbS(W, A) (solid, from cross-validated relaxed￾HAL) against the truth τS,0(W, A) = −B(W, A) (dashed), sliced at W2 = W3 = 0, by treatment arm (A = 0/A = 1, columns) and scenario (rows). The report card reproduces the structure of the enrollment-effect surface, including the arm-specific W1-dependence. Alt text: Grid of line plots. In each panel the learned bias-mode… view at source ↗
Figure 3
Figure 3. Figure 3: The efficiency gain map. Influence-curve gain (6) versus bias magnitude m, by complexity. The gain falls monotonically with magnitude and crosses parity just above m ≈ 1; complexity separates the curves only at large magnitude. Alt text: Line plot of the efficiency gain (vertical axis) against bias magnitude m (horizontal axis) for three bias complexities. All three curves fall monotonically with m and cro… view at source ↗
Figure 3
Figure 3. Figure 3: The efficiency gain map. Influence-curve gain (6) versus bias magnitude m, by complexity. The gain falls monotonically with magnitude and crosses parity just above m ≈ 1; complexity separates the curves only at large magnitude. Alt text: Line plot of the efficiency gain (vertical axis) against bias magnitude m (horizontal axis) for three bias complexities. All three curves fall monotonically with m and cro… view at source ↗
Figure 4
Figure 4. Figure 4: Fixed-truth coverage of the ten selection-aware SEs. 95% CI coverage of the efficiency gain, scored against the locked B = 1000 truth, by bias magnitude m and shape (faceted); the dashed line marks the 0.95 target. The block jackknife (gold) is the only method to reach the target (0.98–1.00); the other nine fall entirely within the shaded min–max envelope (never above 0.87), undercovering across all three … view at source ↗
Figure 4
Figure 4. Figure 4: Fixed-truth coverage of the ten selection-aware SEs. 95% CI coverage of the efficiency gain, scored against the locked B = 1000 truth, by bias magnitude m and shape (faceted); the dashed line marks the 0.95 target. The block jackknife (gold) is the only method to reach the target (0.98–1.00); the other nine fall entirely within the shaded min–max envelope (never above 0.87), undercovering across all three … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

79 extracted references · 9 linked inside Pith

  1. [1]

    and Qiu, Sky and Tarp, Jens Magelund and van der Laan, Lars , title =

    van der Laan, Mark J. and Qiu, Sky and Tarp, Jens Magelund and van der Laan, Lars , title =. Journal of Causal Inference , year =

  2. [4]

    and Rubin, Daniel , title =

    van der Laan, Mark J. and Rubin, Daniel , title =. The International Journal of Biostatistics , year =

  3. [5]

    and Rose, Sherri , title =

    van der Laan, Mark J. and Rose, Sherri , title =. 2011 , isbn =

  4. [6]

    and Rose, Sherri , title =

    van der Laan, Mark J. and Rose, Sherri , title =. 2018 , isbn =

  5. [7]

    , title =

    Zheng, Wenjing and van der Laan, Mark J. , title =. Targeted Learning: Causal Inference for Observational and Experimental Data , editor =. 2011 , doi =

  6. [8]

    , title =

    Benkeser, David and van der Laan, Mark J. , title =. 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA) , pages =. 2016 , doi =

  7. [9]

    , title =

    van der Laan, Mark J. , title =. The International Journal of Biostatistics , year =

  8. [10]

    and Coyle, Jeremy R

    Hejazi, Nima S. and Coyle, Jeremy R. and van der Laan, Mark J. , title =. Journal of Open Source Software , year =

  9. [11]

    and Polley, Eric C

    van der Laan, Mark J. and Polley, Eric C. and Hubbard, Alan E. , title =. Statistical Applications in Genetics and Molecular Biology , year =

  10. [12]

    and Hejazi, Nima S

    Coyle, Jeremy R. and Hejazi, Nima S. and Malenica, Ivana and Phillips, Rachael V. and Sofrygin, Oleg , title =. 2023 , note =

  11. [13]

    and Kherad-Pajouh, Sara and van der Laan, Mark J

    Hubbard, Alan E. and Kherad-Pajouh, Sara and van der Laan, Mark J. , title =. The International Journal of Biostatistics , year =

  12. [14]

    The Econometrics Journal , year =

    Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Duflo, Esther and Hansen, Christian and Newey, Whitney and Robins, James , title =. The Econometrics Journal , year =

  13. [15]

    and Balakrishnan, Sivaraman and Wasserman, Larry , title =

    Kuchibhotla, Arun K. and Balakrishnan, Sivaraman and Wasserman, Larry , title =. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , year =

  14. [16]

    The Annals of Statistics , year =

    Efron, Bradley and Stein, Charles , title =. The Annals of Statistics , year =

  15. [17]

    , title =

    Quenouille, Maurice H. , title =. Biometrika , year =

  16. [18]

    and Klaassen, Chris A

    Bickel, Peter J. and Klaassen, Chris A. J. and Ritov, Ya'acov and Wellner, Jon A. , title =. 1993 , isbn =

  17. [19]

    The American Statistician , year =

    Hines, Oliver and Dukes, Oliver and Diaz-Ordaz, Karla and Vansteelandt, Stijn , title =. The American Statistician , year =

  18. [20]

    Causal Inference Methods for Combining Randomized Trials and Observational Studies: A Review , journal =

    Colnet, B. Causal Inference Methods for Combining Randomized Trials and Observational Studies: A Review , journal =. 2024 , volume =

  19. [21]

    Journal of the Royal Statistical Society: Series B (Statistical Methodology) , year =

    Yang, Shu and Gao, Chenyin and Zeng, Donglin and Wang, Xiaofei , title =. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , year =

  20. [22]

    Journal of the American Statistical Association , year =

    Yang, Shu and Ding, Peng , title =. Journal of the American Statistical Association , year =

  21. [23]

    Rosenman, Evan T. R. and Basse, Guillaume and Owen, Art B. and Baiocchi, Mike , title =. Biometrics , year =

  22. [24]

    2018 , month =

    Framework for. 2018 , month =

  23. [25]

    34) , year =

    21st Century Cures Act (Public Law 114-255, H.R. 34) , year =

  24. [26]

    and Rahman, Mahbubur and Arnold, Benjamin F

    Luby, Stephen P. and Rahman, Mahbubur and Arnold, Benjamin F. and Unicomb, Leanne and others , title =. The Lancet Global Health , year =

  25. [27]

    and Petersen, Maya and van der Laan, Mark , title =

    Dang, Lauren Eyler and Tarp, Jens Magelund and Abrahamsen, Trine Julie and Kvist, Kajsa and Buse, John B. and Petersen, Maya and van der Laan, Mark , title =. Journal of Causal Inference , year =

  26. [28]

    and Katzenstein, David A

    Hammer, Scott M. and Katzenstein, David A. and Hughes, Michael D. and Gundacker, Holly and Schooley, Robert T. and Haubrich, Richard H. and Henry, W. Keith and Lederman, Michael M. and Phair, John P. and Niu, Manette and Hirsch, Martin S. and Merigan, Thomas C. , title =. New England Journal of Medicine , year =

  27. [29]

    and Lu, Xiaomin and Zhang, Min and Davidian, Marie and Tsiatis, Anastasios A

    Juraska, Michal and Gilbert, Peter B. and Lu, Xiaomin and Zhang, Min and Davidian, Marie and Tsiatis, Anastasios A. , title =. 2022 , note =

  28. [30]

    and Robertson, Sarah E

    Dahabreh, Issa J. and Robertson, Sarah E. and Steingrimsson, Jon A. and Stuart, Elizabeth A. and Hern. Extending inferences from a randomized trial to a new target population , journal =. 2020 , volume =

  29. [31]

    and Cole, Stephen R

    Stuart, Elizabeth A. and Cole, Stephen R. and Bradshaw, Catherine P. and Leaf, Philip J. , title =. Journal of the Royal Statistical Society: Series A (Statistics in Society) , year =

  30. [32]

    and Lesko, Catherine R

    Westreich, Daniel and Edwards, Jessie K. and Lesko, Catherine R. and Stuart, Elizabeth and Cole, Stephen R. , title =. American Journal of Epidemiology , year =

  31. [33]

    Journal of the American Statistical Association , year =

    Efron, Bradley , title =. Journal of the American Statistical Association , year =

  32. [34]

    The Annals of Statistics , year =

    Berk, Richard and Brown, Lawrence and Buja, Andreas and Zhang, Kai and Zhao, Linda , title =. The Annals of Statistics , year =

  33. [35]

    and Sun, Dennis L

    Lee, Jason D. and Sun, Dennis L. and Sun, Yuekai and Taylor, Jonathan E. , title =. The Annals of Statistics , year =

  34. [36]

    2014 , journal =

    Optimal Inference After Model Selection , author =. 2014 , journal =. 1410.2597 , archivePrefix =

  35. [37]

    Shao, Jun and Wu, C. F. J. , title =. The Annals of Statistics , year =

  36. [38]

    and Chen, Ming-Hui , title =

    Ibrahim, Joseph G. and Chen, Ming-Hui , title =. Statistical Science , year =

  37. [39]

    Biometrics , year =

    Schmidli, Heinz and Gsteiger, Sandro and Roychoudhury, Satrajit and O'Hagan, Anthony and Spiegelhalter, David and Neuenschwander, Beat , title =. Biometrics , year =

  38. [40]

    and Kinnersley, Nelson and Lindborg, Stacy and Micallef, Sandrine and Roychoudhury, Satrajit and Thompson, Laura , title =

    Viele, Kert and Berry, Scott and Neuenschwander, Beat and Amzal, Billy and Chen, Fang and Enas, Nathan and Hobbs, Brian and Ibrahim, Joseph G. and Kinnersley, Nelson and Lindborg, Stacy and Micallef, Sandrine and Roychoudhury, Satrajit and Thompson, Laura , title =. Pharmaceutical Statistics , year =

  39. [41]

    , title =

    Newey, Whitney K. , title =. Journal of Applied Econometrics , volume =. 1990 , doi =

  40. [42]

    , title =

    Pocock, Stuart J. , title =. Journal of Chronic Diseases , year =

  41. [43]

    , title =

    LaLonde, Robert J. , title =. The American Economic Review , year =

  42. [44]

    and Wahba, Sadek , title =

    Dehejia, Rajeev H. and Wahba, Sadek , title =. Journal of the American Statistical Association , year =

  43. [45]

    and Wahba, Sadek , title =

    Dehejia, Rajeev H. and Wahba, Sadek , title =. The Review of Economics and Statistics , year =

  44. [46]

    21st century cures act (public law 114-255, h.r

    114th United States Congress . 21st century cures act (public law 114-255, h.r. 34), 2016. URL https://www.congress.gov/114/plaws/publ255/PLAW-114publ255.pdf. Enacted December 13, 2016

  45. [47]

    Benkeser and M

    D. Benkeser and M. J. van der Laan. The highly adaptive lasso estimator. In 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA), pages 689--696, 2016. doi:10.1109/DSAA.2016.93

  46. [48]

    R. Berk, L. Brown, A. Buja, K. Zhang, and L. Zhao. Valid post-selection inference. The Annals of Statistics, 41 0 (2): 0 802--837, 2013. doi:10.1214/12-AOS1077

  47. [49]

    Chernozhukov, D

    V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins. Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal, 21 0 (1): 0 C1--C68, 2018. doi:10.1111/ectj.12097

  48. [50]

    Colnet, I

    B. Colnet, I. Mayer, G. Chen, A. Dieng, R. Li, G. Varoquaux, J.-P. Vert, J. Josse, and S. Yang. Causal inference methods for combining randomized trials and observational studies: A review. Statistical Science, 39 0 (1): 0 165--191, 2024. doi:10.1214/23-STS889. arXiv:2011.08047

  49. [51]

    I. J. Dahabreh, S. E. Robertson, J. A. Steingrimsson, E. A. Stuart, and M. A. Hern \'a n. Extending inferences from a randomized trial to a new target population. Statistics in Medicine, 39 0 (14): 0 1999--2014, 2020. doi:10.1002/sim.8426

  50. [52]

    L. E. Dang, J. M. Tarp, T. J. Abrahamsen, K. Kvist, J. B. Buse, M. Petersen, and M. van der Laan. Experiment-selector cross-validated targeted maximum likelihood estimator for hybrid RCT -external data studies. Journal of Causal Inference, 13 0 (1): 0 20240041, 2025. doi:10.1515/jci-2024-0041. arXiv:2210.05802

  51. [53]

    R. H. Dehejia and S. Wahba. Causal effects in nonexperimental studies: Reevaluating the evaluation of training programs. Journal of the American Statistical Association, 94 0 (448): 0 1053--1062, 1999. doi:10.1080/01621459.1999.10473858

  52. [54]

    R. H. Dehejia and S. Wahba. Propensity score-matching methods for nonexperimental causal studies. The Review of Economics and Statistics, 84 0 (1): 0 151--161, 2002. doi:10.1162/003465302317331982

  53. [55]

    Efron and C

    B. Efron and C. Stein. The jackknife estimate of variance. The Annals of Statistics, 9 0 (3): 0 586--596, 1981. doi:10.1214/aos/1176345462

  54. [56]

    Fithian, D

    W. Fithian, D. L. Sun, and J. Taylor. Optimal inference after model selection. arXiv preprint arXiv:1410.2597, 2014. doi:10.48550/arXiv.1410.2597

  55. [57]

    S. M. Hammer, D. A. Katzenstein, M. D. Hughes, H. Gundacker, R. T. Schooley, R. H. Haubrich, W. K. Henry, M. M. Lederman, J. P. Phair, M. Niu, M. S. Hirsch, and T. C. Merigan. A trial comparing nucleoside monotherapy with combination therapy in HIV -infected adults with CD4 cell counts from 200 to 500 per cubic millimeter. New England Journal of Medicine,...

  56. [58]

    A. E. Hubbard, S. Kherad-Pajouh, and M. J. van der Laan. Statistical inference for data adaptive target parameters. The International Journal of Biostatistics, 12 0 (1): 0 3--19, 2016. doi:10.1515/ijb-2015-0013

  57. [59]

    J. G. Ibrahim and M.-H. Chen. Power prior distributions for regression models. Statistical Science, 15 0 (1): 0 46--60, 2000. doi:10.1214/ss/1009212673

  58. [60]

    Juraska, P

    M. Juraska, P. B. Gilbert, X. Lu, M. Zhang, M. Davidian, and A. A. Tsiatis. speff2trial : Semiparametric Efficient Estimation for a Two-Sample Treatment Effect , 2022. URL https://CRAN.R-project.org/package=speff2trial. R package version 1.0.5; includes the ACTG175 dataset

  59. [61]

    A. K. Kuchibhotla, S. Balakrishnan, and L. Wasserman. The HulC : Confidence regions from convex hulls. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 86 0 (3): 0 586--622, 2024. doi:10.1093/jrsssb/qkad134. arXiv:2105.14577

  60. [62]

    R. J. LaLonde. Evaluating the econometric evaluations of training programs with experimental data. The American Economic Review, 76 0 (4): 0 604--620, 1986

  61. [63]

    J. D. Lee, D. L. Sun, Y. Sun, and J. E. Taylor. Exact post-selection inference, with application to the lasso. The Annals of Statistics, 44 0 (3): 0 907--927, 2016. doi:10.1214/15-AOS1371

  62. [64]

    Y. Li, S. Qiu, Z. Wang, and M. van der Laan. Regularized targeted maximum likelihood estimation in highly adaptive lasso implied working models. arXiv preprint arXiv:2506.17214, 2025. URL https://arxiv.org/abs/2506.17214. stat.ME

  63. [65]

    S. P. Luby, M. Rahman, B. F. Arnold, L. Unicomb, et al. Effects of water quality, sanitation, handwashing, and nutritional interventions on diarrhoea and child growth in rural Bangladesh : A cluster randomised controlled trial. The Lancet Global Health, 6 0 (3): 0 e302--e315, 2018. doi:10.1016/S2214-109X(17)30490-4

  64. [66]

    S. J. Pocock. The combination of randomized and historical controls in clinical trials. Journal of Chronic Diseases, 29 0 (3): 0 175--188, 1976. doi:10.1016/0021-9681(76)90044-8

  65. [67]

    M. H. Quenouille. Notes on bias in estimation. Biometrika, 43 0 (3--4): 0 353--360, 1956. doi:10.1093/biomet/43.3-4.353

  66. [68]

    E. T. R. Rosenman, G. Basse, A. B. Owen, and M. Baiocchi. Combining observational and experimental datasets using shrinkage estimators. Biometrics, 79 0 (4): 0 2961--2973, 2023. doi:10.1111/biom.13827. arXiv:2002.06708

  67. [69]

    Schmidli, S

    H. Schmidli, S. Gsteiger, S. Roychoudhury, A. O'Hagan, D. Spiegelhalter, and B. Neuenschwander. Robust meta-analytic-predictive priors in clinical trials with historical control information. Biometrics, 70 0 (4): 0 1023--1032, 2014. doi:10.1111/biom.12242

  68. [70]

    E. A. Stuart, S. R. Cole, C. P. Bradshaw, and P. J. Leaf. The use of propensity scores to assess the generalizability of results from randomized trials. Journal of the Royal Statistical Society: Series A (Statistics in Society), 174 0 (2): 0 369--386, 2011. doi:10.1111/j.1467-985X.2010.00673.x

  69. [71]

    Food and Drug Administration

    U.S. Food and Drug Administration . Framework for FDA's real-world evidence program. Technical report, U.S. Food and Drug Administration, 12 2018. URL https://www.fda.gov/media/120060/download

  70. [72]

    van der Laan, M

    L. van der Laan, M. Carone, A. Luedtke, and M. van der Laan. Adaptive debiased machine learning using data-driven model selection techniques. arXiv preprint arXiv:2307.12544, 2023. URL https://arxiv.org/abs/2307.12544. stat.ME

  71. [73]

    M. J. van der Laan. A generally efficient targeted minimum loss based estimator based on the highly adaptive lasso. The International Journal of Biostatistics, 13 0 (2): 0 20150097, 2017. doi:10.1515/ijb-2015-0097

  72. [74]

    M. J. van der Laan and D. Rubin. Targeted maximum likelihood learning. The International Journal of Biostatistics, 2 0 (1): 0 Article 11, 2006. doi:10.2202/1557-4679.1043

  73. [75]

    M. J. van der Laan, E. C. Polley, and A. E. Hubbard. Super learner. Statistical Applications in Genetics and Molecular Biology, 6 0 (1): 0 Article 25, 2007. doi:10.2202/1544-6115.1309

  74. [76]

    M. J. van der Laan, S. Qiu, J. M. Tarp, and L. van der Laan. Adaptive- TMLE for the average treatment effect based on randomized controlled trial augmented with real-world data. Journal of Causal Inference, 14 0 (1): 0 20240025, 2026. doi:10.1515/jci-2024-0025. arXiv:2405.07186

  75. [77]

    Viele, S

    K. Viele, S. Berry, B. Neuenschwander, B. Amzal, F. Chen, N. Enas, B. Hobbs, J. G. Ibrahim, N. Kinnersley, S. Lindborg, S. Micallef, S. Roychoudhury, and L. Thompson. Use of historical control data for assessing treatment effects in clinical trials. Pharmaceutical Statistics, 13 0 (1): 0 41--54, 2014. doi:10.1002/pst.1589

  76. [78]

    Westreich, J

    D. Westreich, J. K. Edwards, C. R. Lesko, E. Stuart, and S. R. Cole. Transportability of trial results using inverse odds of sampling weights. American Journal of Epidemiology, 186 0 (8): 0 1010--1014, 2017. doi:10.1093/aje/kwx164

  77. [79]

    Yang and P

    S. Yang and P. Ding. Combining multiple observational data sources to estimate causal effects. Journal of the American Statistical Association, 115 0 (531): 0 1540--1554, 2020. doi:10.1080/01621459.2019.1609973

  78. [80]

    S. Yang, C. Gao, D. Zeng, and X. Wang. Elastic integrative analysis of randomised trial and real-world data for treatment heterogeneity estimation. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 85 0 (3): 0 575--596, 2023. doi:10.1093/jrsssb/qkad017. arXiv:2005.10579

  79. [81]

    Zheng and M

    W. Zheng and M. J. van der Laan. Cross-validated targeted minimum-loss-based estimation. In M. J. van der Laan and S. Rose, editors, Targeted Learning: Causal Inference for Observational and Experimental Data, chapter 27, pages 459--474. Springer, New York, 2011. doi:10.1007/978-1-4419-9782-1_27