Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

From Observational Data to Clinical Recommendations: A Causal Framework for Estimating Patient-level Treatment Effects and Learning Policies

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A staged causal pipeline can turn observational hospital data into patient-specific treatment policies that beat current care, demonstrated on diuretic dosing for heart-failure patients with acute kidney injury.

desk verdict A useful framework paper whose case study is undermined by its own admitted SUTVA violation: T=1 mixes dose increase and maintenance, so the headline policy values are not well-defined causal effects. read the letter →

arxiv 2507.11381 v2 pith:6XCQB27S submitted 2025-07-15 stat.ML cs.LGstat.AP

classification stat.MLcs.LGstat.AP
keywords individualizedtreatmenteffectspolicyvalueobservationaldatacausalidentificationdeferralheartfailureacutekidneyinjurytargettrial
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that retrospective hospital data can support safe, patient-specific treatment recommendations, provided the modeling is organized as an explicit causal pipeline rather than a single predictive model. The authors propose a step-by-step framework—called the Target Recommendation System—that covers defining the clinical decision, checking whether causal effects are identifiable, estimating propensity overlap and patient-level treatment effects, deferring recommendations when uncertainty is high, and evaluating the resulting policy on held-out data. In a case study on heart-failure patients who develop acute kidney injury, the recommended policies are estimated to outperform the treatments actually given: return-to-baseline creatinine improves from 22.1% under current care to about 45.8% under the best learned policy, and a gradient-boosted model also reduces 30-day re-hospitalization. The take-home claim is that the value is in process—identification, deferral, and rigorous policy evaluation—not in any particular learning algorithm.

What carries the argument

The load-bearing object is the combined deferral rule of the Target Recommendation System, written as $\mathrm{Rej}'_\theta(x)=1$ when the estimated propensity score falls outside the overlap bounds or when $0\in[\hat\tau_\theta(x),\,\hat\tau^\theta(x)]$ for the CATE uncertainty interval, where CATE is the conditional average treatment effect, the expected individual-level gain from treatment. This rule converts a CATE estimate into a policy that either recommends treatment $\psi(\hat\tau_A(x))$ or defers to the care observed in the data, and the policy’s quality is measured by its policy value $V(\pi)$ estimated with doubly robust and inverse-propensity weighting on held-out data. That combination—restricting recommendations to patients whose effect is both identifiable and reliably signed, then evaluating counterfactual value with bootstrap—is what carries the claim that the learned policies improve on current care.

What would settle it

A randomized three-arm trial in the same patient population—decrease, maintain, or increase loop diuretic dose at the first creatinine rise, with creatinine return to baseline and 30-day rehospitalization as outcomes—would reveal whether the combined 'increase' arm has a single causal effect and whether a learned policy can beat usual care.

Watch

Extended reading notes

Core claim

The central discovery is that a policy learned through this staged causal pipeline can beat the observed clinician behavior in estimated value. In the case study, the decision is what to do with loop diuretic dosage at the first creatinine rise: decrease, versus maintain or increase; the outcome is the percentage return to baseline creatinine (RTB) measured within a week. With estimates that combine outcome and propensity modeling, evaluated on held-out data, ridge and gradient-boosted tree policies reach RTB values of 45.8% and 40.9%, versus 22.1% for the treatment actually observed, and the gradient-boosted policy also achieves a lower 30-day rehospitalization rate than current care. The authors do not claim every patient should get a recommendation: the policy defers for patients outside propensity overlap and for patients whose conditional average treatment effect uncertainty interval includes zero, with deferral meaning the patient receives the historically observed treatment.

Load-bearing premise

The comparison stands on treating 'increase or maintain diuretic dose' as a single well-defined treatment while also assuming every confounder of dosing and recovery is measured; the paper itself concedes that the first condition is violated as currently defined.

Editorial extensions

If this is right

  • A deployed system following the framework would recommend a diuretic dose change for only the subset of patients whose treatment effect is reliably estimated, and would otherwise let clinicians continue usual care.
  • Under the paper's estimates, following the learned policy would roughly double the average return-to-baseline creatinine in the studied population relative to observed treatment.
  • Model-based policies improve on treating everyone the same in bootstrap comparisons, except that they do not beat a simple decrease-for-all policy with statistical significance on the primary outcome alone.
  • The gradient-boosted policy is the one that improves both renal recovery and 30-day rehospitalization, indicating the framework can surface a policy balancing competing clinical goals.
  • The semi-synthetic simulation provides evidence that, when the data-generating assumptions hold, the estimated policy value tracks the true policy value and most learned policies beat current practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because 'maintain or increase' is one arm, the reported gain over current care bundles two different actions; a three-arm analysis could show the benefit comes mostly from one of them or that the arm is not a well-defined treatment at all.
  • The deferral assumption—deferred patients get the historically observed treatment—means the realized value depends on clinicians' behavior not changing for deferred patients; if the system alters practice, the policy value itself shifts, a performative effect the paper flags.
  • In the simulation, the true CATE is constructed partly from the propensity-score direction, so the simulation's success at finding beneficial policies is partly built in; a simulation with clinician-independent CATE would test the pipeline's ability to discover policies that diverge from current practice.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes the Target Recommendation System (TRS), a practical framework for building and validating patient-level treatment recommendation policies from observational health data. The framework is organized into identification, estimation, and validation stages, and it emphasizes deferral under uncertainty (via overlap trimming and causal/statistical uncertainty bounds), semi-synthetic simulation checks, and held-out doubly robust policy evaluation. The framework is applied to a case study of diuretic management for heart-failure patients who develop acute kidney injury, with treatment T=0 defined as lowering diuretic dose and T=1 as maintaining or increasing it, and with return-to-baseline creatinine (RTB) as the primary outcome. The central claim is that the learned policies improve on current care: Table 3 reports DR policy values of 45.8% for Ridge and 40.9% for XGBoost versus 22.1% for the observed treatment.

Significance. If the case-study validation is sound, the paper is a useful contribution: it operationalizes the target-trial logic for individualized policies, gives reproducible evaluative practices (semi-synthetic ground truth, held-out bootstrap policy values, IPW/DR estimators, deferral rules), and includes careful discussion of identification assumptions with clinical partners. The semi-synthetic simulation in Section 3.5 and Appendix C provides an external check against known ground truth, which is a genuine strength. However, the headline result rests on a treatment arm whose definition the paper itself admits violates SUTVA, so the case study does not currently validate the framework's central promise. The framework may still be valuable, but the demonstrated improvement over current care is not yet established.

major comments (3)
  1. [Section 4.1, Assumption 1; Table 3] The T=1 arm is not a well-defined intervention. The text states: "as currently defined, our formulation violates SUTVA as there are two versions of treatment for T=1: increasing or maintaining the dosage are not the same thing." This is not a minor caveat: it breaks Assumption 1, so a single potential outcome Y^1 does not exist for a patient; consistency (Assumption 2) has no well-defined counterfactual to be consistent with; and the CATE tau(x)=E[Y^1-Y^0|X=x] in Eq. (1) is not identified. Consequently, every policy value in Table 3 that ever assigns T=1 estimates a value for a mixture of two distinct interventions, and the headline comparison (Ridge 45.8%, XGBoost 40.9% vs. Doctors 22.1%) cannot be interpreted as the value of a well-defined treatment policy. The paper must either separate 'increase' and 'maintain' into distinct arms, or restrict the analysis to a treatment contrast for which SUTVA holds, and the case-study conclusions need to be reframed accordingly.
  2. [Section 4.4.6; Section 3.8] The causal sensitivity parameter is chosen as exp(0.1) with no sensitivity analysis and no clinical or empirical justification. The deferral rule Rej'_theta in Eq. (4) is central to the framework's safety claims, and the size of the deferral set (139 of 481 patients) and the downstream policy values on the 'Conservative' set (Figs. S.10 and S.11) will depend on this arbitrary Gamma value. The paper should report deferral proportions and policy values over a plausible range of causal uncertainty levels, and ideally relate those levels to measurable proxies of hidden confounding, before claiming that the framework safely handles causal uncertainty.
  3. [Section 4.4.8; Table 4; Fig. 7] The claimed improvement over current care is not statistically distinguished from a simple 'Decrease all' policy. The text states that "we cannot reject the possibility that using a model based policy has the same value as simply decreasing dosages for all patients," and Table 4 shows that the Ridge policy beats 'Decrease' in only 8,058 of 10,000 bootstrap rounds while XGBoost beats it in only 6,339 rounds. Moreover, Fig. 9 shows that the 'Decrease' policy performs worse than current care on 30-day re-hospitalization, so the apparent superiority of XGBoost on RTB is offset by a worse secondary outcome. The claim that the learned policies "improve patient outcomes over the current treatment regime" should be weakened to reflect that (a) the improvement is not significant relative to a non-personalized decrease-all policy, and (b) the multi-outcome trade-off is unresolved. The framework itself can survive this, but the case study's central conclusion needs revision.
minor comments (5)
  1. [Section 4.4.6] The sentence beginning "Using a causal sensitivity parameter of exp(0.1) and a statistical point estimate" should be clarified: it is not clear whether a statistical uncertainty interval was used at all, and the interaction between the statistical and causal components of the deferral rule is not described precisely.
  2. [Section 4.4.7; Table 3] The policies "Increase all" in Table 3 and "Keep/Increase" in Table S.7 refer to the same intervention but use inconsistent labels; unify the terminology across the main text, tables, and figure captions.
  3. [Section 4.1; Eq. (6)] The RTB outcome is described as "percent return to baseline" but Eq. (6) is a ratio, not a percentage; state explicitly whether the reported policy values are percentage points or fractions, and ensure consistency with Table 3 and the outcome tree in Fig. 6.
  4. [Appendix D; Section 3.10] The DR and IPW estimators in Appendix D use a propensity model p* trained on all the data; clarify whether this model is trained on the same data used for policy evaluation, and if so, discuss the implications for overfitting and whether cross-fitting was used.
  5. [General] There are several typographical and formatting issues: "T able" appears in table captions, "Absulte Error" in Table 2, "DrangoNet" in Tables S.7 and S.8, and "V alidation" in the framework diagram. These should be corrected in a final revision.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the case study is validated against a semi-synthetic ground truth and held-out policy evaluation, and the self-citations are not load-bearing.

full rationale

The central derivation chain is self-contained. The framework's output policies are evaluated on held-out test data (train/validation/test split of 1305/322/530) with doubly robust and IPW estimators, and the DR estimator uses a plug-in outcome model and a propensity model; neither is defined in terms of the policy value being reported. The semi-synthetic simulation (Section 4.4.3, Appendix C) provides an external ground truth: potential outcomes are generated from a known linear CATE and the estimated policy values are compared to the true simulation policy values, so the simulation is not a renamed fit. The paper explicitly acknowledges that the 'Increase' arm mixes maintenance and dose increase and thus violates SUTVA (Section 4.1); this is a validity threat to interpreting the policy values as effects of a well-defined treatment, but it is not circularity, because the reported estimates do not reduce by construction to the fitted inputs. Self-citations to Quince [28], B-learner [31], and the earlier heart-failure modeling paper [87] exist, but the headline results are reported on the 'Inclusive' set without Quince-based deferral, and covariate selection from [87] is an input to modeling, not a fitted quantity renamed as a prediction. No load-bearing step equates a prediction with an input by definition.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

The central claims rest on standard causal identification assumptions plus several user-chosen quantities. The most fragile inputs are the overlap thresholds and the sensitivity parameter, which directly determine deferrals and hence policy values; and the acknowledged SUTVA violation in the case study.

free parameters (3)
  • Overlap thresholds (eta_l, eta_h) = 0.21, 0.9
    Chosen based on training data to define common support; directly determines which patients receive recommendations (Section 4.4.1, Eq. (2)).
  • Quince causal sensitivity parameter (Gamma) = exp(0.1)
    User-specified bound on unobserved confounding for the deferral rule; no a priori justification or sensitivity sweep reported (Section 4.4.6).
  • Simulation parameters (lambda, C, variance multiplier) = lambda in [0,1], C chosen to match clinically reasonable average CATE, variance multiplier 1.2
    Algorithm SA1 requires choosing lambda (blend of propensity and random CATE direction) and C (target average effect); results used to conclude data is sufficient (Section C).
assumptions (6)
  • domain assumption Ignorability holds: no unmeasured common causes of treatment and outcome
    Required for CATE/policy identification (Assumption 4, Section 2.2); untestable; mitigated only by expert-elicited covariate list and proxies.
  • domain assumption SUTVA holds: no interference and no treatment versions
    Assumption 1, Section 2.2; the paper admits violation in the case study because T=1 includes both increase and maintain (Section 4.1).
  • domain assumption Consistency and accurate treatment recording
    Assumption 2; the paper asserts treatment allocations are accurately recorded (Section 3.2).
  • domain assumption Outcome RTB is a valid and complete measure of renal benefit within 7 days
    RTB defined in Eq. (6) uses last creatinine within a week; missing outcomes and competing risks are not addressed for the primary outcome.
  • standard math DR/IPW policy value estimators are consistent under the identification assumptions
    Standard doubly robust theory (Dudik et al. 2014); requires correct propensity or outcome model.
  • domain assumption Propensity model (XGBoost) is well calibrated
    Used to define overlap and weights; calibration shown in Fig. 4 but with AUROC 0.698 on validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Observational Data to Clinical Recommendations: A Causal Framework for Estimating Patient-level Treatment Effects and Learning Policies." pith.science (2026). https://pith.science/paper/6XCQB27S

@misc{pith2026250711381,
  author       = {Pith},
  title        = {Pith review of: From Observational Data to Clinical Recommendations: A Causal Framework for Estimating Patient-level Treatment Effects and Learning Policies},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6XCQB27S}},
  note         = {Machine review of arXiv:2507.11381}
}
read the original abstract

We propose a framework for building patient-specific treatment recommendation models, building on the large recent literature on learning patient-level causal models and inspired by the target trial paradigm of Hernan and Robins. We focus on safety and validity, including the crucial issue of causal identification when using observational data. We do not provide a specific model, but rather a way to integrate existing methods and know-how into a practical pipeline. We further provide a real world use-case of treatment optimization for patients with heart failure who develop acute kidney injury during hospitalization. The results suggest our pipeline can improve patient outcomes over the current treatment regime.

Figures

Figures reproduced from arXiv: 2507.11381 by the authors.

Figure 1
Figure 1. The outline of the target system, including identification, estimation and validation steps. Steps 1-2 refer to defining the clinical and causal question of interests and step 5 is designed to validate its feasibility. Steps 3,6,8,10 are for estimating the causal quantities of interest such as the CATE, and establishing a recommendation policy (step 10). Steps 4,7,9,11 are for validating and evaluating the correspon… view at source ↗
Figure 2
Figure 2. A schematic illustration of the different types of causal variables for the question of the causal effect of the treatment T on the outcome Y ; this is a modified version of [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Schematic illustration of propensity overlap, presenting the distribution of propensity scores (x-axis) of patients received treatment T = 1 (blue) or treatment T = 0 (red). A Strong overlap, where minor trimming could be considered. B Mild overlap, where trimming for extreme values is advised. C Non-overlap case, where the researchers should re-consider the research question (i.e. using this data as-is is not advis… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Calibration curve of propensity estimator, using XGBoost, on the validation set. data of patients in the common support region (Section 4.4.1), we simulated potential outcomes for both treatment arms. The generated outcomes were chosen such that the CATE is a linear fu…
Figure 5
Figure 5. Figure 5: Scatter plot of the true policy value versus the estimated policy values, using DR (Fig. 5a) and IPW (Fig. 5b) methods. The graph represents 6 simulation runs, using multiple policies: T-learner of XGboost (“XGB”), Ridge (“RIDGE”), Lasso (“LASSO”), and BART (“BART”), C…
Figure 6
Figure 6. Figure 6: Outcome Tree: A graph representing the mean RTB value for each policy group, using XGBoost T-learner as the basis for the policy. The first split (left) is by actual treatment received. The second split (right) is by whether the policy agreed or disagreed with the actu…
Figure 7
Figure 7. Figure 7: Policy value box-plot, the result of running 10K bootstraps evaluation on held-out data, showing 6 policies: Current: current treatment, Random: randomly assigning treatment at the same propotion as current treatment, Keep/Increase: all patients given “increase”, Decre…
Figure 8
Figure 8. Figure 8: Policy value rank graph, result of running 10K bootstraps evaluation on held-out data, showing two T-learner models: XGBoost (“XGB”) and Ridge (“RIDGE”), compared with “Doctors” treatment policy (“doctors”). Each model “ rank” markers (e.g., “RIDGE Rank”) represent the…
Figure 9
Figure 9. Figure 9: Policy value box-plot, a result of running 10K bootstraps evaluation on held-out data, in terms of re-hospitalization 30 days from decision-point. The DR policy value where estimated using L2-regularized logistic regression estimator. The policies are: Current: current…

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ConfoundingSHAP: Quantifying confounding strength in causal inference

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    ConfoundingSHAP defines a custom Shapley game to attribute confounding strength to individual covariates and uses TabPFN to estimate it scalably without exhaustive refitting.

Reference graph

Works this paper leans on

156 extracted references · 47 canonical work pages · cited by 1 Pith paper

  1. [1]

    Hern´ an and James M

    Miguel A. Hern´ an and James M. Robins. Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available.American Journal of Epidemiology, 183(8):758–764,

  2. [2]

    Alaa, Craig Lambert, and Mihaela van der Schaar

    Ioana Bica, Ahmed M. Alaa, Craig Lambert, and Mihaela van der Schaar. From Real-World Patient Data to Individualized Treatment Effects Using Machine Learning: Current and Future Methods to Address Underlying Challenges.Clinical Pharmacology & Therapeutics, 109(1):87–100, 1 2021. ISSN 1532-6535. doi: 10.1002/CPT.1907

  3. [3]

    Konig, Ruoxuan Xiong, Sadiqa Mahmood, Vera Mucaj, Chetan Bettegowda, Liam Rose, Suzanne Tamang, Adam Sacarny, Brian Caffo, Susan Athey, Elizabeth A

    Michael Powell, Allison Koenecke, James Brian Byrd, Akihiko Nishimura, Maximilian F. Konig, Ruoxuan Xiong, Sadiqa Mahmood, Vera Mucaj, Chetan Bettegowda, Liam Rose, Suzanne Tamang, Adam Sacarny, Brian Caffo, Susan Athey, Elizabeth A. Stuart, and Joshua T. Vogelstein. Ten Rules for Conducting Retrospective Pharmacoepidemiological Analyses: Example COVID-19...

  4. [4]

    Meid, Carmen Ruff, Lucas Wirbka, Felicitas Stoll, Hanna M

    Andreas D. Meid, Carmen Ruff, Lucas Wirbka, Felicitas Stoll, Hanna M. Seidling, Andreas Groll, and Walter E. Haefeli. Using the causal inference framework to support individualized drug treatment decisions based on observational healthcare data.Clinical Epidemiology, 12: 1223–1234, 2020. ISSN 11791349. doi: 10.2147/CLEP.S274466

  5. [5]

    Kent, Ewout Steyerberg, and David Van Klaveren

    David M. Kent, Ewout Steyerberg, and David Van Klaveren. Personalized evidence based medicine: Predictive approaches to heterogeneous treatment effects.BMJ (Online), 363,

  6. [6]

    Tell me something interesting: Clinical utility of machine learning prediction models in the icu.Journal of Biomedical Informatics, page 104107, 2022

    Bar Eini-Porat, Ofra Amir, Danny Eytan, and Uri Shalit. Tell me something interesting: Clinical utility of machine learning prediction models in the icu.Journal of Biomedical Informatics, page 104107, 2022

  7. [7]

    Koopman, Jae S

    Mattia Prosperi, Yi Guo, Matt Sperrin, James S. Koopman, Jae S. Min, Xing He, Shan- nan Rich, Mo Wang, Iain E. Buchan, and Jiang Bian. Causal inference and counterfactual prediction in machine learning for actionable healthcare.Nature Machine Intelligence, 2(7): 369–375, 7 2020. ISSN 25225839. doi: 10.1038/s42256-020-0197-y

  8. [8]

    Abassi, Zaher S

    Jubran Boulos, Wisam Darawsha, Zaid A. Abassi, Zaher S. Azzam, and Doron Aronson. Treatment patterns of patients with acute heart failure who develop acute kidney injury. ESC Heart Failure, 6(1):45–52, 2 2019. ISSN 20555822. doi: 10.1002/ehf2.12364

Show all 156 references
  1. [9]

    Valente, Adriaan A

    Kevin Damman, Mattia A.E. Valente, Adriaan A. Voors, Christopher M. O’Connor, Dirk J. Van Veldhuisen, and Hans L. Hillege. Renal impairment, worsening renal function, and outcome in patients with heart failure: An updated meta-analysis.European Heart Journal, 35(7):455–469, 2 ...

  2. [10]

    Cardiorenal syndrome in decompensated heart fail- ure.Heart, 96(4):255–260, 2010

    WH Wilson Tang and Wilfried Mullens. Cardiorenal syndrome in decompensated heart fail- ure.Heart, 96(4):255–260, 2010

  3. [11]

    Donald B. Rubin. Estimating causal effects of treatments in randomized and nonrandomized studies.Journal of Educational Psychology, 66(5):688–701, 1974. ISSN 00220663. doi: 10. 1037/h0037350. 32

  4. [12]

    Holland and Donald B

    Paul W. Holland and Donald B. Rubin. Causal inference in retrospective studies.ETS Research Report Series, 1987(1):203–231, 6 1987. ISSN 2330-8516. doi: 10.1002/j.2330-8516. 1987.tb00211.x

  5. [13]

    Cambridge university press, illustrate edition, 2009

    Judea Pearl.Causality. Cambridge university press, illustrate edition, 2009. ISBN 9780521895606

  6. [14]

    With great data comes great responsibility: Publishing comparative effec- tiveness research in epidemiology.Epidemiology, 22(3):290–291, 2011

    Miguel A Hern´ an. With great data comes great responsibility: Publishing comparative effec- tiveness research in epidemiology.Epidemiology, 22(3):290–291, 2011. ISSN 10443983. doi: 10.1097/EDE.0b013e3182114039

  7. [15]

    Can we learn individual-level treatment policies from clinical data?Biostatistics, 21(2):359–362, 11 2019

    Uri Shalit. Can we learn individual-level treatment policies from clinical data?Biostatistics, 21(2):359–362, 11 2019. ISSN 1465-4644. doi: 10.1093/biostatistics/kxz043

  8. [16]

    Learning causal effects from observational data in healthcare: A review and summary.Frontiers in Medicine, page 2027, 2022

    Jingpu Shi and Beau Norgeot. Learning causal effects from observational data in healthcare: A review and summary.Frontiers in Medicine, page 2027, 2022

  9. [17]

    Causal Decision Making and Causal Effect Esti- mation Are Not the Same

    Carlos Fern´ andez-Lor ´ ıa and Foster Provost. Causal Decision Making and Causal Effect Esti- mation Are Not the Same... and Why It Matters.arXiv preprint arXiv:2104.04103, 2021

  10. [18]

    MA Hernan and J Robins.Causal Inference: What if.Chapman & Hill/CRC, Boca Raton, 2020

  11. [19]

    Applied causal inference powered by ml and ai.arXiv preprint arXiv:2403.02467, 2024

    Victor Chernozhukov, Christian Hansen, Nathan Kallus, Martin Spindler, and Vasilis Syrgka- nis. Applied causal inference powered by ml and ai.arXiv preprint arXiv:2403.02467, 2024

  12. [20]

    Causal inference: A statistical learning approach, 2024

    Stefan Wager. Causal inference: A statistical learning approach, 2024

  13. [21]

    Nonparametric estimation of average treatment effects under exogeneity: A review.Review of Economics and Statistics, 86(1):4–29, 2004

    Guido W Imbens. Nonparametric estimation of average treatment effects under exogeneity: A review.Review of Economics and Statistics, 86(1):4–29, 2004. ISSN 00346535. doi: 10. 1162/003465304323023651

  14. [22]

    Rosenbaum and Donald B

    Paul R. Rosenbaum and Donald B. Rubin. The central role of the propensity score in ob- servational studies for causal effects.Biometrika, 70(1):41–55, 4 1983. ISSN 00063444. doi: 10.1093/biomet/70.1.41

  15. [23]

    Randomization Analysis of Experimental Data: The Fisher Randomization Test Comment.Journal of the American Statistical Association, 75(371):591–593, 1980

    Donald B Rubin. Randomization Analysis of Experimental Data: The Fisher Randomization Test Comment.Journal of the American Statistical Association, 75(371):591–593, 1980

  16. [24]

    Cole and Miguel A

    Stephen R. Cole and Miguel A. Hern´ an. Constructing inverse probability weights for marginal structural models.American Journal of Epidemiology, 168(6):656–664, 7 2008. ISSN 00029262. doi: 10.1093/aje/kwn164

  17. [25]

    A distributional approach for causal inference using propensity scores.Jour- nal of the American Statistical Association, 101(476):1619–1637, dec 2006

    Zhiqiang Tan. A distributional approach for causal inference using propensity scores.Jour- nal of the American Statistical Association, 101(476):1619–1637, dec 2006. doi: 10.1198/ 016214506000000023

  18. [26]

    Bounds on the conditional and average treatment effect with unobserved confounding factors.The Annals of Statistics, 50(5):2587–2615, 2022

    Steve Yadlowsky, Hongseok Namkoong, Sanjay Basu, John Duchi, and Lu Tian. Bounds on the conditional and average treatment effect with unobserved confounding factors.The Annals of Statistics, 50(5):2587–2615, 2022. 33

  19. [27]

    Interval estimation of individual-level causal effects under unobserved confounding

    Nathan Kallus, Xiaojie Mao, and Angela Zhou. Interval estimation of individual-level causal effects under unobserved confounding. InThe 22nd international conference on artificial intelligence and statistics, pages 2281–2290. PMLR, 2019

  20. [28]

    Quantifying ignorance in individual-level causal-effect estimates under hidden confounding

    Andrew Jesson, S¨ oren Mindermann, Yarin Gal, and Uri Shalit. Quantifying ignorance in individual-level causal-effect estimates under hidden confounding. InInternational Conference on Machine Learning, pages 4829–4838. PMLR, 2021

  21. [29]

    Sensitivity analysis of individual treatment effects: A robust conformal inference approach.arXiv preprint arXiv:2111.12161, 2021

    Ying Jin, Zhimei Ren, and Emmanuel J Cand` es. Sensitivity analysis of individual treatment effects: A robust conformal inference approach.arXiv preprint arXiv:2111.12161, 2021

  22. [30]

    Conformal sensitivity analysis for individual treatment effects.Journal of the American Statistical Association, pages 1–14, 2022

    Mingzhang Yin, Claudia Shi, Yixin Wang, and David M Blei. Conformal sensitivity analysis for individual treatment effects.Journal of the American Statistical Association, pages 1–14, 2022

  23. [31]

    B-learner: Quasi-oracle bounds on heterogeneous causal effects under hidden con- founding.arXiv preprint arXiv:2304.10577, 2023

    Miruna Oprescu, Jacob Dorn, Marah Ghoummaid, Andrew Jesson, Nathan Kallus, and Uri Shalit. B-learner: Quasi-oracle bounds on heterogeneous causal effects under hidden con- founding.arXiv preprint arXiv:2304.10577, 2023

  24. [32]

    Peter C. Austin. An introduction to propensity score methods for reducing the effects of confounding in observational studies.Multivariate Behavioral Research, 46(3):399–424, 5

  25. [33]

    K¨ unzel, Jasjeet S

    S¨ oren R. K¨ unzel, Jasjeet S. Sekhon, Peter J. Bickel, and Bin Yu. Metalearners for estimating heterogeneous treatment effects using machine learning.Proceedings of the National Academy of Sciences of the United States of America, 116(10):4156–4165, 2019. ISSN 10916490. doi:...

  26. [34]

    Quasi-oracle estimation of heterogeneous treatment effects

    Xinkun Nie and Stefan Wager. Quasi-oracle estimation of heterogeneous treatment effects. Biometrika, 108(2):299–319, 5 2021. ISSN 0006-3444

  27. [35]

    Towards optimal doubly robust estimation of heterogeneous causal effects.arXiv preprint arXiv:2004.14497, 2020

    Edward H Kennedy. Towards optimal doubly robust estimation of heterogeneous causal effects.arXiv preprint arXiv:2004.14497, 2020

  28. [36]

    Meta-learners for estimation of causal effects: Finite sample cross-fit perfor- mance.arXiv preprint arXiv:2201.12692, 2022

    Gabriel Okasa. Meta-learners for estimation of causal effects: Finite sample cross-fit perfor- mance.arXiv preprint arXiv:2201.12692, 2022

  29. [37]

    Estimation and Inference of Heterogeneous Treatment Effects using Random Forests.Journal of the American Statistical Association, 113(523):1228–1242, 7 2018

    Stefan Wager and Susan Athey. Estimation and Inference of Heterogeneous Treatment Effects using Random Forests.Journal of the American Statistical Association, 113(523):1228–1242, 7 2018. ISSN 1537274X. doi: 10.1080/01621459.2017.1319839

  30. [38]

    Random Forests.Machine Learning, 45(1):5–32, 2001

    Leo Breiman. Random Forests.Machine Learning, 45(1):5–32, 2001. ISSN 08856125. doi: 10.1023/A:1010933404324

  31. [39]

    Shah, Trevor Hastie, and Robert Tibshirani

    Scott Powers, Junyang Qian, Kenneth Jung, Alejandro Schuler, Nigam H. Shah, Trevor Hastie, and Robert Tibshirani. Some methods for heterogeneous treatment effect estima- tion in high dimensions.Statistics in Medicine, 37(11):1767–1787, 5 2018. ISSN 10970258. doi: 10.1002/sim.7623

  32. [40]

    A targeted maximum likelihood estimator of a causal effect on a bounded continuous outcome.The international journal of biostatistics, 6 (1):Article 26, 2010

    Susan Gruber and Mark J van der Laan. A targeted maximum likelihood estimator of a causal effect on a bounded continuous outcome.The international journal of biostatistics, 6 (1):Article 26, 2010. ISSN 1557-4679. doi: 10.2202/1557-4679.1260. 34

  33. [41]

    van der Laan and Alexander R

    Mark J. van der Laan and Alexander R. Luedtke. Targeted Learning of the Mean Outcome under an Optimal Dynamic Treatment Rule.Journal of Causal Inference, 3(1), 10 2015. ISSN 2193-3677. doi: 10.1515/jci-2013-0022

  34. [42]

    Engelhardt

    Li-Fang Cheng, Bianca Dumitrascu, Michael Zhang, Corey Chivers, Michael Draugelis, Kai Li, and Barbara E. Engelhardt. Patient-Specific Effects of Medication Using Latent Force Models with Gaussian Processes.arXiv preprint arXiv:1906.00226, 2019

  35. [43]

    Estimating individual treatment ef- fect: Generalization bounds and algorithms

    Uri Shalit, Fredrik D Johansson, and David Sontag. Estimating individual treatment ef- fect: Generalization bounds and algorithms. In34th International Conference on Machine Learning, ICML 2017, volume 6, pages 4709–4718, 2017. ISBN 9781510855144

  36. [44]

    Johansson, Nathan Kallus, Uri Shalit, and David Sontag

    Fredrik D. Johansson, Nathan Kallus, Uri Shalit, and David Sontag. Learning Weighted Representations for Generalization Across Designs.arXiv preprint arXiv:1802.08598, 2 2018

  37. [45]

    Causal Effect Inference with Deep Latent-Variable Models

    Christos Louizos, Uri Shalit, Joris M Mooij, David Sontag, Richard Zemel, and Max Welling. Causal Effect Inference with Deep Latent-Variable Models. In I Guyon, U V Luxburg, S Ben- gio, H Wallach, R Fergus, S Vishwanathan, and R Garnett, editors,Advances in Neural Information ...

  38. [46]

    Adapting neural networks for the estimation of treatment effects.arXiv preprint arXiv:1906.02120, 2019

    Claudia Shi, David M Blei, and Victor Veitch. Adapting neural networks for the estimation of treatment effects.arXiv preprint arXiv:1906.02120, 2019. ISSN 23318422

  39. [47]

    McCulloch

    Hugh A Chipman, Edward I George, and Robert E. McCulloch. BART: Bayesian additive regression trees.Annals of Applied Statistics, 4(1):266–298, 2010. ISSN 19326157. doi: 10.1214/09-AOAS285

  40. [48]

    Bayesian nonparametric modeling for causal inference.Journal of Computa- tional and Graphical Statistics, 20(1):217–240, 2011

    Jennifer L Hill. Bayesian nonparametric modeling for causal inference.Journal of Computa- tional and Graphical Statistics, 20(1):217–240, 2011. ISSN 10618600. doi: 10.1198/jcgs.2010. 08162

  41. [49]

    Gaussian processes in machine learning

    Carl Edward Rasmussen. Gaussian processes in machine learning. InSummer school on machine learning, pages 63–71. Springer, 2003

  42. [50]

    A tutorial on conformal prediction.Journal of Machine Learning Research, 9(3), 2008

    Glenn Shafer and Vladimir Vovk. A tutorial on conformal prediction.Journal of Machine Learning Research, 9(3), 2008

  43. [51]

    Conformal inference of counterfactuals and individual treatment effects.Journal of the Royal Statistical Society Series B: Statistical Methodology, 83(5):911–938, 2021

    Lihua Lei and Emmanuel J Cand` es. Conformal inference of counterfactuals and individual treatment effects.Journal of the Royal Statistical Society Series B: Statistical Methodology, 83(5):911–938, 2021

  44. [52]

    Bounds on the conditional and average treatment effect with unobserved confounding factors.arXiv preprint arXiv:1808.09521, 8 2018

    Steve Yadlowsky, Hongseok Namkoong, Sanjay Basu, John Duchi, and Lu Tian. Bounds on the conditional and average treatment effect with unobserved confounding factors.arXiv preprint arXiv:1808.09521, 8 2018

  45. [53]

    Identifying causal-effect inference failure with uncertainty-aware models.Advances in Neural Information Processing Systems, 33, 2020

    Andrew Jesson, S¨ oren Mindermann, Uri Shalit, and Yarin Gal. Identifying causal-effect inference failure with uncertainty-aware models.Advances in Neural Information Processing Systems, 33, 2020. 35

  46. [54]

    Min Qian and Susan A. Murphy. Performance guarantees for individualized treatment rules. The Annals of Statistics, 39(2):1180–1210, 4 2011. ISSN 0090-5364. doi: 10.1214/10-aos864

  47. [55]

    Doubly robust policy eval- uation and optimization.Statistical Science, 29(4):485–511, 2014

    Miroslav Dud ´ ık, Dumitru Erhan, John Langford, and Lihong Li. Doubly robust policy eval- uation and optimization.Statistical Science, 29(4):485–511, 2014

  48. [56]

    Estimators for the value of the optimal dynamic treatment rule with application to criminal justice interventions.The International Journal of Biostatistics, 2022

    Lina M Montoya, Mark J van der Laan, Jennifer L Skeem, and Maya L Petersen. Estimators for the value of the optimal dynamic treatment rule with application to criminal justice interventions.The International Journal of Biostatistics, 2022

  49. [57]

    Dimitris Bertsimas, Agni Orfanoudaki, and Rory B. Weiner. Personalized Treatment for Coronary Artery Disease Patients: A Machine Learning Approach.arXiv preprint arXiv:1910.08483, 10 2019

  50. [58]

    Efficient Policy Learning.arXiv preprint arXiv:1702.02896, 2 2017

    Susan Athey and Stefan Wager. Efficient Policy Learning.arXiv preprint arXiv:1702.02896, 2 2017

  51. [59]

    Policy learning with observational data.Econometrica, 89 (1):133–161, 2021

    Susan Athey and Stefan Wager. Policy learning with observational data.Econometrica, 89 (1):133–161, 2021

  52. [60]

    Balanced policy evaluation and learning

    Nathan Kallus. Balanced policy evaluation and learning. InAdvances in Neural Information Processing Systems, volume 2018-Decem, pages 8895–8906, 2018

  53. [61]

    The optimal dynamic treatment rule superlearner: consid- erations, performance, and application to criminal justice interventions.The International Journal of Biostatistics, 2022

    Lina M Montoya, Mark J van der Laan, Alexander R Luedtke, Jennifer L Skeem, Jeremy R Coyle, and Maya L Petersen. The optimal dynamic treatment rule superlearner: consid- erations, performance, and application to criminal justice interventions.The International Journal of Biost...

  54. [62]

    Teaching statistical inference for causal effects in experiments and observa- tional studies.Journal of Educational and Behavioral Statistics, 29(3):343–367, 2004

    Donald B Rubin. Teaching statistical inference for causal effects in experiments and observa- tional studies.Journal of Educational and Behavioral Statistics, 29(3):343–367, 2004

  55. [63]

    Avoidable flaws in observational analyses: an application to statins and cancer

    Barbra A Dickerman, Xabier Garc ´ ıa-Alb´ eniz, Roger W Logan, Spiros Denaxas, and Miguel A Hern´ an. Avoidable flaws in observational analyses: an application to statins and cancer. Nature medicine, 25(10):1601–1606, 2019

  56. [64]

    Sendak, Joshua D’Arcy, Sehj Kashyap, Michael Gao, Marshall Nichols, Kristin Corey, William Ratliff, and Suresh Balu

    Mark P. Sendak, Joshua D’Arcy, Sehj Kashyap, Michael Gao, Marshall Nichols, Kristin Corey, William Ratliff, and Suresh Balu. A Path for Translation of Machine Learning Products into Healthcare Delivery.EMJ Innovations, 1 2020. ISSN 2513-8634. doi: 10.33590/emjinnov/ 19-00172

  57. [65]

    Use of directed acyclic graphs (dags) to identify confounders in applied health research: review and recommendations.International journal of epidemiology, 50(2):620–632, 2021

    Peter WG Tennant, Eleanor J Murray, Kellyn F Arnold, Laurie Berrie, Matthew P Fox, Sarah C Gadd, Wendy J Harrison, Claire Keeble, Lynsie R Ranker, Johannes Textor, et al. Use of directed acyclic graphs (dags) to identify confounders in applied health research: review and recom...

  58. [66]

    A behavioral model of rational choice.The quarterly journal of economics, pages 99–118, 1955

    Herbert A Simon. A behavioral model of rational choice.The quarterly journal of economics, pages 99–118, 1955

  59. [67]

    Analysis of complex decision-making processes in health care: cognitive approaches to health informatics.Journal of biomedical informatics, 34(5):365–376, 2001

    Andre W Kushniruk. Analysis of complex decision-making processes in health care: cognitive approaches to health informatics.Journal of biomedical informatics, 34(5):365–376, 2001. 36

  60. [68]

    Emerging paradigms of cognition in medical decision-making.Journal of biomedical informatics, 35(1):52–75, 2002

    Vimla L Patel, David R Kaufman, and Jose F Arocha. Emerging paradigms of cognition in medical decision-making.Journal of biomedical informatics, 35(1):52–75, 2002

  61. [69]

    Causal diagrams for empirical research.Biometrika, 82(4):669–688, 1995

    Judea Pearl. Causal diagrams for empirical research.Biometrika, 82(4):669–688, 1995

  62. [70]

    Prognostic value of estimated plasma volume in heart failure.JACC: Heart Failure, 3(11):886–893, 2015

    K´ evin Duarte, Jean-Marie Monnez, Eliane Albuisson, Bertram Pitt, Faiez Zannad, and Patrick Rossignol. Prognostic value of estimated plasma volume in heart failure.JACC: Heart Failure, 3(11):886–893, 2015

  63. [71]

    Instrumental variables as bias amplifiers with general outcome and confounding.Biometrika, 104(2):291–302, 2017

    Peng Ding, TJ VanderWeele, and James M Robins. Instrumental variables as bias amplifiers with general outcome and confounding.Biometrika, 104(2):291–302, 2017

  64. [72]

    Covariate selection

    Brian Sauer, M Alan Brookhart, Jason A Roy, and Tyler J VanderWeele. Covariate selection. InDeveloping a protocol for observational comparative effectiveness research: a user’s guide. Agency for Healthcare Research and Quality (US), 2013

  65. [73]

    The choice of control variables: How causal graphs can inform the decision

    Paul Huenermund, Beyers Louw, and Mikko R¨ onkk¨ o. The choice of control variables: How causal graphs can inform the decision. InAcademy of Management Proceedings, page 294. Academy of Management Briarcliff Manor, NY 10510, 2022. doi: 10.5465/AMBPP.2022.294

  66. [74]

    Tianqi Chen and Carlos Guestrin. XGBoost. InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining - KDD ’16, pages 785– 794, 3 2016. ISBN 9781450342322. doi: 10.1145/2939672.2939785

  67. [75]

    Outcome adaptive lasso: Variable selection for causal inference.Biometrics, 73(4):1111–1122, dec 2017

    Ashkan Ertefaie Susan M Shortreed. Outcome adaptive lasso: Variable selection for causal inference.Biometrics, 73(4):1111–1122, dec 2017. ISSN 0006-341X. doi: 10.1111/biom.12679

  68. [76]

    Propensity score models are better when post-calibrated.arXiv preprint arXiv:2211.01221, 2022

    Rom Gutman, Ehud Karavani, and Yishai Shimoni. Propensity score models are better when post-calibrated.arXiv preprint arXiv:2211.01221, 2022

  69. [77]

    Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods.Advances in large margin classifiers, 10(3):61–74, 1999

    John Platt et al. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods.Advances in large margin classifiers, 10(3):61–74, 1999

  70. [78]

    Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers

    Bianca Zadrozny and Charles Elkan. Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. InIcml, volume 1, pages 609–616, 2001

  71. [79]

    A unified approach to interpreting model predictions

    Scott M Lundberg and Su-In Lee. A unified approach to interpreting model predictions. Advances in neural information processing systems, 30, 2017

  72. [80]

    Causal effects in nonexperimental studies: Reevaluating the evaluation of training programs.Journal of the American statistical Association, 94(448): 1053–1062, 1999

    Rajeev H Dehejia and Sadek Wahba. Causal effects in nonexperimental studies: Reevaluating the evaluation of training programs.Journal of the American statistical Association, 94(448): 1053–1062, 1999

  73. [81]

    Colin B Fogarty, Mark E Mikkelsen, David F Gaieski, and Dylan S Small. Discrete opti- mization for interpretable study populations and randomization inference in an observational study of severe sepsis mortality.Journal of the American Statistical Association, 111(514): 447–458, 2016

  74. [82]

    Dealing with limited overlap in estimation of average treatment effects.Biometrika, 96(1):187–199, 2009

    Richard K Crump, V Joseph Hotz, Guido W Imbens, and Oscar A Mitnik. Dealing with limited overlap in estimation of average treatment effects.Biometrika, 96(1):187–199, 2009. 37

  75. [83]

    Overlap in observational studies with high-dimensional covariates.Journal of Econometrics, 221(2):644– 654, 2021

    Alexander D’Amour, Peng Ding, Avi Feller, Lihua Lei, and Jasjeet Sekhon. Overlap in observational studies with high-dimensional covariates.Journal of Econometrics, 221(2):644– 654, 2021

  76. [84]

    In search of insights, not magic bullets: Towards demystification of the model selection dilemma in heterogeneous treatment effect estimation

    Alicia Curth and Mihaela Van Der Schaar. In search of insights, not magic bullets: Towards demystification of the model selection dilemma in heterogeneous treatment effect estimation. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonat...

  77. [85]

    Really doing great at estimating cate? a critical look at ml benchmarking practices in treatment effect estimation

    Alicia Curth, David Svensson, Jim Weatherall, and Mihaela van der Schaar. Really doing great at estimating cate? a critical look at ml benchmarking practices in treatment effect estimation. InThirty-fifth conference on neural information processing systems datasets and benchma...

  78. [86]

    Biases in electronic health record data due to processes within the healthcare system: retrospective observational study.Bmj, 361, 2018

    Denis Agniel, Isaac S Kohane, and Griffin M Weber. Biases in electronic health record data due to processes within the healthcare system: retrospective observational study.Bmj, 361, 2018

  79. [87]

    What drives performance in machine learning models for predicting heart failure outcome?European Heart Journal - Digital Health, 9 2022

    Rom Gutman, Doron Aronson, Oren Caspi, and Uri Shalit. What drives performance in machine learning models for predicting heart failure outcome?European Heart Journal - Digital Health, 9 2022. doi: 10.1093/EHJDH/ZTAC054

  80. [88]

    Estimating treatment effects with causal forests: An appli- cation.Observational Studies, 5(2):37–51, 2019

    Susan Athey and Stefan Wager. Estimating treatment effects with causal forests: An appli- cation.Observational Studies, 5(2):37–51, 2019

  81. [89]

    An introduction to the augmented inverse propensity weighted estimator.Political analysis, pages 36–56, 2010

    Adam N Glynn and Kevin M Quinn. An introduction to the augmented inverse propensity weighted estimator.Political analysis, pages 36–56, 2010

  82. [90]

    Balanced policy evaluation and learning for right censored data.arXiv preprint arXiv:1911.05728, 2019

    Owen E Leete, Nathan Kallus, Michael G Hudgens, Sonia Napravnik, and Michael R Kosorok. Balanced policy evaluation and learning for right censored data.arXiv preprint arXiv:1911.05728, 2019

  83. [91]

    Experimental evaluation of individualized treatment rules.Journal of the American Statistical Association, 118(541):242–256, 2023

    Kosuke Imai and Michael Lingzhi Li. Experimental evaluation of individualized treatment rules.Journal of the American Statistical Association, 118(541):242–256, 2023

  84. [92]

    Stratification and weighting via the propensity score in estimation of causal treatment effects: A comparative study.Statistics in Medicine, 23 (19):2937–2960, 2004

    Jared K Lunceford and Marie Davidian. Stratification and weighting via the propensity score in estimation of causal treatment effects: A comparative study.Statistics in Medicine, 23 (19):2937–2960, 2004. ISSN 02776715. doi: 10.1002/sim.1903

  85. [93]

    Toward robust policy summarization.Autonomous agents and multi-agent systems, 2019:2081, 2019

    Isaac Lage, Daphna Lifschitz, Finale Doshi-Velez, and Ofra Amir. Toward robust policy summarization.Autonomous agents and multi-agent systems, 2019:2081, 2019

  86. [94]

    Case-based off-policy evaluation using prototype learning

    Anton Matsson and Fredrik D Johansson. Case-based off-policy evaluation using prototype learning. InUncertainty in Artificial Intelligence, pages 1339–1349. PMLR, 2022

  87. [95]

    Theresa A McDonagh, Marco Metra, Marianna Adamo, Roy S Gardner, Andreas Baum- bach, Michael B¨ ohm, Haran Burri, Javed Butler, Jelena ˇCelutkien˙ e, Ovidiu Chioncel, John G F Cleland, Andrew J S Coats, Maria G Crespo-Leiro, Dimitrios Farmakis, Martine Gi- lard, Stephane Heyman...

  88. [96]

    Tricuspid regurgitation in acute heart failure: is there any incremental risk? European Heart Journal - Cardiovascular Imaging, 19(9):993–1001, 9 2018

    Diab Mutlak, Jonathan Lessick, Shehrban Khalil, Sergey Yalonetsky, Yoram Agmon, and Doron Aronson. Tricuspid regurgitation in acute heart failure: is there any incremental risk? European Heart Journal - Cardiovascular Imaging, 19(9):993–1001, 9 2018. ISSN 2047-2404. doi: 10.10...

  89. [97]

    Makhoul, Diab Mutlak, Jonathan Lessick, Shemy Carasso, Shimon Reisner, Yoram Agmon, Robert Dragu, and Zaher S

    Doron Aronson, Wisam Darawsha, Aula Atamna, Marielle Kaplan, Badira F. Makhoul, Diab Mutlak, Jonathan Lessick, Shemy Carasso, Shimon Reisner, Yoram Agmon, Robert Dragu, and Zaher S. Azzam. Pulmonary Hypertension, Right Ventricular Function, and Clinical Outcome in Acute Decomp...

  90. [98]

    Voors, Stefan D

    Piotr Ponikowski, Adriaan A. Voors, Stefan D. Anker, H´ ector Bueno, John G. F. Cleland, An- drew J. S. Coats, Volkmar Falk, Jos´ e Ram´ on Gonz´ alez-Juanatey, Veli-Pekka Harjola, Ewa A. Jankowska, Mariell Jessup, Cecilia Linde, Petros Nihoyannopoulos, John T. Parissis, Burke...

  91. [99]

    A unified approach to interpreting model predictions

    Scott M Lundberg and Su In Lee. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems, volume 2017-Decem, pages 4766–4775, 5 2017

  92. [100]

    A causal roadmap for generating high-quality real-world evidence.arXiv preprint arXiv:2305.06850, 2023

    Lauren E Dang, Susan Gruber, Hana Lee, Issa Dahabreh, Elizabeth A Stuart, Brian D Williamson, Richard Wyss, Iv´ an D ´ ıaz, Debashis Ghosh, Emre Kıcıman, et al. A causal roadmap for generating high-quality real-world evidence.arXiv preprint arXiv:2305.06850, 2023

  93. [101]

    Kravitz, Naihua Duan, and Joel Braslow

    Richard L. Kravitz, Naihua Duan, and Joel Braslow. Evidence-based medicine, heterogeneity of treatment effects, and the trouble with averages.Milbank Quarterly, 82(4):661–687, 12

  94. [102]

    Reporting clinical trial results to inform providers, payers, and consumers.Health Affairs, 24(6):1571–1581, 2005

    Rodney A Hayward, David M Kent, Sandeep Vijan, and Timothy P Hofer. Reporting clinical trial results to inform providers, payers, and consumers.Health Affairs, 24(6):1571–1581, 2005

  95. [103]

    Dahabreh, Rodney Hayward, and David M

    Issa J. Dahabreh, Rodney Hayward, and David M. Kent. Using group data to treat indi- viduals: Understanding heterogeneous treatment effects in the age of precision medicine and patient-centred evidence.International Journal of Epidemiology, 45(6):2184–2193, 12 2016. ISSN 14643...

  96. [104]

    A general statistical framework for subgroup identification and comparative treatment scoring.Biometrics, 73(4):1199–1209,

    Shuai Chen, Lu Tian, Tianxi Cai, and Menggang Yu. A general statistical framework for subgroup identification and comparative treatment scoring.Biometrics, 73(4):1199–1209,

  97. [105]

    Anoke, Sharon Lise Normand, and Corwin M

    Sarah C. Anoke, Sharon Lise Normand, and Corwin M. Zigler. Approaches to treatment effect heterogeneity in the presence of confounding.Statistics in Medicine, 38(15):2797–2815, 7 2019. ISSN 10970258. doi: 10.1002/sim.8143

  98. [106]

    Luedtke and Mark J

    Alexander R. Luedtke and Mark J. Van Der Laan. Evaluating the impact of treating the optimal subgroup.Statistical Methods in Medical Research, 26(4):1630–1640, 8 2017. ISSN 14770334. doi: 10.1177/0962280217708664

  99. [107]

    Yaobin Ling, Pulakesh Upadhyaya, Luyao Chen, Xiaoqian Jiang, and Yejin Kim. Emulate randomized clinical trials using heterogeneous treatment effect estimation for personalized treatments: methodology review and benchmark.Journal of Biomedical Informatics, 137: 104256, 2023

  100. [108]

    The predictive approaches to treatment effect heterogeneity (path) statement.Annals of internal medicine, 172(1):35–45, 2020

    David M Kent, Jessica K Paulus, David Van Klaveren, Ralph D’Agostino, Steve Goodman, Rodney Hayward, John PA Ioannidis, Bray Patrick-Lake, Sally Morton, Michael Pencina, et al. The predictive approaches to treatment effect heterogeneity (path) statement.Annals of internal medi...

  101. [109]

    Dowhy: An end-to-end library for causal inference.arXiv preprint arXiv:2011.04216, 2020

    Amit Sharma and Emre Kıcıman. Dowhy: An end-to-end library for causal inference.arXiv preprint arXiv:2011.04216, 2020. doi: 10.48550/arxiv.2011.04216

  102. [110]

    Treatment Policy Learning in Multiobjective Settings with Fully Observed Outcomes

    Soorajnath Boominathan, Michael Oberst, Helen Zhou, Sanjat Kanjilal, and David Sontag. Treatment Policy Learning in Multiobjective Settings with Fully Observed Outcomes. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 193...

  103. [111]

    Can machine learning from real-world data support drug treatment decisions? a prediction modeling case for direct oral anticoagulants.Medical Decision Making, 42(5):587–598, 2022

    Andreas D Meid, Lucas Wirbka, ARMIN Study Group, Andreas Groll, and Walter E Haefeli. Can machine learning from real-world data support drug treatment decisions? a prediction modeling case for direct oral anticoagulants.Medical Decision Making, 42(5):587–598, 2022

  104. [112]

    Causal models and learning from data: integrating causal modeling and statistical estimation.Epidemiology (Cambridge, Mass.), 25(3):418, 2014

    Maya L Petersen and Mark J van der Laan. Causal models and learning from data: integrating causal modeling and statistical estimation.Epidemiology (Cambridge, Mass.), 25(3):418, 2014

  105. [113]

    Bates, Andrew Auerbach, Peter Schulam, Adam Wright, and Suchi Saria

    David W. Bates, Andrew Auerbach, Peter Schulam, Adam Wright, and Suchi Saria. Reporting and Implementing Interventions Involving Machine Learning and Artificial Intelligence.An- nals of internal medicine, 172(11):S137–S144, 6 2020. ISSN 15393704. doi: 10.7326/M19-0872

  106. [114]

    AI in health and medicine

    Pranav Rajpurkar, Emma Chen, Oishi Banerjee, and Eric J Topol. AI in health and medicine. Nature medicine, 28(1):31–38, 2022

  107. [115]

    Clinical deployment environments: Five pillars of translational machine learning for health.Frontiers in Digital Health, 4:939292, 2022

    Steve Harris, Tim Bonnici, Thomas Keen, Watjana Lilaonitkul, Mark J White, and Nel Swanepoel. Clinical deployment environments: Five pillars of translational machine learning for health.Frontiers in Digital Health, 4:939292, 2022. 40

  108. [116]

    Shifting machine learning for health- care from development to deployment and from models to data.Nature Biomedical Engineer- ing, 6(12):1330–1345, 2022

    Angela Zhang, Lei Xing, James Zou, and Joseph C Wu. Shifting machine learning for health- care from development to deployment and from models to data.Nature Biomedical Engineer- ing, 6(12):1330–1345, 2022

  109. [117]

    Hundreds of AI tools have been built to catch Covid

    Will Douglas Heaven. Hundreds of AI tools have been built to catch Covid. none of them helped.MIT Technology Review. Retrieved October, 6:2021, 2021

  110. [118]

    Causal inference using observational inten- sive care unit data: a scoping review and recommendations for future practice.npj Digital Medicine, 6(1):221, 2023

    JM Smit, JH Krijthe, WMR Kant, JA Labrecque, M Komorowski, DAMPJ Gommers, J van Bommel, MJT Reinders, and ME van Genderen. Causal inference using observational inten- sive care unit data: a scoping review and recommendations for future practice.npj Digital Medicine, 6(1):221, 2023

  111. [119]

    Machine learning for healthcare that matters: Reorienting from technical novelty to equitable impact.PLOS Digital Health, 3(4):e0000474, 2024

    Aparna Balagopalan, Ioana Baldini, Leo Anthony Celi, Judy Gichoya, Liam G McCoy, Tristan Naumann, Uri Shalit, Mihaela van der Schaar, and Kiri L Wagstaff. Machine learning for healthcare that matters: Reorienting from technical novelty to equitable impact.PLOS Digital Health, ...

  112. [120]

    McCradden, Elizabeth A

    Melissa D. McCradden, Elizabeth A. Stephenson, and James A. Anderson. Clinical research underlies ethical integration of healthcare artificial intelligence.Nature Medicine 2020 26:9, 26(9):1325–1326, 9 2020. ISSN 1546-170X. doi: 10.1038/s41591-020-1035-9

  113. [121]

    AI model transferability in healthcare: a sociotechnical perspective.Nature Machine Intelligence, 4(10):807–809, 2022

    Batia Mishan Wiesenfeld, Yin Aphinyanaphongs, and Oded Nov. AI model transferability in healthcare: a sociotechnical perspective.Nature Machine Intelligence, 4(10):807–809, 2022

  114. [122]

    Shalmali Joshi, I˜ nigo Urteaga, Wouter AC van Amsterdam, George Hripcsak, Pierre Elias, Benjamin Recht, No´ emie Elhadad, James Fackler, Mark P Sendak, Jenna Wiens, et al. AI as an intervention: improving clinical outcomes relies on a causal approach to AI development and val...

  115. [123]

    Generalization in medical ai: a perspective on developing scalable models.arXiv preprint arXiv:2311.05418, 2023

    Joachim A Behar, Jeremy Levy, and Leo Anthony Celi. Generalization in medical ai: a perspective on developing scalable models.arXiv preprint arXiv:2311.05418, 2023

  116. [124]

    Causal Inference Through Potential Outcomes and Principal Stratification: Application to Studies with ”Censoring” Due to Death 1.Statistical Science, 21(3):299–309,

    Donald B Rubin. Causal Inference Through Potential Outcomes and Principal Stratification: Application to Studies with ”Censoring” Due to Death 1.Statistical Science, 21(3):299–309,

  117. [125]

    The win ratio: a new approach to the analysis of composite endpoints in clinical trials based on clinical priorities

    Stuart J Pocock, Cono A Ariti, Timothy J Collier, and Duolao Wang. The win ratio: a new approach to the analysis of composite endpoints in clinical trials based on clinical priorities. European heart journal, 33(2):176–182, 2012

  118. [126]

    Rethinking the win ratio: A causal framework for hierarchical outcome analysis.arXiv preprint arXiv:2501.16933, 2025

    Mathieu Even and Julie Josse. Rethinking the win ratio: A causal framework for hierarchical outcome analysis.arXiv preprint arXiv:2501.16933, 2025

  119. [127]

    Optimal dynamic treatment regimes.Journal of the Royal Statistical Society Series B: Statistical Methodology, 65(2):331–355, 2003

    Susan A Murphy. Optimal dynamic treatment regimes.Journal of the Royal Statistical Society Series B: Statistical Methodology, 65(2):331–355, 2003

  120. [128]

    Guidelines for reinforcement learning in healthcare

    Omer Gottesman, Fredrik Johansson, Matthieu Komorowski, Aldo Faisal, David Sontag, Finale Doshi-Velez, and Leo Anthony Celi. Guidelines for reinforcement learning in healthcare. Nature medicine, 25(1):16–18, 2019. 41

  121. [129]

    Predict responsibly: improving fairness and accuracy by learning to defer.Advances in Neural Information Processing Systems, 31, 2018

    David Madras, Toni Pitassi, and Richard Zemel. Predict responsibly: improving fairness and accuracy by learning to defer.Advances in Neural Information Processing Systems, 31, 2018

  122. [130]

    Consistent estimators for learning to defer to an expert

    Hussein Mozannar and David Sontag. Consistent estimators for learning to defer to an expert. InInternational Conference on Machine Learning, pages 7076–7087. PMLR, 2020

  123. [131]

    When to act and when to ask: Policy learning with deferral under hidden confounding

    Marah Ghoummaid and Uri Shalit. When to act and when to ask: Policy learning with deferral under hidden confounding. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  124. [132]

    Confounding-robust policy improvement with human-AI teams.arXiv preprint arXiv:2310.08824, 2023

    Ruijiang Gao and Mingzhang Yin. Confounding-robust policy improvement with human-AI teams.arXiv preprint arXiv:2310.08824, 2023

  125. [133]

    Is the most accurate AI the best teammate? optimizing AI for teamwork.Proceedings of the AAAI Conference on Artificial Intelligence, 35(13):11405–11414, 2021

    Gagan Bansal, Besmira Nushi, Ece Kamar, Eric Horvitz, and Daniel S Weld. Is the most accurate AI the best teammate? optimizing AI for teamwork.Proceedings of the AAAI Conference on Artificial Intelligence, 35(13):11405–11414, 2021

  126. [134]

    How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection.Translational psychiatry, 11(1):108, 2021

    Maia Jacobs, Melanie F Pradier, Thomas H McCoy Jr, Roy H Perlis, Finale Doshi-Velez, and Krzysztof Z Gajos. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection.Translational psychiatry, 11(1):108, 2021

  127. [135]

    Impact of artificial intelligence on pathologists’ decisions: an experiment.Journal of the American Medical Informatics Association, 29(10):1688–1695, 2022

    Julien Meyer, April Khademi, Bernard Tˆ etu, Wencui Han, Pria Nippak, and David Remisch. Impact of artificial intelligence on pathologists’ decisions: an experiment.Journal of the American Medical Informatics Association, 29(10):1688–1695, 2022

  128. [136]

    The clinician and dataset shift in artificial intelligence.New England Journal of Medicine, 385(3):283–286, 2021

    Samuel G Finlayson, Adarsh Subbaswamy, Karandeep Singh, John Bowers, Annabel Kupke, Jonathan Zittrain, Isaac S Kohane, and Suchi Saria. The clinician and dataset shift in artificial intelligence.New England Journal of Medicine, 385(3):283–286, 2021

  129. [137]

    Performative prediction

    Juan Perdomo, Tijana Zrnic, Celestine Mendler-D¨ unner, and Moritz Hardt. Performative prediction. InInternational Conference on Machine Learning, pages 7599–7609. PMLR, 2020

  130. [138]

    An introduction to proximal causal learning.arXiv preprint arXiv:2009.10982, 2020

    Eric J Tchetgen Tchetgen, Andrew Ying, Yifan Cui, Xu Shi, and Wang Miao. An introduction to proximal causal learning.arXiv preprint arXiv:2009.10982, 2020

  131. [139]

    A selective review of negative control methods in epidemiology.Current epidemiology reports, 7(4):190–202, 2020

    Xu Shi, Wang Miao, and Eric Tchetgen Tchetgen. A selective review of negative control methods in epidemiology.Current epidemiology reports, 7(4):190–202, 2020

  132. [140]

    Proximal causal learning of conditional average treatment effects

    Erik Sverdrup and Yifan Cui. Proximal causal learning of conditional average treatment effects. InInternational Conference on Machine Learning, pages 33285–33298. PMLR, 2023

  133. [141]

    Mihai Gheorghiade and Peter S. Pang. Acute Heart Failure Syndromes.Journal of the American College of Cardiology, 53(7):557–573, 2 2009. ISSN 07351097. doi: 10.1016/j.jacc. 2008.10.041

  134. [142]

    Negative trials in critical care: why most research is probably wrong.The Lancet

    John G Laffey and Brian P Kavanagh. Negative trials in critical care: why most research is probably wrong.The Lancet. Respiratory medicine, 6(9):659–660, 9 2018. ISSN 2213-2619. doi: 10.1016/S2213-2600(18)30279-0. 42

  135. [143]

    McMurray, Milton Packer, Akshay S

    John J.V. McMurray, Milton Packer, Akshay S. Desai, Jianjian Gong, Martin P. Lefkowitz, Adel R. Rizkala, Jean L. Rouleau, Victor C. Shi, Scott D. Solomon, Karl Swedberg, and Michael R. Zile. Angiotensin–Neprilysin Inhibition versus Enalapril in Heart Failure.New England Journa...

  136. [144]

    Massie, Christopher M

    Barry M. Massie, Christopher M. O’Connor, Marco Metra, Piotr Ponikowski, John R. Teer- link, Gad Cotter, Beth Davison Weatherley, John G.F. Cleland, Michael M. Givertz, Adriaan Voors, Paul DeLucca, George A. Mansoor, Christina M. Salerno, Daniel M. Bloomfield, and Howard C. Di...

  137. [145]

    O’Connor, R.C

    C.M. O’Connor, R.C. Starling, A.F. Hernandez, P.W. Armstrong, K. Dickstein, V. Hassel- blad, G.M. Heizer, M. Komajda, B.M. Massie, J.J.V. McMurray, M.S. Nieminen, C.J. Reist, J.L. Rouleau, K. Swedberg, K.F. Adams, S.D. Anker, D. Atar, A. Battler, R. Botero, N.R. Bo- hidar, J. ...

  138. [146]

    The impact of transient and persistent acute kidney injury on long-term outcomes after acute myocardial infarction.Kidney International, 76(8):900–906, 10 2009

    Alexander Goldberg, Elena Kogan, Haim Hammerman, Walter Markiewicz, and Doron Aron- son. The impact of transient and persistent acute kidney injury on long-term outcomes after acute myocardial infarction.Kidney International, 76(8):900–906, 10 2009. ISSN 0085-2538. doi: 10.103...

  139. [147]

    Relationship between heart failure treatment and development of worsening renal function among hospitalized patients.American Heart Journal, 147(2):331–338, 2 2004

    Javed Butler, Daniel E Forman, William T Abraham, Stephen S Gottlieb, Evan Loh, Barry M Massie, Christopher M O’Connor, Michael W Rich, Lynne Warner Stevenson, Yongfei Wang, James B Young, and Harlan M Krumholz. Relationship between heart failure treatment and development of w...

  140. [148]

    Kane, Joseph K

    Jesse A. Kane, Joseph K. Kim, Syed Abbas Haidry, Louis Salciccioli, and Jason Lazar. Dis- continuation/dose reduction of angiotensin-converting enzyme inhibitors/angiotensin recep- tor blockers during acute decompensated heart failure in African-American patients with reduced ...

  141. [149]

    Nieminen, Gerasimos S

    Michael B¨ ohm, Andreas Link, Danlin Cai, Markku S. Nieminen, Gerasimos S. Filippatos, Reda Salem, Alain Cohen Solal, Bidan Huang, Robert J. Padley, Matti Kivikko, and Alexan- dre Mebazaa. Beneficial association ofβ-blocker therapy on recovery from severe acute heart failure t...

  142. [156]

    yes”, we are left with the following: Covariates for which the answer tobANDcis “yes

    (a)At what point in the clinical workflow would the clinical team want the recommendation? (b)Does the clinical recommendation time point correspond to the time where the decision is currently made? B Which covariates should be used Creating and validating a causal graph that ...

  143. [2004]

    doi: 10.1111/j.0887-378X.2004.00327.x

    ISSN 0887378X. doi: 10.1111/j.0887-378X.2004.00327.x

  144. [2006]

    doi: 10.1214/088342306000000114

  145. [2011]

    doi: 10.1080/00273171.2011.568786

    ISSN 00273171. doi: 10.1080/00273171.2011.568786

  146. [2016]

    doi: 10.1093/aje/kwv254

    ISSN 14766256. doi: 10.1093/aje/kwv254

  147. [2017]

    doi: 10.1111/biom.12676

    ISSN 15410420. doi: 10.1111/biom.12676

  148. [2018]

    doi: 10.1136/bmj.k4245

    ISSN 17561833. doi: 10.1136/bmj.k4245

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.