REVIEW 4 major objections 5 minor 92 references
Scoping review of methodology for aiding generalisability and transportability of clinical prediction models
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This scoping review claims that development-phase methods for generalisable or transportable clinical prediction models split into data-driven and knowledge-driven families, with no head-to-head evaluation yet.
desk verdict A useful but search-restricted taxonomy of development-stage methods for CPM generalisability and transportability; worth reviewing, but the completeness claims need major revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central analytic object is the two-way classification (data-driven vs knowledge-driven × not tailored vs tailored) constructed from the 18 included papers. Within that taxonomy, the load-bearing technical apparatus is the density ratio $r(x)=P_{target}(x)/P_{source}(x)$ for data-driven tailoring, and, for knowledge-driven methods, the selection diagram and its associated transport formula, e.g., $P(Y=1 \mid X=x, S=0) = \sum_z P(Y=1 \mid X=x, S=1, Z=z)P(Z=z)$ under S-admissibility, together with the graph surgery estimator that turns interventional distributions into observable ones. These devices let a model be transported or made invariant to shifts without re-estimation on the target population.
What would settle it
A replication of the search with additional title and abstract terms such as 'domain adaptation', 'domain generalization', 'invariant learning', 'distribution shift', 'external validity', and 'transfer learning' in the same databases up to September 2023, followed by full-text screening against the review's inclusion and exclusion criteria, would test the taxonomy's completeness. If any resultant paper proposes a development-phase method for generalisable or transportable clinical prediction models that fits neither the data-driven nor the knowledge-driven category, or neither the tailored nor the not-tailored axis, the review's classification is incomplete.
Extended reading notes
Core claim
The paper's central claim is that the emerging literature on development-phase methods for generalisability and transportability of clinical prediction models can be organised by two independent axes: data-driven versus knowledge-driven, and whether the method is tailored to a specific target population. Data-driven methods include training on heterogeneous multi-centre data (internal-external cross-validation), ensemble models, density-ratio or importance weighting, adversarial validation, and measurement-error alignment; knowledge-driven methods include selection diagrams and transport formulas, graph surgery estimators, shortcut-predictor removal, distributionally robust and worst-case risk minimisation with mutable/immutable sets, Markov-blanket or parent-child predictor selection, and counterfactual adjustment for anticausal (diagnostic) tasks. The review states that no included study has comparatively evaluated these methods on shared simulated or real datasets, and it identifies integration of the two families as an open direction.
Load-bearing premise
The search strategy, based on titles containing generalisab*, generalizab*, transportab* or 'dataset shift' plus the Geersing filters, captures the full universe of relevant methodology papers, even though the review itself found that one known seed paper (Bellamy et al.) was missed by the search and only included manually.
Editorial extensions
If this is right
- If the taxonomy is correct, a model developer can choose a method family based on whether covariate data from the target population is available and whether reliable causal knowledge exists.
- Because the review finds that density-ratio weighting corrects covariate shift but not concept shift, these methods cannot rescue models when the predictor-outcome relationship itself changes.
- The absence of comparative evaluation means that, as of the review's search date, no evidence indicates which method family performs best for a given type of dataset shift.
- Knowledge-driven methods all rest on correctly specified causal structures, so misspecification of the selection diagram or the mutable/immutable split is a shared failure mode.
- Integrating data-driven heterogeneity with causal invariant selection is identified by the paper as a promising direction for future methodology.
Reading between the lines
- A direct test of the taxonomy would be to run the 18 methods on a common battery of simulated shifts (covariate, prior, concept, measurement error) and real EHR datasets; the paper's own call for such benchmarks implies that none exists.
- Because one known seed paper (Bellamy et al.) escaped the search and snowballing, the review's map is probably incomplete; adding domain-adaptation and invariant-learning search terms could surface methods that fit neither family.
- The tailored-versus-universal axis suggests a potential bridge to transfer learning: density-ratio weighting is already used in domain adaptation, so formal connections could unify the two literatures.
- A head-to-head comparison might find that knowledge-driven methods degrade gracefully under correct causal assumptions but fail catastrophically when those assumptions are wrong, while data-driven methods trade off performance on any single population for robustness across many.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a scoping review of methodology for improving the generalisability and transportability of clinical prediction models (CPMs). The authors searched MEDLINE, Embase, medRxiv, and arXiv up to September 2023 using title-based search terms (generalisab*, generalizab*, transportab*, 'dataset shift') plus Geersing filters, screened 1,761 records, and included 18 papers. They categorize the identified methods into data-driven versus knowledge-driven families and, within each, whether the method is tailored to a specific target population. The paper summarises assumptions, advantages, disadvantages, and applications of each included method, and offers future research directions such as comparative evaluations on simulated and real data and integration of the two families.
Significance. If the taxonomy is accepted, this review provides a useful map of an emerging and fragmented literature, and it makes a plausible conceptual distinction between methods that exploit data heterogeneity or weighting to improve robustness and methods that rely on causal structure or domain knowledge. The review also clearly identifies the absence of comparative evaluations and the need for methods that require less target-population data. The paper's strengths include a documented search strategy, a PRISMA flow diagram, a structured extraction of assumptions and limitations in supplementary material, and an explicit discussion of search limitations. The main value is as a synthesis and agenda-setting review for clinical prediction methodology; its central claim, however, depends on the completeness and representativeness of the included literature.
major comments (4)
- [Search Strategy; Results; Discussion] The central claim that methodologies for aiding generalisability/transportability fall into two families is load-bearing, yet the evidence for completeness is weak. The search was restricted to titles containing generalisab*, generalizab*, transportab* or 'dataset shift', and the paper concedes that one of the five validation papers (Bellamy et al.) was not found by the search or snowballing, and that Debray et al.'s framework was excluded because it did not match keywords. This means the search has demonstrated imperfect recall even on a hand-picked set; it may have missed large adjacent literatures using terms such as domain shift, domain generalization, distributional robustness, invariant risk minimization, or external validity. The paper should either rerun the search with additional synonym-based terms and a broader validation set, or explicitly reframe the taxonomy as provisional and limited to the included papers rather than as a complete classification of the field.
- [Information sources; Data extraction] Screening and data extraction were conducted by a single author (KP) with no independent second reviewer or adjudication process. For a review whose core output is a classification and summary of included methods, single-reviewer screening is a substantial risk to reliability, as title/abstract decisions determine the inclusion set on which the taxonomy is built. The manuscript should report whether any reliability checks (e.g., dual screening on a subset, or verification of extraction by a second author) were performed, or justify why single-reviewer screening is appropriate for this scoping review.
- [Scope of review; Results; Table 1] The exclusion of transfer learning and adjustment methods is not consistently applied. The scope states that methods that adjust an already-developed model are excluded, but density ratio/importance weighting methods (Gao et al., Steingrimsson et al., Dockès et al.) are included in Table 1 even though they are commonly classified as transfer-learning or domain-adaptation techniques. If these are included because they are used during development, the boundary should be stated more precisely; if they are excluded in other guises, the review risks including members of a family while claiming that family is out of scope, which weakens the exhaustiveness of the two-family taxonomy.
- [Knowledge-driven methodology that are not tailored for target population; Table 1] The manuscript states 'Fewer studies (7 of 19) focused on knowledge-driven methodologies', but the review includes 18 papers, not 19. More substantively, Piccininni et al. and Liu et al. are placed in both the data-driven and knowledge-driven categories, so the two families are not mutually exclusive. The paper should provide explicit criteria for assigning a method to one or both families, and should discuss whether the overlap indicates that the data-driven/knowledge-driven distinction is better viewed as a continuum or as complementary axes rather than a clean dichotomy.
minor comments (5)
- [Information sources] The database name 'medRxiv' is misspelled as 'merRxiv' in the Information sources section.
- [Table 1] Table 1 contains stray timestamps '26/11/2024 12:10:00' in two cells; these artifacts should be removed.
- [Knowledge-driven methodology that are not tailored for target population] The text says '7 of 19'; the correct total is 18 included papers. Also, 'Parent Child - PB' should likely be 'Parent-Child (PC)' for consistency.
- [Knowledge-driven methodology that is tailored for a specific target population] There is a typo, 'Causaility-aware', which should read 'Causality-aware'.
- [Snowballing (Citation search)] The phrase 'To compliment the database search' should be 'To complement the database search'.
Circularity Check
No significant circularity: the taxonomy is a literature synthesis, and the search check is reported transparently.
full rationale
This is a scoping review rather than a derivation, so its central claim—that methodologies aiding generalisability and transportability of clinical prediction models can be grouped into data-driven and knowledge-driven families—is a categorization of the 18 included papers, not a quantity fitted to data and then renamed as a prediction. The search terms were derived from five seed papers and checked against them; while this is a sensitivity exercise rather than an independent validation, the paper does not infer completeness from that check and explicitly reports that one seed paper (Bellamy et al.) was not found by the search and had to be added manually. The self-citations (Martin et al. 2020 for scoping-review search strategy, Sperrin et al. 2022 for tailored validation, and Lin et al. 2021 for causal methods under hypothetical interventions) are contextual and are not load-bearing premises of the taxonomy. No uniqueness theorem from the authors is invoked to forbid alternative classifications, and no method is defined in terms of the taxonomy it is supposed to support. The acknowledged search gap—title-based terms that missed synonym-rich literatures such as transfer learning—is a limitation on completeness, not a circularity, because the taxonomy is transparently stated relative to the included set.
Assumptions & free parameters
assumptions (3)
- domain assumption Search terms generalisab*, generalizab*, transportab*, 'dataset shift' combined with Geersing filters identify the relevant methodology.
- ad hoc to paper The categorization axes (data-driven versus knowledge-driven, tailored versus not tailored) are natural and exhaustive enough for the field.
- ad hoc to paper Methods that only reduce predictive uncertainty are out of scope.
Cite this review
Pith. "Pith review of Scoping review of methodology for aiding generalisability and transportability of clinical prediction models." pith.science (2026). https://pith.science/paper/LQYJQJLX
@misc{pith2026241204275,
author = {Pith},
title = {Pith review of: Scoping review of methodology for aiding generalisability and transportability of clinical prediction models},
year = {2026},
howpublished = {\url{https://pith.science/paper/LQYJQJLX}},
note = {Machine review of arXiv:2412.04275}
}
read the original abstract
Generalisability and transportability of clinical prediction models (CPMs) refer to their ability to maintain predictive performance when applied to new populations. While CPMs may show good generalisability or transportability to a specific new population, it is rare for a CPM to be developed using methods that prioritise good generalisability or transportability. There is an emerging literature of such techniques; therefore, this scoping review aims to summarise the main methodological approaches, assumptions, advantages, disadvantages and future development of methodology aiding the generalisability/transportability. Relevant articles were systematically searched from MEDLINE, Embase, medRxiv, arxiv databases until September 2023 using a predefined set of search terms. Extracted information included methodology description, assumptions, applied examples, advantages and disadvantages. The searches found 1,761 articles; 172 were retained for full text screening; 18 were finally included. We categorised the methodologies according to whether they were data-driven or knowledge-driven, and whether are specifically tailored for target population. Data-driven approaches range from data augmentation to ensemble methods and density ratio weighting, while knowledge-driven strategies rely on causal methodology. Future research could focus on comparison of such methodologies on simulated and real datasets to identify their strengths specific applicability, as well as synthesising these approaches for enhancing their practical usefulness.
Figures
Reference graph
Works this paper leans on
-
[1]
(Validat* OR Predict*.ti. OR Rule*) OR (Predict* AND (Outcome* OR Risk* OR Model*)) OR ((History OR Variable* OR Criteria OR Scor* OR Characteristic* OR Finding* OR Factor*) AND (Predict* OR Model* OR Decision* OR Identif* OR Prognos*)) OR (Decision* AND (Model* OR Clinical* OR Logistic Models)) OR (Prognostic AND (History OR Variable* OR Criteria OR Scor...
-
[2]
( generalisability terms ) AND ( CPM term (1) OR CPM term (2) )
Stratification OR ROC Curve OR Discrimination OR Discriminate OR c-statistic OR c statistic OR Area under the curve OR AUC OR Calibration OR Indices OR Algorithm OR Multivariable Articles were searched based on titles containing terms related to generalisability AND terms in (1) OR (2), i.e. ( generalisability terms ) AND ( CPM term (1) OR CPM term (2) )....
-
[3]
Ensemble modelling [27]
-
[4]
Density ratio/importance weighting [29]
-
[5]
Adversarial validation [35]
-
[6]
Measurement error comparison [29]26/11/2024 12:10:00 Knowledge- driven
2024
-
[7]
Selection diagram [19, 37]
-
[8]
Graph surgery estimator [19]
Show all 92 references
-
[9]
Shortcut predictor [22]
-
[10]
Distributionally robust model [30, 31]
-
[11]
Minimising worst-case risk [30, 31]
-
[12]
Optimal stable and immutable sets [30, 31]
-
[13]
Exposing the model to heterogeneous development data will dilute location- or cluster- specific patterns and this can lead to a more generalisable model [33]
Markov blanket, Parent-child nodes [20, 23, 38]26/11/2024 12:10:00 Causality-aware prediction by counterfactual approach [36] Table 1: Classification of the methodology aiding generalisability and transportability of clinical prediction models Data-driven methodologies that ar...
2022
-
[14]
Updating methods improved the performance of a clinical prediction model in new patients
Janssen KJM, Moons KGM, Kalkman CJ, et al. Updating methods improved the performance of a clinical prediction model in new patients. J Clin Epidemiol 2008; 61: 76–86
2008
-
[15]
Clinical prediction models : a practical approach to development, validation, and updating
Steyerberg EW. Clinical prediction models : a practical approach to development, validation, and updating. Second edition. Cham, Switzerland: Springer, 2019. Epub ahead of print 2019. DOI: 10.1007/978-3-030-16399-0
2019 doi
-
[16]
Riley RD, van der Windt D, Croft P, et al. (eds). Prognosis Research in Healthcare: Concepts, Methods, and Impact. Oxford University Press. Epub ahead of print 1 January 2019. DOI: 10.1093/med/9780198796619.001.0001
2019
-
[17]
Assessing the Generalizability of Prognostic Information
Justice AC, Covinsky KE, Berlin JA. Assessing the Generalizability of Prognostic Information. Ann Intern Med 1999; 130: 515–524
1999
-
[18]
A framework for developing, implementing, and evaluating clinical prediction models in an individual participant data meta-analysis
Debray TPA, Moons KGM, Ahmed I, et al. A framework for developing, implementing, and evaluating clinical prediction models in an individual participant data meta-analysis. Stat Med 2013; 32: 3158–3180
2013
-
[19]
Evaluation of clinical prediction models (part 1): from development to external validation
Collins GS, Dhiman P, Ma J, et al. Evaluation of clinical prediction models (part 1): from development to external validation. BMJ 2024; 384: e074819
2024
-
[20]
Evaluation of clinical prediction models (part 2): how to undertake an external validation study
Riley RD, Archer L, Snell KIE, et al. Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ 2024; 384: e074820
2024
-
[21]
Quiñonero-Candela J, Sugiyama M, Schwaighofer A, et al. (eds). Dataset Shift in Machine Learning. The MIT Press. Epub ahead of print 12 December 2008. DOI: 10.7551/mitpress/9780262170055.001.0001
2008
-
[22]
From development to deployment: dataset shift, causality, and shift - stable models in health AI
Subbaswamy A, Saria S. From development to deployment: dataset shift, causality, and shift - stable models in health AI. Biostatistics 2020; 21: 345–352
2020
-
[23]
The Clinician and Dataset Shift in Artificial Intelligence
Finlayson SG, Subbaswamy A, Singh K, et al. The Clinician and Dataset Shift in Artificial Intelligence. N Engl J Med 2021; 385: 283–286
2021
-
[24]
Prognostic models will be victims of their own success, unless…
Lenert MC, Matheny ME, Walsh CG. Prognostic models will be victims of their own success, unless…. J Am Med Inform Assoc 2019; 26: 1645–1650
2019
-
[25]
A review of statistical updating methods for clinical prediction models
Su T-L, Jaki T, Hickey GL, et al. A review of statistical updating methods for clinical prediction models. Stat Methods Med Res 2018; 27: 185–197
2018
-
[26]
Dynamic models to predict health outcomes: current status and methodological challenges
Jenkins DA, Sperrin M, Martin GP, et al. Dynamic models to predict health outcomes: current status and methodological challenges. Diagn Progn Res 2018; 2: 23
2018
-
[27]
Assessing the Performance of Prediction Models: A Framework for Traditional and Novel Measures
Steyerberg EW, Vickers AJ, Cook NR, et al. Assessing the Performance of Prediction Models: A Framework for Traditional and Novel Measures. Epidemiology 2010; 21: 128
2010
-
[28]
Preventing dataset shift from breaking machine-learning biomarkers
Dockès J, Varoquaux G, Poline J-B. Preventing dataset shift from breaking machine-learning biomarkers. GigaScience 2021; 10: giab055
2021
-
[29]
Aggregating published prediction models with individual participant data: a comparison of different approaches
Debray TPA, Koffijberg H, Vergouwe Y, et al. Aggregating published prediction models with individual participant data: a comparison of different approaches. Stat Med 2012; 31: 2697– 2712
2012
-
[30]
Meta-analysis and aggregation of multiple published prediction models
Debray TPA, Koffijberg H, Nieboer D, et al. Meta-analysis and aggregation of multiple published prediction models. Stat Med 2014; 33: 2341–2362
2014
-
[31]
Dynamic Prediction Modeling Approaches for Cardiac Surgery
Hickey GL, Grant SW, Caiado C, et al. Dynamic Prediction Modeling Approaches for Cardiac Surgery. Circ Cardiovasc Qual Outcomes 2013; 6: 649–658. 22
2013
-
[32]
Dynamic Logistic Regression and Dynamic Model Averaging for Binary Classification
McCormick TH, Raftery AE, Madigan D, et al. Dynamic Logistic Regression and Dynamic Model Averaging for Binary Classification. Biometrics 2012; 68: 23–30
2012
-
[33]
Preventing Failures Due to Dataset Shift: Learning Predictive Models That Transport
Subbaswamy A, Schulam P, Saria S. Preventing Failures Due to Dataset Shift: Learning Predictive Models That Transport. In: Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics. PMLR, pp. 3118–3127
-
[34]
A causal framework for assessing the transportability of clinical prediction models
Fehr J, Piccininni M, Kurth T, et al. A causal framework for assessing the transportability of clinical prediction models. 2022; 2022.03.01.22271617
2022
-
[35]
Toward a framework for the design, implementation, and reporting of methodology scoping reviews
Martin GP, Jenkins DA, Bull L, et al. Toward a framework for the design, implementation, and reporting of methodology scoping reviews. J Clin Epidemiol 2020; 127: 191–197
2020
-
[36]
A structural characterization of shortcut features for prediction
Bellamy D, Hernán MA, Beam A. A structural characterization of shortcut features for prediction. Eur J Epidemiol 2022; 37: 563–568
2022
-
[37]
Directed acyclic graphs and causal thinking in clinical risk prediction modeling
Piccininni M, Konigorski S, Rohmann JL, et al. Directed acyclic graphs and causal thinking in clinical risk prediction modeling. BMC Med Res Methodol 2020; 20: 179
2020
-
[38]
switching
and 3) Shapley value [45]. Regularisation shrinks the predictor coefficients toward zero. It 17 reduces model variance and hence overfitting to development data. However, there is no guarantee that taking shrunken coefficients fitted from development population would improve g...
-
[39]
Developing more generalizable prediction models from pooled studies and large clustered data sets
de Jong VMT, Moons KGM, Eijkemans MJC, et al. Developing more generalizable prediction models from pooled studies and large clustered data sets. Stat Med 2021; 40: 3533–3559
2021
-
[40]
Search Filters for Finding Prognostic and Diagnostic Prediction Studies in Medline to Enhance Systematic Reviews
Geersing G-J, Bouwmeester W, Zuithoff P, et al. Search Filters for Finding Prognostic and Diagnostic Prediction Studies in Medline to Enhance Systematic Reviews. PLOS ONE 2012; 7: e32844
2012
-
[41]
More Generalizable Models For Sepsis Detection Under Covariate Shift
Gao J, Mar PL, Chen G. More Generalizable Models For Sepsis Detection Under Covariate Shift. AMIA Jt Summits Transl Sci Proc AMIA Jt Summits Transl Sci 2021; 2021: 220–228
2021
-
[42]
Learning patient-level prediction models across multiple healthcare databases: evaluation of ensembles for increasing model transportability
Reps J.M., Williams R.D., Schuemie M.J., et al. Learning patient-level prediction models across multiple healthcare databases: evaluation of ensembles for increasing model transportability. BMC Med Inform Decis Mak 2022; 22: 142
2022
-
[43]
How variation in predictor measurement affects the discriminative ability and transportability of a prediction model
Pajouheshnia R, van Smeden M, Peelen LM, et al. How variation in predictor measurement affects the discriminative ability and transportability of a prediction model. J Clin Epidemiol 2019; 105: 136–141
2019
-
[44]
Which Invariance Should We Transfer? A Causal Minimax Learning Approach, http://arxiv.org/abs/2107.01876 (2023, accessed 25 September 2023)
Liu M, Zheng X, Sun X, et al. Which Invariance Should We Transfer? A Causal Minimax Learning Approach, http://arxiv.org/abs/2107.01876 (2023, accessed 25 September 2023)
2023 arXiv
-
[45]
Evaluating Model Robustness and Stability to Dataset Shift, http://arxiv.org/abs/2010.15100 (2021, accessed 25 September 2023)
Subbaswamy A, Adams R, Saria S. Evaluating Model Robustness and Stability to Dataset Shift, http://arxiv.org/abs/2010.15100 (2021, accessed 25 September 2023)
2010 arXiv
-
[46]
Transporting a Prediction Model for Use in a New Target Population
Steingrimsson JA, Gatsonis C, Li B, et al. Transporting a Prediction Model for Use in a New Target Population. Am J Epidemiol 2023; 192: 296–304
2023
-
[47]
Embracing cohort heterogeneity in clinical machine learning development: a step toward generalizable models
Schinkel M, Bennis FC, Boerman AW, et al. Embracing cohort heterogeneity in clinical machine learning development: a step toward generalizable models. Sci Rep 2023; 13: 8363
2023
-
[48]
Instance Weighting for Patient-Specific Risk Stratification Models
Gong JJ, Sundt TM, Rawn JD, et al. Instance Weighting for Patient-Specific Risk Stratification Models. In: Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. New York, NY, USA: Association for Computing Machinery, pp. 369–378. 23
-
[49]
Managing dataset shift by adversarial validation for credit scoring, http://arxiv.org/abs/2112.10078 (2021, accessed 25 September 2023)
Qian H, Wang B, Ma P, et al. Managing dataset shift by adversarial validation for credit scoring, http://arxiv.org/abs/2112.10078 (2021, accessed 25 September 2023)
2021 arXiv
-
[50]
Neto EC. Towards causality-aware predictions in static anticausal machine learning tasks: the linear structural causal model case, http://arxiv.org/abs/2001.03998 (2020, accessed 25 September 2023)
2001 arXiv
-
[51]
Transportability of Causal and Statistical Relations: A Formal Approach
Pearl J, Bareinboim E. Transportability of Causal and Statistical Relations: A Formal Approach. In: 2011 IEEE 11th International Conference on Data Mining Workshops, pp. 540–547
2011
-
[52]
Evaluation of Feature Selection Methods for Preserving Machine Learning Performance in the Presence of Temporal Dataset Shift in Clinical Medicine
Lemmon J., Guo L.L., Posada J., et al. Evaluation of Feature Selection Methods for Preserving Machine Learning Performance in the Presence of Temporal Dataset Shift in Clinical Medicine. Methods Inf Med 2023; 62: 60–70
2023
-
[53]
Targeted validation: validating clinical prediction models in their intended population and setting
Sperrin M, Riley RD, Collins GS, et al. Targeted validation: validating clinical prediction models in their intended population and setting. Diagn Progn Res 2022; 6: 24
2022
-
[54]
Why do probabilistic clinical models fail to transport between sites
Lasko TA, Strobl EV, Stead WW. Why do probabilistic clinical models fail to transport between sites. Npj Digit Med 2024; 7: 1–8
2024
-
[55]
The Book of Why: The New Science of Cause and Effect
Pearl J, Mackenzie D. The Book of Why: The New Science of Cause and Effect. 1st ed. USA: Basic Books, Inc., 2018
2018
-
[56]
Meta-Transportability of Causal Effects: A Formal Approach
Bareinboim E, Pearl J. Meta-Transportability of Causal Effects: A Formal Approach. In: Proceedings of the Sixteenth International Conference on Artificial Intelligence and Statistics . PMLR, pp. 135–143
-
[57]
There is no such thing as a validated prediction model
Van Calster B, Steyerberg EW, Wynants L, et al. There is no such thing as a validated prediction model. BMC Med 2023; 21: 70
2023
-
[58]
External validation of clinical prediction models using big datasets from e-health records or IPD meta-analysis: opportunities and challenges
Riley RD, Ensor J, Snell KIE, et al. External validation of clinical prediction models using big datasets from e-health records or IPD meta-analysis: opportunities and challenges. BMJ 2016; 353: i3140
2016
-
[59]
A feature selection method based on Shapley values robust to concept shift in regression, http://arxiv.org/abs/2304.14774 (2023, accessed 25 September 2023)
Sebastián C, González-Guillén CE. A feature selection method based on Shapley values robust to concept shift in regression, http://arxiv.org/abs/2304.14774 (2023, accessed 25 September 2023)
2023 arXiv
-
[60]
Transportability of Trial Results Using Inverse Odds of Sampling Weights
Westreich D., Edwards J.K., Lesko C.R., et al. Transportability of Trial Results Using Inverse Odds of Sampling Weights. Am J Epidemiol 2017; 186: 1010–1014
2017
-
[61]
Robust Estimation of Encouragement Design Intervention Effects Transported Across Sites
Rudolph KE, Laan MJ. Robust Estimation of Encouragement Design Intervention Effects Transported Across Sites. J R Stat Soc Ser B Stat Methodol 2017; 79: 1509–1525
2017
-
[62]
Extending inferences from a randomized trial to a new target population
Dahabreh IJ, Robertson SE, Steingrimsson JA, et al. Extending inferences from a randomized trial to a new target population. Stat Med 2020; 39: 1999–2014
2020
-
[63]
A scoping review of causal methods enabling predictions under hypothetical interventions
Lin L, Sperrin M, Jenkins DA, et al. A scoping review of causal methods enabling predictions under hypothetical interventions. Diagn Progn Res 2021; 5: 3
2021
-
[64]
A Review of Generalizability and Transportability
Degtiar I, Rose S. A Review of Generalizability and Transportability. Annu Rev Stat Its Appl 2023; 10: 501–524
2023
-
[65]
An Overview of Current Methods for Real-World Applications to Generalize or Transport Clinical Trial Findings to Target Populations of Interest
Ling AY, Montez-Rath ME, Carita P, et al. An Overview of Current Methods for Real-World Applications to Generalize or Transport Clinical Trial Findings to Target Populations of Interest. Epidemiology; 10.1097/EDE.0000000000001633. 24
-
[66]
A Survey on Transfer Learning
Pan SJ, Yang Q. A Survey on Transfer Learning. IEEE Trans Knowl Data Eng 2010; 22: 1345– 1359
2010
-
[67]
A survey of transfer learning
Weiss K, Khoshgoftaar TM, Wang D. A survey of transfer learning. J Big Data 2016; 3: 9
2016
-
[68]
Causal effect on a target population: A sensitivity analysis to handle missing covariates
Colnet B, Josse J, Varoquaux G, et al. Causal effect on a target population: A sensitivity analysis to handle missing covariates. J Causal Inference 2022; 10: 372–414
2022
- [69]
- [70]
-
[71]
Exploiting Causal Structure for Robust Model Selection in Unsupervised Domain Adaptation
Kyono T, van der Schaar M. Exploiting Causal Structure for Robust Model Selection in Unsupervised Domain Adaptation. IEEE Trans Artif Intell 2021; 2: 494–507
2021
-
[72]
Off-Policy Evaluation and Learning for External Validity under a Covariate Shift
Uehara M, Kato M, Yasui S. Off-Policy Evaluation and Learning for External Validity under a Covariate Shift. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 49–61
-
[73]
A Unified Framework on Generalizability of Clinical Prediction Models
Wan B, Caffo B, Vedula SS. A Unified Framework on Generalizability of Clinical Prediction Models. Front Artif Intell 2022; 5: 872720
2022
-
[74]
Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available
Hernán MA, Robins JM. Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available. Am J Epidemiol 2016; 183: 758–764. Supplemental table 1: Summary of included 18 papers N o Aut hor s Methodology Description Need target populatio n? Domain knowledge ? Sta...
2016
-
[75]
Direct methods - KMM and RuLSIF
-
[76]
Using density ratio weighting is not possible to correct concept shift
Indirect methods - LR with L1 and L2 penalty, and Random Forests Yes, note that outcome data is required only for the indirect methods No The distribution shift is compatible with covariate shift Estimation of density ratio with both parametric and non- parametric models are u...
-
[77]
Positivity: this condition necessitates that for every covariate pattern present in the target population, there is a positive probability of encountering the same pattern in the development population. The methodology aids tailoring a prediction model from the development pop...
-
[78]
The applicability of the methods may be limited if there is a lack of access to covariate data from the target population
-
[79]
The set of covariates necessary to satisfy the conditional independence assumption might be extensive than the set actually used in the model
-
[80]
The authors used simulated data to test the performance of the prediction models
The approach did not address the practical issues such as missing data and measurement error. The authors used simulated data to test the performance of the prediction models. Data was generated using a linear model with heteroscedastic errors and simulated membership for deve...
2001
-
[81]
covariate shift Assessing multiple feature selection methods for transportability across several clinical outcomes - Only used MIMIC-IV dataset for evaluation
Remove and Retrain (ROAR): identified importance predictors with a permutation- based feature importance score and removed, then retrained the model No (requires only historical data) No (except for causal predictor selection) No explicitly mentioned - it is inferred that the ...
2008
-
[82]
Clustering the training dataset to mimick real-world heterogeneity and use ICEV for those clusters to perform stepwise predictor selection
-
[83]
The choices of candidate predictors should be guided by clinical knowledge
From stepwise IECV, aggregate performance metrics of all clusters using meta-analytic methods to select the best model that can generalise across heterogenous dataset No 1. The choices of candidate predictors should be guided by clinical knowledge. Heterogeneity of training da...
-
[84]
Generalisability may come at a cost of improving one predictive performance (such as discrimination), but deteriorating another one
-
[85]
Focusing solely on predictor selection may have an impact on model's precision
-
[86]
In their prior work, the author developed site-specific XGBoost models for three healthcare systems (KPSC, IH, UWM) using over 200 candidate features
It is still possible that the model requires recalibration on the target settings. In their prior work, the author developed site-specific XGBoost models for three healthcare systems (KPSC, IH, UWM) using over 200 candidate features. The top 20 features, each with importance v...
-
[87]
The selection diagram is correctly specified
-
[88]
The interventinal distribution is identifiable
-
[89]
(prior probability shift)
Predicting variables is mutable. (prior probability shift)
-
[90]
Simulating data from DGP by pre- defined DAG and generate different environments by vary mutable variables
-
[91]
N o Aut hor s Methodology Description Need target populatio n? Domain knowledge ? Stated assumptions Reported advantages Reported limitations Experiments 1 5 Pearl et al
Using UCI Bike sharing dataset to predict the number of hourly bike rentals in different seasons by assuming possible DGP as a DAG including mutable variables. N o Aut hor s Methodology Description Need target populatio n? Domain knowledge ? Stated assumptions Reported advanta...
-
[92]
The covariate C need to satisfy S- adminissibilty, i.e. y ⫫ x | C - Mainly addresses the causal effect of a single variable X and does not cover prediction scenarios with multiple predictors - Lacks cases where the covariate set C overlaps with the predictor set X, which is co...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.