Pith. sign in

REVIEW 4 major objections 5 minor 92 references

Scoping review of methodology for aiding generalisability and transportability of clinical prediction models

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This scoping review claims that development-phase methods for generalisable or transportable clinical prediction models split into data-driven and knowledge-driven families, with no head-to-head evaluation yet.

desk verdict A useful but search-restricted taxonomy of development-stage methods for CPM generalisability and transportability; worth reviewing, but the completeness claims need major revision. read the letter →

arxiv 2412.04275 v1 pith:LQYJQJLX submitted 2024-12-05 stat.ME

classification stat.ME
keywords clinicalpredictionmodelsgeneralisabilitytransportabilitydatasetshiftcausalinferencescopingreviewdensityratioweightingselectiondiagram
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This scoping review tries to establish that methodology for making clinical prediction models generalisable or transportable at the development stage falls into two distinct families: data-driven methods that enhance robustness through heterogeneous data, ensemble learning, or density-ratio weighting, and knowledge-driven methods that encode causal structure through selection diagrams, graph surgery, invariant predictors, or counterfactual adjustment. The review also classifies each method by whether it requires data from the specific target population. If the taxonomy is right, developers can choose an approach based on whether they have target covariates and causal knowledge, and researchers can see that the two families have not yet been compared on common benchmarks. The paper argues this matters because most models are developed without prioritising transportability, and a better map of methods would reduce the need for post-hoc updating.

What carries the argument

The central analytic object is the two-way classification (data-driven vs knowledge-driven × not tailored vs tailored) constructed from the 18 included papers. Within that taxonomy, the load-bearing technical apparatus is the density ratio $r(x)=P_{target}(x)/P_{source}(x)$ for data-driven tailoring, and, for knowledge-driven methods, the selection diagram and its associated transport formula, e.g., $P(Y=1 \mid X=x, S=0) = \sum_z P(Y=1 \mid X=x, S=1, Z=z)P(Z=z)$ under S-admissibility, together with the graph surgery estimator that turns interventional distributions into observable ones. These devices let a model be transported or made invariant to shifts without re-estimation on the target population.

What would settle it

A replication of the search with additional title and abstract terms such as 'domain adaptation', 'domain generalization', 'invariant learning', 'distribution shift', 'external validity', and 'transfer learning' in the same databases up to September 2023, followed by full-text screening against the review's inclusion and exclusion criteria, would test the taxonomy's completeness. If any resultant paper proposes a development-phase method for generalisable or transportable clinical prediction models that fits neither the data-driven nor the knowledge-driven category, or neither the tailored nor the not-tailored axis, the review's classification is incomplete.

Watch

Extended reading notes

Core claim

The paper's central claim is that the emerging literature on development-phase methods for generalisability and transportability of clinical prediction models can be organised by two independent axes: data-driven versus knowledge-driven, and whether the method is tailored to a specific target population. Data-driven methods include training on heterogeneous multi-centre data (internal-external cross-validation), ensemble models, density-ratio or importance weighting, adversarial validation, and measurement-error alignment; knowledge-driven methods include selection diagrams and transport formulas, graph surgery estimators, shortcut-predictor removal, distributionally robust and worst-case risk minimisation with mutable/immutable sets, Markov-blanket or parent-child predictor selection, and counterfactual adjustment for anticausal (diagnostic) tasks. The review states that no included study has comparatively evaluated these methods on shared simulated or real datasets, and it identifies integration of the two families as an open direction.

Load-bearing premise

The search strategy, based on titles containing generalisab*, generalizab*, transportab* or 'dataset shift' plus the Geersing filters, captures the full universe of relevant methodology papers, even though the review itself found that one known seed paper (Bellamy et al.) was missed by the search and only included manually.

Editorial extensions

If this is right

  • If the taxonomy is correct, a model developer can choose a method family based on whether covariate data from the target population is available and whether reliable causal knowledge exists.
  • Because the review finds that density-ratio weighting corrects covariate shift but not concept shift, these methods cannot rescue models when the predictor-outcome relationship itself changes.
  • The absence of comparative evaluation means that, as of the review's search date, no evidence indicates which method family performs best for a given type of dataset shift.
  • Knowledge-driven methods all rest on correctly specified causal structures, so misspecification of the selection diagram or the mutable/immutable split is a shared failure mode.
  • Integrating data-driven heterogeneity with causal invariant selection is identified by the paper as a promising direction for future methodology.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the taxonomy would be to run the 18 methods on a common battery of simulated shifts (covariate, prior, concept, measurement error) and real EHR datasets; the paper's own call for such benchmarks implies that none exists.
  • Because one known seed paper (Bellamy et al.) escaped the search and snowballing, the review's map is probably incomplete; adding domain-adaptation and invariant-learning search terms could surface methods that fit neither family.
  • The tailored-versus-universal axis suggests a potential bridge to transfer learning: density-ratio weighting is already used in domain adaptation, so formal connections could unify the two literatures.
  • A head-to-head comparison might find that knowledge-driven methods degrade gracefully under correct causal assumptions but fail catastrophically when those assumptions are wrong, while data-driven methods trade off performance on any single population for robustness across many.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript is a scoping review of methodology for improving the generalisability and transportability of clinical prediction models (CPMs). The authors searched MEDLINE, Embase, medRxiv, and arXiv up to September 2023 using title-based search terms (generalisab*, generalizab*, transportab*, 'dataset shift') plus Geersing filters, screened 1,761 records, and included 18 papers. They categorize the identified methods into data-driven versus knowledge-driven families and, within each, whether the method is tailored to a specific target population. The paper summarises assumptions, advantages, disadvantages, and applications of each included method, and offers future research directions such as comparative evaluations on simulated and real data and integration of the two families.

Significance. If the taxonomy is accepted, this review provides a useful map of an emerging and fragmented literature, and it makes a plausible conceptual distinction between methods that exploit data heterogeneity or weighting to improve robustness and methods that rely on causal structure or domain knowledge. The review also clearly identifies the absence of comparative evaluations and the need for methods that require less target-population data. The paper's strengths include a documented search strategy, a PRISMA flow diagram, a structured extraction of assumptions and limitations in supplementary material, and an explicit discussion of search limitations. The main value is as a synthesis and agenda-setting review for clinical prediction methodology; its central claim, however, depends on the completeness and representativeness of the included literature.

major comments (4)
  1. [Search Strategy; Results; Discussion] The central claim that methodologies for aiding generalisability/transportability fall into two families is load-bearing, yet the evidence for completeness is weak. The search was restricted to titles containing generalisab*, generalizab*, transportab* or 'dataset shift', and the paper concedes that one of the five validation papers (Bellamy et al.) was not found by the search or snowballing, and that Debray et al.'s framework was excluded because it did not match keywords. This means the search has demonstrated imperfect recall even on a hand-picked set; it may have missed large adjacent literatures using terms such as domain shift, domain generalization, distributional robustness, invariant risk minimization, or external validity. The paper should either rerun the search with additional synonym-based terms and a broader validation set, or explicitly reframe the taxonomy as provisional and limited to the included papers rather than as a complete classification of the field.
  2. [Information sources; Data extraction] Screening and data extraction were conducted by a single author (KP) with no independent second reviewer or adjudication process. For a review whose core output is a classification and summary of included methods, single-reviewer screening is a substantial risk to reliability, as title/abstract decisions determine the inclusion set on which the taxonomy is built. The manuscript should report whether any reliability checks (e.g., dual screening on a subset, or verification of extraction by a second author) were performed, or justify why single-reviewer screening is appropriate for this scoping review.
  3. [Scope of review; Results; Table 1] The exclusion of transfer learning and adjustment methods is not consistently applied. The scope states that methods that adjust an already-developed model are excluded, but density ratio/importance weighting methods (Gao et al., Steingrimsson et al., Dockès et al.) are included in Table 1 even though they are commonly classified as transfer-learning or domain-adaptation techniques. If these are included because they are used during development, the boundary should be stated more precisely; if they are excluded in other guises, the review risks including members of a family while claiming that family is out of scope, which weakens the exhaustiveness of the two-family taxonomy.
  4. [Knowledge-driven methodology that are not tailored for target population; Table 1] The manuscript states 'Fewer studies (7 of 19) focused on knowledge-driven methodologies', but the review includes 18 papers, not 19. More substantively, Piccininni et al. and Liu et al. are placed in both the data-driven and knowledge-driven categories, so the two families are not mutually exclusive. The paper should provide explicit criteria for assigning a method to one or both families, and should discuss whether the overlap indicates that the data-driven/knowledge-driven distinction is better viewed as a continuum or as complementary axes rather than a clean dichotomy.
minor comments (5)
  1. [Information sources] The database name 'medRxiv' is misspelled as 'merRxiv' in the Information sources section.
  2. [Table 1] Table 1 contains stray timestamps '26/11/2024 12:10:00' in two cells; these artifacts should be removed.
  3. [Knowledge-driven methodology that are not tailored for target population] The text says '7 of 19'; the correct total is 18 included papers. Also, 'Parent Child - PB' should likely be 'Parent-Child (PC)' for consistency.
  4. [Knowledge-driven methodology that is tailored for a specific target population] There is a typo, 'Causaility-aware', which should read 'Causality-aware'.
  5. [Snowballing (Citation search)] The phrase 'To compliment the database search' should be 'To complement the database search'.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the taxonomy is a literature synthesis, and the search check is reported transparently.

full rationale

This is a scoping review rather than a derivation, so its central claim—that methodologies aiding generalisability and transportability of clinical prediction models can be grouped into data-driven and knowledge-driven families—is a categorization of the 18 included papers, not a quantity fitted to data and then renamed as a prediction. The search terms were derived from five seed papers and checked against them; while this is a sensitivity exercise rather than an independent validation, the paper does not infer completeness from that check and explicitly reports that one seed paper (Bellamy et al.) was not found by the search and had to be added manually. The self-citations (Martin et al. 2020 for scoping-review search strategy, Sperrin et al. 2022 for tailored validation, and Lin et al. 2021 for causal methods under hypothetical interventions) are contextual and are not load-bearing premises of the taxonomy. No uniqueness theorem from the authors is invoked to forbid alternative classifications, and no method is defined in terms of the taxonomy it is supposed to support. The acknowledged search gap—title-based terms that missed synonym-rich literatures such as transfer learning—is a limitation on completeness, not a circularity, because the taxonomy is transparently stated relative to the included set.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The review rests on assumptions about search completeness and the usefulness of its organizing categories. It introduces no free parameters and no new entities. The central risk is that the search missed relevant methods, which the paper itself acknowledges.

assumptions (3)
  • domain assumption Search terms generalisab*, generalizab*, transportab*, 'dataset shift' combined with Geersing filters identify the relevant methodology.
    Methods section; the authors rely on this to claim comprehensiveness, and they report one validation paper was missed.
  • ad hoc to paper The categorization axes (data-driven versus knowledge-driven, tailored versus not tailored) are natural and exhaustive enough for the field.
    This is the paper's organizing contribution; it is imposed by the reviewers and not derived from the included papers.
  • ad hoc to paper Methods that only reduce predictive uncertainty are out of scope.
    Defined in the exclusion criteria: 'We aim to find the methodology that deterministically improves generalisability/transportability rather than reducing the predictive uncertainty.'

how reviews work

0 comments
Cite this review

Pith. "Pith review of Scoping review of methodology for aiding generalisability and transportability of clinical prediction models." pith.science (2026). https://pith.science/paper/LQYJQJLX

@misc{pith2026241204275,
  author       = {Pith},
  title        = {Pith review of: Scoping review of methodology for aiding generalisability and transportability of clinical prediction models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LQYJQJLX}},
  note         = {Machine review of arXiv:2412.04275}
}
read the original abstract

Generalisability and transportability of clinical prediction models (CPMs) refer to their ability to maintain predictive performance when applied to new populations. While CPMs may show good generalisability or transportability to a specific new population, it is rare for a CPM to be developed using methods that prioritise good generalisability or transportability. There is an emerging literature of such techniques; therefore, this scoping review aims to summarise the main methodological approaches, assumptions, advantages, disadvantages and future development of methodology aiding the generalisability/transportability. Relevant articles were systematically searched from MEDLINE, Embase, medRxiv, arxiv databases until September 2023 using a predefined set of search terms. Extracted information included methodology description, assumptions, applied examples, advantages and disadvantages. The searches found 1,761 articles; 172 were retained for full text screening; 18 were finally included. We categorised the methodologies according to whether they were data-driven or knowledge-driven, and whether are specifically tailored for target population. Data-driven approaches range from data augmentation to ensemble methods and density ratio weighting, while knowledge-driven strategies rely on causal methodology. Future research could focus on comparison of such methodologies on simulated and real datasets to identify their strengths specific applicability, as well as synthesising these approaches for enhancing their practical usefulness.

Figures

Figures reproduced from arXiv: 2412.04275 by the authors.

Figure 4
Figure 4. Prediction of Y with only variables that reside in Markov Blanket (red dash square). [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

92 extracted references · 73 canonical work pages

  1. [1]

    (Validat* OR Predict*.ti. OR Rule*) OR (Predict* AND (Outcome* OR Risk* OR Model*)) OR ((History OR Variable* OR Criteria OR Scor* OR Characteristic* OR Finding* OR Factor*) AND (Predict* OR Model* OR Decision* OR Identif* OR Prognos*)) OR (Decision* AND (Model* OR Clinical* OR Logistic Models)) OR (Prognostic AND (History OR Variable* OR Criteria OR Scor...

  2. [2]

    ( generalisability terms ) AND ( CPM term (1) OR CPM term (2) )

    Stratification OR ROC Curve OR Discrimination OR Discriminate OR c-statistic OR c statistic OR Area under the curve OR AUC OR Calibration OR Indices OR Algorithm OR Multivariable Articles were searched based on titles containing terms related to generalisability AND terms in (1) OR (2), i.e. ( generalisability terms ) AND ( CPM term (1) OR CPM term (2) )....

  3. [3]

    Ensemble modelling [27]

  4. [4]

    Density ratio/importance weighting [29]

  5. [5]

    Adversarial validation [35]

  6. [6]

    Measurement error comparison [29]26/11/2024 12:10:00 Knowledge- driven

  7. [7]

    Selection diagram [19, 37]

  8. [8]

    Graph surgery estimator [19]

Show all 92 references
  1. [9]

    Shortcut predictor [22]

  2. [10]

    Distributionally robust model [30, 31]

  3. [11]

    Minimising worst-case risk [30, 31]

  4. [12]

    Optimal stable and immutable sets [30, 31]

  5. [13]

    Exposing the model to heterogeneous development data will dilute location- or cluster- specific patterns and this can lead to a more generalisable model [33]

    Markov blanket, Parent-child nodes [20, 23, 38]26/11/2024 12:10:00 Causality-aware prediction by counterfactual approach [36] Table 1: Classification of the methodology aiding generalisability and transportability of clinical prediction models Data-driven methodologies that ar...

  6. [14]

    Updating methods improved the performance of a clinical prediction model in new patients

    Janssen KJM, Moons KGM, Kalkman CJ, et al. Updating methods improved the performance of a clinical prediction model in new patients. J Clin Epidemiol 2008; 61: 76–86

  7. [15]

    Clinical prediction models : a practical approach to development, validation, and updating

    Steyerberg EW. Clinical prediction models : a practical approach to development, validation, and updating. Second edition. Cham, Switzerland: Springer, 2019. Epub ahead of print 2019. DOI: 10.1007/978-3-030-16399-0

  8. [16]

    Riley RD, van der Windt D, Croft P, et al. (eds). Prognosis Research in Healthcare: Concepts, Methods, and Impact. Oxford University Press. Epub ahead of print 1 January 2019. DOI: 10.1093/med/9780198796619.001.0001

  9. [17]

    Assessing the Generalizability of Prognostic Information

    Justice AC, Covinsky KE, Berlin JA. Assessing the Generalizability of Prognostic Information. Ann Intern Med 1999; 130: 515–524

  10. [18]

    A framework for developing, implementing, and evaluating clinical prediction models in an individual participant data meta-analysis

    Debray TPA, Moons KGM, Ahmed I, et al. A framework for developing, implementing, and evaluating clinical prediction models in an individual participant data meta-analysis. Stat Med 2013; 32: 3158–3180

  11. [19]

    Evaluation of clinical prediction models (part 1): from development to external validation

    Collins GS, Dhiman P, Ma J, et al. Evaluation of clinical prediction models (part 1): from development to external validation. BMJ 2024; 384: e074819

  12. [20]

    Evaluation of clinical prediction models (part 2): how to undertake an external validation study

    Riley RD, Archer L, Snell KIE, et al. Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ 2024; 384: e074820

  13. [21]

    Quiñonero-Candela J, Sugiyama M, Schwaighofer A, et al. (eds). Dataset Shift in Machine Learning. The MIT Press. Epub ahead of print 12 December 2008. DOI: 10.7551/mitpress/9780262170055.001.0001

  14. [22]

    From development to deployment: dataset shift, causality, and shift - stable models in health AI

    Subbaswamy A, Saria S. From development to deployment: dataset shift, causality, and shift - stable models in health AI. Biostatistics 2020; 21: 345–352

  15. [23]

    The Clinician and Dataset Shift in Artificial Intelligence

    Finlayson SG, Subbaswamy A, Singh K, et al. The Clinician and Dataset Shift in Artificial Intelligence. N Engl J Med 2021; 385: 283–286

  16. [24]

    Prognostic models will be victims of their own success, unless…

    Lenert MC, Matheny ME, Walsh CG. Prognostic models will be victims of their own success, unless…. J Am Med Inform Assoc 2019; 26: 1645–1650

  17. [25]

    A review of statistical updating methods for clinical prediction models

    Su T-L, Jaki T, Hickey GL, et al. A review of statistical updating methods for clinical prediction models. Stat Methods Med Res 2018; 27: 185–197

  18. [26]

    Dynamic models to predict health outcomes: current status and methodological challenges

    Jenkins DA, Sperrin M, Martin GP, et al. Dynamic models to predict health outcomes: current status and methodological challenges. Diagn Progn Res 2018; 2: 23

  19. [27]

    Assessing the Performance of Prediction Models: A Framework for Traditional and Novel Measures

    Steyerberg EW, Vickers AJ, Cook NR, et al. Assessing the Performance of Prediction Models: A Framework for Traditional and Novel Measures. Epidemiology 2010; 21: 128

  20. [28]

    Preventing dataset shift from breaking machine-learning biomarkers

    Dockès J, Varoquaux G, Poline J-B. Preventing dataset shift from breaking machine-learning biomarkers. GigaScience 2021; 10: giab055

  21. [29]

    Aggregating published prediction models with individual participant data: a comparison of different approaches

    Debray TPA, Koffijberg H, Vergouwe Y, et al. Aggregating published prediction models with individual participant data: a comparison of different approaches. Stat Med 2012; 31: 2697– 2712

  22. [30]

    Meta-analysis and aggregation of multiple published prediction models

    Debray TPA, Koffijberg H, Nieboer D, et al. Meta-analysis and aggregation of multiple published prediction models. Stat Med 2014; 33: 2341–2362

  23. [31]

    Dynamic Prediction Modeling Approaches for Cardiac Surgery

    Hickey GL, Grant SW, Caiado C, et al. Dynamic Prediction Modeling Approaches for Cardiac Surgery. Circ Cardiovasc Qual Outcomes 2013; 6: 649–658. 22

  24. [32]

    Dynamic Logistic Regression and Dynamic Model Averaging for Binary Classification

    McCormick TH, Raftery AE, Madigan D, et al. Dynamic Logistic Regression and Dynamic Model Averaging for Binary Classification. Biometrics 2012; 68: 23–30

  25. [33]

    Preventing Failures Due to Dataset Shift: Learning Predictive Models That Transport

    Subbaswamy A, Schulam P, Saria S. Preventing Failures Due to Dataset Shift: Learning Predictive Models That Transport. In: Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics. PMLR, pp. 3118–3127

  26. [34]

    A causal framework for assessing the transportability of clinical prediction models

    Fehr J, Piccininni M, Kurth T, et al. A causal framework for assessing the transportability of clinical prediction models. 2022; 2022.03.01.22271617

  27. [35]

    Toward a framework for the design, implementation, and reporting of methodology scoping reviews

    Martin GP, Jenkins DA, Bull L, et al. Toward a framework for the design, implementation, and reporting of methodology scoping reviews. J Clin Epidemiol 2020; 127: 191–197

  28. [36]

    A structural characterization of shortcut features for prediction

    Bellamy D, Hernán MA, Beam A. A structural characterization of shortcut features for prediction. Eur J Epidemiol 2022; 37: 563–568

  29. [37]

    Directed acyclic graphs and causal thinking in clinical risk prediction modeling

    Piccininni M, Konigorski S, Rohmann JL, et al. Directed acyclic graphs and causal thinking in clinical risk prediction modeling. BMC Med Res Methodol 2020; 20: 179

  30. [38]

    switching

    and 3) Shapley value [45]. Regularisation shrinks the predictor coefficients toward zero. It 17 reduces model variance and hence overfitting to development data. However, there is no guarantee that taking shrunken coefficients fitted from development population would improve g...

  31. [39]

    Developing more generalizable prediction models from pooled studies and large clustered data sets

    de Jong VMT, Moons KGM, Eijkemans MJC, et al. Developing more generalizable prediction models from pooled studies and large clustered data sets. Stat Med 2021; 40: 3533–3559

  32. [40]

    Search Filters for Finding Prognostic and Diagnostic Prediction Studies in Medline to Enhance Systematic Reviews

    Geersing G-J, Bouwmeester W, Zuithoff P, et al. Search Filters for Finding Prognostic and Diagnostic Prediction Studies in Medline to Enhance Systematic Reviews. PLOS ONE 2012; 7: e32844

  33. [41]

    More Generalizable Models For Sepsis Detection Under Covariate Shift

    Gao J, Mar PL, Chen G. More Generalizable Models For Sepsis Detection Under Covariate Shift. AMIA Jt Summits Transl Sci Proc AMIA Jt Summits Transl Sci 2021; 2021: 220–228

  34. [42]

    Learning patient-level prediction models across multiple healthcare databases: evaluation of ensembles for increasing model transportability

    Reps J.M., Williams R.D., Schuemie M.J., et al. Learning patient-level prediction models across multiple healthcare databases: evaluation of ensembles for increasing model transportability. BMC Med Inform Decis Mak 2022; 22: 142

  35. [43]

    How variation in predictor measurement affects the discriminative ability and transportability of a prediction model

    Pajouheshnia R, van Smeden M, Peelen LM, et al. How variation in predictor measurement affects the discriminative ability and transportability of a prediction model. J Clin Epidemiol 2019; 105: 136–141

  36. [44]

    Which Invariance Should We Transfer? A Causal Minimax Learning Approach, http://arxiv.org/abs/2107.01876 (2023, accessed 25 September 2023)

    Liu M, Zheng X, Sun X, et al. Which Invariance Should We Transfer? A Causal Minimax Learning Approach, http://arxiv.org/abs/2107.01876 (2023, accessed 25 September 2023)

  37. [45]

    Evaluating Model Robustness and Stability to Dataset Shift, http://arxiv.org/abs/2010.15100 (2021, accessed 25 September 2023)

    Subbaswamy A, Adams R, Saria S. Evaluating Model Robustness and Stability to Dataset Shift, http://arxiv.org/abs/2010.15100 (2021, accessed 25 September 2023)

  38. [46]

    Transporting a Prediction Model for Use in a New Target Population

    Steingrimsson JA, Gatsonis C, Li B, et al. Transporting a Prediction Model for Use in a New Target Population. Am J Epidemiol 2023; 192: 296–304

  39. [47]

    Embracing cohort heterogeneity in clinical machine learning development: a step toward generalizable models

    Schinkel M, Bennis FC, Boerman AW, et al. Embracing cohort heterogeneity in clinical machine learning development: a step toward generalizable models. Sci Rep 2023; 13: 8363

  40. [48]

    Instance Weighting for Patient-Specific Risk Stratification Models

    Gong JJ, Sundt TM, Rawn JD, et al. Instance Weighting for Patient-Specific Risk Stratification Models. In: Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. New York, NY, USA: Association for Computing Machinery, pp. 369–378. 23

  41. [49]

    Managing dataset shift by adversarial validation for credit scoring, http://arxiv.org/abs/2112.10078 (2021, accessed 25 September 2023)

    Qian H, Wang B, Ma P, et al. Managing dataset shift by adversarial validation for credit scoring, http://arxiv.org/abs/2112.10078 (2021, accessed 25 September 2023)

  42. [50]

    Neto EC. Towards causality-aware predictions in static anticausal machine learning tasks: the linear structural causal model case, http://arxiv.org/abs/2001.03998 (2020, accessed 25 September 2023)

  43. [51]

    Transportability of Causal and Statistical Relations: A Formal Approach

    Pearl J, Bareinboim E. Transportability of Causal and Statistical Relations: A Formal Approach. In: 2011 IEEE 11th International Conference on Data Mining Workshops, pp. 540–547

  44. [52]

    Evaluation of Feature Selection Methods for Preserving Machine Learning Performance in the Presence of Temporal Dataset Shift in Clinical Medicine

    Lemmon J., Guo L.L., Posada J., et al. Evaluation of Feature Selection Methods for Preserving Machine Learning Performance in the Presence of Temporal Dataset Shift in Clinical Medicine. Methods Inf Med 2023; 62: 60–70

  45. [53]

    Targeted validation: validating clinical prediction models in their intended population and setting

    Sperrin M, Riley RD, Collins GS, et al. Targeted validation: validating clinical prediction models in their intended population and setting. Diagn Progn Res 2022; 6: 24

  46. [54]

    Why do probabilistic clinical models fail to transport between sites

    Lasko TA, Strobl EV, Stead WW. Why do probabilistic clinical models fail to transport between sites. Npj Digit Med 2024; 7: 1–8

  47. [55]

    The Book of Why: The New Science of Cause and Effect

    Pearl J, Mackenzie D. The Book of Why: The New Science of Cause and Effect. 1st ed. USA: Basic Books, Inc., 2018

  48. [56]

    Meta-Transportability of Causal Effects: A Formal Approach

    Bareinboim E, Pearl J. Meta-Transportability of Causal Effects: A Formal Approach. In: Proceedings of the Sixteenth International Conference on Artificial Intelligence and Statistics . PMLR, pp. 135–143

  49. [57]

    There is no such thing as a validated prediction model

    Van Calster B, Steyerberg EW, Wynants L, et al. There is no such thing as a validated prediction model. BMC Med 2023; 21: 70

  50. [58]

    External validation of clinical prediction models using big datasets from e-health records or IPD meta-analysis: opportunities and challenges

    Riley RD, Ensor J, Snell KIE, et al. External validation of clinical prediction models using big datasets from e-health records or IPD meta-analysis: opportunities and challenges. BMJ 2016; 353: i3140

  51. [59]

    A feature selection method based on Shapley values robust to concept shift in regression, http://arxiv.org/abs/2304.14774 (2023, accessed 25 September 2023)

    Sebastián C, González-Guillén CE. A feature selection method based on Shapley values robust to concept shift in regression, http://arxiv.org/abs/2304.14774 (2023, accessed 25 September 2023)

  52. [60]

    Transportability of Trial Results Using Inverse Odds of Sampling Weights

    Westreich D., Edwards J.K., Lesko C.R., et al. Transportability of Trial Results Using Inverse Odds of Sampling Weights. Am J Epidemiol 2017; 186: 1010–1014

  53. [61]

    Robust Estimation of Encouragement Design Intervention Effects Transported Across Sites

    Rudolph KE, Laan MJ. Robust Estimation of Encouragement Design Intervention Effects Transported Across Sites. J R Stat Soc Ser B Stat Methodol 2017; 79: 1509–1525

  54. [62]

    Extending inferences from a randomized trial to a new target population

    Dahabreh IJ, Robertson SE, Steingrimsson JA, et al. Extending inferences from a randomized trial to a new target population. Stat Med 2020; 39: 1999–2014

  55. [63]

    A scoping review of causal methods enabling predictions under hypothetical interventions

    Lin L, Sperrin M, Jenkins DA, et al. A scoping review of causal methods enabling predictions under hypothetical interventions. Diagn Progn Res 2021; 5: 3

  56. [64]

    A Review of Generalizability and Transportability

    Degtiar I, Rose S. A Review of Generalizability and Transportability. Annu Rev Stat Its Appl 2023; 10: 501–524

  57. [65]

    An Overview of Current Methods for Real-World Applications to Generalize or Transport Clinical Trial Findings to Target Populations of Interest

    Ling AY, Montez-Rath ME, Carita P, et al. An Overview of Current Methods for Real-World Applications to Generalize or Transport Clinical Trial Findings to Target Populations of Interest. Epidemiology; 10.1097/EDE.0000000000001633. 24

  58. [66]

    A Survey on Transfer Learning

    Pan SJ, Yang Q. A Survey on Transfer Learning. IEEE Trans Knowl Data Eng 2010; 22: 1345– 1359

  59. [67]

    A survey of transfer learning

    Weiss K, Khoshgoftaar TM, Wang D. A survey of transfer learning. J Big Data 2016; 3: 9

  60. [68]

    Causal effect on a target population: A sensitivity analysis to handle missing covariates

    Colnet B, Josse J, Varoquaux G, et al. Causal effect on a target population: A sensitivity analysis to handle missing covariates. J Causal Inference 2022; 10: 372–414

  61. [69]

    Covariate Balancing Sensitivity Analysis for Extrapolating Randomized Trials across Locations

    Nie X, Imbens G, Wager S. Covariate Balancing Sensitivity Analysis for Extrapolating Randomized Trials across Locations. Epub ahead of print 9 December 2021. DOI: 10.48550/arXiv.2112.04723

  62. [70]

    Federated Adaptive Causal Estimation (FACE) of Target Treatment Effects

    Han L, Hou J, Cho K, et al. Federated Adaptive Causal Estimation (FACE) of Target Treatment Effects. Epub ahead of print 5 October 2023. DOI: 10.48550/arXiv.2112.09313

  63. [71]

    Exploiting Causal Structure for Robust Model Selection in Unsupervised Domain Adaptation

    Kyono T, van der Schaar M. Exploiting Causal Structure for Robust Model Selection in Unsupervised Domain Adaptation. IEEE Trans Artif Intell 2021; 2: 494–507

  64. [72]

    Off-Policy Evaluation and Learning for External Validity under a Covariate Shift

    Uehara M, Kato M, Yasui S. Off-Policy Evaluation and Learning for External Validity under a Covariate Shift. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 49–61

  65. [73]

    A Unified Framework on Generalizability of Clinical Prediction Models

    Wan B, Caffo B, Vedula SS. A Unified Framework on Generalizability of Clinical Prediction Models. Front Artif Intell 2022; 5: 872720

  66. [74]

    Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available

    Hernán MA, Robins JM. Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available. Am J Epidemiol 2016; 183: 758–764. Supplemental table 1: Summary of included 18 papers N o Aut hor s Methodology Description Need target populatio n? Domain knowledge ? Sta...

  67. [75]

    Direct methods - KMM and RuLSIF

  68. [76]

    Using density ratio weighting is not possible to correct concept shift

    Indirect methods - LR with L1 and L2 penalty, and Random Forests Yes, note that outcome data is required only for the indirect methods No The distribution shift is compatible with covariate shift Estimation of density ratio with both parametric and non- parametric models are u...

  69. [77]

    Positivity: this condition necessitates that for every covariate pattern present in the target population, there is a positive probability of encountering the same pattern in the development population. The methodology aids tailoring a prediction model from the development pop...

  70. [78]

    The applicability of the methods may be limited if there is a lack of access to covariate data from the target population

  71. [79]

    The set of covariates necessary to satisfy the conditional independence assumption might be extensive than the set actually used in the model

  72. [80]

    The authors used simulated data to test the performance of the prediction models

    The approach did not address the practical issues such as missing data and measurement error. The authors used simulated data to test the performance of the prediction models. Data was generated using a linear model with heteroscedastic errors and simulated membership for deve...

  73. [81]

    covariate shift Assessing multiple feature selection methods for transportability across several clinical outcomes - Only used MIMIC-IV dataset for evaluation

    Remove and Retrain (ROAR): identified importance predictors with a permutation- based feature importance score and removed, then retrained the model No (requires only historical data) No (except for causal predictor selection) No explicitly mentioned - it is inferred that the ...

  74. [82]

    Clustering the training dataset to mimick real-world heterogeneity and use ICEV for those clusters to perform stepwise predictor selection

  75. [83]

    The choices of candidate predictors should be guided by clinical knowledge

    From stepwise IECV, aggregate performance metrics of all clusters using meta-analytic methods to select the best model that can generalise across heterogenous dataset No 1. The choices of candidate predictors should be guided by clinical knowledge. Heterogeneity of training da...

  76. [84]

    Generalisability may come at a cost of improving one predictive performance (such as discrimination), but deteriorating another one

  77. [85]

    Focusing solely on predictor selection may have an impact on model's precision

  78. [86]

    In their prior work, the author developed site-specific XGBoost models for three healthcare systems (KPSC, IH, UWM) using over 200 candidate features

    It is still possible that the model requires recalibration on the target settings. In their prior work, the author developed site-specific XGBoost models for three healthcare systems (KPSC, IH, UWM) using over 200 candidate features. The top 20 features, each with importance v...

  79. [87]

    The selection diagram is correctly specified

  80. [88]

    The interventinal distribution is identifiable

  81. [89]

    (prior probability shift)

    Predicting variables is mutable. (prior probability shift)

  82. [90]

    Simulating data from DGP by pre- defined DAG and generate different environments by vary mutable variables

  83. [91]

    N o Aut hor s Methodology Description Need target populatio n? Domain knowledge ? Stated assumptions Reported advantages Reported limitations Experiments 1 5 Pearl et al

    Using UCI Bike sharing dataset to predict the number of hourly bike rentals in different seasons by assuming possible DGP as a DAG including mutable variables. N o Aut hor s Methodology Description Need target populatio n? Domain knowledge ? Stated assumptions Reported advanta...

  84. [92]

    The covariate C need to satisfy S- adminissibilty, i.e. y ⫫ x | C - Mainly addresses the causal effect of a single variable X and does not cover prediction scenarios with multiple predictors - Lacks cases where the covariate set C overlaps with the predictor set X, which is co...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.