Pith. sign in

REVIEW 4 major objections 6 minor 98 references

Recent Advances, Applications and Open Challenges in Machine Learning for Health: Reflections from Research Roundtables at ML4H 2024 Symposium

T0 review · 4 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The ML4H 2024 Research Roundtables highlighted diverse challenges and opportunities in applying machine learning to healthcare, with recurring themes of data standardization, deployment gaps, monitoring, and interdisciplinary collaboration.

desk verdict A readable, useful community snapshot of ML4H 2024 roundtables, but the representativeness claim is unverifiable and one empirical assertion in Section 4 is unsupported; treat it as a report, not research. read the letter →

arxiv 2502.06693 v1 pith:E345N7OS submitted 2025-02-10 cs.LG cs.AIcs.CY

classification cs.LGcs.AIcs.CY
keywords machinelearningforhealthML4Hresearchroundtablescommunityperspectivesopenchallengesdatastandardizationclinicalintegrationfairness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This report synthesizes what 13 research roundtables at the ML4H 2024 symposium identified as the most pressing challenges and opportunities in machine learning for health. Across topics from multimodal foundation models to AI in low-resource settings, the authors argue that the community's central problems are data standardization, clinical integration, trust, and equity. If the summaries faithfully capture the discussions, they give a reliable snapshot of where the field currently sees its bottlenecks. The paper is a community-authored reflection, not a study with new empirical results.

What carries the argument

The central object carrying the argument is the research roundtable itself: a structured, in-person discussion session with an invited senior chair, two junior chairs, and a diverse group of attendees. Before the event, chairs drafted an introductory paragraph and up to four discussion questions; after the event, they submitted written summaries that became the paper's primary data. This mechanism converts open-ended conversation into condensed, thematically organized reflections. The paper's credibility depends on the assumption that these written summaries faithfully preserve the discussions, including disagreements and minority views.

What would settle it

A direct check would be to compare the chair-written summaries against audio recordings or a structured survey of all attendees from the same sessions; if an independent survey finds different priorities or records major disagreements that the summaries omit, the representativeness claim would collapse. For instance, collecting anonymous post-session rankings of the topics raised would reveal whether the summaries reflect consensus or the chairs' selection.

Watch

Extended reading notes

Core claim

The central claim is that the ML4H 2024 Research Roundtables highlighted diverse challenges and opportunities in applying machine learning to healthcare, and that these discussions cluster around a few recurring themes. Across 13 tables, the most heavily discussed topic was multimodal foundation models, with attention to incomplete data, temporal alignment, and model specialization. The paper also identifies a persistent gap between model evaluation and clinical deployment, noting that most evaluated models are not deployed while many deployed models are not evaluated. Participants across tables emphasized monitoring, regulation, and the value of 'bilingual' experts who can translate between technical and clinical languages. The paper's main assertion is that these chair summaries represent the key insights and takeaways from the roundtable discussions.

Load-bearing premise

The chairs' written summaries faithfully and completely preserve the roundtable discussions, including minority viewpoints and disagreements.

Editorial extensions

If this is right

  • If accurate, the summaries provide a shared agenda for the ML4H community, highlighting data standardization and deployment gaps as priorities.
  • The recurrence of similar challenges across very different tables, from causality to drug discovery, suggests that structural issues like incentives and infrastructure, not just technical ones, are the real bottlenecks.
  • The call for continuous monitoring and regulation implies that future research should invest in evaluation frameworks and governance, not only in new model architectures.
  • The emphasis on interdisciplinary teams and stakeholder engagement suggests that funding and recognition should shift to support collaborative, translation-oriented work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The convergence on data-standardization challenges, such as the MEDS format discussion, suggests the community could benefit from shared infrastructure and standardized benchmarks, a conclusion the paper only reports as one table's view.
  • The paper's reliance on chair summaries is a potential selection bias; a controlled comparison with recorded discussions or attendee surveys would test whether the reported consensus is representative.
  • The set of roundtable topics chosen by the organizers reveals what the community currently considers timely; tracking topic selection year to year would show how ML4H's priorities evolve.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This manuscript reports on the research roundtables held at the ML4H 2024 symposium. It describes the organizational process, summarizes thirteen roundtable discussions on topics ranging from foundation models and causality to health economics, fairness, and drug discovery, and closes with a synthesis of cross-cutting challenges and lessons learned for future events. The central claim is that the summaries in Sections 3.1–3.13 accurately reflect the roundtable discussions and thereby provide a reliable map of current challenges and opportunities in machine learning for health.

Significance. If taken as a faithful record, the paper is a useful community resource: it condenses a broad set of expert discussions into a compact, well-organized overview, and it documents practical organizational details that could help future ML4H roundtables. The paper is transparent about the event structure, names chairs, and supports many background statements with citations. The paper makes no technical or empirical research claim, so its contribution is descriptive and organizational rather than scientific. Its value therefore depends entirely on the trustworthiness of the summarization process, and that trustworthiness is not established by the manuscript as written. The central claim is defensible, but the evidence behind it needs to be made visible and the over-precise claims need to be qualified.

major comments (4)
  1. [Sections 2 and 5] The central claim that these summaries accurately represent the roundtable discussions rests entirely on chair-written summaries, but the reader is given no access to those summaries and no account of how they were synthesized. Section 2 says only that 'the chairs submitted written summaries that highlight the key insights and takeaways,' and Section 5 confirms that the summaries 'was used to create this document.' There are no transcripts, no attendee counts for most tables, no note-taking protocol, and no statement that participants or chairs reviewed the synthesized text. This is load-bearing because every 'participants noted' claim depends on an unobserved condensation step. The authors should provide the chair summaries as supplementary material or, failing that, describe the synthesis process, state who performed it, and state whether the final text was circulated for verification.
  2. [Sections 3.2, 3.6.1, and 3.6.2] Several quantitative statements are presented with a precision that the summarized-discussion method cannot support. Section 3.2 reports that 'about half of the participants have hands-on experience' with causal inference methods, Section 3.6.1 reports 'estimates suggesting that up to half of the variables might be incorrectly defined or interpreted,' and Section 3.6.2 reports that longitudinal EHR benchmarks are 'used by a maximum of nine studies.' No denominator, measurement instrument, or source is given for any of these numbers, and they read as aggregated statistics rather than reflections of a free-flowing discussion. These claims should either be removed or attributed to specific participants as individual impressions, with explicit qualifications about their anecdotal status.
  3. [Sections 3.4 and 3.10] The manuscript does not consistently separate chair-authored exposition from participant discussion, which makes it impossible for the reader to tell which statements represent group views and which represent the chairs' framing. Section 3.4 contains a tutorial on QALYs/DALYs, NICE, and CDA-AMC, and Section 3.10 contains a 'professional biographies' paragraph; neither is presented as a participant discussion theme. The authors should clearly label tutorial and expository content as such, or move it to a separately marked background section, so that the roundtable discussion summaries are not conflated with the authors' own framing.
  4. [Section 4] The aphorism 'Most evaluated models are not deployed, but many deployed models are not evaluated' is presented as a finding of the roundtables, but it is not attributed to any specific table and no supporting evidence or citation is provided. If the authors wish to retain this sentence, they must attribute it to an identifiable discussion or participant and frame it as an opinion expressed during the sessions rather than as an established result.
minor comments (6)
  1. [Section 5] The sentence 'After the conference, junior and senior chairs provided summaries of discussions, which was used to create this document' has a subject-verb agreement error; it should read 'which were used.'
  2. [Section 3.11.2] The phrase 'if a patient experiences experiences financial misconduct' contains a duplicated word; it should read 'if a patient experiences financial misconduct.'
  3. [Section 3.11.3] The phrase 'the scarsity for long-term ed data' appears to be a typo; it should read 'the scarcity of long-term data.'
  4. [Section 3.6.1] The phrase 'the unbelievable heterogeneity' is informal for a journal-style report; 'substantial heterogeneity' would be more appropriate.
  5. [Section 3.4.3] The text refers to the 'Food and Drug Association'; the correct name is the Food and Drug Administration (FDA).
  6. [Section 3.13] The spelling of 'generalizability' is inconsistent ('generalizabiltiy,' 'generaliazability,' 'generalisability'); the manuscript should use one standard spelling throughout.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a narrative report of roundtable discussions and contains no derivation chain whose conclusions are equivalent to its inputs.

full rationale

The paper is not an empirical or theoretical derivation; it is a structured summary of discussions at ML4H 2024. Section 2 states that 'Following the event, the chairs submitted written summaries that highlight the key insights and takeaways from their discussions,' and Section 5 says these summaries 'was used to create this document.' There is thus a condensation step from live discussions to chair-written summaries to the report, but the report does not claim to derive any quantitative result, fit any parameter, or prove any theorem from assumptions. The central assertion in Section 4 that the roundtables 'highlighted diverse challenges and opportunities' is a descriptive meta-summary, not a prediction or a formal consequence. The prior ML4H roundtable reports cited in Section 1 are contextual references and are not load-bearing for any claim. The limitation that raw discussion records are not exposed makes representativeness difficult to audit from the document alone, and some statements (e.g., 'about half of the participants have hands-on experience' in §3.2 and 'estimates suggesting that up to half of the variables might be incorrectly defined or interpreted' in §3.6.1) lack provenance details; however, unverifiable reporting is an evidential or methodological weakness, not circular reasoning. There is no equation, fitted parameter, or self-citation chain that reduces the output to its input, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The report's reliability rests on unverifiable assumptions about the fidelity and representativeness of the written summaries and of the self-selected participants. There are no free parameters or invented entities because the paper makes no quantitative or mechanistic claims.

assumptions (2)
  • domain assumption Chairs' written summaries faithfully represent the roundtable discussions, including the range of participant views and dissenting opinions.
    Section 2 states: 'Following the event, the chairs submitted written summaries that highlight the key insights and takeaways from their discussions.' The accuracy of the entire report depends on this unverified fidelity.
  • domain assumption The self-selected and organizer-invited roundtable participants are representative of the broader ML4H community's challenges and priorities.
    Section 2 describes topic selection by organizers and participant recruitment without systematic sampling. The report generalizes from these participants to community-wide reflections, which assumes representativeness.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Recent Advances, Applications and Open Challenges in Machine Learning for Health: Reflections from Research Roundtables at ML4H 2024 Symposium." pith.science (2026). https://pith.science/paper/E345N7OS

@misc{pith2026250206693,
  author       = {Pith},
  title        = {Pith review of: Recent Advances, Applications and Open Challenges in Machine Learning for Health: Reflections from Research Roundtables at ML4H 2024 Symposium},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E345N7OS}},
  note         = {Machine review of arXiv:2502.06693}
}
read the original abstract

The fourth Machine Learning for Health (ML4H) symposium was held in person on December 15th and 16th, 2024, in the traditional, ancestral, and unceded territories of the Musqueam, Squamish, and Tsleil-Waututh Nations in Vancouver, British Columbia, Canada. The symposium included research roundtable sessions to foster discussions between participants and senior researchers on timely and relevant topics for the ML4H community. The organization of the research roundtables at the conference involved 13 senior and 27 junior chairs across 13 tables. Each roundtable session included an invited senior chair (with substantial experience in the field), junior chairs (responsible for facilitating the discussion), and attendees from diverse backgrounds with an interest in the session's topic.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

98 extracted references · 51 canonical work pages

  1. [1]

    Amin Adibi, Mohsen Sadatsafavi, and John P. A. Ioannidis. Validation and utility testing of clinical prediction models: Time to change the approach. JAMA, 324 0 (3): 0 235--236, 07 2020. ISSN 0098-7484. doi:10.1001/jama.2020.1230. URL https://doi.org/10.1001/jama.2020.1230

  2. [2]

    Topo-cxr: Chest x-ray tb and pneumonia screening with topological machine learning

    Faisal Ahmed, Brighton Nuwagira, Furkan Torlak, and Baris Coskunuzer. Topo-cxr: Chest x-ray tb and pneumonia screening with topological machine learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2326--2336, 2023

  3. [3]

    Ayers, Adam Poliak, Mark Dredze, Eric C

    John W. Ayers, Adam Poliak, Mark Dredze, Eric C. Leas, Zechariah Zhu, Jessica B. Kelley, Dennis J. Faix, Aaron M. Goodman, Christopher A. Longhurst, Michael Hogarth, and Davey M. Smith. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum . JAMA Internal Medicine, 183 0 (6): 0 589--59...

  4. [4]

    Elena, Dávid P

    Ilyes Batatia, Philipp Benner, Yuan Chiang, Alin M. Elena, Dávid P. Kovács, Janosh Riebesell, Xavier R. Advincula, Mark Asta, Matthew Avaylon, William J. Baldwin, Fabian Berger, Noam Bernstein, Arghya Bhowmik, Samuel M. Blau, Vlad Cărare, James P. Darby, Sandip De, Flaviano Della Pia, Volker L. Deringer, Rokas Elijošius, Zakariya El-Machachi, Fabio Falcio...

  5. [5]

    Beecy, Christopher A

    Ashley N. Beecy, Christopher A. Longhurst, Karandeep Singh, Robert M. Wachter, and Sara G. Murray. The Chief Health AI Officer — An Emerging Role for an Emerging Technology . NEJM AI, 1 0 (7): 0 AIp2400109, June 2024. doi:10.1056/AIp2400109. URL https://ai.nejm.org/doi/abs/10.1056/AIp2400109. Publisher: Massachusetts Medical Society

  6. [6]

    The regulation of clinical artificial intelligence

    David Blumenthal and Bakul Patel. The regulation of clinical artificial intelligence. NEJM AI, 1 0 (8): 0 AIpc2400545, 2024. doi:10.1056/AIpc2400545. URL https://ai.nejm.org/doi/full/10.1056/AIpc2400545

  7. [7]

    On the opportunities and risks of foundation models

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021

  8. [8]

    Guidelines for the economic evaluation of health technologies: Canada—4th edition

    CADTH. Guidelines for the economic evaluation of health technologies: Canada—4th edition. Technical report, CADTH Methods and Guidelines, 2017. URL https://www.cda-amc.ca/guidelines-economic-evaluation-health-technologies-canada-4th-edition

Show all 98 references
  1. [9]

    Mpoxvlm: A vision-language model for diagnosing skin lesions from mpox virus infection

    Xu Cao, Wenqian Ye, Kenny Moise, and Megan Coffee. Mpoxvlm: A vision-language model for diagnosing skin lesions from mpox virus infection. arXiv preprint arXiv:2411.10888, 2024

  2. [10]

    Revolution, interrupted: Why AI has failed to live up to the hype in drug development , 2024

    Joe Castaldo and Sean Silcoff. Revolution, interrupted: Why AI has failed to live up to the hype in drug development , 2024. URL https://www.theglobeandmail.com/business/article-artificial-intelligence-drug-research-hype/

  3. [11]

    Statistics canada: Canadian community health survey, 2024

    CCHS. Statistics canada: Canadian community health survey, 2024. URL https://www23.statcan.gc.ca/imdb/p2SV.pl?Function=getSurvey&SDDS=3226. (accessed Dec 21, 2024)

  4. [12]

    Assessing the usability of gutgpt: A simulation study of an ai clinical decision support system for gastrointestinal bleeding risk, 2023

    Colleen Chan, Kisung You, Sunny Chung, Mauro Giuffrè, Theo Saarinen, Niroop Rajashekar, Yuan Pu, Yeo Eun Shin, Loren Laine, Ambrose Wong, René Kizilcec, Jasjeet Sekhon, and Dennis Shung. Assessing the usability of gutgpt: A simulation study of an ai clinical decision support s...

  5. [13]

    Alfonso, Marisa Cobanaj, Amelia Morel Fiske, Alexander J

    Marie-Laure Charpignon, Joao Matos, Luis Filipe Nakayama, Jack Gallifant, Pia Gabrielle I. Alfonso, Marisa Cobanaj, Amelia Morel Fiske, Alexander J. Gates, Frances Dominique V. Ho, Urvish Jain, Mohammad Kashkooli, Naira Link, Liam G. McCoy, Jonathan Shaffer, and Leo Anthony Ce...

  6. [14]

    xTrimoPGLM: Unified 100B-Scale Pre-trained Transformer for Deciphering the Language of Protein

    Bo Chen, Xingyi Cheng, Pan Li, Yangli-ao Geng, Jing Gong, Shen Li, Zhilei Bei, Xu Tan, Boyan Wang, Xin Zeng, Chiming Liu, Aohan Zeng, Yuxiao Dong, Jie Tang, and Le Song. xTrimoPGLM: Unified 100B-Scale Pre-trained Transformer for Deciphering the Language of Protein . arXiv, pag...

  7. [15]

    AI will boost drug development in 2025: Software-inspired medicines are getting closer to prime time

    Shailesh Chitnis. AI will boost drug development in 2025: Software-inspired medicines are getting closer to prime time . The Economist, nov 2024. URL https://www.economist.com/the-world-ahead/2024/11/20/ai-will-boost-drug-development-in-2025

  8. [16]

    The frontier of simulation-based inference

    Kyle Cranmer, Johann Brehmer, and Gilles Louppe. The frontier of simulation-based inference . Proceedings of the National Academy of Sciences of the United States of America, 117 0 (48): 0 30055--30062, 2020. ISSN 10916490. doi:10.1073/pnas.1912789117

  9. [17]

    The Atlas of AI : Power , Politics , and the Planetary Costs of Artificial Intelligence

    Kate Crawford. The Atlas of AI : Power , Politics , and the Planetary Costs of Artificial Intelligence . Yale University Press, April 2021. ISBN 978-0-300-20957-0. Google-Books-ID: KfodEAAAQBAJ

  10. [18]

    Using machine learning to individualize treatment effect estimation: Challenges and opportunities

    Alicia Curth, Richard W Peck, Eoin McKinney, James Weatherall, and Mihaela van Der Schaar. Using machine learning to individualize treatment effect estimation: Challenges and opportunities. Clinical Pharmacology & Therapeutics, 115 0 (4): 0 710--719, 2024

  11. [19]

    Calibration drift in regression and machine learning models for acute kidney injury

    Sharon E Davis, Thomas A Lasko, Guanhua Chen, Edward D Siew, and Michael E Matheny. Calibration drift in regression and machine learning models for acute kidney injury. Journal of the American Medical Informatics Association, 24 0 (6): 0 1052--1061, 2017

  12. [20]

    Explainable framework for glaucoma diagnosis by image processing and convolutional neural network synergy: Analysis with doctor evaluation

    Omer Deperlioglu, Utku Kose, Deepak Gupta, Ashish Khanna, Fabio Giampaolo, and Giancarlo Fortino. Explainable framework for glaucoma diagnosis by image processing and convolutional neural network synergy: Analysis with doctor evaluation. Future Generation Computer Systems, 129...

  13. [21]

    Conceptualizing Epistemic Oppression

    Kristie Dotson. Conceptualizing Epistemic Oppression . Social Epistemology, 28 0 (2): 0 115--138, April 2014. ISSN 0269-1728. doi:10.1080/02691728.2013.782585. URL https://doi.org/10.1080/02691728.2013.782585. Publisher: Routledge

  14. [22]

    Causal machine learning for predicting treatment outcomes

    Stefan Feuerriegel, Dennis Frauen, Valentyn Melnychuk, Jonas Schweisthal, Konstantin Hess, Alicia Curth, Stefan Bauer, Niki Kilbertus, Isaac S Kohane, and Mihaela van der Schaar. Causal machine learning for predicting treatment outcomes. Nature Medicine, 30 0 (4): 0 958--968, 2024

  15. [23]

    Epistemic Injustice : Power and the Ethics of Knowing

    Miranda Fricker. Epistemic Injustice : Power and the Ethics of Knowing . Oxford University Press, Incorporated, Oxford, UNITED KINGDOM, 2007. ISBN 978-0-19-151930-7. URL http://ebookcentral.proquest.com/lib/ubc/detail.action?docID=694009

  16. [24]

    Bayesian Optimization

    Roman Garnett. Bayesian Optimization . Cambridge University Press, 2022

  17. [25]

    fishing expedition

    Andrew Gelman and Eric Loken. The garden of forking paths: Why multiple comparisons can be a problem, even when there is no “fishing expedition” or “p-hacking” and the research hypothesis was posited ahead of time, 2013. URL http://www.stat.columbia.edu/ gelman/research/unpubl...

  18. [26]

    Interactions in epidemiology: relevance, identification, and estimation

    Sander Greenland. Interactions in epidemiology: relevance, identification, and estimation. Epidemiology, 20 0 (1): 0 14--17, 2009

  19. [27]

    A stock approach to the demand for health

    Michael Grossman. A stock approach to the demand for health. In The Demand for Health: A Theoretical and Empirical Investigation, pages 1--10. NBER, 1972

  20. [28]

    Generalizability of cardiovascular disease clinical prediction models: 158 independent external validations of 104 unique models

    Gaurav Gulati, Jenica Upshaw, Benjamin S Wessler, Riley J Brazil, Jason Nelson, David Van Klaveren, Christine M Lundquist, Jinny G Park, Hannah McGinnes, Ewout W Steyerberg, et al. Generalizability of cardiovascular disease clinical prediction models: 158 independent external ...

  21. [29]

    CZII - CryoET Object Identification , 2024

    Kyle; Harrington, Mohammadreza; Paraan, Anchi; Cheng, Utz Heinrich; Ermel, Saugat; Kandel, Dari; Kimanius, Elizabeth; Montabana, Ariana; Peck, Jonathan; Schwartz, Daniel; Serwas, Hannah; Siems, Feng; Wang, Yue; Yu, Zhuowen; Zhao, Shawn; Zheng, Walter Reade, Maggie; Demkin, Kri...

  22. [30]

    Value judgments in a COVID -19 vaccination model: A case study in the need for public involvement in health-oriented modelling

    Stephanie Harvard, Eric Winsberg, John Symons, and Amin Adibi. Value judgments in a COVID -19 vaccination model: A case study in the need for public involvement in health-oriented modelling. Social Science & Medicine, 286: 0 114323, October 2021. ISSN 0277-9536. doi:10.1016/j....

  23. [31]

    Deeperbind: Enhancing prediction of sequence specificities of dna binding proteins

    Hamid Reza Hassanzadeh and May D Wang. Deeperbind: Enhancing prediction of sequence specificities of dna binding proteins. In 2016 IEEE International conference on bioinformatics and biomedicine (BIBM), pages 178--183. IEEE, 2016

  24. [32]

    Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q

    Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J. Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q. Tran, Jonathan Deaton, Marius Wiggert, Rohil Badkundri, Irhum Shafkat, Jun Gong, Alexander Derry, Raul S. Molina, Neil Thomas, Yousuf Khan, Chetan Mishra, Carolyn K...

  25. [33]

    Stefan Hegselmann, Helen Zhou, Yuyin Zhou, Jennifer Chien, Sujay Nagaraj, Neha Hulkund, Shreyas Bhave, Michael Oberst, Amruta Pai, Caleb Ellington, Wisdom Ikezogwo, Jason Xiaotian Dou, Monica Agrawal, Changye Li, Peniel Argaw, Arpita Biswas, Mehak Gupta, Xinhui Li, Marta Leman...

  26. [34]

    Applied artificial intelligence and trust—the case of autonomous vehicles and medical assistance devices

    Monika Hengstler, Ellen Enkel, and Selina Duelli. Applied artificial intelligence and trust—the case of autonomous vehicles and medical assistance devices. Technological Forecasting and Social Change, 105: 0 105--120, 2016. ISSN 0040-1625. doi:https://doi.org/10.1016/j.techfor...

  27. [35]

    Osborne, and Hans P

    Philipp Hennig, Michael A. Osborne, and Hans P. Kersting. Probabilistic Numerics: Computation as Machine Learning , volume 59. Cambridge University Press, jun 2022. ISBN 9781316681411. doi:10.1017/9781316681411. URL https://www.cambridge.org/core/product/identifier/97813166814...

  28. [36]

    Causability and explainability of artificial intelligence in medicine

    Andreas Holzinger, Georg Langs, Helmut Denk, Kurt Zatloukal, and Heimo M \"u ller. Causability and explainability of artificial intelligence in medicine. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 9 0 (4): 0 e1312, 2019

  29. [37]

    Bryant Howren

    M. Bryant Howren. Quality-Adjusted Life Years (QALYs), pages 1605--1606. Springer New York, New York, NY, 2013. ISBN 978-1-4419-1005-9. doi:10.1007/978-1-4419-1005-9_613. URL https://doi.org/10.1007/978-1-4419-1005-9_613

  30. [38]

    Canada denied visas to dozens of Africans for a big artificial intelligence conference

    Matthew Hutson. Canada denied visas to dozens of Africans for a big artificial intelligence conference. Science, December 2018. ISSN 0036-8075, 1095-9203. doi:10.1126/science.aaw3628. URL doi.org/10.1126/science.aaw3628

  31. [39]

    Ingraham, Max Baranov, Zak Costello, Karl W

    John B. Ingraham, Max Baranov, Zak Costello, Karl W. Barber, Wujie Wang, Ahmed Ismail, Vincent Frappier, Dana M. Lord, Christopher Ng-Thow-Hing, Erik R. Van Vlack , Shan Tie, Vincent Xue, Sarah C. Cowles, Alan Leung, Jo \ a o V. Rodrigues, Claudio L. Morales-Perez, Alex M. Ayo...

  32. [40]

    Random survival forests

    Hemant Ishwaran, Udaya B Kogalur, Eugene H Blackstone, and Michael S Lauer. Random survival forests. The Annals of Applied Statistics, 2 0 (3): 0 841--60, 2008

  33. [41]

    Valley, Ella A

    Sarah Jabbour, David Fouhey, Stephanie Shepard, Thomas S. Valley, Ella A. Kazerooni, Nikola Banovic, Jenna Wiens, and Michael W. Sjoding. Measuring the impact of ai in the diagnosis of hospitalized patients: A randomized clinical vignette survey study. JAMA, 330 0 (23): 0 2275...

  34. [42]

    Recent advances, applications, and open challenges in machine learning for health: Reflections from research roundtables at ml4h 2023 symposium

    Hyewon Jeong, Sarah Jabbour, Yuzhe Yang, Rahul Thapta, Hussein Mozannar, William Jongwon Han, Nikita Mehandru, Michael Wornow, Vladislav Lialin, Xin Liu, et al. Recent advances, applications, and open challenges in machine learning for health: Reflections from research roundta...

  35. [43]

    Conceptualizing machine learning for dynamic information retrieval of electronic health record notes

    Sharon Jiang, Shannon Shen, Monica Agrawal, Barbara Lam, Nicholas Kurtzman, Steven Horng, David R Karger, and David Sontag. Conceptualizing machine learning for dynamic information retrieval of electronic health record notes. In Machine Learning for Healthcare Conference, page...

  36. [44]

    Towards an artificial intelligence framework for data-driven prediction of coronavirus clinical severity

    Xiangao Jiang, Megan Coffee, Anasse Bari, Junzhang Wang, Xinyue Jiang, Jianping Huang, Jichan Shi, Jianyi Dai, Jing Cai, Tianxiao Zhang, et al. Towards an artificial intelligence framework for data-driven prediction of coronavirus clinical severity. Computers, Materials & Cont...

  37. [45]

    Time-dependent roc curve analysis in medical research: current methods and applications

    Adina Najwa Kamarudin, Trevor Cox, and Ruwanthi Kolamunnage-Dona. Time-dependent roc curve analysis in medical research: current methods and applications. BMC Medical Research Methodology, 17: 0 1--19, 2017

  38. [46]

    Evaluating the impact of prediction models: lessons learned, challenges, and recommendations

    Teus H Kappen, Wilton A van Klei, Leo van Wolfswinkel, Cor J Kalkman, Yvonne Vergouwe, and Karel GM Moons. Evaluating the impact of prediction models: lessons learned, challenges, and recommendations. Diagnostic and prognostic research, 2 0 (1): 0 11, 2018

  39. [47]

    Deepsurv: personalized treatment recommender system using a cox proportional hazards deep neural network

    Jared L Katzman, Uri Shaham, Alexander Cloninger, Jonathan Bates, Tingting Jiang, and Yuval Kluger. Deepsurv: personalized treatment recommender system using a cox proportional hazards deep neural network. BMC Medical Research Methodology, 18: 0 1--12, 2018

  40. [48]

    Epistemic Injustice in Generative AI

    Jackie Kay, Atoosa Kasirzadeh, and Shakir Mohamed. Epistemic Injustice in Generative AI . Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 7: 0 684--697, October 2024. ISSN 3065-8365. doi:10.1609/aies.v7i1.31671. URL https://ojs.aaai.org/index.php/AIES/articl...

  41. [49]

    Key challenges for delivering clinical impact with artificial intelligence

    Christopher J Kelly, Alan Karthikesalingam, Mustafa Suleyman, Greg Corrado, and Dominic King. Key challenges for delivering clinical impact with artificial intelligence. BMC medicine, 17: 0 1--9, 2019

  42. [50]

    D. M. Kent, D. van Klaveren, J. K. Paulus, R. D'Agostino, S. Goodman, R. Hayward, J. P. A. Ioannidis, B. Patrick-Lake, S. Morton, M. Pencina, G. Raman, J. S. Ross, H. P. Selker, R. Varadhan, A. Vickers, J. B. Wong, and E. W. Steyerberg. The predictive approaches to treatment e...

  43. [51]

    Foundations of machine learning-based clinical prediction modeling: Part ii—generalization and overfitting

    Julius M Kernbach and Victor E Staartjes. Foundations of machine learning-based clinical prediction modeling: Part ii—generalization and overfitting. Machine Learning in Clinical Neuroscience: Foundations and Applications, pages 15--21, 2022

  44. [52]

    Patient information summarization in clinical settings: scoping review

    Daniel Keszthelyi, Christophe Gaudet-Blavignac, Mina Bjelogrlic, Christian Lovis, et al. Patient information summarization in clinical settings: scoping review. JMIR Medical Informatics, 11 0 (1): 0 e44639, 2023

  45. [53]

    S-synth: Knowledge-based, synthetic generation of skin images

    Andrea Kim, Niloufar Saharkhiz, Elena Sizikova, Miguel Lago, Berkman Sahiner, Jana Delfino, and Aldo Badano. S-synth: Knowledge-based, synthetic generation of skin images. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 734--744...

  46. [54]

    Benjamin Kompa, Jasper Snoek, and Andrew L. Beam. Second opinion needed: communicating uncertainty in medical machine learning. npj Digital Medicine, 4 0 (1): 0 1--6, January 2021. ISSN 2398-6352. doi:10.1038/s41746-020-00367-3. URL https://www.nature.com/articles/s41746-020-00367-3

  47. [55]

    Chembo: Bayesian optimization of small organic molecules with synthesizable recommendations

    Ksenia Korovina, Sailun Xu, Kirthevasan Kandasamy, Willie Neiswanger, Barnabas Poczos, Jeff Schneider, and Eric Xing. Chembo: Bayesian optimization of small organic molecules with synthesizable recommendations. In Silvia Chiappa and Roberto Calandra, editors, Proceedings of th...

  48. [56]

    How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons

    Madura KP Jayatunga , Margaret Ayers, Lotte Bruens, Dhruv Jayanth, and Christoph Meier. How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons . Drug Discovery Today, 29 0 (6): 0 104009, 2024. ISSN 18785832. doi:10.1016/j.drudis.2024.1...

  49. [57]

    Artificial intelligence in clinical medicine: catalyzing a sustainable global healthcare paradigm

    Gokul Krishnan, Shiana Singh, Monika Pathania, Siddharth Gosavi, Shuchi Abhishek, Ashwin Parchani, and Minakshi Dhar. Artificial intelligence in clinical medicine: catalyzing a sustainable global healthcare paradigm. Frontiers in Artificial Intelligence, 6: 0 1227091, 2023

  50. [58]

    Confounding by indication in clinical research

    Demetrios N Kyriacou and Roger J Lewis. Confounding by indication in clinical research. Jama, 316 0 (17): 0 1818--1819, 2016

  51. [59]

    McMahon, Jakob Macke, Kyle Cranmer, Jiaxin Zhang, Haruko Wainwright, Adi Hanuka, Manuela Veloso, Samuel Assefa, Stephan Zheng, and Avi Pfeffer

    Alexander Lavin, David Krakauer, Hector Zenil, Justin Gottschlich, Tim Mattson, Johann Brehmer, Anima Anandkumar, Sanjay Choudry, Kamil Rocki, Atılım Güneş Baydin, Carina Prunkl, Brooks Paige, Olexandr Isayev, Erik Peterson, Peter L. McMahon, Jakob Macke, Kyle Cranmer, Jiaxin ...

  52. [60]

    Deephit: A deep learning approach to survival analysis with competing risks

    Changhee Lee, William Zame, Jinsung Yoon, and Mihaela van der Schaar. Deephit: A deep learning approach to survival analysis with competing risks. Proceedings of the AAAI Conference on Artificial Intelligence, 32 0 (1), Apr. 2018. doi:10.1609/aaai.v32i1.11842. URL https://ojs....

  53. [61]

    Thirsty

    Pengfei Li, Jianyi Yang, Mohammad A. Islam, and Shaolei Ren. Making AI Less " Thirsty ": Uncovering and Addressing the Secret Water Footprint of AI Models , October 2023. URL http://arxiv.org/abs/2304.03271. arXiv:2304.03271 [cs]

  54. [62]

    Lu, Bowen Chen, Drew F

    Ming Y. Lu, Bowen Chen, Drew F. K. Williamson, Richard J. Chen, Melissa Zhao, Aaron K. Chow, Kenji Ikemura, Ahrong Kim, Dimitra Pouli, Ankush Patel, Amr Soliman, Chengkuan Chen, Tong Ding, Judy J. Wang, Georg Gerber, Ivy Liang, Long Phi Le, Anil V. Parwani, Luca L. Weishaupt, ...

  55. [63]

    Sasha Luccioni, Yacine Jernite, and Emma Strubell. Power Hungry Processing : Watts Driving the Cost of AI Deployment ? In Proceedings of the 2024 ACM Conference on Fairness , Accountability , and Transparency , FAccT '24, pages 85--99, New York, NY, USA, June 2024. Association...

  56. [64]

    Bran , Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D

    Andres M. Bran , Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D. White, and Philippe Schwaller. Augmenting large language models with chemistry tools . Nature Machine Intelligence, 6 0 (5): 0 525--535, may 2024. ISSN 2522-5839. doi:10.1038/s42256-024-00832-8. URL https:/...

  57. [65]

    Medgan: optimized generative adversarial network with graph convolutional networks for novel molecule design

    Bruno Macedo, In \^e s Ribeiro Vaz, and Tiago Taveira Gomes. Medgan: optimized generative adversarial network with graph convolutional networks for novel molecule design. Scientific Reports, 14 0 (1): 0 1212, 2024

  58. [66]

    Artificial intelligence in low-and middle-income countries: innovating global health radiology

    Daniel J Mollura, Melissa P Culp, Erica Pollack, Gillian Battino, John R Scheel, Victoria L Mango, Ameena Elahi, Alan Schweitzer, and Farouk Dako. Artificial intelligence in low-and middle-income countries: innovating global health radiology. Radiology, 297 0 (3): 0 513--520, 2020

  59. [67]

    National center for health statistics: National health and nutrition examination survey, 2024

    NHANES. National center for health statistics: National health and nutrition examination survey, 2024. URL https://www.cdc.gov/nchs/nhanes/. (accessed Dec 21, 2024)

  60. [68]

    Nice health technology evaluations: the manual

    NICE. Nice health technology evaluations: the manual. Technical report, National Institute for Health and Care Excellence - Process and Methods, 2022. URL https://www.nice.org.uk/process/pmg36/chapter/introduction-to-health-technology-evaluation

  61. [69]

    Korea national health and nutrition examination survey, 20th anniversary: accomplishments and future directions

    Kyungwon Oh, Yoonjung Kim, Sanghui Kweon, Soyeon Kim, Sungha Yun, Suyeon Park, Yeon-Kyeng Lee, Youngtaek Kim, Ok Park, and Eun Kyeong Jeong. Korea national health and nutrition examination survey, 20th anniversary: accomplishments and future directions. Epidemiology and Health...

  62. [70]

    ‘ Exhausted and insulted’: how harsh visa-application policies are hobbling global research

    Sandra Owusu-Gyamfi. ‘ Exhausted and insulted’: how harsh visa-application policies are hobbling global research. Nature, 627 0 (8005): 0 705--705, March 2024. doi:10.1038/d41586-024-00892-1. URL https://www.nature.com/articles/d41586-024-00892-1. Bandiera\_abtest: a Cg\_type:...

  63. [71]

    Machine learning methods in health economics and outcomes research—the palisade checklist: a good practices report of an ispor task force

    William V Padula, Noemi Kreif, David J Vanness, Blythe Adamson, Juan-David Rueda, Federico Felizzi, Pall Jonsson, Maarten J IJzerman, Atul Butte, and William Crown. Machine learning methods in health economics and outcomes research—the palisade checklist: a good practices repo...

  64. [72]

    The crucial role of interdisciplinary conferences in advancing explainable ai in healthcare

    Ankush U Patel, Qiangqiang Gu, Ronda Esper, Danielle Maeser, and Nicole Maeser. The crucial role of interdisciplinary conferences in advancing explainable ai in healthcare. BioMedInformatics, 4 0 (2): 0 1363--1383, 2024

  65. [73]

    Phillippo, Sofia Dias, A

    David M. Phillippo, Sofia Dias, A. E. Ades, Mark Belger, Alan Brnabic, Alexander Schacht, Daniel Saure, Zbigniew Kadziola, and Nicky J. Welton. Multilevel network meta-regression for population-adjusted treatment comparisons. Journal of the Royal Statistical Society Series A: ...

  66. [74]

    Causal inference and counterfactual prediction in machine learning for actionable healthcare

    Mattia Prosperi, Yi Guo, Matt Sperrin, James S Koopman, Jae S Min, Xing He, Shannan Rich, Mo Wang, Iain E Buchan, and Jiang Bian. Causal inference and counterfactual prediction in machine learning for actionable healthcare. Nature Machine Intelligence, 2 0 (7): 0 369--375, 2020

  67. [75]

    Kizilcec, Loren Laine, Terika Mccall, and Dennis Shung

    Niroop Channa Rajashekar, Yeo Eun Shin, Yuan Pu, Sunny Chung, Kisung You, Mauro Giuffre, Colleen E Chan, Theo Saarinen, Allen Hsiao, Jasjeet Sekhon, Ambrose H Wong, Leigh V Evans, Rene F. Kizilcec, Loren Laine, Terika Mccall, and Dennis Shung. Human-algorithmic interaction usi...

  68. [76]

    Scientific Objectivity

    Julian Reiss and Jan Sprenger. Scientific Objectivity . In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy . Metaphysics Research Lab, Stanford University, winter 2020 edition, 2020. URL https://plato.stanford.edu/archives/win2020/entries/scientific-objectivity/

  69. [77]

    Deep learning in drug design: protein-ligand binding affinity prediction

    Mohammad A Rezaei, Yanjun Li, Dapeng Wu, Xiaolin Li, and Chenglong Li. Deep learning in drug design: protein-ligand binding affinity prediction. IEEE/ACM transactions on computational biology and bioinformatics, 19 0 (1): 0 407--417, 2020

  70. [78]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684--10695, 2022

  71. [79]

    Marginal versus conditional odds ratios when updating risk prediction models

    Mohsen Sadatsafavi, Hamid Tavakoli, and Abdollah Safari. Marginal versus conditional odds ratios when updating risk prediction models. Epidemiology, 33 0 (4): 0 555--558, 2022

  72. [80]

    A practical guide for building collaborations between clinical researchers and engineers: lessons learned from a multidisciplinary patient safety project

    Roshun R Sankaran, Jessica M Ameling, Amy EM Cohn, Cyril M Grum, and Jennifer Meddings. A practical guide for building collaborations between clinical researchers and engineers: lessons learned from a multidisciplinary patient safety project. Journal of patient safety, 17 0 (8...

  73. [81]

    Regularization paths for cox’s proportional hazards model via coordinate descent

    Noah Simon, Jerome Friedman, Trevor Hastie, and Rob Tibshirani. Regularization paths for cox’s proportional hazards model via coordinate descent. Journal of Statistical Software, 39 0 (5): 0 1, 2011

  74. [82]

    Preventing failures due to dataset shift: Learning predictive models that transport, 2019

    Adarsh Subbaswamy, Peter Schulam, and Suchi Saria. Preventing failures due to dataset shift: Learning predictive models that transport, 2019. URL https://arxiv.org/abs/1812.04597

  75. [83]

    The virtual lab: Ai agents design new sars-cov-2 nanobodies with experimental validation

    Kyle Swanson, Wesley Wu, Nash L Bulaong, John E Pak, and James Zou. The virtual lab: Ai agents design new sars-cov-2 nanobodies with experimental validation. bioRxiv, pages 2024--11, 2024

  76. [84]

    Seiji Takeda, Akihiro Kishimoto, Lisa Hamada, Daiju Nakano, and John R. Smith. Foundation Model for Material Science . Proceedings of the AAAI Conference on Artificial Intelligence, 37 0 (13): 0 15376--15383, jun 2023. ISSN 2374-3468. doi:10.1609/aaai.v37i13.26793. URL https:/...

  77. [85]

    Scientifc Background to the Nobel Prize in Chemistry 2024: Computational protein design and protein structure prediction

    The Nobel Committee for Chemistry . Scientifc Background to the Nobel Prize in Chemistry 2024: Computational protein design and protein structure prediction . Technical report, The Royal Swedish Academy of Sciences, 2024. URL https://www.nobelprize.org/uploads/2024/10/advanced...

  78. [86]

    Canadian residents face the longest waits in the world for U

    Elizabeth Thompson. Canadian residents face the longest waits in the world for U . S . visas. CBC News, August 2024. URL https://www.cbc.ca/news/politics/u-s-visas-canadian-residents-delays-1.7301250

  79. [87]

    Tierney, Gregg Gayre, Brian Hoberman, Britt Mattern, Manuel Ballesca, Patricia Kipnis, Vincent Liu, and Kristine Lee

    Aaron A. Tierney, Gregg Gayre, Brian Hoberman, Britt Mattern, Manuel Ballesca, Patricia Kipnis, Vincent Liu, and Kristine Lee. Ambient artificial intelligence scribes to alleviate the burden of clinical documentation. NEJM Catalyst, 5 0 (3): 0 CAT.23.0404, 2024. doi:10.1056/CA...

  80. [88]

    Memorization without overfitting: Analyzing the training dynamics of large language models

    Kushal Tirumala, Aram Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. Memorization without overfitting: Analyzing the training dynamics of large language models. Advances in Neural Information Processing Systems, 35: 0 38274--38290, 2022

  81. [89]

    Food and Drug Administration

    U.S. Food and Drug Administration . Using Artificial Intelligence & Machine Learning in the Development of Drug and Biological Products: Discussion Paper and Request for Feedback . Technical report, U.S. Food and Drug Administration , 2023. URL https://www.fda.gov/media/167973...

  82. [90]

    Macaskill Petra McLernon David J

    Ben Van Calster, David J McLernon, Maarten Van Smeden, Laure Wynants, Ewout W Steyerberg, Topic Group ‘Evaluating diagnostic tests, and prediction models’ of the STRATOS initiative Bossuyt Patrick Collins Gary S. Macaskill Petra McLernon David J. Moons Karel GM Steyerberg Ewou...

  83. [91]

    Economic evaluations of artificial intelligence-based healthcare interventions: a systematic literature review of best practices in their conduct and reporting

    Jai Vithlani, Claire Hawksworth, Jamie Elvidge, Lynda Ayiku, and Dalia Dawoud. Economic evaluations of artificial intelligence-based healthcare interventions: a systematic literature review of best practices in their conduct and reporting. Frontiers in Pharmacology, 14: 0 1220...

  84. [92]

    A Comprehensive Guide to Simulation-based Inference in Computational Biology , 2024

    Xiaoyu Wang, Ryan P Kelly, Adrianne L Jenner, and J David. A Comprehensive Guide to Simulation-based Inference in Computational Biology , 2024

  85. [93]

    Weishaupt, T

    LL. Weishaupt, T. Wang, J. Schamroth, P. Morandini, J. Matos, LM Hampton, J. Gallifant, A. Fiske, N. Dundas, K. David, LA. Celi, A. Carrel, J. Byers, and G. Angelotti. Care phenotypes in critical care. medRxiv, 2025. doi:10.1101/2025.01.24.25320468. URL https://www.medrxiv.org...

  86. [94]

    Boltz-1 Democratizing Biomolecular Interaction Modeling , nov 2024

    Jeremy Wohlwend, Gabriele Corso, Saro Passaro, Mateo Reveiz, Ken Leidal, Wojtek Swiderski, Tally Portnoi, Itamar Chinn, Jacob Silterra, Tommi Jaakkola, and Regina Barzilay. Boltz-1 Democratizing Biomolecular Interaction Modeling , nov 2024. URL http://biorxiv.org/lookup/doi/10...

  87. [95]

    Evaluating treatment benefit predictors using observational data: Contending with identification and confounding bias, 2024

    Yuan Xia, Mohsen Sadatsafavi, and Paul Gustafson. Evaluating treatment benefit predictors using observational data: Contending with identification and confounding bias, 2024. URL https://arxiv.org/abs/2407.05585

  88. [96]

    Unremarkable ai: Fitting intelligent decision support into critical, clinical decision-making processes

    Qian Yang, Aaron Steinfeld, and John Zimmerman. Unremarkable ai: Fitting intelligent decision support into critical, clinical decision-making processes. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, CHI '19, page 1–11, New York, NY, USA, 2019...

  89. [97]

    Spurious correlations in machine learning: A survey

    Wenqian Ye, Guangtao Zheng, Xu Cao, Yunsheng Ma, and Aidong Zhang. Spurious correlations in machine learning: A survey. arXiv preprint arXiv:2402.12715, 2024

  90. [98]

    Zaj a c, Dana Li, Xiang Dai, Jonathan F

    Hubert D. Zaj a c, Dana Li, Xiang Dai, Jonathan F. Carlsen, Finn Kensing, and Tariq O. Andersen. Clinician-facing ai in the wild: Taking stock of the sociotechnical challenges and opportunities for hci. ACM Trans. Comput.-Hum. Interact., 30 0 (2), March 2023. ISSN 1073-0516. d...

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.