Pith. sign in

REVIEW 3 cited by

The Shaky Foundations of Clinical Foundation Models: A Survey of Large Language Models and Foundation Models for EMRs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.12961 v2 pith:6AZF3UHC submitted 2023-03-22 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelsfoundationclinicaldataemrstrainedalphafoldarchitectures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The successes of foundation models such as ChatGPT and AlphaFold have spurred significant interest in building similar models for electronic medical records (EMRs) to improve patient care and hospital operations. However, recent hype has obscured critical gaps in our understanding of these models' capabilities. We review over 80 foundation models trained on non-imaging EMR data (i.e. clinical text and/or structured data) and create a taxonomy delineating their architectures, training data, and potential use cases. We find that most models are trained on small, narrowly-scoped clinical datasets (e.g. MIMIC-III) or broad, public biomedical corpora (e.g. PubMed) and are evaluated on tasks that do not provide meaningful insights on their usefulness to health systems. In light of these findings, we propose an improved evaluation framework for measuring the benefits of clinical foundation models that is more closely grounded to metrics that matter in healthcare.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Clinical Models with Pseudo Data for De-identification

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Continued pretraining on masked and pseudo-replaced MIMIC-III notes yields strong de-identification models, with masked RoBERTa large performing best and pseudo data helping XLM-RoBERTa more than RoBERTa.

  2. Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A 3B model trained with a small SFT warm-up followed by verifiable-reward RL matches or exceeds far larger models on EHR-based medical calculation, trial matching, and diagnosis tasks.

  3. Revolutionizing Radiology Workflow with Factual and Efficient CXR Report Generation

    cs.CV 2025-06 reject novelty 2.0 of 10

    CXR-PathFinder claims to outperform much larger medical vision-language models on chest X-ray reporting, using adversarial fine-tuning with clinician feedback and knowledge graph verification, but the evidence is not ...

Pith tools