Pith. sign in

REVIEW 3 cited by

EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.02028 v3 pith:JNYPD52G submitted 2023-07-05 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords datamodelsehrshotfoundationclinicalmodelpatientsstructured
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While the general machine learning (ML) community has benefited from public datasets, tasks, and models, the progress of ML in healthcare has been hampered by a lack of such shared assets. The success of foundation models creates new challenges for healthcare ML by requiring access to shared pretrained models to validate performance benefits. We help address these challenges through three contributions. First, we publish a new dataset, EHRSHOT, which contains deidentified structured data from the electronic health records (EHRs) of 6,739 patients from Stanford Medicine. Unlike MIMIC-III/IV and other popular EHR datasets, EHRSHOT is longitudinal and not restricted to ICU/ED patients. Second, we publish the weights of CLMBR-T-base, a 141M parameter clinical foundation model pretrained on the structured EHR data of 2.57M patients. We are one of the first to fully release such a model for coded EHR data; in contrast, most prior models released for clinical data (e.g. GatorTron, ClinicalBERT) only work with unstructured text and cannot process the rich, structured data within an EHR. We provide an end-to-end pipeline for the community to validate and build upon its performance. Third, we define 15 few-shot clinical prediction tasks, enabling evaluation of foundation models on benefits such as sample efficiency and task adaptation. Our model and dataset are available via a research data use agreement from our website: https://ehrshot.stanford.edu. Code to reproduce our results are available at our Github repo: https://github.com/som-shahlab/ehrshot-benchmark

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 20 citations worldwide. Full citation record

  1. Large Language Models with Temporal Reasoning for Longitudinal Clinical Summarization and Prediction

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Open-source LLMs still struggle with temporal reasoning in long clinical summaries, and adding prior context or retrieval only partially helps.

  2. Towards Foundation Models for Critical Care Time Series

    cs.LG 2024-11 conditional novelty 5.0 of 10

    The paper introduces a harmonized multi-center critical care time series dataset and transfer benchmark covering nine datasets from three continents, with treatment variables, and compares seven models on early event ...

  3. Holistic Artificial Intelligence in Medicine; improved performance and explainability

    cs.AI 2025-06 conditional novelty 4.0 of 10

    An extension of the HAIM multimodal framework that uses LLM-based retrieval and summarization to improve clinical prediction AUC from 79.9% to 90.3% and to generate document-grounded explanations.

Pith tools