Pith. sign in

REVIEW 2 cited by

meds_reader: A fast and efficient EHR processing library

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.09095 v2 pith:BLLMZAB5 submitted 2024-09-12 cs.LG cs.DB

classification cs.LGcs.DB
keywords medsreaderprocessingefficientdatapipelinesspeedachieving
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The growing demand for machine learning in healthcare requires processing increasingly large electronic health record (EHR) datasets, but existing pipelines are not computationally efficient or scalable. In this paper, we introduce meds_reader, an optimized Python package for efficient EHR data processing that is designed to take advantage of many intrinsic properties of EHR data for improved speed. We then demonstrate the benefits of meds_reader by reimplementing key components of two major EHR processing pipelines, achieving 10-100x improvements in memory, speed, and disk usage. The code for meds_reader can be found at https://github.com/som-shahlab/meds_reader.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A pre-training method that aligns ICU time-series windows with LLM-encoded event summaries via a regularised InfoNCE loss improves downstream predictions and cross-dataset transfer.

  2. FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

    cs.LG 2025-05 conditional novelty 5.0 of 10

    FoMoH benchmarks six structured EHR foundation models on 14 tasks and finds they do not consistently outperform supervised baselines, particularly for rare diseases and low-data regimes.

Pith tools