Pith. sign in

REVIEW 13 cited by

MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.19260 v2 pith:F6MIN7X3 submitted 2024-12-26 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords medicalmedecclinicalnotesbenchmarkcorrectiondetectionerror
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Several studies showed that Large Language Models (LLMs) can answer medical questions correctly, even outperforming the average human score in some medical exams. However, to our knowledge, no study has been conducted to assess the ability of language models to validate existing or generated medical text for correctness and consistency. In this paper, we introduce MEDEC (https://github.com/abachaa/MEDEC), the first publicly available benchmark for medical error detection and correction in clinical notes, covering five types of errors (Diagnosis, Management, Treatment, Pharmacotherapy, and Causal Organism). MEDEC consists of 3,848 clinical texts, including 488 clinical notes from three US hospital systems that were not previously seen by any LLM. The dataset has been used for the MEDIQA-CORR shared task to evaluate seventeen participating systems [Ben Abacha et al., 2024]. In this paper, we describe the data creation methods and we evaluate recent LLMs (e.g., o1-preview, GPT-4, Claude 3.5 Sonnet, and Gemini 2.0 Flash) for the tasks of detecting and correcting medical errors requiring both medical knowledge and reasoning capabilities. We also conducted a comparative study where two medical doctors performed the same task on the MEDEC test set. The results showed that MEDEC is a sufficiently challenging benchmark to assess the ability of models to validate existing or generated notes and to correct medical errors. We also found that although recent LLMs have a good performance in error detection and correction, they are still outperformed by medical doctors in these tasks. We discuss the potential factors behind this gap, the insights from our experiments, the limitations of current evaluation metrics, and share potential pointers for future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KokoroChat: A Japanese Psychological Counseling Dialogue Dataset Collected via Role-Playing by Trained Counselors

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A large Japanese counseling dialogue dataset collected via role-play by trained counselors, with per-dialogue client feedback, improves LLM counseling response generation and evaluation.

  2. MedFact: Benchmarking the Fact-Checking Capabilities of Large Language Models on Chinese Medical Texts

    cs.CL 2025-09 conditional novelty 6.0 of 10

    MedFact, a new Chinese medical fact-checking benchmark, shows LLMs often detect errors but localize them poorly, and more reasoning time triggers over-criticism.

  3. Can Large Language Models Understand As Well As Apply Patent Regulations to Pass a Hands-On Patent Attorney Test?

    cs.CY 2025-07 conditional novelty 6.0 of 10

    Across EQE pre-exam legal questions, OpenAI o1 reached the highest accuracy (0.82), but no tested LLM reached the 0.90 threshold the authors set for passing, and human patent experts found systematic flaws in the mode...

  4. IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new multi-shot video dataset and an instance-prompt video LLM report large gains, but the main benchmark is built by the same authors and the model is not released.

  5. MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The authors introduce a 3D CT-based visual question answering benchmark with six error types and three task levels, and show that current 3D medical MLLMs perform poorly on it.

  6. Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark shows LLMs rarely detect and correct errors in user prompts unless explicitly instructed, and fine-tuning on error-handling examples greatly improves this ability.

  7. TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.

  8. Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family

    cs.CL 2025-06 reject novelty 5.0 of 10

    Round-trip translation quality in GPT-4 and Llama 2 is jointly shaped by training data volume and language distance from English, with orthographic, phylogenetic, syntactic, and geographic distances as the strongest p...

  9. Revisiting Uncertainty Estimation and Calibration of Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Across 80 LLMs on MMLU-Pro, linguistic verbal uncertainty judged by another LLM gives better calibration and error ranking on average than token-probability or numeric self-reported uncertainty, with exceptions.

  10. Empowering Tabular Data Preparation with Language Models: Why and How?

    cs.AI 2025-08 accept novelty 4.0 of 10

    A structured survey synthesizes LM-based tabular data preparation methods into four phases and two enabling strategies, with qualitative assessments and future directions.

  11. A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models

    cs.LG 2025-07 reject novelty 4.0 of 10

    A survey that taxonomizes EHR modeling research into data-centric, architectural, learning-focused, multimodal, and LLM-based categories, with datasets and metrics.

  12. Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Fine-tuning small LLMs on benign data raises harmfulness scores, but those scores vary widely across random seeds, temperatures, and repeated runs, making single-run safety comparisons unreliable.

  13. OSoRA: Output-Dimension and Singular-Value Initialized Low-Rank Adaptation

    cs.CL 2025-05 reject novelty 4.0 of 10

    OSoRA fine-tunes LLMs by updating only singular values and one output-dimension vector, using frozen singular vectors from an SVD of the pretrained weights.

Pith tools