REVIEW 13 cited by
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Several studies showed that Large Language Models (LLMs) can answer medical questions correctly, even outperforming the average human score in some medical exams. However, to our knowledge, no study has been conducted to assess the ability of language models to validate existing or generated medical text for correctness and consistency. In this paper, we introduce MEDEC (https://github.com/abachaa/MEDEC), the first publicly available benchmark for medical error detection and correction in clinical notes, covering five types of errors (Diagnosis, Management, Treatment, Pharmacotherapy, and Causal Organism). MEDEC consists of 3,848 clinical texts, including 488 clinical notes from three US hospital systems that were not previously seen by any LLM. The dataset has been used for the MEDIQA-CORR shared task to evaluate seventeen participating systems [Ben Abacha et al., 2024]. In this paper, we describe the data creation methods and we evaluate recent LLMs (e.g., o1-preview, GPT-4, Claude 3.5 Sonnet, and Gemini 2.0 Flash) for the tasks of detecting and correcting medical errors requiring both medical knowledge and reasoning capabilities. We also conducted a comparative study where two medical doctors performed the same task on the MEDEC test set. The results showed that MEDEC is a sufficiently challenging benchmark to assess the ability of models to validate existing or generated notes and to correct medical errors. We also found that although recent LLMs have a good performance in error detection and correction, they are still outperformed by medical doctors in these tasks. We discuss the potential factors behind this gap, the insights from our experiments, the limitations of current evaluation metrics, and share potential pointers for future research.
Forward citations
Cited by 13 Pith papers
-
KokoroChat: A Japanese Psychological Counseling Dialogue Dataset Collected via Role-Playing by Trained Counselors
A large Japanese counseling dialogue dataset collected via role-play by trained counselors, with per-dialogue client feedback, improves LLM counseling response generation and evaluation.
-
MedFact: Benchmarking the Fact-Checking Capabilities of Large Language Models on Chinese Medical Texts
MedFact, a new Chinese medical fact-checking benchmark, shows LLMs often detect errors but localize them poorly, and more reasoning time triggers over-criticism.
-
Can Large Language Models Understand As Well As Apply Patent Regulations to Pass a Hands-On Patent Attorney Test?
Across EQE pre-exam legal questions, OpenAI o1 reached the highest accuracy (0.82), but no tested LLM reached the 0.90 threshold the authors set for passing, and human patent experts found systematic flaws in the mode...
-
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
A new multi-shot video dataset and an instance-prompt video LLM report large gains, but the main benchmark is built by the same authors and the model is not released.
-
MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports
The authors introduce a 3D CT-based visual question answering benchmark with six error types and three task levels, and show that current 3D medical MLLMs perform poorly on it.
-
Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
A new benchmark shows LLMs rarely detect and correct errors in user prompts unless explicitly instructed, and fine-tuning on error-handling examples greatly improves this ability.
-
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.
-
Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family
Round-trip translation quality in GPT-4 and Llama 2 is jointly shaped by training data volume and language distance from English, with orthographic, phylogenetic, syntactic, and geographic distances as the strongest p...
-
Revisiting Uncertainty Estimation and Calibration of Large Language Models
Across 80 LLMs on MMLU-Pro, linguistic verbal uncertainty judged by another LLM gives better calibration and error ranking on average than token-probability or numeric self-reported uncertainty, with exceptions.
-
Empowering Tabular Data Preparation with Language Models: Why and How?
A structured survey synthesizes LM-based tabular data preparation methods into four phases and two enabling strategies, with qualitative assessments and future directions.
-
A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models
A survey that taxonomizes EHR modeling research into data-centric, architectural, learning-focused, multimodal, and LLM-based categories, with datasets and metrics.
-
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
Fine-tuning small LLMs on benign data raises harmfulness scores, but those scores vary widely across random seeds, temperatures, and repeated runs, making single-run safety comparisons unreliable.
-
OSoRA: Output-Dimension and Singular-Value Initialized Low-Rank Adaptation
OSoRA fine-tunes LLMs by updating only singular values and one output-dimension vector, using frozen singular vectors from an SVD of the pretrained weights.
Discussion (0). Continue with ORCID to comment.