REVIEW 8 cited by
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Advancements in Large Language Models (LLMs) and their increasing use in medical question-answering necessitate rigorous evaluation of their reliability. A critical challenge lies in hallucination, where models generate plausible yet factually incorrect outputs. In the medical domain, this poses serious risks to patient safety and clinical decision-making. To address this, we introduce MedHallu, the first benchmark specifically designed for medical hallucination detection. MedHallu comprises 10,000 high-quality question-answer pairs derived from PubMedQA, with hallucinated answers systematically generated through a controlled pipeline. Our experiments show that state-of-the-art LLMs, including GPT-4o, Llama-3.1, and the medically fine-tuned UltraMedical, struggle with this binary hallucination detection task, with the best model achieving an F1 score as low as 0.625 for detecting "hard" category hallucinations. Using bidirectional entailment clustering, we show that harder-to-detect hallucinations are semantically closer to ground truth. Through experiments, we also show incorporating domain-specific knowledge and introducing a "not sure" category as one of the answer categories improves the precision and F1 scores by up to 38% relative to baselines.
Forward citations
Cited by 8 Pith papers
-
CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series
CLIR-Bench shows generalist and time-series LLMs struggle to ground clinical answers in sparse irregular ICU evidence, with top accuracy near 50% and weak causal evidence use.
-
ADRD-Bench: A Preliminary LLM Benchmark for Alzheimer's Disease and Related Dementias
ADRD-Bench combines 1,352 extracted ADRD exam questions with 149 new caregiving questions and shows that even top LLMs fail on daily-care scenarios.
-
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
A dynamic red-teaming audit reports that 94% of MedQA-correct answers fail under adversarial mutation, with 86% privacy leak rates, 81% bias shift rates, and 66-74% hallucination rates across 15 medical LLMs.
-
MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
MTCMB is a 12-dataset benchmark for evaluating LLMs on Traditional Chinese Medicine knowledge, reasoning, and safety, with results showing models still fail at clinical reasoning and safe prescriptions.
-
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
A 5.8M-pair, literature-grounded medical QA corpus improves RAG accuracy and hallucination detection in LLMs compared with unstructured passage retrieval.
-
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection
Using hallucinated benchmark answers as rejected DPO pairs, ordered by an external fact-checker's grounding score, improves hallucination detection in 1B-3B Llama models.
-
A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models
A survey that taxonomizes EHR modeling research into data-centric, architectural, learning-focused, multimodal, and LLM-based categories, with datasets and metrics.
-
A comprehensive taxonomy of hallucinations in Large Language Models
A survey that organizes LLM hallucination types, causes, benchmarks, and mitigations, and restates the theorem that hallucination is inevitable for computable LLMs.
Discussion (0). Continue with ORCID to comment.