REVIEW 2 cited by
VeriFact: Verifying Facts in LLM-Generated Clinical Text with Electronic Health Records
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Methods to ensure factual accuracy of text generated by large language models (LLM) in clinical medicine are lacking. VeriFact is an artificial intelligence system that combines retrieval-augmented generation and LLM-as-a-Judge to verify whether LLM-generated text is factually supported by a patient's medical history based on their electronic health record (EHR). To evaluate this system, we introduce VeriFact-BHC, a new dataset that decomposes Brief Hospital Course narratives from discharge summaries into a set of simple statements with clinician annotations for whether each statement is supported by the patient's EHR clinical notes. Whereas highest agreement between clinicians was 88.5%, VeriFact achieves up to 92.7% agreement when compared to a denoised and adjudicated average human clinican ground truth, suggesting that VeriFact exceeds the average clinician's ability to fact-check text against a patient's medical record. VeriFact may accelerate the development of LLM-based EHR applications by removing current evaluation bottlenecks.
Forward citations
Cited by 2 Pith papers
-
MedFact: Benchmarking the Fact-Checking Capabilities of Large Language Models on Chinese Medical Texts
MedFact, a new Chinese medical fact-checking benchmark, shows LLMs often detect errors but localize them poorly, and more reasoning time triggers over-criticism.
-
MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning
MedReadCtrl instruction-tunes LLaMA3 to control readability at 12 grade levels, reporting lower readability errors than GPT-4 and higher content scores on unseen clinical simplification.
Discussion (0). Sign in to comment.