Pith. sign in

REVIEW 1 cited by

Almanac: Retrieval-Augmented Language Models for Clinical Medicine

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.01229 v2 pith:GOCKFHJL submitted 2023-03-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords clinicallanguagemodelsalmanaccapabilitieslargemedicineacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-language models have recently demonstrated impressive zero-shot capabilities in a variety of natural language tasks such as summarization, dialogue generation, and question-answering. Despite many promising applications in clinical medicine, adoption of these models in real-world settings has been largely limited by their tendency to generate incorrect and sometimes even toxic statements. In this study, we develop Almanac, a large language model framework augmented with retrieval capabilities for medical guideline and treatment recommendations. Performance on a novel dataset of clinical scenarios (n = 130) evaluated by a panel of 5 board-certified and resident physicians demonstrates significant increases in factuality (mean of 18% at p-value < 0.05) across all specialties, with improvements in completeness and safety. Our results demonstrate the potential for large language models to be effective tools in the clinical decision-making process, while also emphasizing the importance of careful testing and deployment to mitigate their shortcomings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative AI in Medicine

    cs.LG 2024-12 conditional novelty 1.0 of 10

    A stakeholder-based review of generative AI use cases in medicine and the consent, privacy, transparency, hallucination, usability, equity, evaluation, and accountability challenges that stand between prototypes and s...

Pith tools