Pith. sign in

REVIEW 2 cited by

A foundation model for human-AI collaboration in medical literature mining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.16255 v1 pith:IMOQXGKG submitted 2025-01-27 cs.CL

classification cs.CL
keywords leadsliteraturemedicalmodelsclinicaldataexpertexperts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Systematic literature review is essential for evidence-based medicine, requiring comprehensive analysis of clinical trial publications. However, the application of artificial intelligence (AI) models for medical literature mining has been limited by insufficient training and evaluation across broad therapeutic areas and diverse tasks. Here, we present LEADS, an AI foundation model for study search, screening, and data extraction from medical literature. The model is trained on 633,759 instruction data points in LEADSInstruct, curated from 21,335 systematic reviews, 453,625 clinical trial publications, and 27,015 clinical trial registries. We showed that LEADS demonstrates consistent improvements over four cutting-edge generic large language models (LLMs) on six tasks. Furthermore, LEADS enhances expert workflows by providing supportive references following expert requests, streamlining processes while maintaining high-quality results. A study with 16 clinicians and medical researchers from 14 different institutions revealed that experts collaborating with LEADS achieved a recall of 0.81 compared to 0.77 experts working alone in study selection, with a time savings of 22.6%. In data extraction tasks, experts using LEADS achieved an accuracy of 0.85 versus 0.80 without using LEADS, alongside a 26.9% time savings. These findings highlight the potential of specialized medical literature foundation models to outperform generic models, delivering significant quality and efficiency benefits when integrated into expert workflows for medical literature mining.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research

    cs.AI 2025-05 conditional novelty 6.0 of 10

    BioDSA-1K is a large, publication-grounded benchmark for evaluating AI agents on biomedical hypothesis validation, including non-verifiable cases.

  2. Artificial intelligence in drug discovery: A comprehensive review with a case study on hyperuricemia, gout arthritis, and hyperuricemic nephropathy

    cs.AI 2025-07 unverdicted novelty 2.0 of 10

    A literature review of AI/ML in drug discovery, with a case study summarizing network pharmacology and machine learning based target and hit discoveries for gout-related diseases.

Pith tools