Pith. sign in

REVIEW 3 cited by

BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.07171 v3 pith:BF4Y2D27 submitted 2025-01-13 cs.CV cs.CL

classification cs.CVcs.CL
keywords datasetmodelsbiologybiomedicabiomedicalaccessibleacrossarchive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The development of vision-language models (VLMs) is driven by large-scale and diverse multimodal datasets. However, progress toward generalist biomedical VLMs is limited by the lack of annotated, publicly accessible datasets across biology and medicine. Existing efforts are restricted to narrow domains, missing the full diversity of biomedical knowledge encoded in scientific literature. To address this gap, we introduce BIOMEDICA, a scalable, open-source framework to extract, annotate, and serialize the entirety of the PubMed Central Open Access subset into an easy-to-use, publicly accessible dataset. Our framework produces a comprehensive archive with over 24 million unique image-text pairs from over 6 million articles. Metadata and expert-guided annotations are also provided. We demonstrate the utility and accessibility of our resource by releasing BMCA-CLIP, a suite of CLIP-style models continuously pre-trained on the BIOMEDICA dataset via streaming, eliminating the need to download 27 TB of data locally. On average, our models achieve state-of-the-art performance across 40 tasks - spanning pathology, radiology, ophthalmology, dermatology, surgery, molecular biology, parasitology, and cell biology - excelling in zero-shot classification with a 6.56% average improvement (as high as 29.8% and 17.5% in dermatology and ophthalmology, respectively), and stronger image-text retrieval, all while using 10x less compute. To foster reproducibility and collaboration, we release our codebase and dataset for the broader research community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Large Language Models Match the Conclusions of Systematic Reviews?

    cs.CL 2025-05 conditional novelty 7.0 of 10

    On 284 medical questions derived from Cochrane systematic reviews, the best of 24 LLMs, DeepSeek V3, matches expert conclusions 62.40% of the time, and all tested models struggle with uncertain or low-quality evidence.

  2. M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision

    cs.CV 2025-09 conditional novelty 6.0 of 10

    One self-supervised encoder trained on unpaired X-ray, ultrasound, endoscopy, and CT data gives competitive zero-shot retrieval and seems to generalize to unseen MRI tasks.

  3. Unifying Biomedical Vision-Language Expertise: Towards a Generalist Foundation Model via Multi-CLIP Knowledge Distillation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A student CLIP model distilled from nine medical CLIP teachers outperforms its teachers across most of 58 biomedical benchmarks.

Pith tools