Pith. sign in

REVIEW 2 cited by

PMC-CLIP: Contrastive Language-Image Pre-training using Biomedical Documents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07240 v1 pith:2CHUESDI submitted 2023-03-13 cs.CV cs.CLcs.LGcs.MM

classification cs.CVcs.CLcs.LGcs.MM
keywords biomedicalpmc-oaclassificationdatasetimageimage-captionimage-textmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Foundation models trained on large-scale dataset gain a recent surge in CV and NLP. In contrast, development in biomedical domain lags far behind due to data scarcity. To address this issue, we build and release PMC-OA, a biomedical dataset with 1.6M image-caption pairs collected from PubMedCentral's OpenAccess subset, which is 8 times larger than before. PMC-OA covers diverse modalities or diseases, with majority of the image-caption samples aligned at finer-grained level, i.e., subfigure and subcaption. While pretraining a CLIP-style model on PMC-OA, our model named PMC-CLIP achieves state-of-the-art results on various downstream tasks, including image-text retrieval on ROCO, MedMNIST image classification, Medical VQA, i.e. +8.1% R@10 on image-text retrieval, +3.9% accuracy on image classification.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A pipeline that converts ultrasound textbooks and reports into 45,000 reasoning QA/VQA samples improves a 7B multimodal model on self-built ultrasound benchmarks.

  2. Barriers in Integrating Medical Visual Question Answering into Radiology Workflows: A Scoping Review and Clinicians' Insights

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A scoping review and 50-clinician survey find that MedVQA research is poorly aligned with radiology practice, with mostly non-diagnostic questions, missing patient context, and evaluation metrics that do not measure c...

Pith tools