Pith. sign in

REVIEW 23 cited by

MedCLIP: Contrastive Learning from Unpaired Medical Images and Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.10163 v1 pith:OCK6PP3D submitted 2022-10-18 cs.CV cs.CL

classification cs.CVcs.CL
keywords contrastiveimageslearningmedclipdatamedicalnegativesfalse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing vision-text contrastive learning like CLIP aims to match the paired image and caption embeddings while pushing others apart, which improves representation transferability and supports zero-shot prediction. However, medical image-text datasets are orders of magnitude below the general images and captions from the internet. Moreover, previous methods encounter many false negatives, i.e., images and reports from separate patients probably carry the same semantics but are wrongly treated as negatives. In this paper, we decouple images and texts for multimodal contrastive learning thus scaling the usable training data in a combinatorial magnitude with low cost. We also propose to replace the InfoNCE loss with semantic matching loss based on medical knowledge to eliminate false negatives in contrastive learning. We prove that MedCLIP is a simple yet effective framework: it outperforms state-of-the-art methods on zero-shot prediction, supervised classification, and image-text retrieval. Surprisingly, we observe that with only 20K pre-training data, MedCLIP wins over the state-of-the-art method (using around 200K data). Our code is available at https://github.com/RyanWangZf/MedCLIP.

Discussion (0). Sign in to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...

  2. Cross-Contextual Vision-Language Adaptation with LoRA for Personalized Severe Adverse Event Detection in Clinical Wound Monitoring

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Cross-contextual dual-stream LoRA on BiomedCLIP plus multi-signal temporal OOD scoring detects personalized SAEs in longitudinal diabetic foot ulcer images better than unimodal baselines on one clinical trial dataset.

  3. GLAM: Geometry-Guided Local Alignment for Multi-View VLP in Mammography

    cs.CV 2025-09 conditional novelty 6.0 of 10

    GLAM adds geometry-guided local alignment to mammography visual-language pre-training and outperforms prior VLP baselines on EMBED, VinDr, and RSNA.

  4. Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A 3D encoder pretrained with GPT-4V slice captions and partial optimal transport alignment beats vision-only SSL baselines on several medical tasks, but a key evaluation dataset may overlap with pretraining.

  5. Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage method (MpGI) produces a 210K-image, 1.26M-caption remote sensing dataset and state-of-the-art CLIP and CoCa models.

  6. MadCLIP: Few-shot Medical Anomaly Detection with CLIP

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A dual-branch CLIP adaptation with learnable adapters and text prompts, trained with SigLIP loss, achieves state-of-the-art few-shot medical anomaly detection and segmentation.

  7. MVP-CBM:Multi-layer Visual Preference-enhanced Concept Bottleneck Model for Explainable Medical Image Classification

    cs.CV 2025-06 conditional novelty 6.0 of 10

    By modeling which visual layers best explain each diagnostic concept and sparsely fusing multi-layer concept activations, MVP-CBM improves accuracy and interpretability over prior concept bottleneck models on seven me...

  8. IQE-CLIP: Instance-aware Query Embedding for Zero-/Few-shot Anomaly Detection in Medical Domain

    cs.CV 2025-06 conditional novelty 6.0 of 10

    IQE-CLIP improves zero- and few-shot medical anomaly detection by building query embeddings that combine text prompts with visual features from each test image, beating prior CLIP-based methods on six BMAD datasets.

  9. Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval

    cs.CV 2025-08 conditional novelty 5.0 of 10

    PECM combines multi-level prototypes with dual-stream confidence weighting and reports up to 10.17% retrieval gains on radiology datasets, including zero-shot transfer.

  10. LSDM: LLM-Enhanced Spatio-temporal Diffusion Model for Service-Level Mobile Traffic Prediction

    cs.LG 2025-07 conditional novelty 5.0 of 10

    LSDM predicts next-hour mobile traffic per app category by feeding a diffusion model with satellite imagery, POI counts, and LLM-generated text descriptions, outperforming eight baselines on a single real-world dataset.

  11. MCA-RG: Enhancing LLMs with Medical Concept Alignment for Radiology Report Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MCA-RG uses concept alignment, contrastive learning, matching loss, and feature gating to generate radiology reports, reporting SOTA on MIMIC-CXR and CheXpert Plus.

  12. Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SPA uses few-shot spatial adaptation, a diffusion model trained on a hand-provided procedure graph, and test-time mutual agreement to achieve strong surgical phase recognition with minimal labeled data.

  13. MOSCARD -- Causal Reasoning and De-confounding for Multimodal Opportunistic Screening of Cardiovascular Adverse Events

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MOSCARD fuses CXR and ECG with co-attention and claims de-confounded MACE risk prediction, reaching 0.75 internal and 0.837 ED AUC, but causal intervention details are missing.

  14. HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

    cs.CV 2025-02 reject novelty 5.0 of 10

    HealthGPT unifies medical image comprehension and generation in a single autoregressive model using heterogeneous low-rank adaptation, reporting strong benchmark results.

  15. Automated Radiology Report Generation Based on Topic-Keyword Semantic Guidance

    cs.MM 2025-09 conditional novelty 4.0 of 10

    A topic-keyword semantic guidance framework improves automated radiology report generation and reaches state-of-the-art on two public chest X-ray benchmarks.

  16. CXR-CML: Improved zero-shot classification of long-tailed multi-label diseases in Chest X-Rays

    cs.CV 2025-07 reject novelty 4.0 of 10

    A CLIP-based chest X-ray classifier enhanced with GMM clustering and triplet loss reports higher AUC, but it is trained on the target dataset rather than being zero-shot.

  17. MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training

    cs.CV 2025-07 conditional novelty 4.0 of 10

    MaskedCLIP jointly trains a masked autoencoder and a CLIP-style contrastive model on paired and unpaired medical images, improving downstream retinal disease classification.

  18. A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models

    cs.LG 2025-07 reject novelty 4.0 of 10

    A survey that taxonomizes EHR modeling research into data-centric, architectural, learning-focused, multimodal, and LLM-based categories, with datasets and metrics.

  19. Medical-Knowledge Driven Multiple Instance Learning for Classifying Severe Abdominal Anomalies on Prenatal Ultrasound

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A medical-knowledge-driven multiple instance learning framework classifies fetal abdominal anomalies at case level from whole ultrasound examination image pools, without standard plane localization.

  20. Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination

    cs.CV 2025-06 reject novelty 4.0 of 10

    A biomedical vision-language model is pre-trained with a contrastive loss that distinguishes original radiology reports from nine perturbed variants, plus a local attention loss, and is reported to beat ConVIRT and GL...

  21. Advancements in Medical Image Classification through Fine-Tuning Natural Domain Foundation Models

    eess.IV 2025-05 conditional novelty 4.0 of 10

    Fine-tuning recent natural-domain foundation models, especially AIMv2, improves medical image classification accuracy across mammography, skin lesion, retinopathy, and chest X-ray benchmarks.

  22. On the Robustness of Medical Vision-Language Models: Are they Truly Generalizable?

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Medical vision-language models lose accuracy on corrupted images; RobustMedCLIP, a few-shot LoRA-tuned BioMedCLIP, partially restores robustness on the new MediMeta-C benchmark.

  23. Quality Versus Sparsity in Image Recovery by Dictionary Learning Using Iterative Shrinkage

    cs.CV 2025-08 reject novelty 3.0 of 10

    A dictionary learning paper whose abstract claims sparsity does not hurt recovery quality, but whose full text is an unrelated medical retrieval manuscript.

Pith tools