Pith. sign in

REVIEW 11 cited by

UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.10372 v1 pith:WUUI4HEP submitted 2024-12-13 cs.CV

classification cs.CV
keywords medicalimage-textdatasetsmodalitiesmodelsunimedunimed-clipvlms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-Language Models (VLMs) trained via contrastive learning have achieved notable success in natural image tasks. However, their application in the medical domain remains limited due to the scarcity of openly accessible, large-scale medical image-text datasets. Existing medical VLMs either train on closed-source proprietary or relatively small open-source datasets that do not generalize well. Similarly, most models remain specific to a single or limited number of medical imaging domains, again restricting their applicability to other modalities. To address this gap, we introduce UniMed, a large-scale, open-source multi-modal medical dataset comprising over 5.3 million image-text pairs across six diverse imaging modalities: X-ray, CT, MRI, Ultrasound, Pathology, and Fundus. UniMed is developed using a data-collection framework that leverages Large Language Models (LLMs) to transform modality-specific classification datasets into image-text formats while incorporating existing image-text data from the medical domain, facilitating scalable VLM pretraining. Using UniMed, we trained UniMed-CLIP, a unified VLM for six modalities that significantly outperforms existing generalist VLMs and matches modality-specific medical VLMs, achieving notable gains in zero-shot evaluations. For instance, UniMed-CLIP improves over BiomedCLIP (trained on proprietary data) by an absolute gain of +12.61, averaged over 21 datasets, while using 3x less training data. To facilitate future research, we release UniMed dataset, training codes, and models at https://github.com/mbzuai-oryx/UniMed-CLIP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A probabilistic vision-language adapter that models text tokens as Gaussian distributions and uses Mahalanobis distance for uncertainty-weighted alignment, improving medical image segmentation under data scarcity and ...

  2. Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Mammography-specific VLMs lead mean OOD linear-probe performance across 15 datasets, but robustness depends on pretraining objective and is highly dataset-heterogeneous, not on mammography exposure alone.

  3. Diffusion-Based Quality Control of Medical Image Segmentations across Organs

    eess.IV 2025-11 conditional novelty 6.0 of 10

    nnQC uses a latent diffusion model conditioned on the image and the slice position to generate a pseudo-ground-truth mask for automatically scoring segmentation quality across organs and modalities.

  4. Unifying Biomedical Vision-Language Expertise: Towards a Generalist Foundation Model via Multi-CLIP Knowledge Distillation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A student CLIP model distilled from nine medical CLIP teachers outperforms its teachers across most of 58 biomedical benchmarks.

  5. Multimodal Medical Image Binding via Shared Text Embeddings

    eess.IV 2025-06 conditional novelty 6.0 of 10

    Five modality-specific CLIP-like medical models are aligned through a shared, distilled text embedding space, enabling zero-shot cross-modal retrieval and improved few-shot classification without paired image data.

  6. OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A validation-tuned, test-time adaptive ensemble of frozen biomedical vision experts improves classification, segmentation, and multimodal diagnosis across nine datasets without updating expert weights.

  7. TextSLIP: Text Self-Supervised CLIP for Medical Report Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Adding ESimCSE text contrastive learning to CLIP improves medical report generation BLEU scores on brain MRI over standard CLIP by about 1-3 points.

  8. ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics

    cs.LG 2026-04 unverdicted novelty 4.0 of 10

    Standard Conditional Flow Matching loss is a misleading early plateau; physics-informed metrics keep improving, so ScatterPrism and multi-metric diagnostics are needed for kinematic fidelity.

  9. MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding

    cs.CV 2025-06 reject novelty 4.0 of 10

    MedMoE inserts report-conditioned mixture-of-experts into a GLoRIA-style medical vision-language model, reporting accuracy gains on several radiology benchmarks.

  10. On the Robustness of Medical Vision-Language Models: Are they Truly Generalizable?

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Medical vision-language models lose accuracy on corrupted images; RobustMedCLIP, a few-shot LoRA-tuned BioMedCLIP, partially restores robustness on the new MediMeta-C benchmark.

  11. From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.

Pith tools