REVIEW 7 cited by
CLIP in Medical Imaging: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Contrastive Language-Image Pre-training (CLIP), a simple yet effective pre-training paradigm, successfully introduces text supervision to vision models. It has shown promising results across various tasks due to its generalizability and interpretability. The use of CLIP has recently gained increasing interest in the medical imaging domain, serving as a pre-training paradigm for image-text alignment, or a critical component in diverse clinical tasks. With the aim of facilitating a deeper understanding of this promising direction, this survey offers an in-depth exploration of the CLIP within the domain of medical imaging, regarding both refined CLIP pre-training and CLIP-driven applications. In this paper, we (1) first start with a brief introduction to the fundamentals of CLIP methodology; (2) then investigate the adaptation of CLIP pre-training in the medical imaging domain, focusing on how to optimize CLIP given characteristics of medical images and reports; (3) further explore practical utilization of CLIP pre-trained models in various tasks, including classification, dense prediction, and cross-modal tasks; and (4) finally discuss existing limitations of CLIP in the context of medical imaging, and propose forward-looking directions to address the demands of medical imaging domain. Studies featuring technical and practical value are both investigated. We expect this survey will provide researchers with a holistic understanding of the CLIP paradigm and its potential implications. The project page of this survey can also be found on https://github.com/zhaozh10/Awesome-CLIP-in-Medical-Imaging.
Forward citations
Cited by 7 Pith papers
-
On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''
FairCLIP's claimed fairness and performance gains over CLIP do not reproduce on two datasets, and its official implementation diverges from the paper's own formulation.
-
SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
General-purpose video foundation models, adapted with lightweight readout heads, reach state-of-the-art performance on three of five scientific video benchmarks.
-
Multimodal Medical Image Binding via Shared Text Embeddings
Five modality-specific CLIP-like medical models are aligned through a shared, distilled text embedding space, enabling zero-shot cross-modal retrieval and improved few-shot classification without paired image data.
-
Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis
Medical CLIP training with text, clinical, and graph soft labels plus negation hard negatives improves chest X-ray zero-shot and fine-tuned performance.
-
Forecasting Continuous Non-Conservative Dynamical Systems in SO(3)
A neural controlled differential equation model with smooth path filtering is proposed for extrapolating 3D rotation trajectories under non-conservative, noisy conditions.
-
CXR-CML: Improved zero-shot classification of long-tailed multi-label diseases in Chest X-Rays
A CLIP-based chest X-ray classifier enhanced with GMM clustering and triplet loss reports higher AUC, but it is trained on the target dataset rather than being zero-shot.
-
MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding
MedMoE inserts report-conditioned mixture-of-experts into a GLoRIA-style medical vision-language model, reporting accuracy gains on several radiology benchmarks.
Discussion (0). Continue with ORCID to comment.