Aligning Medical Images with General Knowledge from Large Language Models

Dong Zhang; Hao Chen; Kwang-Ting Cheng; Xiao Fang; Yi Lin

arxiv: 2409.00341 · v1 · pith:QYLCLNDVnew · submitted 2024-08-31 · 💻 cs.CV

Aligning Medical Images with General Knowledge from Large Language Models

Xiao Fang , Yi Lin , Dong Zhang , Kwang-Ting Cheng , Hao Chen This is my paper

classification 💻 cs.CV

keywords visuallargepromptlanguagemedicalmodelsanalysisclip

0 comments

read the original abstract

Pre-trained large vision-language models (VLMs) like CLIP have revolutionized visual representation learning using natural language as supervisions, and demonstrated promising generalization ability. In this work, we propose ViP, a novel visual symptom-guided prompt learning framework for medical image analysis, which facilitates general knowledge transfer from CLIP. ViP consists of two key components: a visual symptom generator (VSG) and a dual-prompt network. Specifically, VSG aims to extract explicable visual symptoms from pre-trained large language models, while the dual-prompt network utilizes these visual symptoms to guide the training on two learnable prompt modules, i.e., context prompt and merge prompt, which effectively adapts our framework to medical image analysis via large VLMs. Extensive experimental results demonstrate that ViP can outperform state-of-the-art methods on two challenging datasets.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Boosting Ultrasound Image Classification via Attribute-Guided Dual-Branch Framework
cs.CV 2026-07 conditional novelty 5.0

An attribute-guided dual-branch framework fuses a standard classifier with an interpretable attribute-prior branch to boost ultrasound classification accuracy and explainability.