REVIEW 3 cited by
Few-shot medical image classification with simple shape and texture text descriptors using vision-language models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work, we investigate the usefulness of vision-language models (VLMs) and large language models for binary few-shot classification of medical images. We utilize the GPT-4 model to generate text descriptors that encapsulate the shape and texture characteristics of objects in medical images. Subsequently, these GPT-4 generated descriptors, alongside VLMs pre-trained on natural images, are employed to classify chest X-rays and breast ultrasound images. Our results indicate that few-shot classification of medical images using VLMs and GPT-4 generated descriptors is a viable approach. However, accurate classification requires to exclude certain descriptors from the calculations of the classification scores. Moreover, we assess the ability of VLMs to evaluate shape features in breast mass ultrasound images. We further investigate the degree of variability among the sets of text descriptors produced by GPT-4. Our work provides several important insights about the application of VLMs for medical image analysis.
Forward citations
Cited by 3 Pith papers
-
MultiEYE: Dataset and Benchmark for OCT-Enhanced Retinal Disease Recognition from Fundus Images
A concept-guided distillation method lets a fundus-image model learn from unpaired OCT scans during training, improving retinal disease classification when only fundus photos are available at test time.
-
Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks
LLaVA-generated captions fused with image features via dual cross-attention improve remote sensing scene classification accuracy over the paper's own simple baselines.
-
Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning
HiCA, a hierarchical contrastive fine-tuning method for large vision-language models, is claimed to achieve state-of-the-art few-shot medical image classification, but the paper lacks the experimental detail needed to...
Discussion (0). Continue with ORCID to comment.