Pith. sign in

REVIEW 3 cited by

Few-shot medical image classification with simple shape and texture text descriptors using vision-language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.04005 v1 pith:VLYIWGT2 submitted 2023-08-08 cs.CV

classification cs.CV
keywords descriptorsimagesclassificationmedicalvlmsgpt-4few-shotmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we investigate the usefulness of vision-language models (VLMs) and large language models for binary few-shot classification of medical images. We utilize the GPT-4 model to generate text descriptors that encapsulate the shape and texture characteristics of objects in medical images. Subsequently, these GPT-4 generated descriptors, alongside VLMs pre-trained on natural images, are employed to classify chest X-rays and breast ultrasound images. Our results indicate that few-shot classification of medical images using VLMs and GPT-4 generated descriptors is a viable approach. However, accurate classification requires to exclude certain descriptors from the calculations of the classification scores. Moreover, we assess the ability of VLMs to evaluate shape features in breast mass ultrasound images. We further investigate the degree of variability among the sets of text descriptors produced by GPT-4. Our work provides several important insights about the application of VLMs for medical image analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiEYE: Dataset and Benchmark for OCT-Enhanced Retinal Disease Recognition from Fundus Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A concept-guided distillation method lets a fundus-image model learn from unpaired OCT scans during training, improving retinal disease classification when only fundus photos are available at test time.

  2. Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks

    cs.CV 2024-12 conditional novelty 5.0 of 10

    LLaVA-generated captions fused with image features via dual cross-attention improve remote sensing scene classification accuracy over the paper's own simple baselines.

  3. Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning

    cs.CV 2025-01 reject novelty 4.0 of 10

    HiCA, a hierarchical contrastive fine-tuning method for large vision-language models, is claimed to achieve state-of-the-art few-shot medical image classification, but the paper lacks the experimental detail needed to...

Pith tools