Pith. sign in

REVIEW 2 cited by

A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.20381 v5 pith:RCCX2BSE submitted 2023-10-31 cs.CV cs.AI

classification cs.CVcs.AI
keywords evaluationmedicalanalysisgpt-4vcapabilityquantitativevisualdiscrepancy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work conducts an evaluation of GPT-4V's multimodal capability for medical image analysis, with a focus on three representative tasks of radiology report generation, medical visual question answering, and medical visual grounding. For the evaluation, a set of prompts is designed for each task to induce the corresponding capability of GPT-4V to produce sufficiently good outputs. Three evaluation ways including quantitative analysis, human evaluation, and case study are employed to achieve an in-depth and extensive evaluation. Our evaluation shows that GPT-4V excels in understanding medical images and is able to generate high-quality radiology reports and effectively answer questions about medical images. Meanwhile, it is found that its performance for medical visual grounding needs to be substantially improved. In addition, we observe the discrepancy between the evaluation outcome from quantitative analysis and that from human evaluation. This discrepancy suggests the limitations of conventional metrics in assessing the performance of large language models like GPT-4V and the necessity of developing new metrics for automatic quantitative analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MMedPO weights preference-optimization training samples by clinical relevance scores, combining hallucinated text answers and locally noised lesion images, and reports improved medical VQA and report generation metrics.

  2. The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering

    cs.CV 2025-01 reject novelty 2.0 of 10

    A survey tracing the evolution of visual question answering from 2015 CNN-LSTM models through attention mechanisms, modular networks, vision-language pretraining, and large multimodal models.

Pith tools