Pith. sign in

REVIEW 7 cited by

Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.09909 v3 pith:SRLXYMNN submitted 2023-10-15 cs.CV cs.CL

classification cs.CVcs.CL
keywords gpt-4vmedicaldiagnosisdiseasemultimodalusedanatomyapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Driven by the large foundation models, the development of artificial intelligence has witnessed tremendous progress lately, leading to a surge of general interest from the public. In this study, we aim to assess the performance of OpenAI's newest model, GPT-4V(ision), specifically in the realm of multimodal medical diagnosis. Our evaluation encompasses 17 human body systems, including Central Nervous System, Head and Neck, Cardiac, Chest, Hematology, Hepatobiliary, Gastrointestinal, Urogenital, Gynecology, Obstetrics, Breast, Musculoskeletal, Spine, Vascular, Oncology, Trauma, Pediatrics, with images taken from 8 modalities used in daily clinic routine, e.g., X-ray, Computed Tomography (CT), Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), Digital Subtraction Angiography (DSA), Mammography, Ultrasound, and Pathology. We probe the GPT-4V's ability on multiple clinical tasks with or without patent history provided, including imaging modality and anatomy recognition, disease diagnosis, report generation, disease localisation. Our observation shows that, while GPT-4V demonstrates proficiency in distinguishing between medical image modalities and anatomy, it faces significant challenges in disease diagnosis and generating comprehensive reports. These findings underscore that while large multimodal models have made significant advancements in computer vision and natural language processing, it remains far from being used to effectively support real-world medical applications and clinical decision-making. All images used in this report can be found in https://github.com/chaoyi-wu/GPT-4V_Medical_Evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On a new 65,464-item binary benchmark that swaps one medical term per caption, four medical multimodal models scored 49–62%, close to chance.

  2. Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A 3D encoder pretrained with GPT-4V slice captions and partial optimal transport alignment beats vision-only SSL baselines on several medical tasks, but a key evaluation dataset may overlap with pretraining.

  3. Active Learning for Neurosymbolic Program Synthesis

    cs.PL 2025-08 unverdicted novelty 6.0 of 10

    The abstract claims a new active learning technique, constrained conformal evaluation (tool SmartLabel), that finds the ground-truth program in 98% of benchmarks, but the delivered full text is a different paper, leav...

  4. MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Current medical multimodal models, including GPT-4o and Claude 3.5 Sonnet, fail simple perceptual tasks on medical images that human experts solve almost perfectly.

  5. Are MLMs Trapped in the Visual Room?

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new Reddit-derived sarcasm benchmark and two-tier evaluation show that multimodal models can perceive scenes accurately while still failing to grasp sarcastic intent.

  6. Discrete Prompt Tuning via Recursive Utilization of Black-box Multimodal Large Language Model for Personalized Visual Emotion Recognition

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Recursive generation, evaluation, and refinement of discrete prompts tunes a black-box multimodal LLM to each user, raising personalized visual emotion recognition accuracy on Affection from 40.6% to 44.9%.

  7. SpatialFly: Implicit 3D Prior-Guided Visual Reparameterization for Continuous UAV Vision-and-Language Navigation

    cs.CV 2026-03 unverdicted novelty 4.0 of 10

    SpatialFly reparameterizes 2D visual tokens with implicit geometric priors and reports lower navigation error and higher success than prior UAV VLN systems on unseen splits.

Pith tools