Pith. sign in

REVIEW 9 cited by

GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.03361 v7 pith:FNYYNC3L submitted 2024-08-06 eess.IV cs.CV

classification eess.IVcs.CV
keywords lvlmsmedicalapplicationsgmai-mmbenchvariousbenchmarkbenchmarkscomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is crucial to develop benchmarks to evaluate LVLMs' effectiveness in various medical applications. Current benchmarks are often built upon specific academic literature, mainly focusing on a single domain, and lacking varying perceptual granularities. Thus, they face specific challenges, including limited clinical relevance, incomplete evaluations, and insufficient guidance for interactive LVLMs. To address these limitations, we developed the GMAI-MMBench, the most comprehensive general medical AI benchmark with well-categorized data structure and multi-perceptual granularity to date. It is constructed from 284 datasets across 38 medical image modalities, 18 clinical-related tasks, 18 departments, and 4 perceptual granularities in a Visual Question Answering (VQA) format. Additionally, we implemented a lexical tree structure that allows users to customize evaluation tasks, accommodating various assessment needs and substantially supporting medical AI research and applications. We evaluated 50 LVLMs, and the results show that even the advanced GPT-4o only achieves an accuracy of 53.96%, indicating significant room for improvement. Moreover, we identified five key insufficiencies in current cutting-edge LVLMs that need to be addressed to advance the development of better medical applications. We believe that GMAI-MMBench will stimulate the community to build the next generation of LVLMs toward GMAI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

    cs.AI 2026-07 conditional novelty 6.5 of 10

    On a new benchmark of 5,620 real multimodal online consultations, top LLMs trail the original physicians mainly because they trigger more unsafe or unsupported negative criteria.

  2. Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

    cs.AI 2026-05 conditional novelty 6.0 of 10

    Watermarking medical AI outputs can degrade reasoning, terminology, and image interpretation even when benchmark accuracy stays stable, so accuracy-only evaluations hide clinically important damage.

  3. Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning

    cs.AI 2025-07 conditional novelty 6.0 of 10

    FundusExpert, an 8B ophthalmic MLLM trained on region-grounded cognitive-chain instructions, reports state-of-the-art QA and report-generation results, with a fitted data-scaling exponent of 0.068.

  4. MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book

    cs.AI 2025-06 conditional novelty 6.0 of 10

    MedBookVQA is a new 5,000-question, textbook-derived multimodal benchmark for testing medical AI systems, with labels for imaging modality, body anatomy, and clinical specialty.

  5. ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ZeroBench is a hand-built 100-question visual reasoning benchmark, adversarially filtered so every evaluated frontier LMM scored 0% at release.

  6. AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs

    cs.CV 2026-01 conditional novelty 5.0 of 10

    A cross-modal benchmark derived from MMVP shows VLMs and IGMs share several elementary visual-reasoning failure modes, with IGMs struggling most on explicit attribute-control prompts.

  7. MANBench: Is Your Multimodal Model Smarter than Human?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A new bilingual 1,314-question benchmark finds the best multimodal model scores about 60%, below the average human score of 62%, beating humans only on knowledge and basic image-text tasks.

  8. Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Direct answer-only supervised fine-tuning is the most robust adaptation family on MedFrameQA, beating frozen baselines by ~6 points and outperforming complex variants on seed stability.

  9. A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.

Pith tools