REVIEW 9 cited by
GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is crucial to develop benchmarks to evaluate LVLMs' effectiveness in various medical applications. Current benchmarks are often built upon specific academic literature, mainly focusing on a single domain, and lacking varying perceptual granularities. Thus, they face specific challenges, including limited clinical relevance, incomplete evaluations, and insufficient guidance for interactive LVLMs. To address these limitations, we developed the GMAI-MMBench, the most comprehensive general medical AI benchmark with well-categorized data structure and multi-perceptual granularity to date. It is constructed from 284 datasets across 38 medical image modalities, 18 clinical-related tasks, 18 departments, and 4 perceptual granularities in a Visual Question Answering (VQA) format. Additionally, we implemented a lexical tree structure that allows users to customize evaluation tasks, accommodating various assessment needs and substantially supporting medical AI research and applications. We evaluated 50 LVLMs, and the results show that even the advanced GPT-4o only achieves an accuracy of 53.96%, indicating significant room for improvement. Moreover, we identified five key insufficiencies in current cutting-edge LVLMs that need to be addressed to advance the development of better medical applications. We believe that GMAI-MMBench will stimulate the community to build the next generation of LVLMs toward GMAI.
Forward citations
Cited by 9 Pith papers
-
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
On a new benchmark of 5,620 real multimodal online consultations, top LLMs trail the original physicians mainly because they trigger more unsafe or unsupported negative criteria.
-
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
Watermarking medical AI outputs can degrade reasoning, terminology, and image interpretation even when benchmark accuracy stays stable, so accuracy-only evaluations hide clinically important damage.
-
Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning
FundusExpert, an 8B ophthalmic MLLM trained on region-grounded cognitive-chain instructions, reports state-of-the-art QA and report-generation results, with a fitted data-scaling exponent of 0.068.
-
MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book
MedBookVQA is a new 5,000-question, textbook-derived multimodal benchmark for testing medical AI systems, with labels for imaging modality, body anatomy, and clinical specialty.
-
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
ZeroBench is a hand-built 100-question visual reasoning benchmark, adversarially filtered so every evaluated frontier LMM scored 0% at release.
-
AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs
A cross-modal benchmark derived from MMVP shows VLMs and IGMs share several elementary visual-reasoning failure modes, with IGMs struggling most on explicit attribute-control prompts.
-
MANBench: Is Your Multimodal Model Smarter than Human?
A new bilingual 1,314-question benchmark finds the best multimodal model scores about 60%, below the average human score of 62%, beating humans only on knowledge and basic image-text tasks.
-
Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA
Direct answer-only supervised fine-tuning is the most robust adaptation family on MedFrameQA, beating frozen baselines by ~6 points and outperforming complex variants on seed stability.
-
A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications
A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.
Discussion (0). Continue with ORCID to comment.