Pith. sign in

REVIEW 1 cited by

MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.17613 v1 pith:QGDCLHVK submitted 2025-05-23 cs.AI cs.CLcs.CV

MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation

classification cs.AI cs.CLcs.CV
keywords generationmultimodalevaluationimagemmmgmodelsaudiointerleaved
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex tasks that involve multiple modalities. To address this, we present MMMG, a comprehensive and human-aligned benchmark for multimodal generation across 4 modality combinations (image, audio, interleaved text and image, interleaved text and audio), with a focus on tasks that present significant challenges for generation models, while still enabling reliable automatic evaluation through a combination of models and programs. MMMG encompasses 49 tasks (including 29 newly developed ones), each with a carefully designed evaluation pipeline, and 937 instructions to systematically assess reasoning, controllability, and other key capabilities of multimodal generation models. Extensive validation demonstrates that MMMG is highly aligned with human evaluation, achieving an average agreement of 94.3%. Benchmarking results on 24 multimodal generation models reveal that even though the state-of-the-art model, GPT Image, achieves 78.3% accuracy for image generation, it falls short on multimodal reasoning and interleaved generation. Furthermore, results suggest considerable headroom for improvement in audio generation, highlighting an important direction for future research.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

    cs.AI 2026-07 conditional novelty 7.0

    Frontier agentic AI models consistently fail at computational imaging tasks requiring physics-aware inversion, producing visually plausible but physically incorrect outputs.