REVIEW 4 major objections 3 minor 7 cited by
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
T0 review · 4 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read MMAU-Pro, a 5,305-instance audio benchmark spanning speech, sound, music and mixtures, leaves top AI models at 59.2% accuracy.
desk verdict A large, useful audio benchmark resource, but the validity claims hinge on methodology we can't see from the abstract. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The benchmark itself is the central object. It combines three design choices: (1) wild audio sources intended to avoid familiar benchmark distributions, (2) 49 explicitly defined auditory skills spanning speech, sound, music, and mixtures, and (3) human-generated questions in multiple-choice and open-ended formats engineered to require multi-hop reasoning across long-form audio, spatial cues, and multiple audio clips.
What would settle it
Take a random sample of MMAU-Pro instances, replace each wild audio with a cleanly re-recorded version of the same content, and run the same models; if accuracy on the re-recorded set drops to chance while the original scores remain high, the original results are at least partly driven by training-data memorization rather than audio comprehension.
Extended reading notes
Core claim
MMAU-Pro is a comprehensive, human-curated audio benchmark whose central claim is that current multimodal AI systems do not yet possess general auditory intelligence. The benchmark's instances are drawn from the wild rather than from existing datasets, and each question requires multi-hop reasoning over one or more audio inputs, spanning speech, sound, music, and combinations. Across the 22 models evaluated, top accuracy stands at 59.2%, the next at 51.7%, and multiple categories show performance consistent with random guessing. The paper interprets this as evidence of systematic, specific weaknesses that a holistic audio benchmark is needed to expose.
Load-bearing premise
The human expert questions and the wild audio measure genuine audio understanding, not surface cues or memorized training clips.
Editorial extensions
If this is right
- Models that score well on single audio clips may still fail on long-form, spatial, or multi-audio questions, making these distinct test axes.
- Open-ended questions reduce the value of guessing, so accuracy gaps are more likely to reflect comprehension rather than answer-set luck.
- The 49-skill taxonomy gives the field a shared way to describe where audio models fail beyond overall accuracy.
- Categories near random chance point to concrete development targets for next-generation audio models.
Reading between the lines
- Because the audio is claimed to come from the wild, a contamination check could be built by recording fresh clips and re-testing; if scores jump sharply, some of the reported difficulty may reflect memorization rather than comprehension.
- The multi-audio setting suggests a natural diagnostic: determine whether models actually fuse audio streams or attend to just one, which would be visible as near-chance performance on multi-audio questions.
- If the human questions prove reliable, the open-ended, multi-hop protocol could transfer to evaluations in other modalities, such as video or tactile understanding, where comprehension is multi-step rather than recall.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This abstract-only submission introduces MMAU-Pro, a benchmark claimed to contain 5,305 audio-based question-answer instances across speech, sound, music, and their combinations, spanning 49 unique skills and including both multiple-choice and open-ended formats. The authors report evaluations of 22 multimodal AI models, with state-of-the-art systems (Gemini 2.5 Flash, Audio Flamingo 3) achieving only 59.2% and 51.7% accuracy, and claim the results approach random performance in several categories. The central claim is that the benchmark is 'rigorously curated' and measures 'audio general intelligence' by requiring deliberate multi-hop reasoning over audio sourced 'from the wild.' No methodology, annotation details, quality control, contamination analysis, or statistical inference is reported in the abstract.
Significance. If the central claims are substantiated, MMAU-Pro would fill a genuine gap in audio evaluation: existing benchmarks are typically narrower, and a broad skill taxonomy with wild audio and multi-hop questions could push the field toward more holistic audio understanding. The reported poor performance of several strong models is a plausible and important finding that could guide future research. However, the abstract alone provides no verifiable evidence that the benchmark measures audio comprehension rather than linguistic cues or memorization. The claimed rigor and the validity of the accuracy numbers rest entirely on unstated methodology, making the significance strictly conditional.
major comments (4)
- [Abstract] The abstract's claims of 'human expert-generated' and 'meticulously designed' questions are not supported by any annotation quality evidence: no number of annotators, inter-annotator agreement, or item-level quality checks. More critically, no text-only baseline is reported. Without a condition in which audio is withheld, lexical or surface cues in question stems and options could allow high accuracy without processing audio. Because the benchmark's entire purpose is to measure audio comprehension, this omission is load-bearing: the reported accuracies (e.g., 59.2% for Gemini 2.5 Flash) cannot be interpreted as audio-reasoning performance. Please provide a text-only baseline and show that audio is necessary for successful answering.
- [Abstract] The statement that audio is 'sourced directly from the wild' does not address training-data contamination. For 22 models trained on large-scale web data, clips from the wild may already be in pretraining corpora. Without a contamination analysis (e.g., audio fingerprint matching against training sets, or temporal cutoff checks), the reported accuracies could reflect memorization rather than reasoning. Add a contamination analysis and report results on an uncontaminated subset, if any.
- [Abstract] The reported accuracy numbers lack statistical support. The abstract claims models are 'approaching random performance in multiple categories,' but no confidence intervals, significance tests, or per-category sample sizes are given. With 5,305 total items, category-level subsets may be small, making differences between models or against chance potentially noisy. Report binomial or bootstrap confidence intervals and, where relevant, pairwise significance tests.
- [Abstract] The inclusion of 'open-ended response formats' raises scoring-validity concerns. The abstract does not state how open-ended answers are evaluated. If automatic metrics are used, they need validation against human judgment; if human graders are used, inter-rater reliability is required. Without this, the aggregate accuracy conflates answer-format difficulty with audio-reasoning ability. Describe the scoring protocol and its validation.
minor comments (3)
- [Abstract] The claim of being 'the most comprehensive and rigorously curated benchmark' should be supported with a side-by-side comparison to existing audio benchmarks (e.g., number of instances, skills, formats, audio sources).
- [Abstract] Please define 'multi-hop reasoning' operationally. As written it is a qualitative label with no criterion for what constitutes one hop versus multiple hops.
- [Abstract] The abstract refers to '49 unique skills' and 'multiple complex dimensions' without listing them. Provide the skill list in the abstract or explicitly reference an appendix/supplement.
Circularity Check
No circular derivation: benchmark construction is not tautological.
full rationale
This is an abstract-only submission with no equations, fitted parameters, or derivation chain. The claimed contribution is the construction of a benchmark: 5,305 instances, human expert-generated QA pairs, 49 skills, and evaluation of 22 models. There is no step in which a quantity is defined in terms of another quantity and then rediscovered as a prediction. The human expert-generated questions and the skill taxonomy are inputs to the benchmark, not outputs of the model evaluations. The reported accuracies are measurements on an external test set, and the paper does not claim to derive them from the benchmark design. The only self-referential aspect is that the same authors define the skills and curate the questions, but this is a validity/quality concern (e.g., question ambiguity, lexical cues, contamination) rather than circular reasoning. No self-citation is present in the abstract, and no uniqueness theorem or fitted parameter is invoked. Therefore, no circular step can be exhibited, and the honest score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The 49-skill taxonomy fully covers the space of audio comprehension skills relevant to general intelligence.
- domain assumption Human expert-generated question-answer pairs are accurate, unambiguous, and discriminative.
- domain assumption Audio sourced 'from the wild' is not present in training corpora of evaluated models.
Cite this review
Pith. "Pith review of MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence." pith.science (2026). https://pith.science/paper/FSRTMI5U
@misc{pith2026250813992,
author = {Pith},
title = {Pith review of: MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence},
year = {2026},
howpublished = {\url{https://pith.science/paper/FSRTMI5U}},
note = {Machine review of arXiv:2508.13992}
}
read the original abstract
Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro.
Forward citations
Cited by 7 Pith papers
-
Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models
Open-weight speech language models align poorly with human bouba/kiki judgments on real speech and fail crossmodal sound-to-shape matching, while their visual shape ratings remain near human ceiling.
-
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
A self-play game with a known 'odd listener' converts unlabeled audio contrast pairs into a verifiable reward, improving fine-grained audio reasoning on TREA, MMAU, and MMAR.
-
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
Using audio-aware LLMs to judge event presence and temporal order as DPO rewards improves multi-event text-to-audio instruction following.
-
Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models
Self-distillation of Qwen3-Omni-Thinking on 545k Cogito-Pipe audio reasoning traces yields the best open-source MMAR CoT scores and top-tier challenge ranking.
-
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
A Chinese benchmark built on real human speech evaluates large audio language models across instruction following, knowledge, and robustness, revealing large performance gaps.
-
Raon-Speech Technical Report
A 9B SpeechLM trained on 1.38M hours of English/Korean data plus a full-duplex chat extension trained on 119K hours of time-aligned dialogue outperform same-size audio models on speech tasks and FDB turn-taking metrics.
-
ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
ORCA predicts the distribution of human correctness ratings for open-ended audio QA answers and matches or beats LLM judges while also estimating annotator disagreement.
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.