Pith. sign in

REVIEW 4 major objections 3 minor 7 cited by

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

T0 review · 4 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read MMAU-Pro, a 5,305-instance audio benchmark spanning speech, sound, music and mixtures, leaves top AI models at 59.2% accuracy.

desk verdict A large, useful audio benchmark resource, but the validity claims hinge on methodology we can't see from the abstract. read the letter →

arxiv 2508.13992 v1 pith:FSRTMI5U submitted 2025-08-19 eess.AS cs.SD

classification eess.AScs.SD
keywords audiointelligencebenchmarkmulti-hopreasoningspeechmusicsoundmultimodalAIevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces MMAU-Pro, a benchmark of 5,305 audio question-answer instances written by human experts, covering speech, non-speech sound, music, and combinations. Each instance is built to demand deliberate multi-hop reasoning, with long-form audio, spatial audio, and multiple audio clips per question. Evaluating 22 leading open and closed AI models, the paper reports that the best model reaches only 59.2% accuracy and the runner-up 51.7%, with several skill categories near chance. The authors argue the results reveal a large gap between current models and audio general intelligence, and they position the benchmark as a tool to help close it.

What carries the argument

The benchmark itself is the central object. It combines three design choices: (1) wild audio sources intended to avoid familiar benchmark distributions, (2) 49 explicitly defined auditory skills spanning speech, sound, music, and mixtures, and (3) human-generated questions in multiple-choice and open-ended formats engineered to require multi-hop reasoning across long-form audio, spatial cues, and multiple audio clips.

What would settle it

Take a random sample of MMAU-Pro instances, replace each wild audio with a cleanly re-recorded version of the same content, and run the same models; if accuracy on the re-recorded set drops to chance while the original scores remain high, the original results are at least partly driven by training-data memorization rather than audio comprehension.

Watch

Extended reading notes

Core claim

MMAU-Pro is a comprehensive, human-curated audio benchmark whose central claim is that current multimodal AI systems do not yet possess general auditory intelligence. The benchmark's instances are drawn from the wild rather than from existing datasets, and each question requires multi-hop reasoning over one or more audio inputs, spanning speech, sound, music, and combinations. Across the 22 models evaluated, top accuracy stands at 59.2%, the next at 51.7%, and multiple categories show performance consistent with random guessing. The paper interprets this as evidence of systematic, specific weaknesses that a holistic audio benchmark is needed to expose.

Load-bearing premise

The human expert questions and the wild audio measure genuine audio understanding, not surface cues or memorized training clips.

Editorial extensions

If this is right

  • Models that score well on single audio clips may still fail on long-form, spatial, or multi-audio questions, making these distinct test axes.
  • Open-ended questions reduce the value of guessing, so accuracy gaps are more likely to reflect comprehension rather than answer-set luck.
  • The 49-skill taxonomy gives the field a shared way to describe where audio models fail beyond overall accuracy.
  • Categories near random chance point to concrete development targets for next-generation audio models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the audio is claimed to come from the wild, a contamination check could be built by recording fresh clips and re-testing; if scores jump sharply, some of the reported difficulty may reflect memorization rather than comprehension.
  • The multi-audio setting suggests a natural diagnostic: determine whether models actually fuse audio streams or attend to just one, which would be visible as near-chance performance on multi-audio questions.
  • If the human questions prove reliable, the open-ended, multi-hop protocol could transfer to evaluations in other modalities, such as video or tactile understanding, where comprehension is multi-step rather than recall.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. This abstract-only submission introduces MMAU-Pro, a benchmark claimed to contain 5,305 audio-based question-answer instances across speech, sound, music, and their combinations, spanning 49 unique skills and including both multiple-choice and open-ended formats. The authors report evaluations of 22 multimodal AI models, with state-of-the-art systems (Gemini 2.5 Flash, Audio Flamingo 3) achieving only 59.2% and 51.7% accuracy, and claim the results approach random performance in several categories. The central claim is that the benchmark is 'rigorously curated' and measures 'audio general intelligence' by requiring deliberate multi-hop reasoning over audio sourced 'from the wild.' No methodology, annotation details, quality control, contamination analysis, or statistical inference is reported in the abstract.

Significance. If the central claims are substantiated, MMAU-Pro would fill a genuine gap in audio evaluation: existing benchmarks are typically narrower, and a broad skill taxonomy with wild audio and multi-hop questions could push the field toward more holistic audio understanding. The reported poor performance of several strong models is a plausible and important finding that could guide future research. However, the abstract alone provides no verifiable evidence that the benchmark measures audio comprehension rather than linguistic cues or memorization. The claimed rigor and the validity of the accuracy numbers rest entirely on unstated methodology, making the significance strictly conditional.

major comments (4)
  1. [Abstract] The abstract's claims of 'human expert-generated' and 'meticulously designed' questions are not supported by any annotation quality evidence: no number of annotators, inter-annotator agreement, or item-level quality checks. More critically, no text-only baseline is reported. Without a condition in which audio is withheld, lexical or surface cues in question stems and options could allow high accuracy without processing audio. Because the benchmark's entire purpose is to measure audio comprehension, this omission is load-bearing: the reported accuracies (e.g., 59.2% for Gemini 2.5 Flash) cannot be interpreted as audio-reasoning performance. Please provide a text-only baseline and show that audio is necessary for successful answering.
  2. [Abstract] The statement that audio is 'sourced directly from the wild' does not address training-data contamination. For 22 models trained on large-scale web data, clips from the wild may already be in pretraining corpora. Without a contamination analysis (e.g., audio fingerprint matching against training sets, or temporal cutoff checks), the reported accuracies could reflect memorization rather than reasoning. Add a contamination analysis and report results on an uncontaminated subset, if any.
  3. [Abstract] The reported accuracy numbers lack statistical support. The abstract claims models are 'approaching random performance in multiple categories,' but no confidence intervals, significance tests, or per-category sample sizes are given. With 5,305 total items, category-level subsets may be small, making differences between models or against chance potentially noisy. Report binomial or bootstrap confidence intervals and, where relevant, pairwise significance tests.
  4. [Abstract] The inclusion of 'open-ended response formats' raises scoring-validity concerns. The abstract does not state how open-ended answers are evaluated. If automatic metrics are used, they need validation against human judgment; if human graders are used, inter-rater reliability is required. Without this, the aggregate accuracy conflates answer-format difficulty with audio-reasoning ability. Describe the scoring protocol and its validation.
minor comments (3)
  1. [Abstract] The claim of being 'the most comprehensive and rigorously curated benchmark' should be supported with a side-by-side comparison to existing audio benchmarks (e.g., number of instances, skills, formats, audio sources).
  2. [Abstract] Please define 'multi-hop reasoning' operationally. As written it is a qualitative label with no criterion for what constitutes one hop versus multiple hops.
  3. [Abstract] The abstract refers to '49 unique skills' and 'multiple complex dimensions' without listing them. Provide the skill list in the abstract or explicitly reference an appendix/supplement.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: benchmark construction is not tautological.

full rationale

This is an abstract-only submission with no equations, fitted parameters, or derivation chain. The claimed contribution is the construction of a benchmark: 5,305 instances, human expert-generated QA pairs, 49 skills, and evaluation of 22 models. There is no step in which a quantity is defined in terms of another quantity and then rediscovered as a prediction. The human expert-generated questions and the skill taxonomy are inputs to the benchmark, not outputs of the model evaluations. The reported accuracies are measurements on an external test set, and the paper does not claim to derive them from the benchmark design. The only self-referential aspect is that the same authors define the skills and curate the questions, but this is a validity/quality concern (e.g., question ambiguity, lexical cues, contamination) rather than circular reasoning. No self-citation is present in the abstract, and no uniqueness theorem or fitted parameter is invoked. Therefore, no circular step can be exhibited, and the honest score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim that MMAU-Pro is comprehensive and rigorously curated rests on unverified assumptions about the completeness of the skill taxonomy, the quality of expert annotations, and the absence of data contamination. None of these are evidenced in the abstract.

assumptions (3)
  • domain assumption The 49-skill taxonomy fully covers the space of audio comprehension skills relevant to general intelligence.
    The abstract states the benchmark spans 49 unique skills; if this taxonomy is incomplete or arbitrary, the benchmark's comprehensiveness claim fails. No justification for the taxonomy is given in the abstract.
  • domain assumption Human expert-generated question-answer pairs are accurate, unambiguous, and discriminative.
    The abstract asserts questions are human expert-generated and require multi-hop reasoning; no annotation protocol, quality controls, or inter-annotator agreement is reported in the abstract.
  • domain assumption Audio sourced 'from the wild' is not present in training corpora of evaluated models.
    The abstract highlights wild audio to avoid known distributions; however, contamination can still occur and would inflate or deflate accuracy. No contamination check is described.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence." pith.science (2026). https://pith.science/paper/FSRTMI5U

@misc{pith2026250813992,
  author       = {Pith},
  title        = {Pith review of: MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FSRTMI5U}},
  note         = {Machine review of arXiv:2508.13992}
}
read the original abstract

Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models

    eess.AS 2026-07 conditional novelty 6.5 of 10

    Open-weight speech language models align poorly with human bouba/kiki judgments on real speech and fail crossmodal sound-to-shape matching, while their visual shape ratings remain near human ceiling.

  2. Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A self-play game with a known 'odd listener' converts unlabeled audio contrast pairs into a verifiable reward, improving fine-grained audio reasoning on TREA, MMAU, and MMAR.

  3. Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Using audio-aware LLMs to judge event presence and temporal order as DPO rewards improves multi-event text-to-audio instruction following.

  4. Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models

    eess.AS 2026-04 unverdicted novelty 6.0 of 10

    Self-distillation of Qwen3-Omni-Thinking on 545k Cogito-Pipe audio reasoning traces yields the best open-source MMAR CoT scores and top-tier challenge ranking.

  5. VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

    cs.SD 2025-10 conditional novelty 6.0 of 10

    A Chinese benchmark built on real human speech evaluates large audio language models across instruction following, knowledge, and robustness, revealing large performance gaps.

  6. Raon-Speech Technical Report

    cs.CL 2026-04 conditional novelty 5.5 of 10

    A 9B SpeechLM trained on 1.38M hours of English/Korean data plus a full-duplex chat extension trained on 119K hours of time-aligned dialogue outperform same-size audio models on speech tasks and FDB turn-taking metrics.

  7. ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

    cs.SD 2025-11 conditional novelty 5.0 of 10

    ORCA predicts the distribution of human correctness ratings for open-ended audio QA answers and matches or beats LLM judges while also estimating annotator disagreement.

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.