Pith. sign in

Paper Citation Record · LEDGER

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2508.13992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13992 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:13:42.305221Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3c050704-71de-49be-ad09-3bd994b9a4db · inbound

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents cites this paper.

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:13:42.305221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:13:42.305221Z digest=sha256:bc719bc8a57f0fbcdf537de6135a1acb36518198a42a5b24ccba0c842c083d8b

Observation 630738de-3364-4576-8c49-4a3858100d7a · inbound

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering cites this paper.

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:23.429787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T19:38:23.429787Z digest=sha256:ce5d8ed6e265e2a4caa9c10a71b066a07ee7abe874e28d1d8e8c36c38455c829

Observation 62228609-ff36-4d97-b1fe-632e0f933498 · inbound

Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization cites this paper.

Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:50.371597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:41:23.645440Z digest=sha256:20b2ea30dad26e35c6d768b93360c2e8081171d5d2db5bc674dd0c0d884b14c1

Observation 2c85a7f3-9755-430e-a390-5a9abadc296b · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:11:01.470349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:46be4694a881b56fb4a8f0b096e0162af5dcd806fb8a0e8f78cd7f93939473d9

Observation f7cc0b70-022f-4ca1-a213-d6822ddf94b9 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.269718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:10:03.707886Z digest=sha256:19607b091f58a838dbea61d59243408377e8c686fb53fdc1b88b829a35565487

Observation 5d451620-3915-40e8-a784-61bf883bc12c · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T21:18:46.566338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:18:46.566338Z digest=sha256:31f9e9846febbdf176ddfc2e1a998a319d483da0878b86aec828e7d98078e48d

Observation 8eaed678-2693-4dd2-91fc-152624e51389 · inbound

VoxSafeBench: Not Just What Is Said, but Who, How, and Where cites this paper.

VoxSafeBench: Not Just What Is Said, but Who, How, and Where MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:22.023192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T10:19:28.041282Z digest=sha256:87d18f0e0b2548c1404a2929082b6a0746bbe2c2935ee69d68b2e4a4ebea6dc1

Observation a6acffad-05ba-4093-9044-286feb5b7a2b · inbound

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models cites this paper.

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:14.474889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T05:20:30.823304Z digest=sha256:ffe996c9b8c6b856089a31f10f4fd89bc2374c6a8b5667319b1b6995aa2d5738

Observation 2c35ff54-432d-418c-b216-dc49fbaf54a3 · inbound

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation cites this paper.

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:36.228584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T17:39:38.234052Z digest=sha256:b8eb849dd484b981cfe623f5216fd6f6c06810cc225fc799cf55f8cac3204dac

Observation 51ff07d7-eeee-47dc-ac61-93f1c282e891 · inbound

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models cites this paper.

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:41:26.485335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T14:29:18.348031Z digest=sha256:9a1d009f2190a1e30a8cae8fd50b40301abf1710ef403ea5aad23d363b5e3e20

Observation f82d1dc7-c589-4291-9b7c-7d3c7c483e3a · inbound

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio cites this paper.

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T20:17:04.813458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:15:54.497023Z digest=sha256:4b1eda8e7a342fa0d92d7bf102805fd3f606b990dfae96069c794e86c393d9ea

Observation d3485401-3041-438b-91fc-3b663a6dcae9 · inbound

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio cites this paper.

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:55:30.217657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:50:29.976958Z digest=sha256:87a4b819b53a91a00ae1ef8df5ca3ff27e06b8e618e7059744cfc85c371c2027

Observation 7bdcc89a-5dfc-409e-8b91-172ad3794571 · inbound

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB) cites this paper.

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB) MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:07.386028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:15:02.257046Z digest=sha256:21dfd2e7491f46c1e57c07b412e5c32a4120fa7a7e974fe9954aad960dfc4526

Observation df12a7d6-0bc4-40fa-9432-44a744bc39f7 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.882480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:5e72ffe99ee09f46682c1fab3ce241c0437c30a10a4521878d831f6f4caf03a6

Observation 0f4e6d22-4ce6-4acd-93b9-ee1fdeeaae7a · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:09:24.456660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:a660be11fa6da473f49f162381571bfc22a8aa295d2cc1c829ddc5effeaf4447

Observation 01b27112-de83-423d-b388-3ec646801966 · inbound

Raon-Speech Technical Report cites this paper.

Raon-Speech Technical Report MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T08:22:44.347941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:22:44.347941Z digest=sha256:ed76f92fb418eb75e5ad014e9c975db22399ea6d79198accb1be33e7c4aea92c

Observation 09ccbcf8-05eb-4c71-b591-bc88a95f6dd3 · inbound

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization cites this paper.

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.588897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:45:10.339950Z digest=sha256:591d60a5e493ecb3ddb846f5d2941b38e1d9ac045a7333b5c6a586d8d47a2b80

Observation 107c5050-9c44-47d3-9984-f432ec791da5 · inbound

Audio-Mind: An Auditable Agentic Framework for Audio Understanding cites this paper.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.567135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:536036965e0aa25bd95f921063ceee081d638124766f82300512fa5d14ca9ebd

Observation d7bcb2bb-69a6-48ba-85d5-6edbed107792 · inbound

MOSS-Audio Technical Report cites this paper.

MOSS-Audio Technical Report MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.050583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:5ea0044f1714cc05d90188457a31684ccd89d7bf5ea08b32a46f7f3db46ef05b

Observation b2a49e64-eb0b-47c9-a1cf-8e50004d6314 · inbound

VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track cites this paper.

VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:07:21.319265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:02:19.300441Z digest=sha256:ed83b9d4d77f97bc5e5c716041f9b77d3d0709102fd2a084f7f2add4a2613137

Observation c9bcd5b7-c3dc-42bb-bed5-7161c73f42ea · inbound

A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models cites this paper.

A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:39:01.662599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T23:25:46.349380Z digest=sha256:dbcd98262e1bfb607579d2a42e4e9c4a8e87612a7a0680823dcb8a3753e54d3a

Observation bdac4518-886a-45ec-b611-4374ad04782f · inbound

Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions cites this paper.

Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:00:00.285874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:18:55.383851Z digest=sha256:88783b2290e3d3ccd3b21ed353760661e4176ca071dc049f9ef731c47869b607

Observation bddd3957-e5ba-4efb-827e-f0128e321565 · inbound

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models cites this paper.

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:10:07.952661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T20:36:24.901454Z digest=sha256:f2ab42360137c7ad130f84a2a7b3d1e84c48ff2973576fa6bdb4da27524a760b

Observation 244e041f-d3f7-43e2-9324-10d11def72b9 · inbound

Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models cites this paper.

Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T13:49:26.320341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:49:26.320341Z digest=sha256:9c15b9431f25160333c4306ee22cde088581fa659767e888c95e2730cca608b5

Observation 6725413c-ccf2-44a3-92cf-440a8c763c73 · inbound

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models cites this paper.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.638568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.638568Z digest=sha256:7afacd8629c5990f952c3c620c35fae6a4cee43923e07c68dd8f51eb1a935ece

Observation db2dece3-a474-4384-933b-90fc9997a4e1 · inbound

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning cites this paper.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.893951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.893951Z digest=sha256:9bf2e98d2e4f3bfe01eabae55a27329dbb34acebfd6eef0cc917500729ea2142