Pith. sign in

Paper Citation Record · LEDGER

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2412.02611.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02611 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:56.875342Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:49:30.296211Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c0de47b9-1277-447a-b849-99fe422276a8 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.417668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:a4ac06c490d363731e4a5f6c8df39e296d2614ce957455330b1fbc9c505c8357

Observation 141c6743-16af-4f7a-bccc-bcbb9b7fe1d3 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 260

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.590020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:0705bfa7cfba0dd00db6d08582244fa0ec2e5cdb019bdaea9515ef5a8095cf2d

Observation added75e-72c0-4bd8-8f88-2349c968ab3d · inbound

Qwen2.5-Omni Technical Report cites this paper.

Qwen2.5-Omni Technical Report AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:54:03.340882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:54:03.225439Z digest=sha256:54e0fee993311acba479d749d0c3d4b10fefa0a59fa6cd2fcda6f5781587eade

Observation 75bd58d3-d36d-490a-9111-a79eda5b8246 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:56.875342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:56.875342Z digest=sha256:6e6b0a044a41f120614145f852ccb3c59bd498667bd379838ff20546153cd1ff

Observation f714a50a-37d6-47ff-820f-dc548c484757 · inbound

Learning Sparsity for Effective and Efficient Music Performance Question Answering cites this paper.

Learning Sparsity for Effective and Efficient Music Performance Question Answering AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:41.293828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:50:41.293828Z digest=sha256:4c3a3dc2e74d2c42825b8291acd9debe7d0821fba0de5e96c95b6bcbdde42ca3

Observation 484f6ce8-fdd9-4447-b52b-7e1f6f78b46d · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.674331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.674331Z digest=sha256:29dad97defc0e31357b13510f0fcbd10d070ae8866c04c73c87bf86a5dc62293

Observation 20fc6335-e5a8-4b6b-95a5-7aac665adb73 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.597241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.597241Z digest=sha256:57dcf4c91bc20e95a987b9df92420a3fa98352c9ee0894b60b37450cdceff719

Observation bfc23237-ecb4-4048-85aa-77b5b166f59c · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:45:56.100282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:e7daa7a4e41c20cf436cf7a99c8e19da3f55bc563ec3cadec847b9dc2f889fae

Observation 7e9fd2fa-bc3b-44c4-aef5-f66678c16da2 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.046832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:0dd5297f8b406611e65ac09af0b1afb6ae7d510a203d2f2c37f423161a68c50b

Observation 9bff972c-7da0-4f12-9ad8-98248f38af4e · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.354000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:aefc2a3d6bbf0a301454f45dc30bde74f48e4e361c42dcda45626a7639102c30

Observation a5b971d3-a744-429a-acc7-eb4a60f1a7a9 · inbound

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search cites this paper.

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:27.463941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:21:52.972107Z digest=sha256:573e1f5e44a04e8d2105b26490f5954f39d256d6d0b7fde38759306461ddbf03

Observation 6c7a6a4b-b71b-4f7c-95a7-53a6023051c7 · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.681606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T03:49:58.240883Z digest=sha256:c3801e9fa3a412c09c34c9f878ab8aef61e877d1deda2740537b0b98efb5272e

Observation 6add56f3-eb07-4c52-8e7e-da8f291eee42 · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:09:50.271349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T06:06:20.658030Z digest=sha256:53d9c481e2b0d694d9e7fbd1bec0655d50b3ca5b7195694cfd57650ba9a8be79

Observation 3362eafb-b804-4ffc-9e0c-5044270b27b0 · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:24.991923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:1266db2e3e4271b481542795320ab330e2584f39d138235a87fc40d67c18755a

Observation 9b30d95c-6c4d-472c-819b-5bca454f3202 · inbound

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models cites this paper.

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:46:15.354338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T16:24:43.104448Z digest=sha256:4ef6421e2d798b54087af41aa720a2dc674c39d111f9b021f2bbfa95c050dfa4

Observation 596a8cd0-8b46-4d82-9f49-2728b5ab5c91 · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.319448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:17bc828652f8c5347061b5778fb7c4669daa09e4c2a2c22e578c08cde210ac82

Observation 894cfb0a-48d4-4c8e-82ad-491bcf5ef139 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:30.298437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:92bebe4dd622c4d91ebe1fd085105162c47361afab140771ed9d749734ca7460

Observation 966b6e6d-3c38-41f6-add6-84fb95b73292 · inbound

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs cites this paper.

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:34:19.509749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:25:38.593423Z digest=sha256:32eda9f57544a2fe601ab5d7990553baf7909ffca80090c1f936a95c535ff504