Pith. sign in

Paper Citation Record · LEDGER

AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2410.18325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.18325 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:55.925858Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:49:30.450433Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4a6fa07c-a059-48cd-8ecf-06ad266367d1 · inbound

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs cites this paper.

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:12.105188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T16:57:12.821916Z digest=sha256:cb9556e32cb9b49b5bfd24e534a6044add6623345c96a611b7d52bd8554e2374

Observation de196700-7955-47f1-8588-7ca9e4639a40 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.288687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:0667c45e6df7de21ef29324e49a9d27c20fd0f65a4f6ad1fa455cd341a035f4e

Observation ec44756a-9734-4793-b05a-a27563163f71 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.925858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.925858Z digest=sha256:878b0344413bc453c91d2140b3a2b815159999cbd222a646e6161629bacda2c6

Observation c4f3000a-1d19-4b8c-a5ac-83e29c7b5dbf · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.606246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.606246Z digest=sha256:a1184ccc4c1b505c11e8723cf09a4d34c70f97b4ba54bfefed850458046712be

Observation 89fd54c6-55f7-4915-982e-afac26dd7711 · inbound

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction cites this paper.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.958481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.958481Z digest=sha256:48de95ebc7347469f8be2daa909fb1762ec4847815a428a338477e393950938b

Observation 4c463224-8582-4896-ad08-9f122ff28b64 · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:45:56.166828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:dba9a029b624f6e34b73010b70afad2ca7a28a930fd8d66de98ceb8a70d2e50f

Observation 3129e8e4-f223-4896-8319-e0b11ab759f3 · inbound

Do Audio-Visual Large Language Models Really See and Hear? cites this paper.

Do Audio-Visual Large Language Models Really See and Hear? AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.828018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:56:19.815569Z digest=sha256:ff7ce2fd7bf58edc91d22eedbb471835a2fd3b7af58ae52e818768a51e6f2519

Observation 2577d76c-02bb-4687-a202-108c5ef550f7 · inbound

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning cites this paper.

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:38:02.919961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T17:34:10.089555Z digest=sha256:981419ae191381c1cc859bd655eeb02cc1357024babe05e0b49ad1742e18beab

Observation b7ce29ad-8a3d-4a3d-9735-34d282e58704 · inbound

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models cites this paper.

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.999729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:57:47.356373Z digest=sha256:ba25d26e609bec9d5bfcc1079dc9342c593ada9509ff796bbc28f3722db31cb9

Observation e5ae92b7-c40a-42c6-8ebc-8910b5b50601 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.156140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:f8b9416c6320a3ea1ba72b30a22a33b1bf3416ea5de75577387901eb587c3ec3

Observation 025e6ffd-5637-4a80-b4cb-f288c2b5b0a0 · inbound

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models cites this paper.

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.746341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T07:02:02.752466Z digest=sha256:9689eb67ab246b09cc12f1b6bcd4b79809c2ebd000675c70264d89bf98acab08

Observation 0ae14557-2be6-4270-b536-9391d3b4ea03 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.478938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:159b4740203b81f17ad1cef3d5ef15d10121c57a97023386504e81af8979ef81

Observation 42dd175d-8b13-42fd-a150-27dbe2c49a27 · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.375774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:affc358f41b6e206cd12a3b4e91d8bdbd4d70ee096adf4ea3a0dc000364559e7

Observation 2953b6ce-48c4-4a9a-979c-03cf2709848c · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 122

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.453310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:67ada6160be71c14186547b282e75fbcaff3452bcaabed1de9d0794f4ba2a69f

Observation e7b61daa-5d64-4c78-b4c5-7e78914a2a8c · inbound

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation cites this paper.

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.904423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:15:47.056229Z digest=sha256:324ea76c9a5105ea21071a1d5492d7fe410929c471db9b56b80dae70c8c3dfe3

Observation e58dcf26-9c76-4ee4-8165-d7636a993709 · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.176193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T16:59:50.615875Z digest=sha256:2687ee6773012ab1fe4a527aacd39325d2b7eb972a48598080d60eb31ee9f858

Observation 81a73ab0-5f3d-4fcb-bb73-897a73c00d2e · inbound

Adaptive Perturbation Selection for Contrastive Audio Decoding cites this paper.

Adaptive Perturbation Selection for Contrastive Audio Decoding AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T04:34:51.366730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:34:51.366730Z digest=sha256:ec3492d3b08860f66f35ff8693121984bf258c0423e04da67a560a8dcbf0f714

Observation 4125ddc8-3dcc-4bea-9ca8-136e363c02fe · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.095101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.095101Z digest=sha256:2f2e1bfddcf757d5781bd5fdcc0adc9a6a813a99d35df6750ecf89b5f7b24c03