Pith. sign in

Paper Citation Record · LEDGER

On the Audio Hallucinations in Large Audio-Video Language Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2401.09774.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.09774 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:40:26.080330Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T03:07:08.690461Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 33e0debe-e788-4feb-86ea-003d9f9f5388 · inbound

ADIFF: Explaining audio difference using natural language cites this paper.

ADIFF: Explaining audio difference using natural language On the Audio Hallucinations in Large Audio-Video Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.080330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.080330Z digest=sha256:ce147f245233d1017684313cf0975ac366e89ccfda228127957c07628f0e92b0

Observation 35dd95bc-b865-4e6f-820e-d04d7078e991 · inbound

Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding cites this paper.

Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding On the Audio Hallucinations in Large Audio-Video Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:44:10.480282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:44:10.480282Z digest=sha256:af036c9f62ebbebc6ab17b388c60dcf7f1d6e9e292a08bccd602a49edbb5bb60

Observation aa6800e8-532a-4740-b0e4-4bc2e4687a68 · inbound

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction cites this paper.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction On the Audio Hallucinations in Large Audio-Video Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.546588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.546588Z digest=sha256:ddd9c8147cc3b989c72a96e1f13a25ad5b7adea19b557198f4874896c85df9da

Observation a6d96bd4-9e8b-4978-b4ff-1a663c791d7a · inbound

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models cites this paper.

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models On the Audio Hallucinations in Large Audio-Video Language Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:29.040829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T13:57:47.356373Z digest=sha256:d0528edf9a6c8cb5c374160e83d500e11d5d3386bbb99ed37d1e5990a97b3cfe

Observation 7d356a5d-0561-4b7f-b9d4-8a4ab72afa21 · inbound

Probing Cross-modal Information Hubs in Audio-Visual LLMs cites this paper.

Probing Cross-modal Information Hubs in Audio-Visual LLMs On the Audio Hallucinations in Large Audio-Video Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:30.485437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:05:06.936268Z digest=sha256:d8f845cc4a250760c56c8ad755b88b89e266f3f3ff5ee916ae586d0e092dd1ab

Observation d5d3ea50-aaf8-4e5b-91ae-a3545a8b3d34 · inbound

Probing Cross-modal Information Hubs in Audio-Visual LLMs cites this paper.

Probing Cross-modal Information Hubs in Audio-Visual LLMs On the Audio Hallucinations in Large Audio-Video Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:08.692174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T03:04:56.762735Z digest=sha256:cf8066cb9734bfb081a653cfafd5dd056f0b21f0892b92fe9eb6e0bd9c0f4bba

Observation 3ade6424-1a5d-4c06-a5bf-5b2ab43eb11f · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment On the Audio Hallucinations in Large Audio-Video Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.375070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.375070Z digest=sha256:f25107bd5c46c333aefbe2938c0b6bd2567333a9613faa21bfbd5235cfe782bc

Observation f9308c80-372d-436c-85da-83b247d367a0 · inbound

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding cites this paper.

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding On the Audio Hallucinations in Large Audio-Video Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.544745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.544745Z digest=sha256:5e2c9636e1e632aa68410d36d75ac4b63328c99e9d0b1b5d6537b1c2364c7e9b