Pith. sign in

REVIEW 5 cited by

On the Audio Hallucinations in Large Audio-Video Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.09774 v1 pith:BTNKSU7F submitted 2024-01-18 cs.MM cs.CLcs.CVcs.SDeess.AS

classification cs.MMcs.CLcs.CVcs.SDeess.AS
keywords audiomodelsaudio-videohallucinationhallucinationslanguagelargezero-shot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large audio-video language models can generate descriptions for both video and audio. However, they sometimes ignore audio content, producing audio descriptions solely reliant on visual information. This paper refers to this as audio hallucinations and analyzes them in large audio-video language models. We gather 1,000 sentences by inquiring about audio information and annotate them whether they contain hallucinations. If a sentence is hallucinated, we also categorize the type of hallucination. The results reveal that 332 sentences are hallucinated with distinct trends observed in nouns and verbs for each hallucination type. Based on this, we tackle a task of audio hallucination classification using pre-trained audio-text models in the zero-shot and fine-tuning settings. Our experimental results reveal that the zero-shot models achieve higher performance (52.2% in F1) than the random (40.3%) and the fine-tuning models achieve 87.9%, outperforming the zero-shot models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Training three omnimodel speech systems with partial, audio-derived scaffold clues that are removed at test time cuts no-clue mpWER on overlapping noisy speech from 25–71% to 9–15%.

  2. Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Separating the text condition into a video caption and a visually grounded audio caption, and fusing the diffusion towers with dual cross-attention, gives the reported-best text-to-sounding-video quality and synchroni...

  3. ADIFF: Explaining audio difference using natural language

    cs.SD 2025-02 conditional novelty 6.0 of 10

    This paper proposes the audio difference explanation task, creates two LLM-generated datasets (ACD and CLD) with three explanation tiers, and presents ADIFF, a prefix-tuning model with cross-projection that beats base...

  4. X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

    cs.LG 2026-07 conditional novelty 5.0 of 10

    X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.

  5. Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding

    eess.AS 2025-06 conditional novelty 4.0 of 10

    Audio-Aware Decoding, a contrastive decoding method that uses silent audio as the no-context baseline, reduces object hallucination and improves accuracy across three large audio-language models.

Pith tools