Pith. sign in

REVIEW 4 cited by

Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08402 v1 pith:2DB2NHLM submitted 2024-06-12 eess.AS cs.CLcs.LGcs.SD

classification eess.AScs.CLcs.LGcs.SD
keywords lalmsaudiomodelsobjectdiscriminativehallucinationlargequestions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large audio-language models (LALMs) enhance traditional large language models by integrating audio perception capabilities, allowing them to tackle audio-related tasks. Previous research has primarily focused on assessing the performance of LALMs across various tasks, yet overlooking their reliability, particularly concerning issues like object hallucination. In our study, we introduce methods to assess the extent of object hallucination of publicly available LALMs. Our findings reveal that LALMs are comparable to specialized audio captioning models in their understanding of audio content, but struggle to answer discriminative questions, specifically those requiring the identification of the presence of particular object sounds within an audio clip. This limitation highlights a critical weakness in current LALMs: their inadequate understanding of discriminative queries. Moreover, we explore the potential of prompt engineering to enhance LALMs' performance on discriminative questions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive Perturbation Selection for Contrastive Audio Decoding

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    A learned per-example router over a 105-perturbation audio library improves contrastive decoding for audio-LLM hallucination, with task-dependent best distortions (e.g., reverse audio for temporal order).

  2. Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Separating the text condition into a video caption and a visually grounded audio caption, and fusing the diffusion towers with dual cross-attention, gives the reported-best text-to-sounding-video quality and synchroni...

  3. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

  4. CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    CCHall is the first benchmark that jointly evaluates cross-lingual and cross-modal hallucination detection in multimodal LLMs, and current models perform poorly on it.

Pith tools