Pith. sign in

REVIEW 1 cited by

AAD-LLM: Neural Attention-Driven Auditory Scene Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.16794 v3 pith:W3KMAIXQ submitted 2025-02-24 cs.SD cs.AIcs.CLcs.HCeess.AS

classification cs.SDcs.AIcs.CLcs.HCeess.AS
keywords auditoryaad-llmlistenermodelsperceptionspeakerattention-drivenfirst
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Auditory foundation models, including auditory large language models (LLMs), process all sound inputs equally, independent of listener perception. However, human auditory perception is inherently selective: listeners focus on specific speakers while ignoring others in complex auditory scenes. Existing models do not incorporate this selectivity, limiting their ability to generate perception-aligned responses. To address this, we introduce Intention-Informed Auditory Scene Understanding (II-ASU) and present Auditory Attention-Driven LLM (AAD-LLM), a prototype system that integrates brain signals to infer listener attention. AAD-LLM extends an auditory LLM by incorporating intracranial electroencephalography (iEEG) recordings to decode which speaker a listener is attending to and refine responses accordingly. The model first predicts the attended speaker from neural activity, then conditions response generation on this inferred attentional state. We evaluate AAD-LLM on speaker description, speech transcription and extraction, and question answering in multitalker scenarios, with both objective and subjective ratings showing improved alignment with listener intention. By taking a first step toward intention-aware auditory AI, this work explores a new paradigm where listener perception informs machine listening, paving the way for future listener-centered auditory systems. Demo and code available: https://aad-llm.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beamforming-LLM: What, Where and When Did I Miss?

    eess.AS 2025-09 conditional novelty 4.0 of 10

    A microphone array plus beamforming, Whisper transcription, FAISS retrieval, and GPT-4o-mini are combined to let users query what they missed in multi-speaker conversations.

Pith tools