Pith. sign in

REVIEW 3 cited by

Can Large Language Models Understand Spatial Audio?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07914 v2 pith:GL4NGUGE submitted 2024-06-12 cs.SD eess.AS

classification cs.SDeess.AS
keywords audiospatialllmsspeechcircenvironmentslanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This paper explores enabling large language models (LLMs) to understand spatial information from multichannel audio, a skill currently lacking in auditory LLMs. By leveraging LLMs' advanced cognitive and inferential abilities, the aim is to enhance understanding of 3D environments via audio. We study 3 spatial audio tasks: sound source localization (SSL), far-field speech recognition (FSR), and localisation-informed speech extraction (LSE), achieving notable progress in each task. For SSL, our approach achieves an MAE of $2.70^{\circ}$ on the Spatial LibriSpeech dataset, substantially surpassing the prior benchmark of about $6.60^{\circ}$. Moreover, our model can employ spatial cues to improve FSR accuracy and execute LSE by selectively attending to sounds originating from a specified direction via text prompts, even amidst overlapping speech. These findings highlight the potential of adapting LLMs to grasp physical audio concepts, paving the way for LLM-based agents in 3D environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dual-BEATs: Unlocking Zero-Shot Stereo Audio Perception in Audio Large Language Models via Dithering

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Uncorrelated dither noise lets dual frozen BEATs encoders preserve inter-channel amplitude differences across LLM normalizers, yielding up to 97% left/center/right accuracy and zero-shot spatial generalization.

  2. MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses

    eess.AS 2025-07 reject novelty 5.0 of 10

    MMW combines a Mamba-based Mix Block, a Frame Diarization Mamba layer, and multi-scale GRPO to reduce side-talk interference in Whisper ASR, reporting WER as low as 3.71%.

  3. Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A speech large language model trained on beamformed multi-channel audio performs directional speech recognition and source localization across 12 discrete angles on simulated smart glasses data.

Pith tools