Pith. sign in

REVIEW 7 cited by

BAT: Learning to Reason about Spatial Sounds with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.01591 v3 pith:UV5AU5KY submitted 2024-02-02 eess.AS cs.AIcs.CLcs.SD

classification eess.AScs.AIcs.CLcs.SD
keywords spatialsoundaudiomodelreasoninglanguageperceptionsounds
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating Spatial-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT's superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Probing Spatial Structure in Pretrained Audio Representations

    cs.SD 2026-06 unverdicted novelty 7.0 of 10

    Introduces SARL benchmark showing pretrained audio encoders encode source-level spatial factors more readily than room-level factors, with patterns shaped by input configuration and training paradigm.

  2. SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

    cs.CV 2025-06 conditional novelty 7.0 of 10

    SAVVY-Bench tests audio-visual LLMs on dynamic 3D spatial questions, and the SAVVY pipeline, combining visual tracks with spatial audio and global mapping, lifts Gemini-2.5-pro accuracy from 50.9% to 58.0%.

  3. Dual-BEATs: Unlocking Zero-Shot Stereo Audio Perception in Audio Large Language Models via Dithering

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Uncorrelated dither noise lets dual frozen BEATs encoders preserve inter-channel amplitude differences across LLM normalizers, yielding up to 97% left/center/right accuracy and zero-shot spatial generalization.

  4. In-the-wild Audio Spatialization with Flexible Text-guided Localization

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.

  5. MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses

    eess.AS 2025-07 reject novelty 5.0 of 10

    MMW combines a Mamba-based Mix Block, a Frame Diarization Mamba layer, and multi-scale GRPO to reduce side-talk interference in Whisper ASR, reporting WER as low as 3.71%.

  6. Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A speech large language model trained on beamformed multi-channel audio performs directional speech recognition and source localization across 12 discrete angles on simulated smart glasses data.

  7. ASAudio: A Survey of Advanced Spatial Audio Research

    eess.AS 2025-08 unverdicted novelty 3.0 of 10

    A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.

Pith tools