Pith. sign in

REVIEW 11 cited by

GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11768 v1 pith:4UZILHES submitted 2024-06-17 cs.SD cs.AIcs.CLeess.AS

classification cs.SDcs.AIcs.CLeess.AS
keywords audiocomplexgamareasoningunderstandingabilitiesaudio-languagecapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Perceiving and understanding non-speech sounds and non-verbal speech is essential to making decisions that help us interact with our surroundings. In this paper, we propose GAMA, a novel General-purpose Large Audio-Language Model (LALM) with Advanced Audio Understanding and Complex Reasoning Abilities. We build GAMA by integrating an LLM with multiple types of audio representations, including features from a custom Audio Q-Former, a multi-layer aggregator that aggregates features from multiple layers of an audio encoder. We fine-tune GAMA on a large-scale audio-language dataset, which augments it with audio understanding capabilities. Next, we propose CompA-R (Instruction-Tuning for Complex Audio Reasoning), a synthetically generated instruction-tuning (IT) dataset with instructions that require the model to perform complex reasoning on the input audio. We instruction-tune GAMA with CompA-R to endow it with complex reasoning abilities, where we further add a soft prompt as input with high-level semantic evidence by leveraging event tags of the input audio. Finally, we also propose CompA-R-test, a human-labeled evaluation dataset for evaluating the capabilities of LALMs on open-ended audio question-answering that requires complex reasoning. Through automated and expert human evaluations, we show that GAMA outperforms all other LALMs in literature on diverse audio understanding tasks by margins of 1%-84%. Further, GAMA IT-ed on CompA-R proves to be superior in its complex reasoning and instruction following capabilities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

    cs.CL 2026-04 conditional novelty 6.5 of 10

    Current Omni LLMs extract modality-specific cues yet fail to integrate them for accurate multicontext safety judgments, performing better on physical than social/illegal risks.

  2. Empowering Long-form Omni-modal Understanding with Robust Audio Perception

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Decoupled audio-visual caption and CoT-QA datasets plus two-stage fine-tuning measurably strengthen auditory perception and cross-modal reasoning in a 7B omni-modal LLM.

  3. EvA: An Evidence-First Audio Understanding Paradigm for LALMs

    cs.SD 2026-03 conditional novelty 6.0 of 10

    Preserving multi-scale non-speech evidence via hierarchical aggregation and non-compressive time-aligned fusion measurably lifts LALM perception more than reasoning.

  4. The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents

    cs.CR 2025-07 unverdicted novelty 6.0 of 10

    A multi-agent audio-language model framework can automatically profile private attributes, such as age, health, and income, directly from general audio recordings.

  5. Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Adversarial audio noise, optimized with audio augmentations, can trigger and distort the behavior of audio-based LLMs both digitally and when played through the air.

  6. FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation

    eess.AS 2026-01 conditional novelty 5.0 of 10

    A hierarchical Q-Former compresses speech to about 1.67 tokens/sec, enabling hour-long audio processing with near-linear memory scaling and competitive benchmark scores.

  7. Improving Audio Event Recognition with Consistency Regularization

    cs.SD 2025-09 conditional novelty 5.0 of 10

    Consistency regularization improves audio event recognition on AudioSet by about 2 mAP, both in supervised and semi-supervised settings.

  8. From Sound to Sight: Towards AI-authored Music Videos

    cs.SD 2025-08 conditional novelty 5.0 of 10

    This paper presents two off-the-shelf model pipelines (CLAP or LALM, an LLM, and a text-to-video model) for generating music videos from arbitrary songs, validated by a preliminary five-participant user study with mod...

  9. MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses

    eess.AS 2025-07 reject novelty 5.0 of 10

    MMW combines a Mamba-based Mix Block, a Frame Diarization Mamba layer, and multi-scale GRPO to reduce side-talk interference in Whisper ASR, reporting WER as low as 3.71%.

  10. CoLMbo: Speaker Language Model for Descriptive Profiling

    cs.CL 2025-06 reject novelty 5.0 of 10

    CoLMbo pairs a fixed speaker encoder with a small language model to write descriptive profiles from voice, reporting high zero-shot accuracy for age, gender, ethnicity, and dialect.

  11. BoSS: Beyond-Semantic Speech

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Current spoken-language models perform poorly on a new five-task evaluation of beyond-semantic speech signals, including dialect, emotion, age, and non-verbal cues.

Pith tools