Pith. sign in

REVIEW 9 cited by

SpeechVerse: A Large-scale Generalizable Audio Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.08295 v3 pith:SJFBAEMT submitted 2024-05-14 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords tasksmodellanguagemodelsspeechspeechverseaudiobaselines
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have shown incredible proficiency in performing tasks that require semantic understanding of natural language instructions. Recently, many works have further expanded this capability to perceive multimodal audio and text inputs, but their capabilities are often limited to specific fine-tuned tasks such as automatic speech recognition and translation. We therefore develop SpeechVerse, a robust multi-task training and curriculum learning framework that combines pre-trained speech and text foundation models via a small set of learnable parameters, while keeping the pre-trained models frozen during training. The models are instruction finetuned using continuous latent representations extracted from the speech foundation model to achieve optimal zero-shot performance on a diverse range of speech processing tasks using natural language instructions. We perform extensive benchmarking that includes comparing our model performance against traditional baselines across several datasets and tasks. Furthermore, we evaluate the model's capability for generalized instruction following by testing on out-of-domain datasets, novel prompts, and unseen tasks. Our empirical experiments reveal that our multi-task SpeechVerse model is even superior to conventional task-specific baselines on 9 out of the 11 tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity

    cs.SD 2026-02 conditional novelty 6.0 of 10

    RA-QA converts 11 public respiratory-audio datasets into 9M template-generated QA pairs and shows current audio-language models score near zero on clinical task accuracy.

  2. EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

    cs.CV 2025-12 conditional novelty 6.0 of 10

    EchoingPixels prunes audio-visual LLM input tokens jointly across modalities and re-tunes RoPE frequencies so that 5–20% of tokens retain roughly full-model performance.

  3. Your Spending Needs Attention: Modeling Financial Habits with Transformers

    cs.IR 2025-07 conditional novelty 6.0 of 10

    A causal transformer pre-trained with next-token prediction on tokenized bank transactions, fused end-to-end with tabular features, lifts recommendation test AUC by 1.25% relative over a LightGBM baseline at Nubank.

  4. Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Adversarial audio noise, optimized with audio augmentations, can trigger and distort the behavior of audio-based LLMs both digitally and when played through the air.

  5. Unlocking Speech Instruction Data Potential with Query Rewriting

    cs.AI 2025-07 conditional novelty 5.0 of 10

    A multi-LLM rewriting and multi-agent validation pipeline makes text-to-speech synthesized speech instruction data far more usable and improves downstream speech instruction following.

  6. Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Multi-task behavior imitation with speech-text interleaving improves speech LLM generalization on prompts and zero-shot tasks using only paired speech and transcripts.

  7. SparQLe: Speech Queries to Text Translation Through LLMs

    cs.CL 2025-02 conditional novelty 5.0 of 10

    SparQLe uses a Q-Former adapter to bridge frozen HuBERT speech features to Llama-3, achieving BERTScore gains over IWSLT 2022 baselines on English-French and zero-shot English-German speech translation.

  8. TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation

    cs.CL 2025-08 conditional novelty 4.0 of 10

    Adding task-specific learned vectors to acoustic embeddings lets a transducer ASR model train on partially labeled data, matching or beating the fully labeled TokenVerse baseline on most tasks.

  9. EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition

    eess.AS 2025-08 reject novelty 4.0 of 10

    EmoSLLM, a LoRA-fine-tuned 3B LLM with an audio mapper, beats several 7B speech-text LLMs on MSP-Podcast emotion recognition, but only when given the ground-truth transcript at inference.

Pith tools