Pith. sign in

REVIEW 13 cited by

Recent Advances in Speech Language Models: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03751 v4 pith:XD7UAYH7 submitted 2024-10-01 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speechmodelslanguagespeechlmssurveycapabilitiesgithubpipeline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in text-based interactions. However, natural human interaction often relies on speech, necessitating a shift towards voice-based models. A straightforward approach to achieve this involves a pipeline of ``Automatic Speech Recognition (ASR) + LLM + Text-to-Speech (TTS)", where input speech is transcribed to text, processed by an LLM, and then converted back to speech. Despite being straightforward, this method suffers from inherent limitations, such as information loss during modality conversion, significant latency due to the complex pipeline, and error accumulation across the three stages. To address these issues, Speech Language Models (SpeechLMs) -- end-to-end models that generate speech without converting from text -- have emerged as a promising alternative. This survey paper provides the first comprehensive overview of recent methodologies for constructing SpeechLMs, detailing the key components of their architecture and the various training recipes integral to their development. Additionally, we systematically survey the various capabilities of SpeechLMs, categorize their evaluation metrics, and discuss the challenges and future research directions in this rapidly evolving field. The GitHub repository is available at https://github.com/dreamtheater123/Awesome-SpeechLM-Survey

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

    eess.AS 2026-03 unverdicted novelty 7.0 of 10

    FLAIR enables spoken dialogue AI to conduct continuous latent reasoning while perceiving speech through recursive latent embeddings and an ELBO-based finetuning objective.

  2. Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A new spoken math benchmark, Spoken-MQA, shows that current speech-based AI models reason poorly from spoken math input, especially for arithmetic and knowledge-heavy problems.

  3. Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models

    cs.SD 2025-09 conditional novelty 6.0 of 10

    Speech DF Arena standardizes audio deepfake detection benchmarking across 14 datasets and 15 systems, showing that most open-source detectors have high error rates on out-of-domain attacks.

  4. OpusLM: A Family of Open Unified Speech Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new open family of speech-language models trained on public data matches or beats prior systems across ASR, TTS, and text benchmarks.

  5. Intelligibility of Text-to-Speech Systems for Mathematical Expressions

    eess.AS 2025-06 conditional novelty 6.0 of 10

    State-of-the-art text-to-speech models are often unintelligible when reading mathematical expressions aloud, with accuracy varying sharply by expression category and model.

  6. Voice of a Continent: Mapping Africa's Speech Technology Frontier

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark and fine-tuned Simba models improve speech recognition, synthesis, and language identification across 61 African languages, but the claimed state of the art lacks comparisons to prior task-specific systems.

  7. MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

    cs.CL 2026-07 conditional novelty 5.0 of 10

    MEUSLI, a released family of linear projectors, enables open-source LLM-based ASR across 28 European languages and supports few-hour adaptation to unseen languages and extra speech tasks.

  8. Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Quantizing an intermediate layer of a pretrained audio model with residual vector quantization, and finetuning the model with task and codebook losses, preserves ASR and audio classification accuracy at bitrates near ...

  9. Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Step-Audio-AQAA, a 130B end-to-end audio language model using dual-codebook tokens, text-audio interleaving, masked DPO and weight merging, is claimed to outperform Kimi-Audio and Qwen-Omni on the authors' StepEval-Au...

  10. Speechless: Speech Instruction Training Without Speech for Low Resource Languages

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Fine-tuning an LLM on text instructions converted to Whisper semantic tokens enables it to understand spoken instructions at inference, bypassing TTS and speech instruction data.

  11. Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts

    cs.CL 2025-09 conditional novelty 4.0 of 10

    AI voice interviewers are already fit for closed-ended surveys and partially for open-ended interviews, but transcription, emotion, and probing weaknesses limit qualitative use.

  12. Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    Under matched settings, continuous SSL speech features generally outperform discrete tokens on six spoken language understanding tasks in SpeechLLMs.

  13. GR-LLMs: Recent Advances in Generative Recommendation Based on Large Language Models

    cs.IR 2025-07 unverdicted novelty 3.0 of 10

    A survey of LLM-based generative recommendation systems, covering application settings, training pipelines, industrial deployment challenges, and future directions.

Pith tools