REVIEW 13 cited by
Recent Advances in Speech Language Models: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in text-based interactions. However, natural human interaction often relies on speech, necessitating a shift towards voice-based models. A straightforward approach to achieve this involves a pipeline of ``Automatic Speech Recognition (ASR) + LLM + Text-to-Speech (TTS)", where input speech is transcribed to text, processed by an LLM, and then converted back to speech. Despite being straightforward, this method suffers from inherent limitations, such as information loss during modality conversion, significant latency due to the complex pipeline, and error accumulation across the three stages. To address these issues, Speech Language Models (SpeechLMs) -- end-to-end models that generate speech without converting from text -- have emerged as a promising alternative. This survey paper provides the first comprehensive overview of recent methodologies for constructing SpeechLMs, detailing the key components of their architecture and the various training recipes integral to their development. Additionally, we systematically survey the various capabilities of SpeechLMs, categorize their evaluation metrics, and discuss the challenges and future research directions in this rapidly evolving field. The GitHub repository is available at https://github.com/dreamtheater123/Awesome-SpeechLM-Survey
Forward citations
Cited by 13 Pith papers
-
The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning
FLAIR enables spoken dialogue AI to conduct continuous latent reasoning while perceiving speech through recursive latent embeddings and an ELBO-based finetuning objective.
-
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
A new spoken math benchmark, Spoken-MQA, shows that current speech-based AI models reason poorly from spoken math input, especially for arithmetic and knowledge-heavy problems.
-
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
Speech DF Arena standardizes audio deepfake detection benchmarking across 14 datasets and 15 systems, showing that most open-source detectors have high error rates on out-of-domain attacks.
-
OpusLM: A Family of Open Unified Speech Language Models
A new open family of speech-language models trained on public data matches or beats prior systems across ASR, TTS, and text benchmarks.
-
Intelligibility of Text-to-Speech Systems for Mathematical Expressions
State-of-the-art text-to-speech models are often unintelligible when reading mathematical expressions aloud, with accuracy varying sharply by expression category and model.
-
Voice of a Continent: Mapping Africa's Speech Technology Frontier
A new benchmark and fine-tuned Simba models improve speech recognition, synthesis, and language identification across 61 African languages, but the claimed state of the art lacks comparisons to prior task-specific systems.
-
MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond
MEUSLI, a released family of linear projectors, enables open-source LLM-based ASR across 28 European languages and supports few-hour adaptation to unseen languages and extra speech tasks.
-
Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
Quantizing an intermediate layer of a pretrained audio model with residual vector quantization, and finetuning the model with task and codebook losses, preserves ASR and audio classification accuracy at bitrates near ...
-
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
Step-Audio-AQAA, a 130B end-to-end audio language model using dual-codebook tokens, text-audio interleaving, masked DPO and weight merging, is claimed to outperform Kimi-Audio and Qwen-Omni on the authors' StepEval-Au...
-
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
Fine-tuning an LLM on text instructions converted to Whisper semantic tokens enables it to understand spoken instructions at inference, bypassing TTS and speech instruction data.
-
Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts
AI voice interviewers are already fit for closed-ended surveys and partially for open-ended interviews, but transcription, emotion, and probing weaknesses limit qualitative use.
-
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
Under matched settings, continuous SSL speech features generally outperform discrete tokens on six spoken language understanding tasks in SpeechLLMs.
-
GR-LLMs: Recent Advances in Generative Recommendation Based on Large Language Models
A survey of LLM-based generative recommendation systems, covering application settings, training pipelines, industrial deployment challenges, and future directions.
Discussion (0). Continue with ORCID to comment.