REVIEW 20 cited by
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios. Classic end-to-end models fused with extra language models perform well, but mainly in data matching scenarios and are gradually approaching a bottleneck. In this work, we introduce Seed-ASR, a large language model (LLM) based speech recognition model. Seed-ASR is developed based on the framework of audio conditioned LLM (AcLLM), leveraging the capabilities of LLMs by inputting continuous speech representations together with contextual information into the LLM. Through stage-wise large-scale training and the elicitation of context-aware capabilities in LLM, Seed-ASR demonstrates significant improvement over end-to-end models on comprehensive evaluation sets, including multiple domains, accents/dialects and languages. Additionally, Seed-ASR can be further deployed to support specific needs in various scenarios without requiring extra language models. Compared to recently released large ASR models, Seed-ASR achieves 10%-40% reduction in word (or character, for Chinese) error rates on Chinese and English public test sets, further demonstrating its powerful performance.
Forward citations
Cited by 20 Pith papers
-
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos
InteracVid delivers 454K livestream-derived context-query-response triplets, pairing real or LLM-reconstructed chat triggers with real audio-video reactions, and shows fine-tuning gains on genuine queries.
-
Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR
A distilled multilingual ASR student, trained by on-policy distillation from language-specialized RL teachers, outperforms the best individual teacher and several larger open-source models.
-
Context-Aware ASR for Mandarin Technical Lectures
Self-built lecture glossaries from first-pass ASR raise technical-term recall across five backbones while holding or lowering CER on a new Mandarin AI/ML lecture benchmark.
-
LLMs and Speech: Integration vs. Combination
With matched data and sizes, CTC+LLM shallow fusion beats tight speech-LLM integration on in-domain ASR, while prefix LLMs win average WER on out-of-domain HuggingFace sets.
-
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.
-
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
An end-to-end simultaneous speech-to-speech translation model with voice cloning, trained with a two-stage reinforcement learning reward scheme, reports high accuracy and low latency on the authors' RealSI benchmark.
-
Improving Contextual ASR via Multi-grained Fusion with Large Language Models
A multi-grained fusion method that jointly uses token-level and phrase-level scores from ASR and LLM improves keyword recognition in contextual ASR.
-
CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models
An LLM-based ASR system that jointly performs overlapping-speech recognition and rare-word biasing, with a CTC-stage filter that trims large biasing lists, beats the tested baselines on LibriMix and AMI.
-
Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR
EThai-ASR combines a self-refined Zipformer encoder with a Thai LLM and reports SOTA CER on Thai test sets plus a cosine-similarity frame pruning that gives 1.5-2.1x speedups in some modes.
-
ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition
A 4B-parameter LLM ASR system with five multi-token-prediction branches reports 2.97% CER Chinese, 3.68% WER English, 3.70% long-form WER, and a 0.0053 real-time factor, but the acceptance-rate calculation and ablatio...
-
Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database
Joint fine-tuning on seven dysarthric speakers' data reduced per-speaker character error rates by up to 13.15 percentage points compared to single-speaker fine-tuning on the CDSD corpus.
-
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
Adding cross-utterance audio context to Conformer-Transducer ASR reduces WER/CER by 0.5 to 1.1 absolute points on four benchmarks, and a splicing-based batch scheme cuts training time by up to about 19%.
-
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.
-
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
NIM4-ASR delivers SOTA ASR performance on public benchmarks using a 2.3B-parameter LLM with multi-stage training, real-time streaming, and million-scale hotword customization via RAG.
-
Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition
An LLM-based ASR error correction framework with noise-adaptive encoding and dynamic multi-modal fusion reports WER gains, but its fusion weights require ground-truth text at inference.
-
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
SpecASR accelerates LLM-based ASR by 3.04x-3.79x over autoregressive decoding using adaptive draft lengths, draft token recycling, and sparse token trees, but the speedups are simulated from Whisper proxy models rathe...
-
The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge
Combining dual encoders, LID-routed MoE LoRA, and CTC prompts yields top challenge results for multilingual conversational ASR and speech diarization.
-
PMF-CEC: Phoneme-augmented Multimodal Fusion for Context-aware ASR Error Correction with Error-specific Selective Decoding
PMF-CEC combines text and phoneme embeddings with a confidence-based edit filter, improving rare-word and homophone correction in ASR outputs across five benchmarks.
-
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
Combining larger LLM decoders, preceding-text context, and iterative self-correction reduces character error rate on CNVSRC.Single Chinese visual speech recognition from 49.32% to 38.18%.
-
Large Language models for Time Series Analysis: Techniques, Applications, and Challenges
A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.
Discussion (0). Continue with ORCID to comment.