Pith. sign in

REVIEW 20 cited by

Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04675 v2 pith:UKU6VZHO submitted 2024-07-05 eess.AS cs.SD

classification eess.AScs.SD
keywords seed-asrspeechmodelslanguagemodelrecognitionscenariosaccents
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios. Classic end-to-end models fused with extra language models perform well, but mainly in data matching scenarios and are gradually approaching a bottleneck. In this work, we introduce Seed-ASR, a large language model (LLM) based speech recognition model. Seed-ASR is developed based on the framework of audio conditioned LLM (AcLLM), leveraging the capabilities of LLMs by inputting continuous speech representations together with contextual information into the LLM. Through stage-wise large-scale training and the elicitation of context-aware capabilities in LLM, Seed-ASR demonstrates significant improvement over end-to-end models on comprehensive evaluation sets, including multiple domains, accents/dialects and languages. Additionally, Seed-ASR can be further deployed to support specific needs in various scenarios without requiring extra language models. Compared to recently released large ASR models, Seed-ASR achieves 10%-40% reduction in word (or character, for Chinese) error rates on Chinese and English public test sets, further demonstrating its powerful performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

    cs.CV 2026-08 conditional novelty 7.0 of 10

    InteracVid delivers 454K livestream-derived context-query-response triplets, pairing real or LLM-reconstructed chat triggers with real audio-video reactions, and shows fine-tuning gains on genuine queries.

  2. Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A distilled multilingual ASR student, trained by on-policy distillation from language-specialized RL teachers, outperforms the best individual teacher and several larger open-source models.

  3. Context-Aware ASR for Mandarin Technical Lectures

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Self-built lecture glossaries from first-pass ASR raise technical-term recall across five backbones while holding or lowering CER on a new Mandarin AI/ML lecture benchmark.

  4. LLMs and Speech: Integration vs. Combination

    eess.AS 2026-03 unverdicted novelty 6.0 of 10

    With matched data and sizes, CTC+LLM shallow fusion beats tight speech-LLM integration on in-domain ASR, while prefix LLMs win average WER on out-of-domain HuggingFace sets.

  5. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  6. Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An end-to-end simultaneous speech-to-speech translation model with voice cloning, trained with a two-stage reinforcement learning reward scheme, reports high accuracy and low latency on the authors' RealSI benchmark.

  7. Improving Contextual ASR via Multi-grained Fusion with Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A multi-grained fusion method that jointly uses token-level and phrase-level scores from ASR and LLM improves keyword recognition in contextual ASR.

  8. CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models

    eess.AS 2025-05 conditional novelty 6.0 of 10

    An LLM-based ASR system that jointly performs overlapping-speech recognition and rare-word biasing, with a CTC-stage filter that trims large biasing lists, beats the tested baselines on LibriMix and AMI.

  9. Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR

    cs.SD 2025-05 conditional novelty 6.0 of 10

    EThai-ASR combines a self-refined Zipformer encoder with a Thai LLM and reports SOTA CER on Thai test sets plus a cosine-similarity frame pruning that gives 1.5-2.1x speedups in some modes.

  10. ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

    cs.SD 2026-07 reject novelty 5.0 of 10

    A 4B-parameter LLM ASR system with five multi-token-prediction branches reports 2.97% CER Chinese, 3.68% WER English, 3.70% long-form WER, and a 0.0053 real-time factor, but the acceptance-rate calculation and ablatio...

  11. Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Joint fine-tuning on seven dysarthric speakers' data reduced per-speaker character error rates by up to 13.15 percentage points compared to single-speaker fine-tuning on the CDSD corpus.

  12. Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems

    eess.AS 2025-08 conditional novelty 5.0 of 10

    Adding cross-utterance audio context to Conformer-Transducer ASR reduces WER/CER by 0.5 to 1.1 absolute points on four benchmarks, and a splicing-based batch scheme cuts training time by up to about 19%.

  13. From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.

  14. NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR

    eess.AS 2026-04 unverdicted novelty 4.0 of 10

    NIM4-ASR delivers SOTA ASR performance on public benchmarks using a 2.3B-parameter LLM with multi-stage training, real-time streaming, and million-scale hotword customization via RAG.

  15. Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition

    cs.SD 2025-09 reject novelty 4.0 of 10

    An LLM-based ASR error correction framework with noise-adaptive encoding and dynamic multi-modal fusion reports WER gains, but its fusion weights require ground-truth text at inference.

  16. SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding

    eess.AS 2025-07 reject novelty 4.0 of 10

    SpecASR accelerates LLM-based ASR by 3.04x-3.79x over autoregressive decoding using adaptive draft lengths, draft token recycling, and sparse token trees, but the speedups are simulated from Whisper proxy models rathe...

  17. The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Combining dual encoders, LID-routed MoE LoRA, and CTC prompts yields top challenge results for multilingual conversational ASR and speech diarization.

  18. PMF-CEC: Phoneme-augmented Multimodal Fusion for Context-aware ASR Error Correction with Error-specific Selective Decoding

    eess.AS 2025-05 conditional novelty 4.0 of 10

    PMF-CEC combines text and phoneme embeddings with a confidence-based edit filter, improving rare-word and homophone correction in ASR outputs across five benchmarks.

  19. Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Combining larger LLM decoders, preceding-text context, and iterative self-correction reduces character error rate on CNVSRC.Single Chinese visual speech recognition from 49.32% to 38.18%.

  20. Large Language models for Time Series Analysis: Techniques, Applications, and Challenges

    cs.LG 2025-05 reject novelty 3.0 of 10

    A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.

Pith tools