Pith. sign in

REVIEW 31 cited by

Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.01037 v3 pith:GYDIN6CX submitted 2023-03-02 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords modellanguagesspeechmultilingualrecognitionacrossautomaticdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the Universal Speech Model (USM), a single large model that performs automatic speech recognition (ASR) across 100+ languages. This is achieved by pre-training the encoder of the model on a large unlabeled multilingual dataset of 12 million (M) hours spanning over 300 languages, and fine-tuning on a smaller labeled dataset. We use multilingual pre-training with random-projection quantization and speech-text modality matching to achieve state-of-the-art performance on downstream multilingual ASR and speech-to-text translation tasks. We also demonstrate that despite using a labeled training set 1/7-th the size of that used for the Whisper model, our model exhibits comparable or better performance on both in-domain and out-of-domain speech recognition tasks across many languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 112 citations worldwide. Full citation record

  1. OLMoASR: Open Models and Data for Training Robust Speech Recognition Models

    cs.SD 2025-08 conditional novelty 7.0 of 10

    An open 1M-hour English speech dataset plus Whisper-architecture models trained on it match Whisper's word error rates on short and long-form benchmarks.

  2. Gemma 4 Technical Report

    cs.CL 2026-07 accept novelty 6.0 of 10

    Gemma 4 open multimodal models (dense + MoE) with thinking mode, encoder-free 12B path, and KV/memory optimizations leap prior Gemma and rival larger open models on STEM, multimodal, long-context, and Arena benchmarks.

  3. UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations

    eess.AS 2026-04 unverdicted novelty 6.0 of 10

    UniPASE extends the PASE framework with DeWavLM-Omni to convert degraded speech into high-fidelity, low-hallucination audio across sampling rates via phonetic enhancement, acoustic adaptation, and multi-rate vocoding.

  4. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  5. CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese

    cs.CL 2025-08 conditional novelty 6.0 of 10

    CAMOES is a new open benchmark and model collection for European Portuguese ASR, cutting word error rate by about 35% over the best zero-shot model.

  6. Geolocation-Aware Robust Spoken Language Identification

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Auxiliary geolocation prediction with injected conditioning signals improves dialect and accent robustness in SSL-based spoken language identification, reaching new SOTA on FLEURS and ML-SUPERB 2.0.

  7. LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness

    eess.AS 2025-08 conditional novelty 6.0 of 10

    LCS-CTC, a phoneme recognizer trained with similarity-aware LCS alignment masks constraining CTC, outperforms vanilla CTC on all reported PER, WPER, boundary-loss, and articulatory metrics.

  8. Identifying Hearing Difficulty Moments in Conversational Audio

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Prompted Gemini audio models detect hearing difficulty moments in conversation audio with F1 0.87, beating ASR hotword and Wav2Vec baselines.

  9. OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder

    cs.SD 2025-07 conditional novelty 6.0 of 10

    OpenBEATs releases the BEATs audio pretraining pipeline, trains 300M-parameter models on 20k hours of multi-domain audio, and reports strong results across 25 datasets, including bioacoustics and reasoning tasks.

  10. NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A real-world continual learning benchmark for multilingual ASR built from 3,250 hours of Indian language speech shows that no current CL method performs consistently across language- and domain-incremental scenarios.

  11. A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition

    cs.SD 2025-06 conditional novelty 6.0 of 10

    An iterative segmentation-based self-training method for Whisper improved long dysarthric speech recognition and achieved second place in both WER and SemScore at the SAP Challenge.

  12. Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A quality audit of Common Voice, FLEURS, and VoxPopuli finds serious micro- and macro-level data defects concentrated in less institutionalized languages, with a case study of Taiwanese Southern Min.

  13. Early Attentive Sparsification Accelerates Neural Speech Transcription

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Attention-based early audio-token sparsification at 40-60% sparsity accelerates Whisper ASR up to 1.6x with under 1% relative WER loss, across ten model variants, with no fine-tuning.

  14. OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Freezing OWSM v3.1 and adding dynamic-vocabulary biasing modules improves rare-word recognition and reduces real-time factor on LibriSpeech 100.

  15. GigaAM: Efficient Self-Supervised Learner for Speech Recognition

    eess.AS 2025-06 conditional novelty 6.0 of 10

    GigaAM, a Russian ASR model family pretrained with CTC-teacher cluster targets, reports roughly 50% lower WER than Whisper-large-v3 on three Russian benchmarks.

  16. NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    NTPP models dual-channel spoken dialogue by predicting both speakers' next speech tokens as a pair, achieving speaker-independent full-duplex generation in a decoder-only transformer.

  17. The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence

    cs.CL 2025-05 conditional novelty 6.0 of 10

    In 150k-hour speech-to-text training, a sub-exponential learning-rate warmup prevents divergence, while a faster warmup only speeds early convergence and does not improve the final model.

  18. OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A new open suite of 13 multilingual speech models, up to 18B parameters, yields empirical scaling laws for ASR and speech translation performance.

  19. A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition

    eess.AS 2026-03 conditional novelty 5.5 of 10

    On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.

  20. MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Grouping 495 languages into roughly 16 clusters and routing speech to group-specific LoRA experts improves multilingual ASR error rates over dense and random baselines.

  21. GigaAM Multilingual: Foundation Model for Underrepresented Languages

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Cluster-balanced HuBERT-style pre-training on 2M hours plus domain-aware fine-tuning yields a compact encoder that outperforms larger open multilingual ASR models on Kazakh, Kyrgyz and Uzbek.

  22. Towards Improved Speech Recognition through Optimized Synthetic Data Generation

    eess.AS 2025-08 conditional novelty 5.0 of 10

    A fine-tuned TTS with Whisper-based filtering generates synthetic Quebec French speech that trains ASR models well when combined with 10-60 hours of real audio, though a 13-14% WER gap to real-data training remains.

  23. Efficient Multilingual ASR Finetuning via LoRA Language Experts

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Monolingual LoRA language experts, combined by weighted merging (MoLE) or layer-wise knowledge distillation, improve Whisper-based multilingual ASR by about 10-15% relative WER over a plain multilingual LoRA baseline.

  24. A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Fine-tuning Whisper-large-v2 on 10,000 hours of synthesized Mandarin plus small real English/code-switching sets yields Twister, cutting mixed error rate by up to 56% on code-switching and 19% on Taiwanese Mandarin.

  25. OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

    cs.CL 2025-05 conditional novelty 5.0 of 10

    OWSM v4 models, trained on a cleaned 166k-hour multilingual YODAS subset, beat prior open OWSM models and are competitive with Whisper and MMS on several benchmarks.

  26. Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A Data2Vec2 speech encoder pre-trained on 300,000 hours of unlabeled Chinese dialect speech, connected to a small Qwen LLM via a linear projector and fine-tuned in four stages, sets a new state of the art on Chinese d...

  27. DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A single speech encoder trained via ASR-aware distillation with variable attention masking performs competitively in both streaming and full-context modes at 200M and 2B scale.

  28. VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A 68M-parameter Vietnamese ASR model, pretrained on 70,000 hours of unlabeled audio and fine-tuned on 50 hours of labels, reports average WER 8.31, beating Whisper Large-v3 and commercial systems.

  29. Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A spoken large language model learns to judge its own transcription difficulty and routes only hard speech to a stronger ASR model, cutting cost and improving word error rate.

  30. DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

    cs.CL 2026-07 reject novelty 4.0 of 10

    Open w2v-BERT ASR base models for 27 African languages, with a two-step annealing recipe and prefix-frame language conditioning.

  31. Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach

    eess.AS 2025-05 conditional novelty 4.0 of 10

    Llama-SMoP-DEDR, a sparse mixture of projectors with modality-specific experts and routers, lowers word error rate for LLM-based AVSR on LRS3, mainly with smaller LLMs.

Pith tools