Pith. sign in

REVIEW 12 cited by

The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.09344 v1 pith:N7Y5CIIT submitted 2021-11-17 cs.LG stat.ML

classification cs.LGstat.ML
keywords dataspeechdatasetundercollectioncommercialenglishlicensed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The People's Speech is a free-to-download 30,000-hour and growing supervised conversational English speech recognition dataset licensed for academic and commercial usage under CC-BY-SA (with a CC-BY subset). The data is collected via searching the Internet for appropriately licensed audio data with existing transcriptions. We describe our data collection methodology and release our data collection system under the Apache 2.0 license. We show that a model trained on this dataset achieves a 9.98% word error rate on Librispeech's test-clean test set.Finally, we discuss the legal and ethical issues surrounding the creation of a sizable machine learning corpora and plans for continued maintenance of the project under MLCommons's sponsorship.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Swivuriso: The South African Next Voices Multilingual Speech Dataset

    cs.CL 2025-12 conditional novelty 7.0 of 10

    Swivuriso provides a 3,000-hour, seven-language, scripted-and-unscripted South African speech corpus with domain coverage in agriculture, healthcare, and general topics.

  2. Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Two small modules enable continuous latent test-time refinement on a frozen ASR backbone, cutting error on hard speech under a 500-utterance regime where fine-tuning, LoRA and prompt tuning all regress.

  3. TTS-1 Technical Report

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new TTS system family combines pre-training, supervised fine-tuning, and GRPO reinforcement learning with a 48 kHz codec to produce multilingual speech with in-context voice cloning.

  4. Ming-Omni: A Unified Multimodal Model for Perception and Generation

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A single model with modality-specific routing processes image, text, audio, and video inputs and generates text, speech, and images, with public benchmarks reported across all of these abilities.

  5. An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Smaller discrete vocabularies (k = 125 to 1,000), WavLM units, and larger models give the lowest negative log-likelihood in speech language model pre-training.

  6. Group Relative Policy Optimization for Speech Recognition

    eess.AS 2025-09 conditional novelty 5.0 of 10

    Applying GRPO with rule-based rewards to LLM-based ASR improves WER by up to 18.4% relative and reduces hallucination errors on unseen acoustic conditions.

  7. OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

    cs.CL 2025-05 conditional novelty 5.0 of 10

    OWSM v4 models, trained on a cleaned 166k-hour multilingual YODAS subset, beat prior open OWSM models and are competitive with Whisper and MMS on several benchmarks.

  8. Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use

    cs.CL 2025-05 conditional novelty 5.0 of 10

    The Loquacious Set is a curated 25,000-hour English ASR corpus combining six open datasets, with commercial-ready licenses and conformer baselines that reach 4.6% WER on LibriSpeech test-other.

  9. DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A single speech encoder trained via ASR-aware distillation with variable attention masking performs competitively in both streaming and full-context modes at 200M and 2B scale.

  10. Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit

    cs.CL 2025-06 conditional novelty 4.0 of 10

    LoRA fine-tuning plus custom text normalization reduces Whisper word error rate on cockpit pilot speech from 68.49% to 26.26%.

  11. Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning

    cs.CL 2025-06 conditional novelty 4.0 of 10

    The IT-IST IWSLT 2025 submission shows that a 1.5B language model with a speech encoder can do reasonable ASR after alignment, but struggles with ST and SQA.

  12. Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Whale, a 1.87B-parameter ASR model combining w2v-BERT and E-Branchformer, reports 2.4% WER on Librispeech test-clean and 3.4% CER on CSJ eval3, beating Whisper large-v3 and OWSM v3.1 on those benchmarks.

Pith tools