Pith. sign in

REVIEW 4 cited by

WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.03370 v5 pith:WECHKIVZ submitted 2021-10-07 cs.SD cs.CL

classification cs.SDcs.CL
keywords speechtesthoursrecognitionwenetspeechcandidatescorpusdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect the data from YouTube and Podcast, which covers a variety of speaking styles, scenarios, domains, topics, and noisy conditions. An optical character recognition (OCR) based method is introduced to generate the audio/text segmentation candidates for the YouTube data on its corresponding video captions, while a high-quality ASR transcription system is used to generate audio/text pair candidates for the Podcast data. Then we propose a novel end-to-end label error detection approach to further validate and filter the candidates. We also provide three manually labelled high-quality test sets along with WenetSpeech for evaluation -- Dev for cross-validation purpose in training, Test_Net, collected from Internet for matched test, and Test\_Meeting, recorded from real meetings for more challenging mismatched test. Baseline systems trained with WenetSpeech are provided for three popular speech recognition toolkits, namely Kaldi, ESPnet, and WeNet, and recognition results on the three test sets are also provided as benchmarks. To the best of our knowledge, WenetSpeech is the current largest open-sourced Mandarin speech corpus with transcriptions, which benefits research on production-level speech recognition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 32 citations worldwide. Full citation record

  1. Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Fine-tuned MMS outperforms XLS-R on fieldwork ASR with less than one hour of training data, while XLS-R reaches parity beyond one hour.

  2. Ming-Omni: A Unified Multimodal Model for Perception and Generation

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A single model with modality-specific routing processes image, text, audio, and video inputs and generates text, speech, and images, with public benchmarks reported across all of these abilities.

  3. Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

    cs.AI 2025-06 conditional novelty 5.0 of 10

    Stream-Omni uses CTC-based layer-dimension mapping to align speech with text, achieving vision, speech, and text interaction in one 8B model trained on 23,000 hours of speech.

  4. Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge

    cs.SD 2025-05 conditional novelty 5.0 of 10

    Training a neural speech enhancer on pseudo labels derived from aligned close-talk recordings improves far-field meeting ASR, reaching second place in the MISP-Meeting Challenge.

Pith tools