Pith. sign in

REVIEW 7 cited by

GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11546 v2 pith:HXWMX2MY submitted 2024-06-17 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords speechgigaspeechcorpusdatalow-resourcelanguagesmodelspipeline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The evolution of speech technology has been spurred by the rapid increase in dataset sizes. Traditional speech models generally depend on a large amount of labeled training data, which is scarce for low-resource languages. This paper presents GigaSpeech 2, a large-scale, multi-domain, multilingual speech recognition corpus. It is designed for low-resource languages and does not rely on paired speech and text data. GigaSpeech 2 comprises about 30,000 hours of automatically transcribed speech, including Thai, Indonesian, and Vietnamese, gathered from unlabeled YouTube videos. We also introduce an automated pipeline for data crawling, transcription, and label refinement. Specifically, this pipeline involves Whisper for initial transcription, MMS for forced alignment, and multi-dimensional filtering for data quality assurance. A modified Noisy Student Training is developed to further refine flawed pseudo labels iteratively, thereby enhancing model performance. Experimental results on our manually transcribed evaluation set and two public test sets from Common Voice and FLEURS confirm our corpus's high quality and broad applicability. Notably, ASR models trained on GigaSpeech 2 can reduce the word error rate for Thai, Indonesian, and Vietnamese on our challenging and realistic YouTube test set by 25% to 40% compared to Whisper large-v3, with merely 10% model parameters. Furthermore, our ASR models trained on GigaSpeech 2 yield superior performance compared to commercial services. We hope that our newly introduced corpus and pipeline will open a new avenue for low-resource speech recognition and significantly facilitate research in this area.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    The authors built and released the largest open-source Cantonese speech corpus (21,800 hours, 10 domains, rich metadata), and show that models trained on it match or beat existing speech recognition and synthesis systems.

  2. Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR

    cs.SD 2025-05 conditional novelty 6.0 of 10

    EThai-ASR combines a self-refined Zipformer encoder with a Thai LLM and reports SOTA CER on Thai test sets plus a cosine-similarity frame pruning that gives 1.5-2.1x speedups in some modes.

  3. OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

    cs.CL 2025-05 conditional novelty 5.0 of 10

    OWSM v4 models, trained on a cleaned 166k-hour multilingual YODAS subset, beat prior open OWSM models and are competitive with Whisper and MMS on several benchmarks.

  4. VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A 68M-parameter Vietnamese ASR model, pretrained on 70,000 hours of unlabeled audio and fine-tuned on 50 hours of labels, reports average WER 8.31, beating Whisper Large-v3 and commercial systems.

  5. The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Combining dual encoders, LID-routed MoE LoRA, and CTC prompts yields top challenge results for multilingual conversational ASR and speech diarization.

  6. ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A three-stage iterative LoRA training recipe (Focus, Feed Back, Fix) is applied to Whisper-large-v3 and Qwen2-Audio, reporting WER reductions on a multilingual ASR benchmark, with the gains attributed to the iterative...

  7. Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge

    eess.AS 2025-08 conditional novelty 2.0 of 10

    A frozen Whisper-large-v3 encoder plus a trainable adaptor plus LoRA-adapted Qwen2.5-7B achieves 9.83% WER/CER on 11-language conversational ASR and third place in MLC-SLM 2025 Track 1.

Pith tools