REVIEW 4 cited by
The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Code-switching (CS) is a common phenomenon and recognizing CS speech is challenging. But CS speech data is scarce and there' s no common testbed in relevant research. This paper describes the design and main outcomes of the ASRU 2019 Mandarin-English code-switching speech recognition challenge, which aims to improve the ASR performance in Mandarin-English code-switching situation. 500 hours Mandarin speech data and 240 hours Mandarin-English intra-sentencial CS data are released to the participants. Three tracks were set for advancing the AM and LM part in traditional DNN-HMM ASR system, as well as exploring the E2E models' performance. The paper then presents an overview of the results and system performance in the three tracks. It turns out that traditional ASR system benefits from pronunciation lexicon, CS text generating and data augmentation. In E2E track, however, the results highlight the importance of using language identification, building-up a rational set of modeling units and spec-augment. The other details in model training and method comparsion are discussed.
Forward citations
Cited by 4 Pith papers
-
Context-Aware ASR for Mandarin Technical Lectures
Self-built lecture glossaries from first-pass ASR raise technical-term recall across five backbones while holding or lowering CER on a new Mandarin AI/ML lecture benchmark.
-
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
An open, resource-lean speech understanding LLM trained on 50,500 hours matches or beats larger industry models on several Chinese benchmarks, with caveats in its internal evaluation.
-
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
A cross-attention based mixture-of-experts architecture with language bias from a language diarization decoder improves Mandarin-English code-switching ASR, achieving state-of-the-art results on SEAME, ASRU200, and AS...
-
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
Adding a language-identification block trained with non-peaky CTC and injecting the resulting language posteriors reduces mixed-error rate on Mandarin-English SEAME by about 0.5 to 0.8 percent absolute over the D-MoE ...
Discussion (0). Continue with ORCID to comment.