REVIEW 2 cited by
Streaming End-to-End Multilingual Speech Recognition with Joint Language Identification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language identification is critical for many downstream tasks in automatic speech recognition (ASR), and is beneficial to integrate into multilingual end-to-end ASR as an additional task. In this paper, we propose to modify the structure of the cascaded-encoder-based recurrent neural network transducer (RNN-T) model by integrating a per-frame language identifier (LID) predictor. RNN-T with cascaded encoders can achieve streaming ASR with low latency using first-pass decoding with no right-context, and achieve lower word error rates (WERs) using second-pass decoding with longer right-context. By leveraging such differences in the right-contexts and a streaming implementation of statistics pooling, the proposed method can achieve accurate streaming LID prediction with little extra test-time cost. Experimental results on a voice search dataset with 9 language locales shows that the proposed method achieves an average of 96.2% LID prediction accuracy and the same second-pass WER as that obtained by including oracle LID in the input.
Forward citations
Cited by 2 Pith papers
-
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
Multilingual fine-tuning with Dutch and German data plus a language-identification token yields small word error rate gains for Frisian, while dialectal speech errors remain about twice as high as standard speech.
-
On the use of Performer and Agent Attention for Spoken Language Identification
Replacing standard self-attention with performer attention in the pooling layer of a language identification model improves average accuracy on VoxPopuli, FLEURS, and VoxLingua, while agent attention is comparable and...
Discussion (0). Continue with ORCID to comment.