Pith. sign in

REVIEW 2 cited by

Streaming End-to-End Multilingual Speech Recognition with Joint Language Identification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.06058 v1 pith:A52WL46V submitted 2022-09-13 eess.AS cs.CL

classification eess.AScs.CL
keywords languagestreamingachievedecodingend-to-endidentificationmethodmultilingual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language identification is critical for many downstream tasks in automatic speech recognition (ASR), and is beneficial to integrate into multilingual end-to-end ASR as an additional task. In this paper, we propose to modify the structure of the cascaded-encoder-based recurrent neural network transducer (RNN-T) model by integrating a per-frame language identifier (LID) predictor. RNN-T with cascaded encoders can achieve streaming ASR with low latency using first-pass decoding with no right-context, and achieve lower word error rates (WERs) using second-pass decoding with longer right-context. By leveraging such differences in the right-contexts and a streaming implementation of statistics pooling, the proposed method can achieve accurate streaming LID prediction with little extra test-time cost. Experimental results on a voice search dataset with 9 language locales shows that the proposed method achieves an average of 96.2% LID prediction accuracy and the same second-pass WER as that obtained by including oracle LID in the input.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Multilingual fine-tuning with Dutch and German data plus a language-identification token yields small word error rate gains for Frisian, while dialectal speech errors remain about twice as high as standard speech.

  2. On the use of Performer and Agent Attention for Spoken Language Identification

    eess.AS 2025-02 conditional novelty 4.0 of 10

    Replacing standard self-attention with performer attention in the pooling layer of a language identification model improves average accuracy on VoxPopuli, FLEURS, and VoxLingua, while agent attention is comparable and...

Pith tools