Pith. sign in

REVIEW 2 cited by

Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12611 v1 pith:C3SKFQXI submitted 2024-06-18 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords languagerecognitionspeechadaptationencoderlanguage-specificlanguagesmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end multilingual speech recognition models handle multiple languages through a single model, often incorporating language identification to automatically detect the language of incoming speech. Since the common scenario is where the language is already known, these models can perform as language-specific by using language information as prompts, which is particularly beneficial for attention-based encoder-decoder architectures. However, the Connectionist Temporal Classification (CTC) approach, which enhances recognition via joint decoding and multi-task training, does not normally incorporate language prompts due to its conditionally independent output tokens. To overcome this, we introduce an encoder prompting technique within the self-conditioned CTC framework, enabling language-specific adaptation of the CTC model in a zero-shot manner. Our method has shown to significantly reduce errors by 28% on average and by 41% on low-resource languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Multilingual fine-tuning with Dutch and German data plus a language-identification token yields small word error rate gains for Frisian, while dialectal speech errors remain about twice as high as standard speech.

  2. Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Whale, a 1.87B-parameter ASR model combining w2v-BERT and E-Branchformer, reports 2.4% WER on Librispeech test-clean and 3.4% CER on CSJ eval3, beating Whisper large-v3 and OWSM v3.1 on those benchmarks.

Pith tools