Pith. sign in

REVIEW 1 cited by

Code Switched and Code Mixed Speech Recognition for Indic languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.16578 v2 pith:VX55CZDE submitted 2022-03-30 cs.CL eess.AS

classification cs.CLeess.AS
keywords multilingualcodelanguagesindiclanguagerecognitionspeechswitched
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training multilingual automatic speech recognition (ASR) systems is challenging because acoustic and lexical information is typically language specific. Training multilingual system for Indic languages is even more tougher due to lack of open source datasets and results on different approaches. We compare the performance of end to end multilingual speech recognition system to the performance of monolingual models conditioned on language identification (LID). The decoding information from a multilingual model is used for language identification and then combined with monolingual models to get an improvement of 50% WER across languages. We also propose a similar technique to solve the Code Switched problem and achieve a WER of 21.77 and 28.27 over Hindi-English and Bengali-English respectively. Our work talks on how transformer based ASR especially wav2vec 2.0 can be applied in developing multilingual ASR and code switched ASR for Indic languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

    cs.CL 2026-08 conditional novelty 7.0 of 10

    SurakshaEval, a 2,968-prompt safety benchmark in 10 Indic languages plus English, shows that 27 current LLMs pass safety checks far less often in Indic scripts than in English.

Pith tools