REVIEW 2 cited by
Code-Switching without Switching: Language Agnostic End-to-End Speech Translation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a) a Language Agnostic end-to-end Speech Translation model (LAST), and b) a data augmentation strategy to increase code-switching (CS) performance. With increasing globalization, multiple languages are increasingly used interchangeably during fluent speech. Such CS complicates traditional speech recognition and translation, as we must recognize which language was spoken first and then apply a language-dependent recognizer and subsequent translation component to generate the desired target language output. Such a pipeline introduces latency and errors. In this paper, we eliminate the need for that, by treating speech recognition and translation as one unified end-to-end speech translation problem. By training LAST with both input languages, we decode speech into one target language, regardless of the input language. LAST delivers comparable recognition and speech translation accuracy in monolingual usage, while reducing latency and error rate considerably when CS is observed.
Forward citations
Cited by 2 Pith papers
-
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
A survey of 110 SimulST papers shows most systems rely on unrealistic human pre-segmented audio and inconsistent terminology, and it offers a taxonomy and recommendations to fix both.
-
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
PIER is a WER variant restricted to tagged points of interest and is proposed as a more honest evaluation of code-switched ASR.
Discussion (0). Continue with ORCID to comment.