REVIEW 13 cited by
Unsupervised Cross-lingual Representation Learning for Speech Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages. We build on wav2vec 2.0 which is trained by solving a contrastive task over masked latent speech representations and jointly learns a quantization of the latents shared across languages. The resulting model is fine-tuned on labeled data and experiments show that cross-lingual pretraining significantly outperforms monolingual pretraining. On the CommonVoice benchmark, XLSR shows a relative phoneme error rate reduction of 72% compared to the best known results. On BABEL, our approach improves word error rate by 16% relative compared to a comparable system. Our approach enables a single multilingual speech recognition model which is competitive to strong individual models. Analysis shows that the latent discrete speech representations are shared across languages with increased sharing for related languages. We hope to catalyze research in low-resource speech understanding by releasing XLSR-53, a large model pretrained in 53 languages.
Forward citations
Cited by 13 Pith papers
-
Less is More: Modality-Decoupling for General AIGC Audio-Video Detection
A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.
-
Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR
Assamese and Hindi, selected by acoustic and typological similarity to Warlpiri, cut Whisper WER/CER most; acoustic similarity best predicts fine-tuning gains, inventory/typology zero-shot.
-
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.
-
Prominence-aware automatic speech recognition for conversational speech
A wav2vec2 speech recognizer was extended with word-level prominence labels, yielding simultaneous transcription and prominence annotation of conversational Austrian German with accuracy near the detector-only system.
-
Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models
A hybrid decoding scheme with a fast TDT draft decoder and selective transformer patches matches baseline word error rate while cutting decoder latency roughly threefold.
-
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning
A quality audit of Common Voice, FLEURS, and VoxPopuli finds serious micro- and macro-level data defects concentrated in less institutionalized languages, with a case study of Taiwanese Southern Min.
-
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
Vedavani is a new 54-hour benchmark for Vedic Sanskrit poetry ASR, and fine-tuned IndicWhisper reaches about 22% word error rate on the Devanagari test set.
-
ZIPA: A family of efficient models for multilingual phone recognition
ZIPA models, trained with Zipformer backbones on a new 17,132-hour IPA-labeled corpus, achieve state-of-the-art multilingual phone recognition with fewer parameters than prior systems.
-
Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems
A new evaluation protocol exhaustively pairs 164 speech synthesizers with nine bona fide speech types and reports max-pooled EERs, revealing larger failures than pooled averages show.
-
Leveraging Large Language Models for Spontaneous Speech-Based Suicide Risk Detection
A multimodal system combining acoustic, textual, and LLM-extracted features reports 74% accuracy for adolescent suicide risk detection on the SW1 challenge test set.
-
UD-KSL Treebank v1.3: A semi-automated framework for aligning XPOS-extracted units with UPOS tags
A semi-automated XPOS-to-UPOS alignment pipeline for the UD-KSL treebank is reported to improve tagging and some parsing accuracy, but the comparison is confounded because the test gold labels also change between conditions.
-
SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
The authors present SwitchLingua, a large multilingual code-switching text and audio dataset, and SAER, a semantic-aware error metric for code-switching ASR evaluation.
-
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
Adding task-specific learned vectors to acoustic embeddings lets a transducer ASR model train on partially labeled data, matching or beating the fully labeled TokenVerse baseline on most tasks.
Discussion (0). Continue with ORCID to comment.