Pith. sign in

REVIEW 13 cited by

Unsupervised Cross-lingual Representation Learning for Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.13979 v2 pith:KJKTX77O submitted 2020-06-24 cs.CL cs.LGcs.SDeess.AS

classification cs.CLcs.LGcs.SDeess.AS
keywords speechlanguagesmodelcross-lingualpretrainingrepresentationsacrossapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages. We build on wav2vec 2.0 which is trained by solving a contrastive task over masked latent speech representations and jointly learns a quantization of the latents shared across languages. The resulting model is fine-tuned on labeled data and experiments show that cross-lingual pretraining significantly outperforms monolingual pretraining. On the CommonVoice benchmark, XLSR shows a relative phoneme error rate reduction of 72% compared to the best known results. On BABEL, our approach improves word error rate by 16% relative compared to a comparable system. Our approach enables a single multilingual speech recognition model which is competitive to strong individual models. Analysis shows that the latent discrete speech representations are shared across languages with increased sharing for related languages. We hope to catalyze research in low-resource speech understanding by releasing XLSR-53, a large model pretrained in 53 languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.

  2. Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Assamese and Hindi, selected by acoustic and typological similarity to Warlpiri, cut Whisper WER/CER most; acoustic similarity best predicts fine-tuning gains, inventory/typology zero-shot.

  3. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  4. Prominence-aware automatic speech recognition for conversational speech

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A wav2vec2 speech recognizer was extended with word-level prominence labels, yielding simultaneous transcription and prominence annotation of conversational Austrian German with accuracy near the detector-only system.

  5. Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models

    eess.AS 2025-08 conditional novelty 6.0 of 10

    A hybrid decoding scheme with a fast TDT draft decoder and selective transformer patches matches baseline word error rate while cutting decoder latency roughly threefold.

  6. Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A quality audit of Common Voice, FLEURS, and VoxPopuli finds serious micro- and macro-level data defects concentrated in less institutionalized languages, with a case study of Taiwanese Southern Min.

  7. Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Vedavani is a new 54-hour benchmark for Vedic Sanskrit poetry ASR, and fine-tuned IndicWhisper reaches about 22% word error rate on the Devanagari test set.

  8. ZIPA: A family of efficient models for multilingual phone recognition

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ZIPA models, trained with Zipformer backbones on a new 17,132-hour IPA-labeled corpus, achieve state-of-the-art multilingual phone recognition with fewer parameters than prior systems.

  9. Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems

    cs.SD 2025-09 reject novelty 5.0 of 10

    A new evaluation protocol exhaustively pairs 164 speech synthesizers with nine bona fide speech types and reports max-pooled EERs, revealing larger failures than pooled averages show.

  10. Leveraging Large Language Models for Spontaneous Speech-Based Suicide Risk Detection

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A multimodal system combining acoustic, textual, and LLM-extracted features reports 74% accuracy for adolescent suicide risk detection on the SW1 challenge test set.

  11. UD-KSL Treebank v1.3: A semi-automated framework for aligning XPOS-extracted units with UPOS tags

    cs.CL 2025-06 reject novelty 5.0 of 10

    A semi-automated XPOS-to-UPOS alignment pipeline for the UD-KSL treebank is reported to improve tagging and some parsing accuracy, but the comparison is confounded because the test gold labels also change between conditions.

  12. SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset

    cs.CL 2025-05 reject novelty 5.0 of 10

    The authors present SwitchLingua, a large multilingual code-switching text and audio dataset, and SAER, a semantic-aware error metric for code-switching ASR evaluation.

  13. TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation

    cs.CL 2025-08 conditional novelty 4.0 of 10

    Adding task-specific learned vectors to acoustic embeddings lets a transducer ASR model train on partially labeled data, matching or beating the fully labeled TokenVerse baseline on most tasks.

Pith tools