REVIEW 30 cited by
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0. We train models with up to 2B parameters on nearly half a million hours of publicly available speech audio in 128 languages, an order of magnitude more public data than the largest known prior work. Our evaluation covers a wide range of tasks, domains, data regimes and languages, both high and low-resource. On the CoVoST-2 speech translation benchmark, we improve the previous state of the art by an average of 7.4 BLEU over 21 translation directions into English. For speech recognition, XLS-R improves over the best known prior work on BABEL, MLS, CommonVoice as well as VoxPopuli, lowering error rates by 14-34% relative on average. XLS-R also sets a new state of the art on VoxLingua107 language identification. Moreover, we show that with sufficient model size, cross-lingual pretraining can outperform English-only pretraining when translating English speech into other languages, a setting which favors monolingual pretraining. We hope XLS-R can help to improve speech processing tasks for many more languages of the world.
Forward citations
Cited by 30 Pith papers
-
Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry
A four-backbone SSL ensemble achieves near-perfect deepfake detection in the ImageCLEF 2026 track, while the same team's generated audio ranks first in the generation track, revealing an OR-versus-AND asymmetry betwee...
-
Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition
Pre-adapting multilingual ASR models on linguistically related languages does not yield practically meaningful target-language gains once one hour of target data is used.
-
An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures
Controlled non-speech interventions cause the largest detection-cost spikes, confirming non-speech structure as the dominant confound-driven shortcut in ASVspoof-trained XLS-R + RawGAT-ST models.
-
Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study
Across six pooling heads and six frozen SSL backbones on English and Mandarin depression speech, a third of configurations collapse to single-class prediction, so backbone- and seed-robustness should be first-class ev...
-
SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection
SONAR improves audio deepfake detection by explicitly aligning low- and high-frequency representations for real speech and repelling them for fakes, setting new benchmark EERs on ASVspoof 2021 and in-the-wild data.
-
CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
CAMOES is a new open benchmark and model collection for European Portuguese ASR, cutting word error rate by about 35% over the best zero-shot model.
-
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.
-
On Barriers to Archival Audio Processing
On archival radio audio, Whisper V3 identifies languages with 91.3% accuracy, but speaker embeddings lose similarity across ages and languages, making speaker recognition unreliable for indexing.
-
Word stress in self-supervised speech models: A cross-linguistic comparison
Stress classifiers trained on Wav2vec 2.0 embeddings distinguish stressed and unstressed syllables in Dutch, English, German, Polish and Hungarian, with language-specific structure that separates fixed and variable st...
-
RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning
PCA of repeated synthesized utterances with fixed inputs can reveal controllable prosodic features that can be enrolled as new prompts via fine-tuning.
-
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning
A quality audit of Common Voice, FLEURS, and VoxPopuli finds serious micro- and macro-level data defects concentrated in less institutionalized languages, with a case study of Taiwanese Southern Min.
-
Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages
Fine-tuned MMS outperforms XLS-R on fieldwork ASR with less than one hour of training data, while XLS-R reaches parity beyond one hour.
-
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
A federated fine-tuning method prunes 90% of attention heads, weights updates by attention importance, and selects clients by loss gap, cutting communication 1.8x and training compute 3.9x with under 2% accuracy drop.
-
ZIPA: A family of efficient models for multilingual phone recognition
ZIPA models, trained with Zipformer backbones on a new 17,132-hour IPA-labeled corpus, achieve state-of-the-art multilingual phone recognition with fewer parameters than prior systems.
-
PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems
PSRB, a 10.4-hour Persian benchmark built from 3,372 clips and 756 speakers, evaluates ten ASR models and introduces SW-WER, showing that systems are far weaker on regional accents, children's speech, and informal aud...
-
Towards Digital Preservation of Efik: TTS for a Low-Resource African Language
First end-to-end Efik TTS baseline: a 3-hour single-speaker corpus and four fine-tuned models, with MMS-TTS best at MOS 3.80±0.63 but residual tonal errors.
-
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
Adding dataset identity as an auxiliary task or adversarial label improves aggregate audio-deepfake detection EER on the 2025 Speech DeepFake Arena benchmark.
-
Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR
Language-balanced gradient projection plus experience replay yields near-zero average forgetting when adapting Whisper-large-v3 to low-resource languages while preserving target plasticity.
-
GigaAM Multilingual: Foundation Model for Underrepresented Languages
Cluster-balanced HuBERT-style pre-training on 2M hours plus domain-aware fine-tuning yields a compact encoder that outperforms larger open multilingual ASR models on Kazakh, Kyrgyz and Uzbek.
-
Generalizable Audio Spoofing Detection using Non-Semantic Representations
Frozen non-semantic TRILLson embeddings with a lightweight backend beat prior spoofing detectors on out-of-domain datasets while staying competitive in-domain.
-
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
A compressed HuBERT can be trained with the original masked-prediction objective using k-means labels from the teacher, beating feature-distillation methods on four SUPERB tasks.
-
From Sharpness to Better Generalization for Speech Deepfake Detection
Sharpness, a measure of loss sensitivity to weight perturbations, correlates with speech deepfake detection error on unseen data, and Sharpness-Aware Minimization reduces both sharpness and error in most settings.
-
Joint ASR and Speaker Role Tagging with Serialized Output Training
Fine-tuning Whisper with serialized output training and role-specific tokens produces role-aware transcripts in one pass, cutting multi-talker word error rate by 10 to 40 percent versus a WavLM CTC baseline.
-
Towards Generalized Source Tracing for Codec-Based Deepfake Speech
SASTNet, which fuses Whisper semantic features with Wav2Vec2 and AudioMAE acoustic features, improves source tracing for codec-based deepfake speech on CodecFake+, while exposing that prior models overfit to silence.
-
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
R-VC performs zero-shot voice conversion in two sampling steps while transferring the target speaker's rhythm, matching or exceeding prior systems in naturalness and intelligibility.
-
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
ADD-GP, a Gaussian Process classifier with XLS-R speech embeddings, adapts to unseen TTS models with as few as 5 samples and achieves state-of-the-art low error rates on the new LibriFake benchmark.
-
Teffic-Audio: Tell Fact from Fiction
A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.
-
DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages
Open w2v-BERT ASR base models for 27 African languages, with a two-step annealing recipe and prefix-frame language conditioning.
-
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
Swapping the MLP projector for a GR-KAN layer in XLSR-Conformer reduces equal error rates on ASVspoof 2021 LA and DF, reaching 0.70% EER on the variable-length LA set.
-
Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact
A broad survey arguing that AGI requires modular, memory-augmented, embodied architectures rather than scaled-up token prediction, with a brief proposal to decompose intelligence into five components.
Discussion (0). Sign in to comment.