REVIEW 40 cited by
Seamless: Multilingual Expressive and Streaming Speech Translation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large-scale automatic speech translation systems today lack key features that help machine-mediated communication feel seamless when compared to human-to-human dialogue. In this work, we introduce a family of models that enable end-to-end expressive and multilingual translations in a streaming fashion. First, we contribute an improved version of the massively multilingual and multimodal SeamlessM4T model-SeamlessM4T v2. This newer model, incorporating an updated UnitY2 framework, was trained on more low-resource language data. SeamlessM4T v2 provides the foundation on which our next two models are initiated. SeamlessExpressive enables translation that preserves vocal styles and prosody. Compared to previous efforts in expressive speech research, our work addresses certain underexplored aspects of prosody, such as speech rate and pauses, while also preserving the style of one's voice. As for SeamlessStreaming, our model leverages the Efficient Monotonic Multihead Attention mechanism to generate low-latency target translations without waiting for complete source utterances. As the first of its kind, SeamlessStreaming enables simultaneous speech-to-speech/text translation for multiple source and target languages. To ensure that our models can be used safely and responsibly, we implemented the first known red-teaming effort for multimodal machine translation, a system for the detection and mitigation of added toxicity, a systematic evaluation of gender bias, and an inaudible localized watermarking mechanism designed to dampen the impact of deepfakes. Consequently, we bring major components from SeamlessExpressive and SeamlessStreaming together to form Seamless, the first publicly available system that unlocks expressive cross-lingual communication in real-time. The contributions to this work are publicly released and accessible at https://github.com/facebookresearch/seamless_communication
Forward citations
Cited by 40 Pith papers
-
Simultaneous Speech-to-Speech Translation Without Aligned Data
Hibiki-Zero performs simultaneous speech-to-speech translation without word-level aligned data, using sentence-level supervision plus GRPO reinforcement learning with BLEU-based process rewards, and reports state-of-t...
-
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
TTS-CtrlNet adds time-varying emotion control to a frozen flow-matching TTS model using a ControlNet-style trainable copy, improving emotion similarity metrics while preserving the base model's voice cloning.
-
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
A new genderless-to-English benchmark shows that fine-tuning mBART-50 on carefully curated examples cuts gender stereotyping and pronoun-reasoning errors, beating larger proprietary systems on that benchmark.
-
Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models
Fine-tuning a pre-trained speech transformer detects primary stress at 99% word accuracy on Croatian and Serbian and 89% on related Chakavian and Slovenian, with only 500 training words needed.
-
The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages
NaijaVoices is a 1,838-hour, 5,455-speaker speech-text corpus for Igbo, Hausa, and Yoruba whose use in fine-tuning cuts Word Error Rates by 42-76% relative to unadapted baselines.
-
ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality
ToxicTone provides a 52k-clip Mandarin spoken-toxicity dataset with form and source labels, and a multimodal detector that outperforms off-the-shelf text baselines.
-
SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation
Emotion embedding similarities are unsuitable for zero-shot evaluation of emotional expressiveness in speech generation due to confounding by non-emotional acoustic features.
-
Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text
Fine-tuning speech translation on adversarial text embeddings in an aligned speech-text space transfers inflectional robustness to speech: ~3 BLEU average gain on adversarially inflected audio, with no adversarial spe...
-
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
The authors release a rule-based Quran Phonetic Script, an 890-hour expert recitation dataset, and a multi-head CTC model that achieves 0.16% average phoneme error rate on held-out reciters.
-
CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
CAMOES is a new open benchmark and model collection for European Portuguese ASR, cutting word error rate by about 35% over the best zero-shot model.
-
Geolocation-Aware Robust Spoken Language Identification
Auxiliary geolocation prediction with injected conditioning signals improves dialect and accent robustness in SSL-based spoken language identification, reaching new SOTA on FLEURS and ML-SUPERB 2.0.
-
The Prosody of Emojis
Speakers systematically modify prosody when reading sentences with different emojis, and listeners can recover the intended emoji from prosody alone, with larger semantic differences yielding larger prosodic shifts.
-
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
VN-MTEB is a new 41-dataset Vietnamese benchmark for text embeddings, built by machine-translating MTEB datasets with embedding-based and LLM-based quality filters.
-
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
An end-to-end simultaneous speech-to-speech translation model with voice cloning, trained with a two-stage reinforcement learning reward scheme, reports high accuracy and low latency on the authors' RealSI benchmark.
-
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
A unified large speech-language model uses speech chain-of-thought to jointly perform segmentation, generation-policy decisions, and streaming translation.
-
Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
A late-fusion model that combines ASR-transcribed lyrics and speech embeddings detects AI-written lyrics from audio alone, achieving 94.9% recall in-domain and staying robust to attacks.
-
It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems
End-to-end speech translation systems translate idioms worse than text-based systems, frequently producing literal or incorrect outputs, across German and Russian to English.
-
PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems
PSRB, a 10.4-hour Persian benchmark built from 3,372 clips and 756 speakers, evaluates ten ASR models and introduces SW-WER, showing that systems are far weaker on regional accents, children's speech, and informal aud...
-
SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation
SeqPO-SiMT uses sequential policy optimization with a combined quality-and-latency reward to improve simultaneous machine translation, beating supervised fine-tuning on six En-Zh and Zh-En datasets.
-
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
A new open suite of 13 multilingual speech models, up to 18B parameters, yields empirical scaling laws for ASR and speech translation performance.
-
FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
FlashRT's agent harness converts reference multimodal pipelines into optimized multi-GPU deployments, reporting ~70x latency cuts and up to 3.6x throughput gains across five applications on B200 and MI355X.
-
X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
An open, modular cascaded system (streaming ASR + MT + prompt-conditioned TTS) preserves speaker identity in long-form multi-speaker translation, at higher latency and slightly lower translation quality than proprietary APIs.
-
GigaAM Multilingual: Foundation Model for Underrepresented Languages
Cluster-balanced HuBERT-style pre-training on 2M hours plus domain-aware fine-tuning yields a compact encoder that outperforms larger open multilingual ASR models on Kazakh, Kyrgyz and Uzbek.
-
NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
The NADI 2025 shared task introduces a standardized speech benchmark for eight Arabic dialects and reports best results of 79.8% dialect ID accuracy, 35.68 WER for ASR, and 55 WER for diacritic restoration.
-
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
ProMode learns a prosody embedding from partially masked audio and text, improving F0 and energy prediction over baseline style encoders.
-
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
A FreeVC-style conditional VAE with mHuBERT-147 discrete units, mixed-style layer normalization, an augmentation similarity loss, and F0 cross-attention reports better emotion transfer and less source leakage than thr...
-
SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
The authors present SwitchLingua, a large multilingual code-switching text and audio dataset, and SAER, a semantic-aware error metric for code-switching ASR evaluation.
-
ALAS: An Automatic Latent Alignment Score for Audio Language Models
ALAS is a reference-based score for audio-text alignment in speech LLMs, computed from frozen hidden states and a Whisper-derived alignment path, with no training or fitted classifier.
-
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.
-
Teffic-Audio: Tell Fact from Fiction
A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.
-
AMECxSV: Adaptive Metadata-Driven Embedding-Fusion Calibration for X-Lingual Speaker Verification
A metadata-conditioned MLP backend fused with fixed speaker-verification scores reduces EER and improves calibration in X-lingual trials when language and duration cues are available.
-
On the Contribution of Lexical Features to Speech Emotion Recognition
Using Whisper transcriptions plus DeBERTa text features, a lexical-only pipeline beats acoustic-only models on MELD speech emotion recognition: 51.5% vs 49.3% weighted F1.
-
Optimal Multi-Task Learning at Regularization Horizon for Speech Translation Task
Combining consistency regularization, R-drop, and the MT loss weight into a single scalar 'total regularization' predicts speech translation quality, and tuning near its optimum yields near-SOTA BLEU on MuST-C.
-
Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition
An LLM-based ASR error correction framework with noise-adaptive encoding and dynamic multi-modal fusion reports WER gains, but its fusion weights require ground-truth text at inference.
-
Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning
The IT-IST IWSLT 2025 submission shows that a 1.5B language model with a speech encoder can do reasonable ASR after alignment, but struggles with ST and SQA.
-
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
A prompt-based method lets an LLM-based TTS system synthesize speech with mixed emotions in user-specified proportions without training on mixed-emotion data.
-
GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task
Fine-tuning SeamlessM4T-v2 directly for end-to-end speech translation is competitive, and ASR-encoder initialization adds about 1 to 5 BLEU for languages unseen by the base model.
-
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
In a three-stage end-to-end spoken language model, experience replay (mixing old data into later training) was the most effective mitigation against catastrophic forgetting, greatly outperforming model merging and LoR...
Discussion (0). Sign in to comment.