Pith. sign in

REVIEW 40 cited by

Seamless: Multilingual Expressive and Streaming Speech Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.05187 v1 pith:DXZLYYXT submitted 2023-12-08 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords translationexpressivefirstseamlessspeechcommunicationmodelsmultilingual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale automatic speech translation systems today lack key features that help machine-mediated communication feel seamless when compared to human-to-human dialogue. In this work, we introduce a family of models that enable end-to-end expressive and multilingual translations in a streaming fashion. First, we contribute an improved version of the massively multilingual and multimodal SeamlessM4T model-SeamlessM4T v2. This newer model, incorporating an updated UnitY2 framework, was trained on more low-resource language data. SeamlessM4T v2 provides the foundation on which our next two models are initiated. SeamlessExpressive enables translation that preserves vocal styles and prosody. Compared to previous efforts in expressive speech research, our work addresses certain underexplored aspects of prosody, such as speech rate and pauses, while also preserving the style of one's voice. As for SeamlessStreaming, our model leverages the Efficient Monotonic Multihead Attention mechanism to generate low-latency target translations without waiting for complete source utterances. As the first of its kind, SeamlessStreaming enables simultaneous speech-to-speech/text translation for multiple source and target languages. To ensure that our models can be used safely and responsibly, we implemented the first known red-teaming effort for multimodal machine translation, a system for the detection and mitigation of added toxicity, a systematic evaluation of gender bias, and an inaudible localized watermarking mechanism designed to dampen the impact of deepfakes. Consequently, we bring major components from SeamlessExpressive and SeamlessStreaming together to form Seamless, the first publicly available system that unlocks expressive cross-lingual communication in real-time. The contributions to this work are publicly released and accessible at https://github.com/facebookresearch/seamless_communication

Discussion (0). Sign in to comment.

Forward citations

Cited by 40 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 41 citations worldwide. Full citation record

  1. Simultaneous Speech-to-Speech Translation Without Aligned Data

    cs.CL 2026-02 conditional novelty 7.0 of 10

    Hibiki-Zero performs simultaneous speech-to-speech translation without word-level aligned data, using sentence-level supervision plus GRPO reinforcement learning with BLEU-based process rewards, and reports state-of-t...

  2. TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

    cs.SD 2025-07 conditional novelty 7.0 of 10

    TTS-CtrlNet adds time-varying emotion control to a frozen flow-matching TTS model using a ControlNet-style trainable copy, improving emotion similarity metrics while preserving the base model's voice cloning.

  3. Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A new genderless-to-English benchmark shows that fine-tuning mBART-50 on carefully curated examples cuts gender stereotyping and pronoun-reasoning errors, beating larger proprietary systems on that benchmark.

  4. Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models

    eess.AS 2025-05 accept novelty 7.0 of 10

    Fine-tuning a pre-trained speech transformer detects primary stress at 99% word accuracy on Croatian and Serbian and 89% on related Chakavian and Slovenian, with only 500 training words needed.

  5. The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

    cs.CL 2025-05 conditional novelty 7.0 of 10

    NaijaVoices is a 1,838-hour, 5,455-speaker speech-text corpus for Igbo, Hausa, and Yoruba whose use in fine-tuning cuts Word Error Rates by 42-76% relative to unadapted baselines.

  6. ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality

    eess.AS 2025-05 conditional novelty 7.0 of 10

    ToxicTone provides a 52k-clip Mandarin spoken-toxicity dataset with form and source labels, and a multimodal detector that outperforms off-the-shelf text baselines.

  7. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.

  8. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  9. The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

    eess.AS 2026-04 unverdicted novelty 6.0 of 10

    Emotion embedding similarities are unsuitable for zero-shot evaluation of emotional expressiveness in speech generation due to confounding by non-emotional acoustic features.

  10. Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Fine-tuning speech translation on adversarial text embeddings in an aligned speech-text space transfers inflectional robustness to speech: ~3 BLEU average gain on adversarially inflected audio, with no adversarial spe...

  11. Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning

    eess.AS 2025-08 conditional novelty 6.0 of 10

    The authors release a rule-based Quran Phonetic Script, an 890-hour expert recitation dataset, and a multi-head CTC model that achieves 0.16% average phoneme error rate on held-out reciters.

  12. CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese

    cs.CL 2025-08 conditional novelty 6.0 of 10

    CAMOES is a new open benchmark and model collection for European Portuguese ASR, cutting word error rate by about 35% over the best zero-shot model.

  13. Geolocation-Aware Robust Spoken Language Identification

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Auxiliary geolocation prediction with injected conditioning signals improves dialect and accent robustness in SSL-based spoken language identification, reaching new SOTA on FLEURS and ML-SUPERB 2.0.

  14. The Prosody of Emojis

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Speakers systematically modify prosody when reading sentences with different emojis, and listeners can recover the intended emoji from prosody alone, with larger semantic differences yielding larger prosodic shifts.

  15. VN-MTEB: Vietnamese Massive Text Embedding Benchmark

    cs.CL 2025-07 conditional novelty 6.0 of 10

    VN-MTEB is a new 41-dataset Vietnamese benchmark for text embeddings, built by machine-translating MTEB datasets with embedding-based and LLM-based quality filters.

  16. Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An end-to-end simultaneous speech-to-speech translation model with voice cloning, trained with a two-stage reinforcement learning reward scheme, reports high accuracy and low latency on the authors' RealSI benchmark.

  17. StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A unified large speech-language model uses speech chain-of-thought to jointly perform segmentation, generation-policy decisions, and streaming translation.

  18. Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A late-fusion model that combines ASR-transcribed lyrics and speech embeddings detects AI-written lyrics from audio alone, achieving 94.9% recall in-domain and staying robust to attacks.

  19. It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems

    cs.CL 2025-06 conditional novelty 6.0 of 10

    End-to-end speech translation systems translate idioms worse than text-based systems, frequently producing literal or incorrect outputs, across German and Russian to English.

  20. PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems

    eess.AS 2025-05 conditional novelty 6.0 of 10

    PSRB, a 10.4-hour Persian benchmark built from 3,372 clips and 756 speakers, evaluates ten ASR models and introduces SW-WER, showing that systems are far weaker on regional accents, children's speech, and informal aud...

  21. SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SeqPO-SiMT uses sequential policy optimization with a combined quality-and-latency reward to improve simultaneous machine translation, beating supervised fine-tuning on six En-Zh and Zh-En datasets.

  22. OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A new open suite of 13 multilingual speech models, up to 18B parameters, yields empirical scaling laws for ASR and speech translation performance.

  23. FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

    cs.LG 2026-07 conditional novelty 5.0 of 10

    FlashRT's agent harness converts reference multimodal pipelines into optimized multi-GPU deployments, reporting ~70x latency cuts and up to 3.6x throughput gains across five applications on B200 and MI355X.

  24. X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

    eess.AS 2026-07 conditional novelty 5.0 of 10

    An open, modular cascaded system (streaming ASR + MT + prompt-conditioned TTS) preserves speaker identity in long-form multi-speaker translation, at higher latency and slightly lower translation quality than proprietary APIs.

  25. GigaAM Multilingual: Foundation Model for Underrepresented Languages

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Cluster-balanced HuBERT-style pre-training on 2M hours plus domain-aware fine-tuning yields a compact encoder that outperforms larger open multilingual ASR models on Kazakh, Kyrgyz and Uzbek.

  26. NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task

    cs.CL 2025-09 conditional novelty 5.0 of 10

    The NADI 2025 shared task introduces a standardized speech benchmark for eight Arabic dialects and reports best results of 79.8% dialect ID accuracy, 35.68 WER for ASR, and 55 WER for diacritic restoration.

  27. ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs

    eess.AS 2025-08 unverdicted novelty 5.0 of 10

    ProMode learns a prosody embedding from partially masked audio and text, improving F0 and energy prediction over baseline style encoders.

  28. Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A FreeVC-style conditional VAE with mHuBERT-147 discrete units, mixed-style layer normalization, an augmentation similarity loss, and F0 cross-attention reports better emotion transfer and less source leakage than thr...

  29. SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset

    cs.CL 2025-05 reject novelty 5.0 of 10

    The authors present SwitchLingua, a large multilingual code-switching text and audio dataset, and SAER, a semantic-aware error metric for code-switching ASR evaluation.

  30. ALAS: An Automatic Latent Alignment Score for Audio Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    ALAS is a reference-based score for audio-text alignment in speech LLMs, computed from frozen hidden states and a Whisper-derived alignment path, with no training or fitted classifier.

  31. From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.

  32. Teffic-Audio: Tell Fact from Fiction

    cs.SD 2026-07 conditional novelty 4.0 of 10

    A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.

  33. AMECxSV: Adaptive Metadata-Driven Embedding-Fusion Calibration for X-Lingual Speaker Verification

    eess.AS 2026-07 accept novelty 4.0 of 10

    A metadata-conditioned MLP backend fused with fixed speaker-verification scores reduces EER and improves calibration in X-lingual trials when language and duration cues are available.

  34. On the Contribution of Lexical Features to Speech Emotion Recognition

    eess.AS 2025-09 conditional novelty 4.0 of 10

    Using Whisper transcriptions plus DeBERTa text features, a lexical-only pipeline beats acoustic-only models on MELD speech emotion recognition: 51.5% vs 49.3% weighted F1.

  35. Optimal Multi-Task Learning at Regularization Horizon for Speech Translation Task

    cs.CL 2025-09 conditional novelty 4.0 of 10

    Combining consistency regularization, R-drop, and the MT loss weight into a single scalar 'total regularization' predicts speech translation quality, and tuning near its optimum yields near-SOTA BLEU on MuST-C.

  36. Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition

    cs.SD 2025-09 reject novelty 4.0 of 10

    An LLM-based ASR error correction framework with noise-adaptive encoding and dynamic multi-modal fusion reports WER gains, but its fusion weights require ground-truth text at inference.

  37. Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning

    cs.CL 2025-06 conditional novelty 4.0 of 10

    The IT-IST IWSLT 2025 submission shows that a 1.5B language model with a speech encoder can do reasonable ASR after alignment, but struggles with ST and SQA.

  38. Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions

    eess.AS 2025-06 conditional novelty 4.0 of 10

    A prompt-based method lets an LLM-based TTS system synthesize speech with mixed emotions in user-specified proportions without training on mixed-emotion data.

  39. GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Fine-tuning SeamlessM4T-v2 directly for end-to-end speech translation is competitive, and ASR-encoder initialization adds about 1 to 5 BLEU for languages unseen by the base model.

  40. Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models

    cs.CL 2025-05 conditional novelty 4.0 of 10

    In a three-stage end-to-end spoken language model, experience replay (mixing old data into later training) was the most effective mitigation against catastrophic forgetting, greatly outperforming model merging and LoR...

Pith tools