Pith. sign in

REVIEW 4 cited by

fairseq S2T: Fast Speech-to-Text Modeling with fairseq

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.05171 v2 pith:NYGUT64W submitted 2020-10-11 cs.CL eess.AS

classification cs.CLeess.AS
keywords fairseqmodelsspeech-to-textend-to-endexampleslearningmodelingspeech
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce fairseq S2T, a fairseq extension for speech-to-text (S2T) modeling tasks such as end-to-end speech recognition and speech-to-text translation. It follows fairseq's careful design for scalability and extensibility. We provide end-to-end workflows from data pre-processing, model training to offline (online) inference. We implement state-of-the-art RNN-based, Transformer-based as well as Conformer-based models and open-source detailed training recipes. Fairseq's machine translation models and language models can be seamlessly integrated into S2T workflows for multi-task learning or transfer learning. Fairseq S2T documentation and examples are available at https://github.com/pytorch/fairseq/tree/master/examples/speech_to_text.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The paper releases Teochew-Wild, the first publicly available Teochew speech corpus with orthographic and pinyin annotations, and shows it supports ASR and TTS training.

  2. A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    The paper presents a unit-based direct speech-to-speech translation system and a paired English-Spanish movie dataset, claiming better preservation of paralinguistic information while maintaining translation quality.

  3. Transferable Adversarial Attacks on Audio Deepfake Detection

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A transferable GAN-based attack that preserves transcription and perceptual quality can substantially degrade current audio deepfake detection systems.

  4. IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Fine-tuning SeamlessM4T on 20 hours of Bhojpuri-Hindi data with tuned hyperparameters and SpecAugment reaches 36.4 dev BLEU but only 9.9 test BLEU in the IWSLT 2025 low-resource task.

Pith tools