REVIEW 4 cited by
fairseq S2T: Fast Speech-to-Text Modeling with fairseq
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce fairseq S2T, a fairseq extension for speech-to-text (S2T) modeling tasks such as end-to-end speech recognition and speech-to-text translation. It follows fairseq's careful design for scalability and extensibility. We provide end-to-end workflows from data pre-processing, model training to offline (online) inference. We implement state-of-the-art RNN-based, Transformer-based as well as Conformer-based models and open-source detailed training recipes. Fairseq's machine translation models and language models can be seamlessly integrated into S2T workflows for multi-task learning or transfer learning. Fairseq S2T documentation and examples are available at https://github.com/pytorch/fairseq/tree/master/examples/speech_to_text.
Forward citations
Cited by 4 Pith papers
-
Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations
The paper releases Teochew-Wild, the first publicly available Teochew speech corpus with orthographic and pinyin annotations, and shows it supports ASR and TTS training.
-
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation
The paper presents a unit-based direct speech-to-speech translation system and a paired English-Spanish movie dataset, claiming better preservation of paralinguistic information while maintaining translation quality.
-
Transferable Adversarial Attacks on Audio Deepfake Detection
A transferable GAN-based attack that preserves transcription and perceptual quality can substantially degrade current audio deepfake detection systems.
-
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
Fine-tuning SeamlessM4T on 20 hours of Bhojpuri-Hindi data with tuned hyperparameters and SpecAugment reaches 36.4 dev BLEU but only 9.9 test BLEU in the IWSLT 2025 low-resource task.
Discussion (0). Continue with ORCID to comment.