Pith. sign in

REVIEW 1 cited by

Sequence-to-Sequence Models Can Directly Translate Foreign Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1703.08581 v2 pith:OIAWA55I submitted 2017-03-24 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords speechmodelssequence-to-sequencelanguagerecognitiontrainingtranslationarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a recurrent encoder-decoder deep neural network architecture that directly translates speech in one language into text in another. The model does not explicitly transcribe the speech into text in the source language, nor does it require supervision from the ground truth source language transcription during training. We apply a slightly modified sequence-to-sequence with attention architecture that has previously been used for speech recognition and show that it can be repurposed for this more complex task, illustrating the power of attention-based models. A single model trained end-to-end obtains state-of-the-art performance on the Fisher Callhome Spanish-English speech translation task, outperforming a cascade of independently trained sequence-to-sequence speech recognition and machine translation models by 1.8 BLEU points on the Fisher test set. In addition, we find that making use of the training data in both languages by multi-task training sequence-to-sequence speech translation and recognition models with a shared encoder network can improve performance by a further 1.4 BLEU points.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Speech to Speech Translation with Translatotron: A State of the Art Review

    cs.CL 2025-02 reject

    A survey of Translatotron speech-to-speech translation models that asserts, without evidence, that Translatotron 3 is the best choice for low-resource African languages.

Pith tools