Pith. sign in

REVIEW 2 cited by

Unified Speech-Text Pre-training for Speech Translation and Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.05409 v1 pith:AIUJ6Z4R submitted 2022-04-11 cs.CL

classification cs.CL
keywords speechtextrecognitiontranslationmethodpre-trainingsubtasksupervised
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four self-supervised and supervised subtasks for cross modality learning. A self-supervised speech subtask leverages unlabelled speech data, and a (self-)supervised text to text subtask makes use of abundant text training data. Two auxiliary supervised speech tasks are included to unify speech and text modeling space. Our contribution lies in integrating linguistic information from the text corpus into the speech pre-training. Detailed analysis reveals learning interference among subtasks. Two pre-training configurations for speech translation and recognition, respectively, are presented to alleviate subtask interference. Our experiments show the proposed method can effectively fuse speech and text information into one model. It achieves between 1.7 and 2.3 BLEU improvement above the state of the art on the MuST-C speech translation dataset and comparable WERs to wav2vec 2.0 on the Librispeech speech recognition task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

    cs.SD 2025-06 conditional novelty 6.0 of 10

    XY-Tokenizer is a 1 kbps dual-channel speech codec that reports simultaneously strong text alignment and high speaker similarity, comparable to specialized codecs at similar bitrates.

  2. CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing

    eess.AS 2024-12 conditional novelty 6.0 of 10

    CA-SSLR injects condition-aware language and speaker embeddings into a frozen SSL encoder via lightweight FiLM-style adapters, improving ASR, LID, and SV performance and transfer.

Pith tools