Pith. sign in

REVIEW 2 cited by

VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.14321 v5 pith:Z3BRWJYS submitted 2024-01-25 eess.AS cs.SD

classification eess.AScs.SD
keywords decoder-onlyvall-tadaptationarchitecturegenerativemodelsmonotonicrelative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent TTS models with decoder-only Transformer architecture, such as SPEAR-TTS and VALL-E, achieve impressive naturalness and demonstrate the ability for zero-shot adaptation given a speech prompt. However, such decoder-only TTS models lack monotonic alignment constraints, sometimes leading to hallucination issues such as mispronunciation, word skipping and repeating. To address this limitation, we propose VALL-T, a generative Transducer model that introduces shifting relative position embeddings for input phoneme sequence, explicitly indicating the monotonic generation process while maintaining the architecture of decoder-only Transformer. Consequently, VALL-T retains the capability of prompt-based zero-shot adaptation and demonstrates better robustness against hallucinations with a relative reduction of 28.3% in the word error rate.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SpeakStream: Streaming Text-to-Speech with Interleaved Data

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A decoder-only TTS trained on force-aligned interleaved text-speech chunks generates audio after a few words, achieving ~30ms TTS latency and WER comparable to non-streaming.

  2. Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding

    cs.SD 2025-05 conditional novelty 4.0 of 10

    Masked self-distillation training plus semi-autoregressive decoding improves RNN-T keyword spotting recall at low false alarm rates, especially in noisy conditions.

Pith tools