Pith. sign in

REVIEW 7 cited by

Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.00654 v1 pith:WIDG2KLU submitted 2024-06-02 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords humanspeechsubjectivetrainingevaluationsfeedbackzero-shotevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, text-to-speech (TTS) technology has witnessed impressive advancements, particularly with large-scale training datasets, showcasing human-level speech quality and impressive zero-shot capabilities on unseen speakers. However, despite human subjective evaluations, such as the mean opinion score (MOS), remaining the gold standard for assessing the quality of synthetic speech, even state-of-the-art TTS approaches have kept human feedback isolated from training that resulted in mismatched training objectives and evaluation metrics. In this work, we investigate a novel topic of integrating subjective human evaluation into the TTS training loop. Inspired by the recent success of reinforcement learning from human feedback, we propose a comprehensive sampling-annotating-learning framework tailored to TTS optimization, namely uncertainty-aware optimization (UNO). Specifically, UNO eliminates the need for a reward model or preference data by directly maximizing the utility of speech generations while considering the uncertainty that lies in the inherent variability in subjective human speech perception and evaluations. Experimental results of both subjective and objective evaluations demonstrate that UNO considerably improves the zero-shot performance of TTS models in terms of MOS, word error rate, and speaker similarity. Additionally, we present a remarkable ability of UNO that it can adapt to the desired speaking style in emotional TTS seamlessly and flexibly.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech

    eess.AS 2025-08 conditional novelty 6.0 of 10

    MPO improves TTS alignment by constructing multi-dimensional preference pairs and adding cross-entropy regularization to DPO, yielding better intelligibility, speaker similarity, and prosody.

  2. TTS-1 Technical Report

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new TTS system family combines pre-training, supervised fine-tuning, and GRPO reinforcement learning with a 48 kHz codec to produce multilingual speech with in-context voice cloning.

  3. DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.

  4. Differentiable Reward Optimization for LLM based TTS system

    cs.SD 2025-07 conditional novelty 6.0 of 10

    DiffRO optimizes codec-based TTS models directly on differentiable token-level rewards, improving WER and enabling zero-shot emotion control.

  5. Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

    cs.SD 2025-06 conditional novelty 6.0 of 10

    DCAR dynamically schedules chunk-wise token prediction in AR TTS, improving WER by up to 72.27% relative and speeding up inference by up to 2.89x over next-token baselines.

  6. VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.

  7. Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation

    cs.SD 2025-08 conditional novelty 5.0 of 10

    PER-based preference optimization (DPO, PPO, GRPO) reduces lyric-to-song hallucination in an audio language model, with the largest gains from DPO plus reject sampling.

Pith tools