Pith. sign in

REVIEW 4 cited by

Hierarchical Generative Modeling for Controllable Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.07217 v2 pith:SPK2H62U submitted 2018-10-16 cs.CL cs.LGcs.SDeess.AS

classification cs.CLcs.LGcs.SDeess.AS
keywords modelattributescontrolcontrollablelatentlevelspeakingspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes a neural sequence-to-sequence text-to-speech (TTS) model which can control latent attributes in the generated speech that are rarely annotated in the training data, such as speaking style, accent, background noise, and recording conditions. The model is formulated as a conditional generative model based on the variational autoencoder (VAE) framework, with two levels of hierarchical latent variables. The first level is a categorical variable, which represents attribute groups (e.g. clean/noisy) and provides interpretability. The second level, conditioned on the first, is a multivariate Gaussian variable, which characterizes specific attribute configurations (e.g. noise level, speaking rate) and enables disentangled fine-grained control over these attributes. This amounts to using a Gaussian mixture model (GMM) for the latent distribution. Extensive evaluation demonstrates its ability to control the aforementioned attributes. In particular, we train a high-quality controllable TTS model on real found data, which is capable of inferring speaker and style attributes from a noisy utterance and use it to synthesize clean speech with controllable speaking style.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Counterfactual gradient edits to a pretrained TTS model's encoder activations can control prosody and correct mispronunciations at inference time, at least on Tacotron 2.

  2. Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces

    eess.AS 2026-08 reject novelty 5.0 of 10

    An interactive genetic algorithm tunes arousal-valence coordinates per listener in emotional TTS, and personalized or culture-specific coordinates beat a generic U.S.-average baseline in small A/B tests.

  3. DarkStream: real-time speech anonymization with low latency

    eess.AS 2025-09 conditional novelty 5.0 of 10

    DarkStream combines a causal content encoder with limited lookahead, k-means quantization, and a GAN-based pseudo-speaker embedding to anonymize speech in real time with near-chance speaker-verification error rates.

  4. A Methodology for Controlling the Emotional Expressiveness in Synthetic Speech -- a Deep Learning approach

    eess.AS 2019-07 unverdicted novelty 3.0 of 10

    A methodology is proposed for emotional text-to-speech using emotional data collection, transfer-learning-based annotation of expressiveness features, and fine-tuning of a neutral TTS model.

Pith tools