Pith. sign in

REVIEW 2 cited by

NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.09407 v1 pith:IYQKFMII submitted 2022-11-17 cs.SD eess.AS

classification cs.SDeess.AS
keywords voicesynthesisanalysisaudionansytrainingapplicationsbackbone
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, most of the voice synthesis models still require a large number of audio data paired with annotated labels (e.g., text transcription and music score) for training. To this end, we propose a unified framework of synthesizing and manipulating voice signals from analysis features, dubbed NANSY++. The backbone network of NANSY++ is trained in a self-supervised manner that does not require any annotations paired with audio. After training the backbone network, we efficiently tackle four voice applications - i.e. voice conversion, text-to-speech, singing voice synthesis, and voice designing - by partially modeling the analysis features required for each task. Extensive experiments show that the proposed framework offers competitive advantages such as controllability, data efficiency, and fast training convergence, while providing high quality synthesis. Audio samples: tinyurl.com/8tnsy3uc.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch

    eess.AS 2025-07 conditional novelty 6.0 of 10

    SLASH adds DSP-derived absolute pitch objectives, including direct spectrogram generation from F0, to self-supervised pitch estimation and beats DSP and SSL baselines on MIR-1K.

  2. DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.

Pith tools