Pith. sign in

REVIEW 1 cited by

Highly Controllable Diffusion-based Any-to-Any Voice Conversion Model with Frame-level Prosody Feature

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.03364 v1 pith:5546DJN6 submitted 2023-09-06 cs.SD eess.AS

classification cs.SDeess.AS
keywords prosodymodelenergyframe-levelspeakingspeechvoiceany-to-any
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a highly controllable voice manipulation system that can perform any-to-any voice conversion (VC) and prosody modulation simultaneously. State-of-the-art VC systems can transfer sentence-level characteristics such as speaker, emotion, and speaking style. However, manipulating the frame-level prosody, such as pitch, energy and speaking rate, still remains challenging. Our proposed model utilizes a frame-level prosody feature to effectively transfer such properties. Specifically, pitch and energy trajectories are integrated in a prosody conditioning module and then fed alongside speaker and contents embeddings to a diffusion-based decoder generating a converted speech mel-spectrogram. To adjust the speaking rate, our system includes a self-supervised model based post-processing step which allows improved controllability. The proposed model showed comparable speech quality and improved intelligibility compared to a SOTA approach. It can cover a varying range of fundamental frequency (F0), energy and speed modulation while maintaining converted speech quality.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Fast-VGAN is a lightweight GAN-based voice converter that explicitly controls F0, phoneme timing, and intensity, achieving near-perfect intelligibility and competitive speaker similarity on a small test set.

Pith tools