A diffusion voice conversion model, conditioned on clean speaker embeddings and HuBERT content features, is applied after a generative speech restorer to achieve state-of-the-art-comparable speech quality.
Our method leverages a two- stage approach, incorporating a generative speech restoration (GSR) model at the front end, followed by a VC module
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
A diffusion voice conversion model, conditioned on clean speaker embeddings and HuBERT content features, is applied after a generative speech restorer to achieve state-of-the-art-comparable speech quality.