Pith. sign in

REVIEW 8 cited by

Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.13821 v2 pith:MVQGO6KI submitted 2021-09-28 cs.SD cs.LGstat.ML

classification cs.SDcs.LGstat.ML
keywords voiceconversiondiffusiongeneralone-shotqualitysynthesistarget
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Voice conversion is a common speech synthesis task which can be solved in different ways depending on a particular real-world scenario. The most challenging one often referred to as one-shot many-to-many voice conversion consists in copying the target voice from only one reference utterance in the most general case when both source and target speakers do not belong to the training dataset. We present a scalable high-quality solution based on diffusion probabilistic modeling and demonstrate its superior quality compared to state-of-the-art one-shot voice conversion approaches. Moreover, focusing on real-time applications, we investigate general principles which can make diffusion models faster while keeping synthesis quality at a high level. As a result, we develop a novel Stochastic Differential Equations solver suitable for various diffusion model types and generative tasks as shown through empirical studies and justify it by theoretical analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

    eess.AS 2026-08 conditional novelty 7.0 of 10

    A 260-hour emotional deepfake benchmark spanning 21 attack systems shows state-of-the-art speech deepfake detectors degrade badly on emotionally expressive and LALM-based spoofing.

  2. What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Underrepresented gender in training suffers higher deepfake-detection error; WavLM gaps stay large under balance, and all post-hoc calibrations leave the EER gap fixed at 1.317 pp.

  3. StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

    cs.MM 2025-06 conditional novelty 6.0 of 10

    StarVC is an autoregressive voice conversion model that generates text tokens before acoustic tokens, improving linguistic fidelity while retaining speaker similarity.

  4. Approximate Borderline Sampling using Granular-Ball for Classification Tasks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GBABS samples only approximate borderline points detected via non-overlapping granular balls, reporting better classifier accuracy and noise robustness than GB-based and standard sampling baselines.

  5. MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection

    cs.SD 2025-09 conditional novelty 5.0 of 10

    MoLEx combines LoRA adapters with a top-K expert router inside a frozen WavLM model, achieving 5.56% EER on ASVSpoof 5 without augmentation.

  6. ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A rectified-flow voice conversion model with speaker feature fusion achieves zero-shot conversion in one sampling step with quality close to 30-step diffusion baselines.

  7. Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching

    eess.AS 2025-06 conditional novelty 5.0 of 10

    R-VC performs zero-shot voice conversion in two sampling steps while transferring the target speaker's rhythm, matching or exceeding prior systems in naturalness and intelligibility.

  8. EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

    cs.SD 2025-05 reject novelty 4.0 of 10

    EZ-VC combines discrete units from a multilingual self-supervised encoder (Xeus) with an F5-TTS flow-matching decoder to achieve zero-shot any-to-any voice conversion, without text labels or multiple disentangling encoders.

Pith tools