Pith. sign in

REVIEW 2 cited by

VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.06674 v2 pith:D4EISXR7 submitted 2024-04-10 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords frameworkattributesmodelspeakerspeechspeech-to-speechtimbrevoiceshop
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present VoiceShop, a novel speech-to-speech framework that can modify multiple attributes of speech, such as age, gender, accent, and speech style, in a single forward pass while preserving the input speaker's timbre. Previous works have been constrained to specialized models that can only edit these attributes individually and suffer from the following pitfalls: the magnitude of the conversion effect is weak, there is no zero-shot capability for out-of-distribution speakers, or the synthesized outputs exhibit undesirable timbre leakage. Our work proposes solutions for each of these issues in a simple modular framework based on a conditional diffusion backbone model with optional normalizing flow-based and sequence-to-sequence speaker attribute-editing modules, whose components can be combined or removed during inference to meet a wide array of tasks without additional model finetuning. Audio samples are available at \url{https://voiceshopai.github.io}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment

    eess.AS 2025-07 conditional novelty 6.0 of 10

    SemAlignVC strips source-speaker timbre by aligning a speech semantic encoder to BERT text embeddings, then resynthesizes the content conditioned only on a target voice reference.

  2. Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

    cs.SD 2025-02 conditional novelty 6.0 of 10

    Vevo achieves zero-shot timbre, accent, and emotion imitation by using VQ-VAE codebook size on HuBERT features to create content and content-style tokens.

Pith tools