Pith. sign in

REVIEW 8 cited by

OpenVoice: Versatile Instant Voice Cloning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.01479 v6 pith:ZDIOIGTM submitted 2023-12-03 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords voiceopenvoicecloninglanguagesmassive-speakerreferencespeakerstyles
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce OpenVoice, a versatile voice cloning approach that requires only a short audio clip from the reference speaker to replicate their voice and generate speech in multiple languages. OpenVoice represents a significant advancement in addressing the following open challenges in the field: 1) Flexible Voice Style Control. OpenVoice enables granular control over voice styles, including emotion, accent, rhythm, pauses, and intonation, in addition to replicating the tone color of the reference speaker. The voice styles are not directly copied from and constrained by the style of the reference speaker. Previous approaches lacked the ability to flexibly manipulate voice styles after cloning. 2) Zero-Shot Cross-Lingual Voice Cloning. OpenVoice achieves zero-shot cross-lingual voice cloning for languages not included in the massive-speaker training set. Unlike previous approaches, which typically require extensive massive-speaker multi-lingual (MSML) dataset for all languages, OpenVoice can clone voices into a new language without any massive-speaker training data for that language. OpenVoice is also computationally efficient, costing tens of times less than commercially available APIs that offer even inferior performance. To foster further research in the field, we have made the source code and trained model publicly accessible. We also provide qualitative results in our demo website. OpenVoice has been used by more than 2M users worldwide as the voice engine of MyShell.ai

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tell me Habibi, is it Real or Fake?

    cs.CV 2025-05 conditional novelty 7.0 of 10

    ArEnAV, the first large-scale Arabic-English code-switched audio-visual deepfake dataset, makes current state-of-the-art detectors fail much more than on monolingual data.

  2. Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil

    eess.AS 2026-07 conditional novelty 6.0 of 10

    State-of-the-art audio deepfake detectors severely degrade on Brazilian Portuguese political speech, and the main source of performance gaps is the synthesis method, not demographic traits.

  3. SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.

  4. Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A cross-modal watermarking method embeds authentic speech into video frames, enabling recovery of the original audio and localization of tampered segments after voice cloning or lip-sync manipulation.

  5. Dataset of News Articles with Provenance Metadata for Media Relevance Assessment

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark dataset and two tasks let researchers test whether AI systems can judge if a news image's recorded location and date match the article, with current chatbots scoring 64-81% on location but 42-58% on date.

  6. StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

    cs.MM 2025-06 conditional novelty 6.0 of 10

    StarVC is an autoregressive voice conversion model that generates text tokens before acoustic tokens, improving linguistic fidelity while retaining speaker similarity.

  7. De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Existing voice-protection perturbations succeed only against naive attackers; a phoneme-guided purification-refinement pipeline restores cloneability of protected speech for most VC models.

  8. DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation

    cs.SD 2025-06 reject novelty 5.0 of 10

    DS-TTS adds a second MFCC-based style encoder and a length-adaptive variance adapter to a StyleSpeech-style TTS model, reporting higher speaker similarity but not lower WER than two strong baselines.

Pith tools