Pith. sign in

REVIEW 3 cited by

HLTCOE JHU Submission to the Voice Privacy Challenge 2024

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.08913 v2 pith:EAUDONAR submitted 2024-09-13 eess.AS cs.LG

classification eess.AScs.LG
keywords systemsvoiceconversionbetterchallengeincludingmethodprivacy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-to-speech (TTS) based systems including Whisper-VITS. We found that while voice conversion systems better preserve emotional content, they struggle to conceal speaker identity in semi-white-box attack scenarios; conversely, TTS methods perform better at anonymization and worse at emotion preservation. Finally, we propose a random admixture system which seeks to balance out the strengths and weaknesses of the two category of systems, achieving a strong EER of over 40% while maintaining UAR at a respectable 47%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SegReConcat: A Data Augmentation Method for Voice Anonymization Attack

    cs.SD 2025-08 conditional novelty 6.0 of 10

    SegReConcat, a word-shuffle-and-concatenate augmentation, improves attacker speaker verification against five of seven voice anonymization systems in the VPAC 2024 benchmark.

  2. Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A context-dependent encoding of phoneme durations identifies speakers far better than average-duration vectors and remains effective on anonymized speech without retraining on anonymized data.

  3. EASY: Emotion-aware Speaker Anonymization via Factorized Distillation

    eess.AS 2025-05 conditional novelty 5.0 of 10

    EASY separates speaker identity, linguistic content, and emotion through sequential factorized distillation, and reports better privacy and emotion preservation than prior VoicePrivacy 2024 systems.

Pith tools