Pith. sign in

Paper Citation Record · LEDGER

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2505.16691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16691 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:58:58.610208Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:31:12.289334Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0f6a7af-550c-4223-a65f-d92468978971 · outbound

This paper cites YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.253851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.253851Z digest=sha256:a038226f369267bfd1fd693bbb4ace207d89b560d341b2171bd68f54c915a508

Observation bf85f6d3-9313-4d5c-9dff-88dffc891119 · outbound

This paper cites Towards Robust Speech Representation Learning for Thousands of Languages.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Towards Robust Speech Representation Learning for Thousands of Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.363483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.363483Z digest=sha256:7b6a76b316fb33050e7d9a82259721110333c6c4e6b3c56667292847cfcc0adb

Observation 13a8820e-17dd-4d1a-95bf-a9e3cef8c4c1 · outbound

This paper cites Diff-HierVC: Diffusion-based Hierarchical Voice Conversion with Robust Pitch Generation and Masked Prior for Zero-shot Speaker Adaptation.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Diff-HierVC: Diffusion-based Hierarchical Voice Conversion with Robust Pitch Generation and Masked Prior for Zero-shot Speaker Adaptation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.506395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.506395Z digest=sha256:bc64f21a98cfcd60b9310f64470de3bd9fe33ea1d2f72cfe37401edb9c045918

Observation 1adbc139-7c13-48a6-839f-2e209a6be972 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.651626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.651626Z digest=sha256:0c65d523969cc2687cf8b2701767d1da8952ad343148d3f2062c7c59514970a9

Observation 5ad2fdbb-073f-4f29-a04d-568c9628d2a4 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.747744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.747744Z digest=sha256:10a7ca70911de57f5a6600ba6f647d8b0051595e23c1422dfb905ce2ac70834e

Observation e23d19c3-e4d0-4c5b-9b03-a0f3341ab500 · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.060521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.060521Z digest=sha256:cef2d00211e0831dbfdcef79b1906d07d620afe8b0b450f9906edcfb81ed377b

Observation 228b822e-6e8d-4a88-ad92-ce19c9c4def3 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.464868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.464868Z digest=sha256:789af2d53390d71bbd80015076ebc14c33974ce13006f4f720484afb9358f3c8

Observation 1c534e74-3d06-4aae-92a9-1ff10d953b22 · outbound

This paper cites SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:58:59.160738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:58:57.566166Z digest=sha256:c55a9cccb62592f5be9298a261ec7c20585abe1f80d3013d1d893d19664f3f3b

Observation 0b9c0f93-9ca5-4759-9b6e-400d1b5d827d · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Zero-shot Voice Conversion with Diffusion Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.726231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.726231Z digest=sha256:6ef7ff942d99f4cc9320563f4adf84d2cd5101f960e765c564b1d6c80c1dcaf4

Observation b0bae56f-0cdf-4a18-b815-56dfb596753d · outbound

This paper cites In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:59:00.048756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:58:57.886020Z digest=sha256:f8c7287c01d65af7c3f9aa7aeb330b6906a0094c81e935d74d10a1e07a555f6e

Observation 47aa4c1a-aabe-44ef-9e16-7d104fc9573b · outbound

This paper cites Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.976936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.976936Z digest=sha256:668e065604cc218d12b1e94e9b322ac6eff2f4781e6dce0af448acec3349b1eb

Observation e4d833d4-08ac-4fa7-8743-bdc62c644824 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.106744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.106744Z digest=sha256:0669a0a252134f5cddbb33ab91fb36e52d7c1cb9500ce002cf0e0f1ed69744dc

Observation ee506292-b81c-45b4-8312-f35493fea5df · outbound

This paper cites VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.237399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.237399Z digest=sha256:4ee358a965d0d44a9e6f0842126630ad9403253b2d55da96632294cc13c815aa

Observation 493a615c-a7b9-4057-bea5-f2569222ea15 · outbound

This paper cites StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.358652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.358652Z digest=sha256:25bcfc817dbf74629fca2fce3999b3da78babe313976838298ac49b35832f95f

Observation 74f6896e-0381-4b69-8a07-941b1548c3bc · outbound

This paper cites Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.610208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.610208Z digest=sha256:81de5998fa979ca329ed82ed566d8e04017d43d75439a1ac19ce7fc179a56dc8

Observation 495d6f42-2df5-42fd-a52b-50529f4c5772 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.472965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.472965Z digest=sha256:633f33d2170ea02011e889a67318f4418d7321c020ad0bcc548fef4e002e110f

Observation 49016cc0-d922-442c-819a-db771d2bb707 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Common Voice: A Massively-Multilingual Speech Corpus

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:55.919987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:55.919987Z digest=sha256:728ed43964b0105d180aeb2b9365dc86bba1ef2bc0969689720369c26ba0c3bb

Observation 3efc365d-704a-4714-8d27-54e70b726133 · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.191969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.191969Z digest=sha256:e31c09423040623bb80deaf89d6c0ffc12c23037fdc2be8834a66fa513555e43

Observation 8ffef98b-e581-4b4d-b210-3f777faee46b · outbound

This paper cites Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.149069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.149069Z digest=sha256:29896dea5e3cdfb537721a63c7cb973c3ff1d39188d9890657dfa5c18f0f15c6

Observation ba9d2b6d-03b8-49d0-ad90-c6c82a1dff5c · outbound

This paper cites Voice Conversion With Just Nearest Neighbors.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Voice Conversion With Just Nearest Neighbors

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.015567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.015567Z digest=sha256:9efcd196ac9a26bf6e621b1923301ebbdb3d19857145c0af347be578cd03ebf5

Observation 56111811-5b5e-4c57-af0d-eb834e641a8d · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.891964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.891964Z digest=sha256:5b839be5375000b2083eef476b59861a763268e4673a0be709dd7269d3f87c3a

Observation 27f11978-8d2c-4b53-bb74-d756d3c3387b · outbound

This paper cites AdaptVC: High Quality Voice Conversion with Adaptive Learning.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion AdaptVC: High Quality Voice Conversion with Adaptive Learning

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:58:59.436990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:58:57.316479Z digest=sha256:2a88eb919a4eb929b6280cbb4ad567a4ac273cb08cc08e88f5e9e44697d6d309

Pith citing papers

Observation ca3a5311-a37c-4fe1-a337-b0f3b844df9c · inbound

Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil cites this paper.

Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T00:31:12.289334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:31:12.289334Z digest=sha256:38921cf327e54a3237030a88b1401a17e41a5aaa6a45a0d6b051cce1edb9d46b