Pith. sign in

Paper Citation Record · LEDGER

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2505.16691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16691 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:58:58.610208Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:31:12.289334Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0f6a7af-550c-4223-a65f-d92468978971 · outbound

This paper cites YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.253851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.253851Z digest=sha256:3a00661fb7434ace730c53b1ed81f0c1368ba66ebb6f7aa5879785ccc7b77e62

Observation bf85f6d3-9313-4d5c-9dff-88dffc891119 · outbound

This paper cites Towards Robust Speech Representation Learning for Thousands of Languages.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Towards Robust Speech Representation Learning for Thousands of Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.363483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.363483Z digest=sha256:05dcf4f2a3e145f6c73c0f8a638ef99e3a97c6940678c51cbcdba570e7502f7a

Observation 13a8820e-17dd-4d1a-95bf-a9e3cef8c4c1 · outbound

This paper cites Diff-HierVC: Diffusion-based Hierarchical Voice Conversion with Robust Pitch Generation and Masked Prior for Zero-shot Speaker Adaptation.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Diff-HierVC: Diffusion-based Hierarchical Voice Conversion with Robust Pitch Generation and Masked Prior for Zero-shot Speaker Adaptation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.506395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.506395Z digest=sha256:601fd124a007f223e0e6aed907bf256c52487bc8a60ee02eb830f0a6c212693a

Observation 1adbc139-7c13-48a6-839f-2e209a6be972 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.651626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.651626Z digest=sha256:85aa45297a3a235066ad906ee22cc18df589f5f5778ffec58331265a62e15c66

Observation 5ad2fdbb-073f-4f29-a04d-568c9628d2a4 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.747744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.747744Z digest=sha256:1f08c872936455d889b66f2f31b83b1c65ceacca85f5382c5b86d60cc3de07e8

Observation e23d19c3-e4d0-4c5b-9b03-a0f3341ab500 · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.060521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.060521Z digest=sha256:c4b22f6b48fe61354be2e45b8df75225f81ed4657d3b2fa1b907615a74347941

Observation 228b822e-6e8d-4a88-ad92-ce19c9c4def3 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.464868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.464868Z digest=sha256:68fe6c7b7389bc6d0f440be2b48681251bc220fc73e3129227546959cf41c5ea

Observation 1c534e74-3d06-4aae-92a9-1ff10d953b22 · outbound

This paper cites SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:58:59.160738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:58:57.566166Z digest=sha256:f65c870f1ecef28fcdf36024e785bf130177f73e66c9291dfa6e08e40b9106a8

Observation 0b9c0f93-9ca5-4759-9b6e-400d1b5d827d · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Zero-shot Voice Conversion with Diffusion Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.726231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.726231Z digest=sha256:36902ef80ed1a9a0d3a68b103cf3d8ef25cc47e11c3727c3e6d7186b761542de

Observation b0bae56f-0cdf-4a18-b815-56dfb596753d · outbound

This paper cites In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:59:00.048756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:58:57.886020Z digest=sha256:9cc8b7aa366ed281c55d2144535f464e10e55fa368fc9b146de7b19ec809019a

Observation 47aa4c1a-aabe-44ef-9e16-7d104fc9573b · outbound

This paper cites Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.976936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.976936Z digest=sha256:20940f7c3ca8a6fa471bddd5960cafc678094a59c66dd6b9f46a02ed99ae2f3d

Observation e4d833d4-08ac-4fa7-8743-bdc62c644824 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.106744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.106744Z digest=sha256:8220d16b5744ad3b328ede9a99e6d623e11bd6d95a285524945b7722733e7d57

Observation ee506292-b81c-45b4-8312-f35493fea5df · outbound

This paper cites VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.237399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.237399Z digest=sha256:f1f056ed5f01e1ea45039ecb20414e5c5d76b137e239fee4a7ced58bbb111948

Observation 493a615c-a7b9-4057-bea5-f2569222ea15 · outbound

This paper cites StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.358652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.358652Z digest=sha256:aefeceafb869e53e778c2c1f55961e70361a71ff4750fa9d69ca212c64cf6b00

Observation 74f6896e-0381-4b69-8a07-941b1548c3bc · outbound

This paper cites Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.610208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.610208Z digest=sha256:96122b5f718d51a25d503a6b960d471f5bf0322b14eafc1848fd6288734128c6

Observation 495d6f42-2df5-42fd-a52b-50529f4c5772 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.472965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.472965Z digest=sha256:03cf8e458fba51501496c90fa49e6a17df6e093c5590c4b4000a7da90815e71d

Observation 49016cc0-d922-442c-819a-db771d2bb707 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Common Voice: A Massively-Multilingual Speech Corpus

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:55.919987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:55.919987Z digest=sha256:9876f02734d578b6e92f5e1b83dba01aa064e06ed4c7a8e13bb1d50b15a3ebed

Observation 3efc365d-704a-4714-8d27-54e70b726133 · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.191969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.191969Z digest=sha256:128752642b91401d8d450f5fb0fa236740bf04b7d812ffba54e2b1af8c2cebdb

Observation 8ffef98b-e581-4b4d-b210-3f777faee46b · outbound

This paper cites Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.149069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.149069Z digest=sha256:0305d087453e0caf97e220a2121f3ab5b8af10f4d60851728e8243c0a9f20486

Observation ba9d2b6d-03b8-49d0-ad90-c6c82a1dff5c · outbound

This paper cites Voice Conversion With Just Nearest Neighbors.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Voice Conversion With Just Nearest Neighbors

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.015567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.015567Z digest=sha256:e537b21812b203d0cb53876cc7429080b01a9a8ac25b4fc35c13a88e44446380

Observation 56111811-5b5e-4c57-af0d-eb834e641a8d · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:56.891964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:56.891964Z digest=sha256:07d497b9206773a9c9b92b92339909c502c454ea8a2a77fc405679d236eb8a57

Observation 27f11978-8d2c-4b53-bb74-d756d3c3387b · outbound

This paper cites AdaptVC: High Quality Voice Conversion with Adaptive Learning.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion AdaptVC: High Quality Voice Conversion with Adaptive Learning

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:58:59.436990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:58:57.316479Z digest=sha256:98c9f05dc1c9b06236a3d2fd45247811c55e2a691a00e5d1cd3eb7273390970a

Pith citing papers

Observation ca3a5311-a37c-4fe1-a337-b0f3b844df9c · inbound

Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil cites this paper.

Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T00:31:12.289334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:31:12.289334Z digest=sha256:ee37afa2c877711b3f69305432cef9b659930ccce7e372d1c937cfdba8eae289