Pith. sign in

Paper Citation Record · LEDGER

GenVC: Self-Supervised Zero-Shot Voice Conversion

As of 9 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 1 inbound Pith citation observation for arXiv:2502.04519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04519 v2

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:30:27.279088Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T12:21:48.698333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy55
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71c1e84d-e24a-42b8-a8bb-bf8fafaa6425 · outbound

This paper cites AutoVC: Zero-Shot V oice Style Transfer with Only Autoencoder Loss,.

GenVC: Self-Supervised Zero-Shot Voice Conversion AutoVC: Zero-Shot V oice Style Transfer with Only Autoencoder Loss,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.786485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.923708Z digest=sha256:3d569be1907fad1065161cf5d3b28e778474b983cd897a51674f6441ba5e83a4

Observation 57116bc2-e5d8-408b-95e9-061b836bf82b · outbound

This paper cites GAZEV: GAN-Based Zero-Shot V oice Conversion Over Non-Parallel Speech Corpus,.

GenVC: Self-Supervised Zero-Shot Voice Conversion GAZEV: GAN-Based Zero-Shot V oice Conversion Over Non-Parallel Speech Corpus,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.770998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.928673Z digest=sha256:0b5996b9b1b01f0823c759e61cdfb518c6f4b9dc3505e2b2ce481bd7afcfc9d8

Observation df37264a-c034-4dc6-b631-090ebe334686 · outbound

This paper cites SIG-VC: A Speaker Information Guided Zero-Shot V oice Conversion System for Both Human Beings and Machines,.

GenVC: Self-Supervised Zero-Shot Voice Conversion SIG-VC: A Speaker Information Guided Zero-Shot V oice Conversion System for Both Human Beings and Machines,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.756520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.933327Z digest=sha256:b814113c41d9dc1b4860ddf83f1f88d54e541ac91aac35cd20017524879c0a2a

Observation 3295ab77-01da-4aaa-aecf-fabe67f77fe5 · outbound

This paper cites An Overview of V oice Conversion and Its Challenges: From Statistical Modeling to Deep Learning,.

GenVC: Self-Supervised Zero-Shot Voice Conversion An Overview of V oice Conversion and Its Challenges: From Statistical Modeling to Deep Learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.741973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.938863Z digest=sha256:fb2ade5ee1d949b5deadb87f3296a9cb805fd6ca7d020a3ae451bd4e4296c7e7

Observation 64126789-70a5-402e-adbe-610c91a006a8 · outbound

This paper cites Prosodic Features for Speaker Veri- fication,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Prosodic Features for Speaker Veri- fication,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.726198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.943697Z digest=sha256:93f6c20569f30db157c2e2196f926e9fe7c1081ed7df52efbe9ac813d90c7bca

Observation 476bdcb9-b688-495a-b790-e32c5dbce96d · outbound

This paper cites FreeVC: Towards High-Quality Text- Free One-Shot V oice Conversion,.

GenVC: Self-Supervised Zero-Shot Voice Conversion FreeVC: Towards High-Quality Text- Free One-Shot V oice Conversion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.711160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.948572Z digest=sha256:d0eaf23f823e9c35ae09e233e4c31871e1225cc0683f5ba969e80a93f25964f1

Observation 538bf111-925b-44b5-8b3a-4e36cf41c6e4 · outbound

This paper cites The Database and Benchmark For the Source Speaker Tracing Challenge 2024,.

GenVC: Self-Supervised Zero-Shot Voice Conversion The Database and Benchmark For the Source Speaker Tracing Challenge 2024,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.697041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.953894Z digest=sha256:c974a30df484faa74e4652440afe224ffbe456b9728d77e4cca2e12d49f572c3

Observation de760ea2-d4a8-4d40-a297-0217ed2dfa9c · outbound

This paper cites NeuralVC: Any-to-Any V oice Conver- sion Using Neural Networks Decoder For Real-Time V oice Conversion,.

GenVC: Self-Supervised Zero-Shot Voice Conversion NeuralVC: Any-to-Any V oice Conver- sion Using Neural Networks Decoder For Real-Time V oice Conversion,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.681244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.958510Z digest=sha256:050ee1fb90b668338283cafead32c3e82b611432b59b5e68bdddd4d34d2461c2

Observation 06902400-f3bf-48cb-9471-d92026eb4610 · outbound

This paper cites Identifying Source Speakers for V oice Conversion based Spoofing Attacks on Speaker Verification Systems,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Identifying Source Speakers for V oice Conversion based Spoofing Attacks on Speaker Verification Systems,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.665812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.962975Z digest=sha256:7c0028746940ac6ba3f51d7734d5f56f51a60e2f362c4f29673e22b87fcf9eaf

Observation 93174113-860c-4897-b38a-40e4f4548454 · outbound

This paper cites Privacy Versus Emotion Preservation Trade-Offs in Emotion-Preserving Speaker Anonymization,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Privacy Versus Emotion Preservation Trade-Offs in Emotion-Preserving Speaker Anonymization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.650766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.967484Z digest=sha256:42a3dbb7387cd5d3a2c21c290fc4b274d3d8eaf7776aada7417b8b0c9f29e0ff

Observation ae623d40-0a59-4f2f-8eb9-5e9288abd926 · outbound

This paper cites Zero-Shot V oice Conversion with Adjusted Speaker Embeddings and Simple Acoustic Features,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Zero-Shot V oice Conversion with Adjusted Speaker Embeddings and Simple Acoustic Features,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.637204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.972138Z digest=sha256:7367f6793a1116466be9eb0c7a60f70b7bca941abf03214cce20b21e45999ca3

Observation 9de9da97-adc0-4f55-9dfa-719cfdff486d · outbound

This paper cites YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,.

GenVC: Self-Supervised Zero-Shot Voice Conversion YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.622650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.976480Z digest=sha256:0ea0c5b473eca59f6ee75637fabd1161816e61f5f650397eed584c2e1e17b03a

Observation bbf74823-d13d-4fd4-ba91-14c7037c60c9 · outbound

This paper cites Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.607510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.980919Z digest=sha256:48196f9ce6e65d6527349d6f936d2d0b58ba21d82715ae1d464efdf4ba816366

Observation 283487b7-4969-4266-b073-6e2aedeb86a4 · outbound

This paper cites NANSY++: Unified V oice Synthesis with Neural Analysis and Synthesis,.

GenVC: Self-Supervised Zero-Shot Voice Conversion NANSY++: Unified V oice Synthesis with Neural Analysis and Synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.592358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.985319Z digest=sha256:cbb89aa86fb16f327db6a66909b20a7d66f8e466adfdf266951309a08f4ef5d3

Observation 4e19c775-5283-49ee-b401-6287a4341f94 · outbound

This paper cites Better speech synthesis through scaling.

GenVC: Self-Supervised Zero-Shot Voice Conversion Better speech synthesis through scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:26.989835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:26.989835Z digest=sha256:0b17f799ce17599c3ee0d02ec776bc74b2c0b457685b656fe2dc282d4e2919f3

Observation 4c70f03d-fd70-4823-82d4-4c41599edd4f · outbound

This paper cites LM-VC: Zero-Shot V oice Conversion via Speech Generation Based on Language Models,.

GenVC: Self-Supervised Zero-Shot Voice Conversion LM-VC: Zero-Shot V oice Conversion via Speech Generation Based on Language Models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.577280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.995136Z digest=sha256:69e29cb84016c468e51ab115f4fc94a963f4aa50c7fe5485bac77b5c36f1b128

Observation 26d1928a-3ae2-404b-83e9-36498cc8c7cb · outbound

This paper cites AudioGen: Textually Guided Audio Generation,.

GenVC: Self-Supervised Zero-Shot Voice Conversion AudioGen: Textually Guided Audio Generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.561580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:26.999490Z digest=sha256:a326f5d04b5d432d8824d44591563e8532b5aafe7f3bdf140fe3f3b83fe20389

Observation a7b5bfff-4835-4fca-9494-232bfbd952d1 · outbound

This paper cites Towards audio language modeling -- an overview.

GenVC: Self-Supervised Zero-Shot Voice Conversion Towards audio language modeling -- an overview

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.004124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.004124Z digest=sha256:b04efb9801f3481f409bdd0619772d46caaf479d2aa3570efa362aa18620b3be

Observation d04fdc0e-588b-44ab-8ec1-331074856051 · outbound

This paper cites AudioLM: A Language Modeling Approach to Audio Generation,.

GenVC: Self-Supervised Zero-Shot Voice Conversion AudioLM: A Language Modeling Approach to Audio Generation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.545081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.009093Z digest=sha256:d75590d5cf4891293a1c5dfb2359ecc8441cfe9f622b23916c2a4adba3d70205

Observation 206007ce-f73a-4dd6-94e2-9d81f43803f1 · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec,.

GenVC: Self-Supervised Zero-Shot Voice Conversion SoundStream: An End-to-End Neural Audio Codec,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.529707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.013711Z digest=sha256:a474a9276ee4619d3aee75e585fb2746d505f72f7ca784c3746845e70ec6ce07

Observation 883ded18-f916-4399-98d9-a62732267a98 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.513984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.018358Z digest=sha256:aa5c69bb1167464dcb350b8287776311e351541b0f90bc1455f43ec19fa204b4

Observation f221b66a-5587-46ab-8e7c-6fc5c67c127e · outbound

This paper cites V oiceCraft: Zero-shot speech editing and text-to-speech in the wild,.

GenVC: Self-Supervised Zero-Shot Voice Conversion V oiceCraft: Zero-shot speech editing and text-to-speech in the wild,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.498449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.023068Z digest=sha256:aa37a9fbb50c45d33a7925e1f27ed1e56c1d100a806b66e1d538b6c91bf397c5

Observation c5f122a8-dc71-4a6f-8a70-27c097e0ed36 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

GenVC: Self-Supervised Zero-Shot Voice Conversion CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.027812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.027812Z digest=sha256:dd16702bee393ddf802264f4941f4aa16c35f3c0054675ba7c2cc7e63be864b9

Observation 3d79ea97-d09b-43cd-ab5b-75a2af4c6e69 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

GenVC: Self-Supervised Zero-Shot Voice Conversion CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.032838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.032838Z digest=sha256:fa83d4441863eadb3307dca46722d22b1475f87e8ca58aade21c752311d6682e

Observation 87c2e8c0-a97c-4489-87a2-d93220f7b470 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

GenVC: Self-Supervised Zero-Shot Voice Conversion Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.038490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.038490Z digest=sha256:380ad7c302c293f723494c0b79f3171772b968b92356e81f0f85845798396ee2

Observation d06cc4bc-1dc0-4259-bfc5-8c0e0d9f6c57 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.482999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.043717Z digest=sha256:ab8036f87748a72c3ab0eb7c94c7942b73a1896b01e3c5b738a088d1b2ad12ba

Observation 3c7bba64-613b-4124-b151-b901c08c3609 · outbound

This paper cites High Fidelity Neural Audio Compression,.

GenVC: Self-Supervised Zero-Shot Voice Conversion High Fidelity Neural Audio Compression,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.048979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.048979Z digest=sha256:9baaf8bd3e3232f0618b02f018cac9a4345499bdf73f137b29a3bd6a52dcad6f

Observation 4584d383-7a8d-4aa6-adec-bfdb1e2a0d79 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

GenVC: Self-Supervised Zero-Shot Voice Conversion Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.053709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.053709Z digest=sha256:917cd46bd2dc94bfc8b90b966cb12102d8b33be26e25c296807798a832deb453

Observation 2e6b6df4-1043-4cc5-8546-9e5b37c4b2af · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

GenVC: Self-Supervised Zero-Shot Voice Conversion VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.058372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.058372Z digest=sha256:d7853bd8230625f55160b36b2c25aea749f2fec0df7fbbc791cb332ab52e4201

Observation de3460c9-6464-4313-8607-a64dd04f09bd · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

GenVC: Self-Supervised Zero-Shot Voice Conversion CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.062791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.062791Z digest=sha256:3a09a681df13ab1f45e0e7e9332e628e6cb0b6743c609adae9a591606ceb6b98

Observation dc341715-b904-4c60-a1c7-d36307712231 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer,.

GenVC: Self-Supervised Zero-Shot Voice Conversion MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.457575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.067170Z digest=sha256:d9bfaa62a9212ddca6d2db408420ced664729a3767395e4a5a9787e756a57f60

Observation f7010b0d-f354-4d29-956a-7f9edc90c6e6 · outbound

This paper cites Simple and Controllable Music Generation,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Simple and Controllable Music Generation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.442369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.071252Z digest=sha256:53e2e49527a874a9d8ce2f3388985a1d73ef233296fb62de916fcc49815bf384

Observation 89e60a65-2387-4e5f-aa28-4397eca88b5d · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

GenVC: Self-Supervised Zero-Shot Voice Conversion Moshi: a speech-text foundation model for real-time dialogue

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.075348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.075348Z digest=sha256:bf72849256a897300b68c0f94916f091c515040fb8fb8f37ca2a24b68a0118a7

Observation 017cc5e1-ca13-49a4-b0eb-c714a05140fb · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,.

GenVC: Self-Supervised Zero-Shot Voice Conversion XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.427706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.079438Z digest=sha256:ea6eaedf96d0c275638b037e2d1860261db4343b58d4a417016726b4d434d21e

Observation cf02f2e7-e96a-4201-ba1b-a3c5f0534f42 · outbound

This paper cites StreamV oice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot V oice Conversion,.

GenVC: Self-Supervised Zero-Shot Voice Conversion StreamV oice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot V oice Conversion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.413035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.084111Z digest=sha256:3e6b965d3422df4e0991c2fd19300565ebca6ea534ec5fa9e5c01aa2e92ad8dc

Observation e316393a-9101-4af5-8dc3-630a32aa095a · outbound

This paper cites Vevo: Control- lable Zero-Shot V oice Imitation with Self-Supervised Disentanglement,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Vevo: Control- lable Zero-Shot V oice Imitation with Self-Supervised Disentanglement,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.397784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.088859Z digest=sha256:30d0ca296f443241e42264d4365fd5ab70b9b5ca2e564b8972fc46332aa0f374

Observation b59e1905-62e8-44d2-91c3-ab4a9795d566 · outbound

This paper cites The VoicePrivacy 2024 Challenge Evaluation Plan.

GenVC: Self-Supervised Zero-Shot Voice Conversion The VoicePrivacy 2024 Challenge Evaluation Plan

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.094409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.094409Z digest=sha256:d0194edaaaa0a86b697df4e59d6b2d7a8d6b0c9aa8439cefd19a1ff182534465

Observation e8faea19-f188-4800-ae26-702bcc22c750 · outbound

This paper cites Self-Supervised Speech Representations are More Pho- netic than Semantic,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Self-Supervised Speech Representations are More Pho- netic than Semantic,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.382328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.099365Z digest=sha256:fc64be652e84b6867508dd1841770739659e4224d28be6749ea72adbf120f536

Observation 53d349fd-6a88-4d19-92ee-87da1b6168dc · outbound

This paper cites S2VC: A Frame- work for Any-to-Any V oice Conversion with Self-Supervised Pretrained Representations,.

GenVC: Self-Supervised Zero-Shot Voice Conversion S2VC: A Frame- work for Any-to-Any V oice Conversion with Self-Supervised Pretrained Representations,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.366410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.103917Z digest=sha256:d72068569958a3325ad2af517494a214b7997c9d8bb690ef725ca13b511d6c54

Observation 6ab9a73c-379e-4d93-9b55-975ec331535f · outbound

This paper cites S3PRL-VC: Open-Source V oice Conversion Framework with Self-Supervised Speech Representations,.

GenVC: Self-Supervised Zero-Shot Voice Conversion S3PRL-VC: Open-Source V oice Conversion Framework with Self-Supervised Speech Representations,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.349822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.108491Z digest=sha256:3e10646fc335a8218b661f83c30ab701a308c75058128853262eefc00802624b

Observation a16894b2-ec46-47e3-81ce-2c0ba66f305e · outbound

This paper cites Neural Discrete Representation Learning,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Neural Discrete Representation Learning,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.331698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.113146Z digest=sha256:51981e285555f1eb6fe80749fed2305e66ab93e6fab07d14cdd913aa03f33af7

Observation 30a846b8-d196-4590-949b-9d622174ba4c · outbound

This paper cites Attention is All you Need,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Attention is All you Need,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.313805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.117803Z digest=sha256:4fb14dbe374afbcf13f78256b3d22dbae5856d4b9839d08c00e57669a087bdfb

Observation f26daa61-61b2-4ba1-bac6-0a2fead6521e · outbound

This paper cites Flamingo: A Visual Language Model for Few-Shot Learning,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Flamingo: A Visual Language Model for Few-Shot Learning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.298619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.122721Z digest=sha256:2de64d3720a60553fec06a652b2e3586ec925a606828f7c61841762bfecbddfd

Observation 7288cac4-2c4e-4d43-afa4-b221ae96f723 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers,.

GenVC: Self-Supervised Zero-Shot Voice Conversion NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.282304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.127444Z digest=sha256:003820eee04e32e5231826add938bc2e0ae0964b8a8963121fceb243b5e4d3a4

Observation 4fd34713-3c3d-48c1-a1fe-7319b6d7fe23 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,.

GenVC: Self-Supervised Zero-Shot Voice Conversion HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.266549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.131978Z digest=sha256:0397fb1b1afb6dad7bbc032fd67a16bff48463f6f6eca490a6c70c6ffb425a8f

Observation 77a7b491-dd80-4067-af6a-a7683c040642 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech,.

GenVC: Self-Supervised Zero-Shot Voice Conversion LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.250786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.136838Z digest=sha256:5fe6b005614777ca021b26179413224c3d18d68f9de51148f1f6f2fbdc6aeeb6

Observation 21fcf77c-46f6-4cff-8ffd-0aed48e9c707 · outbound

This paper cites Common V oice: A Massively-Multilingual Speech Corpus,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Common V oice: A Massively-Multilingual Speech Corpus,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.142012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.142012Z digest=sha256:5744549560b58bc280920c7157a9df55f932fb80d56194c96418fbd08c3629bd

Observation ddf739fd-5418-4521-b8e7-fd27cd1818d8 · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research,.

GenVC: Self-Supervised Zero-Shot Voice Conversion MLS: A Large-Scale Multilingual Dataset for Speech Research,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.146476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.146476Z digest=sha256:51521c5839ec4dc34f65c82eb7831f7b23e81b30af49516b10fc11ce309d450f

Observation d3d017ef-8948-4d91-9071-165606f983ca · outbound

This paper cites CMU ARCTIC Databases for Speech Synthesis,.

GenVC: Self-Supervised Zero-Shot Voice Conversion CMU ARCTIC Databases for Speech Synthesis,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.211890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.151196Z digest=sha256:3b0d756fa36dab79421fbf6725b661d84d8f8c307479fce159f18629303d36e9

Observation 48910e9d-6469-4166-9e87-0251622825f1 · outbound

This paper cites The EMIME Mandarin Bilingual Database,.

GenVC: Self-Supervised Zero-Shot Voice Conversion The EMIME Mandarin Bilingual Database,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.195744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.155742Z digest=sha256:e981008bed31082e6a0f9d285f2e31368d3349163d110f20d607ca1a5c6acd32

Observation f0fc4360-e347-43b2-b436-a216be764e81 · outbound

This paper cites Librispeech: An ASR Corpus Based on Public Domain Audio Books,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Librispeech: An ASR Corpus Based on Public Domain Audio Books,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.179990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.160700Z digest=sha256:adcbd71c4f12ddf99bc9590d36d4b86a87eb655748abb06fdfe13fc476fc17e7

Observation d3f548f3-a30c-4472-82a9-f95c79720c30 · outbound

This paper cites V oxCeleb: A Large-Scale Speaker Identification Dataset,.

GenVC: Self-Supervised Zero-Shot Voice Conversion V oxCeleb: A Large-Scale Speaker Identification Dataset,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.163119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.165129Z digest=sha256:dd357158d614e3c06873c15b4164679019926e28eeb0fb433335597364ae71ff

Observation 1eac42d6-8919-4482-86c3-ed284ed169d8 · outbound

This paper cites V oxCeleb2: Deep Speaker Recognition,.

GenVC: Self-Supervised Zero-Shot Voice Conversion V oxCeleb2: Deep Speaker Recognition,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.148092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.169658Z digest=sha256:698dec1749ff3ebc6ede2a2c73d9b8d4ab0055ee42ca6c093d35da96210e0a5f

Observation d2f1e9d4-05c6-4b38-b143-3270e332a156 · outbound

This paper cites ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers,.

GenVC: Self-Supervised Zero-Shot Voice Conversion ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.132448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.174264Z digest=sha256:1aa0af281124a7077fffbebbab8e34ae086b8d8311a3841faa1130f22768183b

Observation 3e0a27f5-c905-4ff7-bf7b-7f44d10c0a7b · outbound

This paper cites Language Models Are Unsupervised Multitask Learners,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Language Models Are Unsupervised Multitask Learners,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.116030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.178754Z digest=sha256:7266a79d0f2f3cedbaff0dcf2cf2ae9d2238ed3ea7950b54ba5290dddc28dca1

Observation 1f6270ec-ed66-4316-94af-48eef01aa0de · outbound

This paper cites Multi-Scale Sub-Band Constant-Q Transform Discriminator for High-Fidelity V ocoder,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Multi-Scale Sub-Band Constant-Q Transform Discriminator for High-Fidelity V ocoder,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.099154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.183079Z digest=sha256:ed5284e1907c11a5eb389cc37fb60bdf6c2e6854113b74a17c57f8779e73a567

Observation f4122fc2-ac26-4c4d-ad20-377b459695d2 · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre- Training for Full Stack Speech Processing,.

GenVC: Self-Supervised Zero-Shot Voice Conversion WavLM: Large-Scale Self-Supervised Pre- Training for Full Stack Speech Processing,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.187548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.187548Z digest=sha256:ffc7f2632a78dfb606b328d99308fe7370833a0a4c7c302422ad92ed3c63f1ab

Observation 65b9d42d-a89e-4127-83db-e8c489112a3e · outbound

This paper cites The T05 System for The V oiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech,.

GenVC: Self-Supervised Zero-Shot Voice Conversion The T05 System for The V oiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.071545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.192099Z digest=sha256:1a82d139d4bf80b9b43f220b8aac3704104135c92055f030baa506ab0c215e7b

Observation 0908d85b-52f9-4b4e-a5a8-e5168c5ac5d8 · outbound

This paper cites Bias and Statistical Significance in Evaluating Speech Synthesis with Mean Opinion Scores,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Bias and Statistical Significance in Evaluating Speech Synthesis with Mean Opinion Scores,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.056193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.196870Z digest=sha256:18ab89286cb7092254bbe9f04225deb090eda1c312c263360d2d11ca6109cd38

Observation 73d23412-d759-4870-b8a1-8e9756cca0da · outbound

This paper cites Good Practices for Evaluation of Synthesized Speech,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Good Practices for Evaluation of Synthesized Speech,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.201566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.201566Z digest=sha256:7ea23db91a3ecce91dbc637f1408b8ec00bdcba489e03372f8223529440b926b

Observation f2f18524-0dae-410b-90d5-34ab5563589d · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,.

GenVC: Self-Supervised Zero-Shot Voice Conversion ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.039189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.206324Z digest=sha256:cabade517b306a1a260bbb5c6528c810c315d88821ee7a55f6d32b8b829df636

Observation 98089578-ca5b-461d-adeb-6d92b72c1c06 · outbound

This paper cites Scaling Laws for Neural Language Models.

GenVC: Self-Supervised Zero-Shot Voice Conversion Scaling Laws for Neural Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.210988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.210988Z digest=sha256:084226e48add99b9f9400cc58ac13b34b6aea92b29dcbd6e49e4e40fea7c0940

Observation 7f04a0d9-cb12-4ceb-9e0f-57512511c733 · outbound

This paper cites Deep Residual Learning for Image Recognition,.

GenVC: Self-Supervised Zero-Shot Voice Conversion Deep Residual Learning for Image Recognition,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.024116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.215698Z digest=sha256:6b414009a552834a0e17f65180798ba9938d6fe5b29d79ea234f25834f302d3b

Observation 2c692cd9-628e-45a2-97e9-f3fb3be45556 · outbound

This paper cites The first two encoder convolutional layers upsample the input dimensionality to the DV AE hidden dimension of 1024.

GenVC: Self-Supervised Zero-Shot Voice Conversion The first two encoder convolutional layers upsample the input dimensionality to the DV AE hidden dimension of 1024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:28.008877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.220159Z digest=sha256:af5e99328d8608e9f7fe5d67a44c1ad283d0d41a46fc773d49373590cb2f27a5

Observation a61af82e-6815-44e1-b56d-b39355c0d6b8 · outbound

This paper cites The learned queries attends to all input frames, transforming them into fixed- length representations.

GenVC: Self-Supervised Zero-Shot Voice Conversion The learned queries attends to all input frames, transforming them into fixed- length representations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.993036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.225587Z digest=sha256:ec31fe3d4f2667a17382474cc921d59535dfe4f1bcb549fd0aba874c503bc9dd

Observation b770d178-6fed-4271-937f-a964500f4b1f · outbound

This paper cites 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5 and 5 are allowed on the half-point scale.

GenVC: Self-Supervised Zero-Shot Voice Conversion 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5 and 5 are allowed on the half-point scale

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.979268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.229649Z digest=sha256:040f722b14daedbddcfc1a20ed2b538c5033ced856ad240fe77e89a15b4575ba

Observation 68f8817e-affe-45a1-81d0-288b84771820 · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.965523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.234565Z digest=sha256:eb606ff4bd511c05f6fe061abcb7385392baef2a44661f1a0376762b6fd0841e

Observation bcc41a1f-95d0-4356-8017-833399b18f64 · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.951564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.238691Z digest=sha256:b9fa79988c98e6ff5e027bc3a94863eb56885aff0648000cb0dcd5b70daa9942

Observation 2e924d13-7264-4f1f-a34e-8a8377acfc2b · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.937513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.243255Z digest=sha256:c17acd50c368b12abe126f6e96dc9a20d7a3764796fad9f32fa9ae8725f54a05

Observation 9531351d-38d9-4636-9ab3-127abfa78418 · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.922807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.248147Z digest=sha256:95f75f55cf40725bbc5253440b476ef7582a8dc3b244c4b6a8f629dd86ca1eb2

Observation d6f96f98-5b1f-4c1d-b957-4431079b99a0 · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.907383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.252499Z digest=sha256:7239aef8a2dcd693c40a66df4f6721cc4fb744131a57c0d1ac30be3c777a09dd

Observation 0a08ca7b-02ba-4ae9-8f00-e10c947f5196 · outbound

This paper cites Use a scale from 1 (Bad) to 5 (Excellent), with increments of 0.5.

GenVC: Self-Supervised Zero-Shot Voice Conversion Use a scale from 1 (Bad) to 5 (Excellent), with increments of 0.5

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.891406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.256544Z digest=sha256:08b9d91a930f13cdee1877eb901568c733411c3ac8cea8ebd08f1850ecfbb911

Observation c38cca02-6785-4edc-9958-72c691439cb4 · outbound

This paper cites an unresolved cited work.

GenVC: Self-Supervised Zero-Shot Voice Conversion Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:30:27.874432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.260729Z digest=sha256:3d48f8bd84eca4d8b5068cef196cd26bec4098c9b8e2dde001f47dfb596bc884

Observation 649cbab6-ba49-4ae5-9f27-1b4e48a71dfd · outbound

This paper cites For example, the speakers may differ in gender or pitch (e.g., a high-pitched female voice vs.

GenVC: Self-Supervised Zero-Shot Voice Conversion For example, the speakers may differ in gender or pitch (e.g., a high-pitched female voice vs

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.858754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.265152Z digest=sha256:5761837528dbfddcd310672da77d8e581f41d87370a0b98f7eee0f6cd591d36e

Observation ec066b31-e376-4c15-a4a1-43c746ce1885 · outbound

This paper cites It’s clear the speakers are the same gender, but their voices are distinctly different, and their speaking styles differ somewhat.

GenVC: Self-Supervised Zero-Shot Voice Conversion It’s clear the speakers are the same gender, but their voices are distinctly different, and their speaking styles differ somewhat

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.841437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.270024Z digest=sha256:8dc260feea988691533e9c601930638a38226f9e7bab9e5f4980099e2082c676

Observation ecda95be-b217-4955-b8f7-8c711b69230b · outbound

This paper cites The speaker voices are close, and the speaking styles largely match.

GenVC: Self-Supervised Zero-Shot Voice Conversion The speaker voices are close, and the speaking styles largely match

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.824860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.274365Z digest=sha256:d06c6c2ac926b6d6319a6351021f2a5b2fa910c71fed9c8922fe421073e3e856

Observation 98c69fa1-f049-493a-a3cd-c5e23bbccbb9 · outbound

This paper cites The timbre of the speakers is the same, their speaking styles match perfectly, and the environmental background is similar, including any acoustic noise.

GenVC: Self-Supervised Zero-Shot Voice Conversion The timbre of the speakers is the same, their speaking styles match perfectly, and the environmental background is similar, including any acoustic noise

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:30:27.809162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:30:27.279088Z digest=sha256:d4cc41e12328079a3f1c005e8ca30804120c12283f6c4d95917558b7b5f511c4

Pith citing papers

Observation 7860c175-f3de-4ccb-9b1b-5cb692542da4 · inbound

Universal Speech Content Factorization cites this paper.

Universal Speech Content Factorization GenVC: Self-Supervised Zero-Shot Voice Conversion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T12:21:48.698333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:21:48.698333Z digest=sha256:451437419135f6f157f8cad7d8a4faa3098c4ee92658582617f3cd68f9b25039