Pith. sign in

Paper Citation Record · LEDGER

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

As of 8 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 7 inbound Pith citation observations for arXiv:2502.07243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07243 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:26:19.369265Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:21:59.575812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:38.691071Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1234a951-ef09-4289-a4bb-218485f1202d · outbound

This paper cites Neural discrete representation learning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural discrete representation learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.088321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.088321Z digest=sha256:ad7c194d768bdd7202a8ae5fdd5a2bd42dfc56bc3a4acb3ac374398afb43ccd7

Observation 9621dbf3-776a-4578-87e4-6d5cc7130f2d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.092101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.092101Z digest=sha256:fc34f7a3a7d3c92bfaf7a9785cd7fee31e34cda52c09ea49890d7d2bf2fa9dae

Observation cdcc9f2f-cdc6-46b8-8c46-646cb2c70740 · outbound

This paper cites An overview of voice conversion systems.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement An overview of voice conversion systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.095768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.095768Z digest=sha256:178447d6eaeeff705f4f00f430ab3904a0fd9bbc6c15ea23a963474f99d7a2c3

Observation c384d50e-393a-4c38-9e76-05d4c1c59737 · outbound

This paper cites An overview of voice con- version and its challenges: From statistical modeling to deep learning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement An overview of voice con- version and its challenges: From statistical modeling to deep learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.099143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.099143Z digest=sha256:8ed4f700363947b470bf004a7e6aa56ff73fbe2435d031d35714b7706f2c31ed

Observation 2ef66d9b-3efe-4cc5-9fdb-32a19f32fb3b · outbound

This paper cites Foreign accent conversion in computer assisted pronunciation training.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Foreign accent conversion in computer assisted pronunciation training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.102330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.102330Z digest=sha256:42e42e3cc394cb914196b2659bfc8a6bcdd4a33a24309726c9b8b89e613dcf45

Observation 9f2d5106-5bcc-4539-92f2-5c08a4174d82 · outbound

This paper cites L2-ARCTIC: A non-native english speech corpus.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement L2-ARCTIC: A non-native english speech corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.105470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.105470Z digest=sha256:edfc857f8492a453854c2c2765bc6632fa8695f7ea3177144e715fd4d86d60bb

Observation 84010b50-fc81-409d-b51c-3e124781da88 · outbound

This paper cites Emotional voice conversion: Theory, databases and ESD.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Emotional voice conversion: Theory, databases and ESD

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.108882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.108882Z digest=sha256:fb11dd6326010d36662ab96f82343d6124d0fbe8980c55dd4e547bf008b72d91

Observation 8f4141dd-7b92-4bc3-a3ed-6c8dade41c33 · outbound

This paper cites Neural Text-to-Speech Synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural Text-to-Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.112021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.112021Z digest=sha256:1609629d2f6169313909525abc1d4c412aa9dd61d6365b4ba3648838a59fac97

Observation bbf8051d-f030-4f39-8e3f-6990403e10e7 · outbound

This paper cites Converting foreign accent speech without a reference.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Converting foreign accent speech without a reference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.114754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.114754Z digest=sha256:37ed85f098abbd209c5cfed5334c84cd78035a0d9d02eb9422ac8c25aa3ea49b

Observation b5ddc4a1-f58f-452b-9dab-64b5718a6ba1 · outbound

This paper cites Sahidullah, Aur ´elien Bellet, Marc Tom- masi, and Emmanuel Vincent.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Sahidullah, Aur ´elien Bellet, Marc Tom- masi, and Emmanuel Vincent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.117635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.117635Z digest=sha256:c966f1d557ef5d1705db0c706ea5d17d491d5e10a2982ef3750b7b1da80066a3

Observation 05a4d2f1-477d-49e4-9574-6aa2484f14a0 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.120659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.120659Z digest=sha256:64a38144260d5f1950c69eb327c0341617782be791768cdd1b33c68cf8a1e966

Observation ab9f82f4-1ab7-4ad4-82d3-3aab0706d44f · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.124096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.124096Z digest=sha256:a9c08c29edf08efd9d2d54d9927c4c6f66bbfa7c3d12d6da3b9b0e2a9330758e

Observation 11b705f3-918e-4255-a214-222525ab2f0e · outbound

This paper cites Maskgct: Zero-shot text-to- speech with masked generative codec transformer.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Maskgct: Zero-shot text-to- speech with masked generative codec transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.127638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.127638Z digest=sha256:9121e65e6d4fe636c1f08355a970c92148175a168434678597b6fe83d68e6dc2

Observation 07ff5752-9eaa-4c8d-a501-508a918d48fe · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.130748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.130748Z digest=sha256:51ccb7100fc3ec713e117305202ac24ca37bce8515f84d0bf95452d7b872cc1d

Observation 6a2d694b-e80e-42a1-9d7c-b13db2eb9f4b · outbound

This paper cites Speech resynthesis from discrete disentan- gled self-supervised representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Speech resynthesis from discrete disentan- gled self-supervised representations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.133715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.133715Z digest=sha256:760aaaaf5a16fec90a63657a41ec0ce4c99836bb059f11fb90ec9a9eeac4d2fd

Observation 6e9e5b21-1953-45c9-9c85-9baf1e1642fb · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.136859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.136859Z digest=sha256:7000a8429fe66274d1c30eaee402d800c4cb43dec158980bec9750792918307c

Observation 53b580d5-50d5-4831-b8b8-aab0451b14bc · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.140473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.140473Z digest=sha256:f21cbac2211a193b2a0bf753502aca0ff5dc3b027c1f241356474ecc8c531555

Observation 48cc366c-9511-4c1d-b45c-122c4c30fd9f · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:20.024364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.143740Z digest=sha256:bccb8972deda823c651457b6d20fddc316f7f9ee2deff56cc051ce4b51fd85c5

Observation b16bd86c-730f-47cb-a719-d59fc2465f63 · outbound

This paper cites Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:20.015767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.146950Z digest=sha256:63342287bf735e320edb376a0088cf883bce4a65f29655950f6420652aa7f639

Observation b4e0abf5-deb8-4310-a6aa-de5bb23ebfa5 · outbound

This paper cites VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.150092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.150092Z digest=sha256:962eb6ba207793d3f71315528f88ebb1a242bc232173ec24d42a23d81ae6ae77

Observation 6e1cff98-0b03-451c-8ab3-a3b8d7261506 · outbound

This paper cites Convert and speak: Zero-shot accent conversion with minimum supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Convert and speak: Zero-shot accent conversion with minimum supervision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:20.006904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.153241Z digest=sha256:4c93db89674242f4dce52205cb3ef559e571a82e87b1c00d76676e3e9d41d224

Observation d63fbd8f-c1e4-4798-99a1-4a417dabffc5 · outbound

This paper cites Au- tovc: Zero-shot voice style transfer with only autoencoder loss.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Au- tovc: Zero-shot voice style transfer with only autoencoder loss

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.998020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.156236Z digest=sha256:1d09f63285d075b957563a59e326f012381847d3bf19df5bdb2463912733545e

Observation be8810d0-db54-481e-abb2-683a25f231af · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.158875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.158875Z digest=sha256:78b5e5a2896cef9171d60dcb48f9925a5b824391a7cd73246e088c31b4b25bb6

Observation 712c1ffd-5225-45d3-bd18-6aba84d44a47 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.162381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.162381Z digest=sha256:b56f6dbfaa8b0995586dce76242c0f6b30f7775095af0a0d74db63cd403ee101

Observation 54dc202b-f995-44b5-9425-fc2300fe4d46 · outbound

This paper cites Better speech synthesis through scaling.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Better speech synthesis through scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.165401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.165401Z digest=sha256:ac7d83f2fced49a1be1a85ead3aea70bd488e894de0dbbf609887fc27d8971a7

Observation 8dd7fa07-0c42-4fbc-9577-7def92b7ad11 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.168802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.168802Z digest=sha256:1767322e0369f65ce84cc0078ced06b3c6282f00ff2fb469ac60492047454b2e

Observation 9bf9cba5-4724-401d-84a2-5dc3a569a6c9 · outbound

This paper cites V oicebox: Text- guided multilingual universal speech generation at scale.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oicebox: Text- guided multilingual universal speech generation at scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.172375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.172375Z digest=sha256:d67beb73fd85649af1a8309e4ad58effaff84bb43f0980d95bf89e9bc0e14313

Observation 8fc2728d-5e2a-48df-ab02-34693677b7c2 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.984343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.175536Z digest=sha256:e04cdda44494efdf63aee6ede0381bb9e053c6d75dc3f8179f208998fe2b28c4

Observation ff1696aa-260f-49fe-be8d-3e0e641aa098 · outbound

This paper cites V oice- preserving zero-shot multiple accent conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oice- preserving zero-shot multiple accent conversion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.975383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.178687Z digest=sha256:0e172be7bda472cdddea822764720ec74de84eb3a6d4c51b351f2829199f7e80

Observation 9d8b0e25-1f48-424f-a484-258480a55ff0 · outbound

This paper cites Schuller, and Haizhou Li.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Schuller, and Haizhou Li

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.966785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.181727Z digest=sha256:eb3293b78ff39a39c7943cd94650e8df1ac5a658cc65972453ff673018f2e203

Observation e4bf249b-c83e-41c0-bb00-de4227ca1539 · outbound

This paper cites PA VITS: exploring prosody-aware VITS for end-to-end emotional voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement PA VITS: exploring prosody-aware VITS for end-to-end emotional voice conversion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.958284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.184577Z digest=sha256:c648489a660c91aa924df260b548e7f8419017764e465ea4875a600dbfc4d39c

Observation cb915b3c-08b9-4090-8a15-d8a50bf05a15 · outbound

This paper cites Transfer the linguistic representa- tions from TTS to accent conversion with non-parallel data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Transfer the linguistic representa- tions from TTS to accent conversion with non-parallel data

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.949309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.187544Z digest=sha256:8d33d01fca40bc97b67953f967a57e525a18e3108643968c60593dbf5954e6f6

Observation a52ec863-cd90-4e3b-97da-87c94d9798dc · outbound

This paper cites U-style: Cascading u-nets with multi-level speaker and style modeling for zero-shot voice cloning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement U-style: Cascading u-nets with multi-level speaker and style modeling for zero-shot voice cloning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.939715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.190541Z digest=sha256:5a9a8e7f5f02557cdbbaca8bdf3f1028e49e4ec5652387abded8457a44b8f6da

Observation 32c230e2-b170-4780-855b-e79bba4ccccb · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.193725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.193725Z digest=sha256:c808efc69cb8aea1a2b563cb7252c7af8debf33844181695a6490fdb4d180395

Observation 6baec9ef-0a8b-435f-909f-835a1966cab0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement LLaMA: Open and Efficient Foundation Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.196997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.196997Z digest=sha256:26727e7871e21f728fa102db32f647b285feeb5e52c5892618d8361ce2be1f2b

Observation ffb292e5-e1b2-4a92-b961-b30131fcd04a · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.200113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.200113Z digest=sha256:1097d7fbe4fc127b49245b3e98ba013496396089a30b31f61ead4ffcbef13ddd

Observation f59cd892-b39e-4fb8-b11f-f5882fcac97e · outbound

This paper cites Scalable diffusion models with transformers.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Scalable diffusion models with transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.920850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.202942Z digest=sha256:0cc4e90d7254ac4080716812bdae13c6ddbff8cdfd8c9ea399afa06ab5a7d5fc

Observation f9b274d1-daa9-4b5a-85f6-5411061b25d3 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.205829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.205829Z digest=sha256:717ecf9db260a55c66477160dee7c3904fdc92e017f8bc5982975e1652ce7623

Observation 02e7d8be-5b9d-4760-aba5-1cb3debb871c · outbound

This paper cites One-shot voice conversion by vector quantization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement One-shot voice conversion by vector quantization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.912239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.209096Z digest=sha256:d0720c8bd498e45899804f88336085d05d4a12d67a73f074b111efcdceeeac61

Observation fff07ec3-4590-4663-93c6-cc10cca13954 · outbound

This paper cites Unsupervised learning of disentangled speech content and style representation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unsupervised learning of disentangled speech content and style representation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.903838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.211965Z digest=sha256:0262b459f785f10c7641e6c3d99171a40074e2a2e1e59a8473beb38cef0f7741

Observation 0b2c7c32-6e36-488c-843a-6e54cf7b2674 · outbound

This paper cites Text- less speech-to-speech translation on real data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Text- less speech-to-speech translation on real data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.895972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.215144Z digest=sha256:9507bf5fee153f371913e362ee6b3abc26889538f15948ba39e553c19db3467d

Observation f2dd42d1-ebc7-41b9-9bc5-b15b453909e6 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.888168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.218507Z digest=sha256:b7765c749c53338657c49a08cdf5458675857877f873ac249a6bfe10026075d8

Observation cce790f6-3be1-4ef6-8e45-4148352c4a91 · outbound

This paper cites A comparative study of self-supervised speech representation based voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement A comparative study of self-supervised speech representation based voice conversion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.880694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.221950Z digest=sha256:88fa593ec3087e279c5f9038c320b6a938df12f06214c2991dd66ffac01cf7cd

Observation 7e0f126a-39a0-4447-90e0-55372264de06 · outbound

This paper cites Leveraging diverse semantic-based audio pretrained models for singing voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Leveraging diverse semantic-based audio pretrained models for singing voice conversion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.872428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.225095Z digest=sha256:48594f6d9ab47a35ad8b8199be7b9590d31ef1dc86fd23f8e7e6afa8d1de67e9

Observation c6debe6e-8d3c-4c3c-9386-b7b172cb9bbb · outbound

This paper cites Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.864166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.228370Z digest=sha256:e2e0253835d593c79d770abeae3fec9f92f4b1d615c33609649eee6810203586

Observation b21ccbbb-9ffe-4d95-b1a1-946d7527d093 · outbound

This paper cites Stargan-vc: non- parallel many-to-many voice conversion using star generative adversarial networks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Stargan-vc: non- parallel many-to-many voice conversion using star generative adversarial networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.855792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.231460Z digest=sha256:bbf0054feb72339f355baa52ae63d65b2f2c3fc573c00b0a3b64c0c6c4042d3a

Observation 7ade7335-7583-48a0-bb39-d1564a094f0a · outbound

This paper cites Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.847440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.234763Z digest=sha256:e5ba8b010071e4b3298a3fce1f836c0457bd446230ea3c3806647abd4213be77

Observation e579a085-1702-4f69-8e7a-d3a41f7d7b54 · outbound

This paper cites Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.839236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.238179Z digest=sha256:632b08adabeeec16cd168e73d8afe770d89a1ddb83a5a8bd0ca2e7d07db8dadf

Observation 0117c55c-388b-4907-b30b-5a76b2feb2ed · outbound

This paper cites Tts-guided training for accent conversion without parallel data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Tts-guided training for accent conversion without parallel data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.830985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.241259Z digest=sha256:fd73d6bf8ec526b41820fa27beed2b12e23247d3eae5a2704d96f28dcf484912

Observation efd90fa0-40a8-41eb-84b7-b3095e3c6b4f · outbound

This paper cites End-to-end accent conversion without using native utterances.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement End-to-end accent conversion without using native utterances

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.822725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.244394Z digest=sha256:e5fdb092dcba8db27c30a1d52f6590f0f833068fc79c7336cc54a51c0af08f81

Observation 2333a5bb-d67a-49cb-a334-c7bb3dee15cc · outbound

This paper cites Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.814630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.247887Z digest=sha256:1bf5646e708a95a003aaba0386d4b58335d8253e7d74bd265bbfbc058fab4924

Observation e933a787-f410-4603-b343-601a48a25b3f · outbound

This paper cites LM-VC: zero-shot voice conversion via speech generation based on language models.IEEE Signal Process.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement LM-VC: zero-shot voice conversion via speech generation based on language models.IEEE Signal Process

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.806051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.251307Z digest=sha256:af8bcc94399367b365c844fdbc168e81bde8b8ed9b679a3fd64fb09378576951

Observation 49e1554d-4460-4704-a07a-acb5915ffb54 · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.254583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.254583Z digest=sha256:a28ee8cec1d13edbf4188c819dd6f8963237537bb8f48575d37c4a92642f119a

Observation 616a3740-187f-4e40-9ae3-7ad81e2b7bca · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement A comparison of discrete and soft speech units for improved voice conversion

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.797639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.258005Z digest=sha256:e5353d693205d6e83701ef63af0450b988a5666da732eaaed40cd281248788b4

Observation 7f46fa70-3f63-4b79-80e6-1882ea7259c6 · outbound

This paper cites SEF-VC: speaker embedding free zero-shot voice conversion with cross attention.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement SEF-VC: speaker embedding free zero-shot voice conversion with cross attention

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.789762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.260858Z digest=sha256:056754455ec60e09fe0c3b1195f4459f9eabceebfc76608f1a8e8be9d3f370f2

Observation 019df985-128f-4a50-9923-fc61b67370ae · outbound

This paper cites Neu- ral analysis and synthesis: Reconstructing speech from self-supervised representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neu- ral analysis and synthesis: Reconstructing speech from self-supervised representations

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.782063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.263434Z digest=sha256:a17e75c41edefae960e78c0fbe582126e165ab06cce1ebad0df124d265558a58

Observation 8d86c200-b042-4825-9fd0-b3f4f953501e · outbound

This paper cites NANSY++: unified voice synthesis with neural analysis and synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement NANSY++: unified voice synthesis with neural analysis and synthesis

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.774318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.265980Z digest=sha256:60fe370298cfda2885be68e74100751df9efbdcb2ce88779b73dd21c79189c73

Observation def82cc8-225b-4949-9917-0a981577767f · outbound

This paper cites Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottle- necks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottle- necks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.766187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.268522Z digest=sha256:a14b1a097e7c4430d20407b05b4dc02b60accd1ce525cc51e01bd6d860ae95d2

Observation f0e34687-921c-4eed-87e3-61b6123eb262 · outbound

This paper cites Cox, Mark Hasegawa-Johnson, and Shiyu Chang.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Cox, Mark Hasegawa-Johnson, and Shiyu Chang

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.757329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.271441Z digest=sha256:6648f45abc627102940fe2d110734800e8e15c47dc41d0a50a0e18d5f4b42104

Observation e0381617-1b1b-4451-9a2e-b8ba001f3301 · outbound

This paper cites Multi-speaker expressive speech synthesis via multiple factors decoupling.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Multi-speaker expressive speech synthesis via multiple factors decoupling

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.748797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.274184Z digest=sha256:d8d6c033cb149921a4af6eb8a9c485eac009f4458e18cdd64c240dd4bad7267d

Observation efdb67c7-5795-43ff-878b-ab42b2040bc5 · outbound

This paper cites CLUB: A contrastive log-ratio upper bound of mutual information.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CLUB: A contrastive log-ratio upper bound of mutual information

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.739916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.276826Z digest=sha256:481455f88fd2beccbe1bd2b3d6cd0300031f34fc4e3474c4f7c4d86c946571e3

Observation 44448173-2d8b-4098-ae9b-fd6b8fe82fd4 · outbound

This paper cites Repcodec: A speech representation codec for speech tokenization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Repcodec: A speech representation codec for speech tokenization

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.731258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.279359Z digest=sha256:a50085994055abcb1ecc4061483287558069ce433009bd82dc50b3bf351413eb

Observation ee6f654c-e8bf-43d3-92ba-67f5aa894cb3 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Soundstream: An end-to-end neural audio codec

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.281773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.281773Z digest=sha256:68c44aeec0d1f9deaf367737eb1d18ee7abbe8c3197c870e97e600eee852accd

Observation 3a7660a3-b4de-4b7a-852e-2f8cbe5b0437 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Wavlm: Large-scale self- supervised pre-training for full stack speech processing

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.717869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.284565Z digest=sha256:665ef8fb5ba2cebb1c1182cf35720accdb09bc5f642f280f50a86d4fc9e8c501

Observation 04a83850-4d53-4914-b52d-5c58321399a7 · outbound

This paper cites ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.709198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.287428Z digest=sha256:89313d9ee5ed848eca643f6da86f589d023d9502f62de2b709d6c4859ccee2a6

Observation cb1b2cf2-77a4-4c63-8175-dbb8f8527d1a · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BERT: pre-training of deep bidirectional transformers for language understanding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.700442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.290292Z digest=sha256:d6e72b85275ed0bf600c8f98324ec55e825177983d48b6ea79894e68b295f653

Observation 1caf02e0-b17b-4418-ab8a-847aa87da851 · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Bigvgan: A universal neural vocoder with large-scale training

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.687235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.296292Z digest=sha256:54497d86b4524b8305a0e3342a1dca6bc99db61b042d9ccbd3f18c7fdff2eb6c

Observation 97818603-d8b1-4009-b416-4ffbedb85cdf · outbound

This paper cites Tyers, and Gregor Weber.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Tyers, and Gregor Weber

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.678910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.299049Z digest=sha256:afaa6f287becf476c39c2bd3f20ef393c229ac9b39f50065b0b20e5161114e2c

Observation 7a190bd8-a972-40ec-b49f-710944c8370f · outbound

This paper cites The singing voice conversion challenge 2023.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement The singing voice conversion challenge 2023

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.671037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.301962Z digest=sha256:1a9a0d66270f7e987ca97c0eb569740a151597dcde73d98f6d3e00736a2c155a

Observation 41e9396b-97de-4d23-847c-cb904661bfc9 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Robust speech recognition via large-scale weak supervision

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.662748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.304779Z digest=sha256:802430b6312aea84aca22a5a6edfe22e0f67ebaac7b25031e9976c9692016b78

Observation f129dbbf-ba2e-4d4d-b419-611e0c960a33 · outbound

This paper cites Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.654569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.307589Z digest=sha256:587e6d96b4d0c97a38a8a96f0539271b1984fb3713ea8b6c47868653fcd98e64

Observation d69f8ee8-2721-4124-a54a-686965e816ef · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.647207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.310413Z digest=sha256:ff607723dd4307a02d8a0476845d144a7f8091220807a47fe3a00f54a449d156

Observation 08f94af4-4fca-4428-8029-ea535e194cae · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.639684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.313298Z digest=sha256:07e57fc036c6df87121cf3f04407b4c4393065d8ef1de227e7fb942938b4e49b

Observation d2b6dc8d-aa5c-4335-a1a2-5d2e740ad999 · outbound

This paper cites Libri-light: A benchmark for ASR with limited or no supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Libri-light: A benchmark for ASR with limited or no supervision

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.631954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.316297Z digest=sha256:539ebb8d1a6f3d98e5305d2ffd70b6c309bbf01f8952ede13f4ce105c7537d9a

Observation 91b2c8cf-ee67-4ac5-a632-a16b75a59656 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Librispeech: An ASR corpus based on public domain audio books

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.622739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.319210Z digest=sha256:5863aec4b7571dcc729a3b54eb18cb802cea2a01adec4333149d767d0fcf45ff

Observation cc567952-405a-4834-803e-a1f8637b199a · outbound

This paper cites Amphion: An open-source audio, music and speech generation toolkit.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Amphion: An open-source audio, music and speech generation toolkit

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.613920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.322124Z digest=sha256:c3e25ba0d37730093b97b9ac67880a312c7f69c7ee54c2fe3d00b5e21edee3fb

Observation 35599e59-9f1a-431a-b55f-394a892cbc4b · outbound

This paper cites V oicecraft: Zero-shot speech editing and text-to-speech in the wild.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oicecraft: Zero-shot speech editing and text-to-speech in the wild

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.605013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.324971Z digest=sha256:ad1cfcc1dcc46c7e47871c549967c440b2ab95e1f9bc5f97ae2571f11a932a2f

Observation d4d9660f-dee0-4fd3-836c-94376ab66820 · outbound

This paper cites Overview of the Amphion Toolkit (v0.2).

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Overview of the Amphion Toolkit (v0.2)

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.327889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.327889Z digest=sha256:da625e0be42c60a37091fe315378af2c7d66f15e6b9040502d7b6a48318e4f12

Observation f007b9ae-b6b3-4469-a20f-6ff698ee87a3 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.596064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.331052Z digest=sha256:b9b1400a6fafe13df49b80de82182a06cbc6c6fbae6e76691603b4754ca6ba2c

Observation 8ea99c5d-0d73-46ab-a699-2c5b5fd2d259 · outbound

This paper cites Decoupled weight decay regularization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Decoupled weight decay regularization

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.586922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.333766Z digest=sha256:349581e42b79fe73639178eb89c8f7d413161892e1d18cd860ea92aa389dca86

Observation f4437809-355c-4f4e-9839-d4877cffa94a · outbound

This paper cites Kingma and Jimmy Ba.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Kingma and Jimmy Ba

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.336676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.336676Z digest=sha256:7d7a43bb0711c8520ba36a649e8a59a45673c2dca9501da87795a4fd25563286

Observation ef3c1bcf-7fce-40a9-a5af-1a4ec8b76b51 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Classifier-Free Diffusion Guidance

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.339612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.339612Z digest=sha256:d35207a6aef873daccfe2958828331a2e5764348af94461f2fe384a5949c0277

Observation 71139633-d1f4-49e7-9afe-ed8423788c6e · outbound

This paper cites Scaling speech technology to 1, 000+ languages.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Scaling speech technology to 1, 000+ languages

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.573407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.342687Z digest=sha256:d524aad9da29f33cf3c4c66e465a6ec7e22943dfcb34ed29768031d7f360851a

Observation fab7ebe3-0ec2-4843-ae1e-2de7f9ddbfb5 · outbound

This paper cites Conditional variational autoencoder with adver- sarial learning for end-to-end text-to-speech.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Conditional variational autoencoder with adver- sarial learning for end-to-end text-to-speech

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.564600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.345544Z digest=sha256:bea1ec20184250dff33b2a816f9abbdcd4a01f0a49d7cc8d21586ed00fb43d1b

Observation f0be682c-099a-45b8-9852-3d4430a7214b · outbound

This paper cites Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.555736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.348402Z digest=sha256:bbf406eeb5ef80cf0a30fa8b64407aa24d8454a987116e52662f2fb093abbc29

Observation ed0f2eb6-4b91-4dec-9f18-5b4edbf627ae · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.351240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.351240Z digest=sha256:7a20a4c6214754ece216408660b0e22a94008defd32792023c1fc48a9f115a5c

Observation d2eb39b9-9991-480f-9f29-ec8fef4334f2 · outbound

This paper cites BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.541564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.354382Z digest=sha256:ab227de44178c9baed653f36190f373d5a87eea3b03afd81e51e0e60d0fc781a

Observation c7454f28-2d70-4bc5-869a-936e0d8b3b9b · outbound

This paper cites CSTR VCTK Corpus: En- glish multi-speaker corpus for cstr voice cloning toolkit (version 0.92).

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CSTR VCTK Corpus: En- glish multi-speaker corpus for cstr voice cloning toolkit (version 0.92)

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.532562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.357239Z digest=sha256:f204b3de72fb7b3b28c2ad9235cbc12d622c9b9d21b090fbf0e56ad25f62e0e3

Observation 0b0598f3-d986-426f-9a83-05f99e54586e · outbound

This paper cites High Fidelity Neural Audio Compression.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement High Fidelity Neural Audio Compression

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.360299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.360299Z digest=sha256:fe178b72bd4ac46604b8272e5af97d683155a6eadef9949073c1f69f53f541c0

Observation 050584c9-9979-4017-a8a1-00ce234971ad · outbound

This paper cites MLS: A large-scale multilingual dataset for speech research.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement MLS: A large-scale multilingual dataset for speech research

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.523449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.363655Z digest=sha256:535ae15de6c8050e57b267b4b6c8525e72315c36c7835d12acb5a239650391ef

Observation b8aac263-107c-4a80-9934-b90cdf3a6e72 · outbound

This paper cites Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.515151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.366424Z digest=sha256:c9cc533b603c4030436f8b58f3f266f42d0e7788e797d2ebfdb492452a7769a4

Observation 06cef095-3980-4f5f-b5c1-4fef04f485f6 · outbound

This paper cites bit” vs. “bet.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement bit” vs. “bet

Reference 92

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T13:26:19.505856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:26:19.369265Z digest=sha256:024e4b547991e251cb407d45d11dc0d40d854eb3d9e801a53806383e0e4992b1

Observation 7687b7be-def6-4817-b4ce-463f6fe28086 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 4186

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.293571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.293571Z digest=sha256:8086c5e68be34e3aaeb1fc66146ebe34d8816703109c6357fad932deadd9b56e

Pith citing papers

Observation abde2652-5844-45b5-bc1b-6b14c8aafb41 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.575812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.575812Z digest=sha256:2fcc17264403cf8f25fa412299cba1d4e79f0489b7366e61dc71b0bf73e275a0

Observation 64546f18-8417-4c2a-b48b-0557f261eb3f · inbound

Entropy-based Coarse and Compressed Semantic Speech Representation Learning cites this paper.

Entropy-based Coarse and Compressed Semantic Speech Representation Learning Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:07.305867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:36:07.305867Z digest=sha256:e21c927df0c7ac49f0224677804a07132cb2e90c8d6f998910f734de8e77d01b

Observation 80a7b18a-b5da-4ab4-997e-a3a9b98d8558 · inbound

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck cites this paper.

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:50:51.780513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:51:38.030059Z digest=sha256:a261734930f5c2a4a12608d5c560d463597cc9edb377cf65885e8ec842449284

Observation 148ee6a2-99f8-4893-a50c-40d16a870387 · inbound

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora cites this paper.

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:01.383468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:57:07.894455Z digest=sha256:84faf7fdbe85ddfec9285301895ef724311e6cde3787c9f5e05b53849ae3cd8b

Observation aef88493-6709-4379-9fe4-f2919aaf5837 · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:39:38.692633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:bd26c40c33596fe5d850b109c0a6d8789b98651aea4237591f25e3835b47b405

Observation 6c019909-a819-4aef-b337-348e77d88fbb · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.092101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:a9383dcc9caaaf76e5726ff46e9578aa9219b978ae926f58da925b12ec03f50b

Observation cffc3372-92ee-43d3-99f7-c5a05a951d70 · inbound

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization cites this paper.

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T22:31:42.563915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:31:42.563915Z digest=sha256:43e18c9127673be15edea110012c6e1dd701c5bb69b3f8fd9d3098d668ff1ee3