Pith. sign in

Paper Citation Record · LEDGER

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

As of 10 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 7 inbound Pith citation observations for arXiv:2502.07243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07243 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:26:19.369265Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:21:59.575812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:38.691071Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1234a951-ef09-4289-a4bb-218485f1202d · outbound

This paper cites Neural discrete representation learning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural discrete representation learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.088321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.088321Z digest=sha256:bc74e8b1a753c6eafd8f715ae8c5852914a578f70b88a523713b6101c756cb69

Observation 9621dbf3-776a-4578-87e4-6d5cc7130f2d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.092101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.092101Z digest=sha256:a8db9b2864e8c3e304a920fe11c6ad378e2712f59457772c596b9a3cbd86a536

Observation cdcc9f2f-cdc6-46b8-8c46-646cb2c70740 · outbound

This paper cites An overview of voice conversion systems.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement An overview of voice conversion systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.095768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.095768Z digest=sha256:e19f6b30880661a46ae725fa8f8279d1f44da4de4c6c229d2ca341175a9aac49

Observation c384d50e-393a-4c38-9e76-05d4c1c59737 · outbound

This paper cites An overview of voice con- version and its challenges: From statistical modeling to deep learning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement An overview of voice con- version and its challenges: From statistical modeling to deep learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.099143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.099143Z digest=sha256:bfecaef24b179fb5852c4d115e79b69cd3ab154ff69ed092e8d02093cb087289

Observation 2ef66d9b-3efe-4cc5-9fdb-32a19f32fb3b · outbound

This paper cites Foreign accent conversion in computer assisted pronunciation training.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Foreign accent conversion in computer assisted pronunciation training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.102330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.102330Z digest=sha256:ae24414368f848eb4e8cbf56919dd0c46754244df21bf039f8e900a49d237ae7

Observation 9f2d5106-5bcc-4539-92f2-5c08a4174d82 · outbound

This paper cites L2-ARCTIC: A non-native english speech corpus.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement L2-ARCTIC: A non-native english speech corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.105470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.105470Z digest=sha256:659ab8785a9f9c761f21a92a624ed25f5738bf73809c0bc87bc64ba4b880e5a9

Observation 84010b50-fc81-409d-b51c-3e124781da88 · outbound

This paper cites Emotional voice conversion: Theory, databases and ESD.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Emotional voice conversion: Theory, databases and ESD

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.108882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.108882Z digest=sha256:bcdf43d37d406a9bed0ddf3258056874f73f62d96e69c5b38b4f6718a8773b71

Observation 8f4141dd-7b92-4bc3-a3ed-6c8dade41c33 · outbound

This paper cites Neural Text-to-Speech Synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural Text-to-Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.112021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.112021Z digest=sha256:39bbe6af22dae263ce6ca2393b0c6f94f47cdc3d42c89dcba3374fa23b27c93e

Observation bbf8051d-f030-4f39-8e3f-6990403e10e7 · outbound

This paper cites Converting foreign accent speech without a reference.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Converting foreign accent speech without a reference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.114754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.114754Z digest=sha256:4ce9c5ad99ec5b3bbd4f617f345cee585b6f954d61a22b78e36de9efdb4eaf1a

Observation b5ddc4a1-f58f-452b-9dab-64b5718a6ba1 · outbound

This paper cites Sahidullah, Aur ´elien Bellet, Marc Tom- masi, and Emmanuel Vincent.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Sahidullah, Aur ´elien Bellet, Marc Tom- masi, and Emmanuel Vincent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.117635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.117635Z digest=sha256:4602060dd4f9eb7a260b6124955307d197d9d638a4f68dc74ef788d268978404

Observation 05a4d2f1-477d-49e4-9574-6aa2484f14a0 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.120659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.120659Z digest=sha256:3221286dd3c35bdb05c6cfc5f7bcf9d0c5aa02b1b63d15f1eb2f0b25298da14c

Observation ab9f82f4-1ab7-4ad4-82d3-3aab0706d44f · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.124096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.124096Z digest=sha256:e287484e547dda8ebad352bfc0ca1dd55707107599892febaf4cd0586bc6e834

Observation 11b705f3-918e-4255-a214-222525ab2f0e · outbound

This paper cites Maskgct: Zero-shot text-to- speech with masked generative codec transformer.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Maskgct: Zero-shot text-to- speech with masked generative codec transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.127638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.127638Z digest=sha256:a6bbc353a12a280c9dcf49de3635c1479b60aa3d1a6808e393469e9f9d50f44c

Observation 07ff5752-9eaa-4c8d-a501-508a918d48fe · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.130748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.130748Z digest=sha256:25975864563580a438a597fce7040d32f51983fdb3b728c954e758d7f23e3cef

Observation 6a2d694b-e80e-42a1-9d7c-b13db2eb9f4b · outbound

This paper cites Speech resynthesis from discrete disentan- gled self-supervised representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Speech resynthesis from discrete disentan- gled self-supervised representations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.133715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.133715Z digest=sha256:8e9622415424204bcccbcd397040a227d9b6bf0472ada412769a7c62705b1fcf

Observation 6e9e5b21-1953-45c9-9c85-9baf1e1642fb · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.136859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.136859Z digest=sha256:65a417afdf6fd71ecf67d9e488eb1e16b44e14f6100074d5d3216058a49c3560

Observation 53b580d5-50d5-4831-b8b8-aab0451b14bc · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.140473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.140473Z digest=sha256:9e814a35b62d14c3725adb4562b48e5f52ce731ce175511761ba06b61f06bf63

Observation 48cc366c-9511-4c1d-b45c-122c4c30fd9f · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:20.024364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.143740Z digest=sha256:ed6e7b3215a6dd7e8fc034ede544fd035d583f3c8689d7da23fc96734e0a8db9

Observation b16bd86c-730f-47cb-a719-d59fc2465f63 · outbound

This paper cites Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:20.015767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.146950Z digest=sha256:5f5f58da09b092c1435471ee50ff821d1afeefa77230b8e3123288450c3bf310

Observation b4e0abf5-deb8-4310-a6aa-de5bb23ebfa5 · outbound

This paper cites VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.150092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.150092Z digest=sha256:fb81d015a88b168c67fa21b828e10a58e8b3a396264b78a922dd5973260a6c4c

Observation 6e1cff98-0b03-451c-8ab3-a3b8d7261506 · outbound

This paper cites Convert and speak: Zero-shot accent conversion with minimum supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Convert and speak: Zero-shot accent conversion with minimum supervision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:20.006904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.153241Z digest=sha256:3db39b7f67e5b0cec27d048052cb766110f68412468d44039994fc8b22c027e2

Observation d63fbd8f-c1e4-4798-99a1-4a417dabffc5 · outbound

This paper cites Au- tovc: Zero-shot voice style transfer with only autoencoder loss.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Au- tovc: Zero-shot voice style transfer with only autoencoder loss

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.998020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.156236Z digest=sha256:139baf4f5a440c7e2b32b2e738728a85aae6eb889a306483954cb11fce75cec5

Observation be8810d0-db54-481e-abb2-683a25f231af · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.158875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.158875Z digest=sha256:94b92c2c01dd13910e075767aeb5402eb106a3f25300766953e38f3439b2903e

Observation 712c1ffd-5225-45d3-bd18-6aba84d44a47 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.162381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.162381Z digest=sha256:df82bd587f03672fb17471562624e28fc27bc17fea02523c8d97ab4a32561c3c

Observation 54dc202b-f995-44b5-9425-fc2300fe4d46 · outbound

This paper cites Better speech synthesis through scaling.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Better speech synthesis through scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.165401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.165401Z digest=sha256:1423561de31df12e1209750d0a200e7bb98ddcd46d9b4940d3bf9ee0850a1437

Observation 8dd7fa07-0c42-4fbc-9577-7def92b7ad11 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.168802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.168802Z digest=sha256:002838c1af1969d4b753b7590c51933eeb1caf12b8591b7c9f107a506f38c6bc

Observation 9bf9cba5-4724-401d-84a2-5dc3a569a6c9 · outbound

This paper cites V oicebox: Text- guided multilingual universal speech generation at scale.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oicebox: Text- guided multilingual universal speech generation at scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.172375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.172375Z digest=sha256:c78b9d02e903ef309ff4c31b68fdc4e04f821ff490b5225ffad6aa97b4383af3

Observation 8fc2728d-5e2a-48df-ab02-34693677b7c2 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.984343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.175536Z digest=sha256:60bf0b1a3f42700f6e1f65bae2156a053e65a67da0d2c9e27a93855b23d5db66

Observation ff1696aa-260f-49fe-be8d-3e0e641aa098 · outbound

This paper cites V oice- preserving zero-shot multiple accent conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oice- preserving zero-shot multiple accent conversion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.975383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.178687Z digest=sha256:06816b63c466015b9d539bea8b87ad8304aed09e2f7cda3b7ad336a6f11b6ef7

Observation 9d8b0e25-1f48-424f-a484-258480a55ff0 · outbound

This paper cites Schuller, and Haizhou Li.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Schuller, and Haizhou Li

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.966785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.181727Z digest=sha256:2a5cca11947dd0446149d0894d43b7bc0f1395fb7fb194b471d20935c60cff7e

Observation e4bf249b-c83e-41c0-bb00-de4227ca1539 · outbound

This paper cites PA VITS: exploring prosody-aware VITS for end-to-end emotional voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement PA VITS: exploring prosody-aware VITS for end-to-end emotional voice conversion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.958284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.184577Z digest=sha256:802e7fc59e6babf3c83f10e2cc305df44c51e527a49bc833485d47d248573a6e

Observation cb915b3c-08b9-4090-8a15-d8a50bf05a15 · outbound

This paper cites Transfer the linguistic representa- tions from TTS to accent conversion with non-parallel data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Transfer the linguistic representa- tions from TTS to accent conversion with non-parallel data

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.949309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.187544Z digest=sha256:43c8393683d82f7d463a39031df26535b40cbf0c1be52d2a542bfa421a60a4f1

Observation a52ec863-cd90-4e3b-97da-87c94d9798dc · outbound

This paper cites U-style: Cascading u-nets with multi-level speaker and style modeling for zero-shot voice cloning.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement U-style: Cascading u-nets with multi-level speaker and style modeling for zero-shot voice cloning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.939715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.190541Z digest=sha256:6bc85a0d292d4ddcf41e36681fcf1865788b63779d55888f0d025e591fb28212

Observation 32c230e2-b170-4780-855b-e79bba4ccccb · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.193725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.193725Z digest=sha256:740d50492cab30091232ddfaabe04eafac9e4cd857a40c75743acfbb494e899b

Observation 6baec9ef-0a8b-435f-909f-835a1966cab0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement LLaMA: Open and Efficient Foundation Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.196997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.196997Z digest=sha256:9863627c22568f6093d8319b02a4fe981d703ed03b91dad06bdcbbdb8bcb13e6

Observation ffb292e5-e1b2-4a92-b961-b30131fcd04a · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.200113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.200113Z digest=sha256:ca1dd70af696ff09863891c7180990c2c3fc5e995b9411803ea5431dc671a5bc

Observation f59cd892-b39e-4fb8-b11f-f5882fcac97e · outbound

This paper cites Scalable diffusion models with transformers.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Scalable diffusion models with transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.920850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.202942Z digest=sha256:dd13311bda61ba9e030120455830e5a224d6c0b01551b0a4d5170342723a3677

Observation f9b274d1-daa9-4b5a-85f6-5411061b25d3 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.205829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.205829Z digest=sha256:07e69173bfabe1b7825373eb03d46529f9880199fa47a5b4f0c7425abe6a9b3b

Observation 02e7d8be-5b9d-4760-aba5-1cb3debb871c · outbound

This paper cites One-shot voice conversion by vector quantization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement One-shot voice conversion by vector quantization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.912239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.209096Z digest=sha256:e0d185e7aa77e7e7d56aa85e67bfcd7aab77c671cb8489dc12d6de5a4a0007a9

Observation fff07ec3-4590-4663-93c6-cc10cca13954 · outbound

This paper cites Unsupervised learning of disentangled speech content and style representation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unsupervised learning of disentangled speech content and style representation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.903838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.211965Z digest=sha256:a541570df1041f233373dd74f7690e53fc0052feb40b87bf430594416f4a9f93

Observation 0b2c7c32-6e36-488c-843a-6e54cf7b2674 · outbound

This paper cites Text- less speech-to-speech translation on real data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Text- less speech-to-speech translation on real data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.895972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.215144Z digest=sha256:67044e2751e52ed1dfbeecb719b32858a4c1a239afbe70fae4c2758ec2937f17

Observation f2dd42d1-ebc7-41b9-9bc5-b15b453909e6 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.888168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.218507Z digest=sha256:a6ded04c91b98abc9d979fba7a1ce07246e9ab81c192b596790c9abc65a81ff6

Observation cce790f6-3be1-4ef6-8e45-4148352c4a91 · outbound

This paper cites A comparative study of self-supervised speech representation based voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement A comparative study of self-supervised speech representation based voice conversion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.880694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.221950Z digest=sha256:0075c70191aeefc633c6a2f0d63523fa832fde04ea70d42b1da49b80a1ef8a17

Observation 7e0f126a-39a0-4447-90e0-55372264de06 · outbound

This paper cites Leveraging diverse semantic-based audio pretrained models for singing voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Leveraging diverse semantic-based audio pretrained models for singing voice conversion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.872428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.225095Z digest=sha256:e65671e9e8d7e9f93049f98deb244d2bbd68944bdd48a9e3e7969792d53b0f8d

Observation c6debe6e-8d3c-4c3c-9386-b7b172cb9bbb · outbound

This paper cites Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.864166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.228370Z digest=sha256:0136c842b76da459643eeaf63f9c38f8b9fc680d53b25c0d0d07c0bbbcfbc9fa

Observation b21ccbbb-9ffe-4d95-b1a1-946d7527d093 · outbound

This paper cites Stargan-vc: non- parallel many-to-many voice conversion using star generative adversarial networks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Stargan-vc: non- parallel many-to-many voice conversion using star generative adversarial networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.855792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.231460Z digest=sha256:f48faf07e34bfd2b68f34f7471930f5f225b9c7f5fd5c8441503aa5118bb79df

Observation 7ade7335-7583-48a0-bb39-d1564a094f0a · outbound

This paper cites Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.847440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.234763Z digest=sha256:d063c9e8ac9bda816df38521025b0394b745b1d39202f9baaf01a3fb62362404

Observation e579a085-1702-4f69-8e7a-d3a41f7d7b54 · outbound

This paper cites Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.839236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.238179Z digest=sha256:3320e737672b996d096a64a17dd86b1882380b7c25d4972e749457308ebf06a9

Observation 0117c55c-388b-4907-b30b-5a76b2feb2ed · outbound

This paper cites Tts-guided training for accent conversion without parallel data.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Tts-guided training for accent conversion without parallel data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.830985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.241259Z digest=sha256:9875003a1eb7129a9b1f4a5a28b79e1dca99c55389031bdda99964e0a6b1e1aa

Observation efd90fa0-40a8-41eb-84b7-b3095e3c6b4f · outbound

This paper cites End-to-end accent conversion without using native utterances.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement End-to-end accent conversion without using native utterances

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.822725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.244394Z digest=sha256:6ba7604f7012f12230d70e9918a99489971091eca180a0ae6c6009bbcc9f45f0

Observation 2333a5bb-d67a-49cb-a334-c7bb3dee15cc · outbound

This paper cites Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.814630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.247887Z digest=sha256:27dc467dfe347974f0f7a2b44e231c3b32b37267e5704e1ccee8c56a62193844

Observation e933a787-f410-4603-b343-601a48a25b3f · outbound

This paper cites LM-VC: zero-shot voice conversion via speech generation based on language models.IEEE Signal Process.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement LM-VC: zero-shot voice conversion via speech generation based on language models.IEEE Signal Process

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.806051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.251307Z digest=sha256:bb45e4076a2cd3f426c0d77198ef21c6a5d2a20ded7dd765e7705cad35346a36

Observation 49e1554d-4460-4704-a07a-acb5915ffb54 · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.254583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.254583Z digest=sha256:95c220c0f063eb8ab40cfc00e030bd98a70a9e550c5e46e6f626d470760dbb22

Observation 616a3740-187f-4e40-9ae3-7ad81e2b7bca · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement A comparison of discrete and soft speech units for improved voice conversion

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.797639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.258005Z digest=sha256:0e27ae0d269e8eabbc9788ec5d27b1f02d8ff9f17f1ca40ad90b41add2056e96

Observation 7f46fa70-3f63-4b79-80e6-1882ea7259c6 · outbound

This paper cites SEF-VC: speaker embedding free zero-shot voice conversion with cross attention.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement SEF-VC: speaker embedding free zero-shot voice conversion with cross attention

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.789762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.260858Z digest=sha256:6df34fb5dc5829b93ea18eec0f8990085d443723af49ddbdf2ca4c07e10ba5e2

Observation 019df985-128f-4a50-9923-fc61b67370ae · outbound

This paper cites Neu- ral analysis and synthesis: Reconstructing speech from self-supervised representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neu- ral analysis and synthesis: Reconstructing speech from self-supervised representations

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.782063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.263434Z digest=sha256:46f2b4320c7bda05e387688559abea06144cea85763f4e687470068cc51bcfe0

Observation 8d86c200-b042-4825-9fd0-b3f4f953501e · outbound

This paper cites NANSY++: unified voice synthesis with neural analysis and synthesis.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement NANSY++: unified voice synthesis with neural analysis and synthesis

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.774318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.265980Z digest=sha256:27b2e574d3b47a4210801d1bba56f8be51527bf7b80c97f1deaee532895565f2

Observation def82cc8-225b-4949-9917-0a981577767f · outbound

This paper cites Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottle- necks.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottle- necks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.766187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.268522Z digest=sha256:4366479072740ddd44c6826a97b77a3bc63e71297ab47e0afaff16aa4c384728

Observation f0e34687-921c-4eed-87e3-61b6123eb262 · outbound

This paper cites Cox, Mark Hasegawa-Johnson, and Shiyu Chang.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Cox, Mark Hasegawa-Johnson, and Shiyu Chang

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.757329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.271441Z digest=sha256:7b4d1107134c8264915732282d08de3e61cbc99bf7884c11ec1918a071bb9e19

Observation e0381617-1b1b-4451-9a2e-b8ba001f3301 · outbound

This paper cites Multi-speaker expressive speech synthesis via multiple factors decoupling.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Multi-speaker expressive speech synthesis via multiple factors decoupling

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.748797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.274184Z digest=sha256:aacc9daee333f58e10e664a23cddfeff48245ba5fc0d406d41c628f9fb63ba59

Observation efdb67c7-5795-43ff-878b-ab42b2040bc5 · outbound

This paper cites CLUB: A contrastive log-ratio upper bound of mutual information.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CLUB: A contrastive log-ratio upper bound of mutual information

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.739916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.276826Z digest=sha256:0bcd025ec652c1d281a370a809e717ac1b33a26b89ffa2d5214373aa962a60c2

Observation 44448173-2d8b-4098-ae9b-fd6b8fe82fd4 · outbound

This paper cites Repcodec: A speech representation codec for speech tokenization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Repcodec: A speech representation codec for speech tokenization

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.731258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.279359Z digest=sha256:b62efc8882516ad9a52ab7093f4c0948b1914652795bb72c8a227121a116d9ff

Observation ee6f654c-e8bf-43d3-92ba-67f5aa894cb3 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Soundstream: An end-to-end neural audio codec

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.281773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.281773Z digest=sha256:47eae18f35d71706d1c12c5ed03847fa41c20d3b5f79a64266e5a4dd7c3dd246

Observation 3a7660a3-b4de-4b7a-852e-2f8cbe5b0437 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Wavlm: Large-scale self- supervised pre-training for full stack speech processing

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.717869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.284565Z digest=sha256:9f8a250ed75010d75231b17a58f2a62cbb8a498558e08d8b2408d21ca7195ed4

Observation 04a83850-4d53-4914-b52d-5c58321399a7 · outbound

This paper cites ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.709198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.287428Z digest=sha256:9889df8623d922d5f7487ae26aa7dd3034d70aa1c9ce16db59e641e901f01c3e

Observation cb1b2cf2-77a4-4c63-8175-dbb8f8527d1a · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BERT: pre-training of deep bidirectional transformers for language understanding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.700442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.290292Z digest=sha256:c93e712e433c4eb65052db3d5f0f24ce9a0c3105d043b719ed5b96df3398e63e

Observation 1caf02e0-b17b-4418-ab8a-847aa87da851 · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Bigvgan: A universal neural vocoder with large-scale training

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.687235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.296292Z digest=sha256:f6d43e91f9f77e8d91e80420bbae3c7a1a791c11e9a4108ba28641a466a7749b

Observation 97818603-d8b1-4009-b416-4ffbedb85cdf · outbound

This paper cites Tyers, and Gregor Weber.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Tyers, and Gregor Weber

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.678910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.299049Z digest=sha256:d51194da8d1d4ff4f0d4f68db6702ec79196a0053289ce1d1d1582dcc83092f7

Observation 7a190bd8-a972-40ec-b49f-710944c8370f · outbound

This paper cites The singing voice conversion challenge 2023.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement The singing voice conversion challenge 2023

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.671037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.301962Z digest=sha256:27af3b7a0fb1bf454f97cf6cb3cd979b1eb6c1ffdb965549062324e57486294f

Observation 41e9396b-97de-4d23-847c-cb904661bfc9 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Robust speech recognition via large-scale weak supervision

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.662748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.304779Z digest=sha256:a7c0c0a891a100571a7749c7e18719796103ef03905e7daa29a7420ed9f51543

Observation f129dbbf-ba2e-4d4d-b419-611e0c960a33 · outbound

This paper cites Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.654569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.307589Z digest=sha256:f061a25e3e30a910fced29175819729d3d5eab972f2cba893e776b1dcb065e15

Observation d69f8ee8-2721-4124-a54a-686965e816ef · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.647207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.310413Z digest=sha256:c4b260f7ec99731db8d81a7c5058d891d8d3f09c65329e28d98ada2dd93210a3

Observation 08f94af4-4fca-4428-8029-ea535e194cae · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:26:19.639684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.313298Z digest=sha256:7246e7b50d51c281b80e7ca5c92649564298af80996ca074715f95a96da0b46a

Observation d2b6dc8d-aa5c-4335-a1a2-5d2e740ad999 · outbound

This paper cites Libri-light: A benchmark for ASR with limited or no supervision.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Libri-light: A benchmark for ASR with limited or no supervision

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.631954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.316297Z digest=sha256:e3dbcf28696b90c09573cad393d2d82e7d64eaef176ae678681b74bd3a56b8d9

Observation 91b2c8cf-ee67-4ac5-a632-a16b75a59656 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Librispeech: An ASR corpus based on public domain audio books

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.622739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.319210Z digest=sha256:791b92a4566ad6933863565cf2dfe60c819c141d00329fe91a3aaf8b733573ca

Observation cc567952-405a-4834-803e-a1f8637b199a · outbound

This paper cites Amphion: An open-source audio, music and speech generation toolkit.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Amphion: An open-source audio, music and speech generation toolkit

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.613920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.322124Z digest=sha256:c66f1c1df29d6d4cfe39ff52594018d9652030804000c490f6b42361bd440971

Observation 35599e59-9f1a-431a-b55f-394a892cbc4b · outbound

This paper cites V oicecraft: Zero-shot speech editing and text-to-speech in the wild.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oicecraft: Zero-shot speech editing and text-to-speech in the wild

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.605013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.324971Z digest=sha256:101cbd74d2fbd0190d35563a1d2fd9aa8ce4f03d51863a0383394e192e6790b6

Observation d4d9660f-dee0-4fd3-836c-94376ab66820 · outbound

This paper cites Overview of the Amphion Toolkit (v0.2).

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Overview of the Amphion Toolkit (v0.2)

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.327889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.327889Z digest=sha256:667312a639d47002b2e594a21e1f698a4da6ae6acd757e8e0650d09053a59183

Observation f007b9ae-b6b3-4469-a20f-6ff698ee87a3 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.596064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.331052Z digest=sha256:f24d85fc0b18f5970d44c719bc64d0c978480822df9fb02d5ca7103e7b692932

Observation 8ea99c5d-0d73-46ab-a699-2c5b5fd2d259 · outbound

This paper cites Decoupled weight decay regularization.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Decoupled weight decay regularization

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.586922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.333766Z digest=sha256:9768952d3838dc819b38ea1549f3ae5aed18d93a2ee397ee3b1b22722417fdd0

Observation f4437809-355c-4f4e-9839-d4877cffa94a · outbound

This paper cites Kingma and Jimmy Ba.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Kingma and Jimmy Ba

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.336676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.336676Z digest=sha256:843ed96465a933140b0162b695cce105d1eca5ff00da0430548c20ac794ae094

Observation ef3c1bcf-7fce-40a9-a5af-1a4ec8b76b51 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Classifier-Free Diffusion Guidance

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.339612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.339612Z digest=sha256:03e627f982d6c1d2023b827d6db44812d7b36a6f51abdb5baeb79bee17cc1481

Observation 71139633-d1f4-49e7-9afe-ed8423788c6e · outbound

This paper cites Scaling speech technology to 1, 000+ languages.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Scaling speech technology to 1, 000+ languages

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.573407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.342687Z digest=sha256:69c0796455897c15e162219d04414864d8ea31fdceef0c7808408fd85747caf1

Observation fab7ebe3-0ec2-4843-ae1e-2de7f9ddbfb5 · outbound

This paper cites Conditional variational autoencoder with adver- sarial learning for end-to-end text-to-speech.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Conditional variational autoencoder with adver- sarial learning for end-to-end text-to-speech

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.564600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.345544Z digest=sha256:db4707dbdf5d54e20d0eeae66cc9d57870830e7b3a656fd87018a71f448b8146

Observation f0be682c-099a-45b8-9852-3d4430a7214b · outbound

This paper cites Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.555736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.348402Z digest=sha256:c931cb9b35b4f4b38a099c586ab384c4a4a13d48abe1a3720360cb1a7fbf984b

Observation ed0f2eb6-4b91-4dec-9f18-5b4edbf627ae · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.351240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.351240Z digest=sha256:64a633ced8a9e3aa1adc42c4261795c9776b49ed38dc70a18c5031491f4cd7f4

Observation d2eb39b9-9991-480f-9f29-ec8fef4334f2 · outbound

This paper cites BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.541564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.354382Z digest=sha256:d902bad88e384b082ee1c47d51fa0ac41a166926349487d82eb095711d8235b9

Observation c7454f28-2d70-4bc5-869a-936e0d8b3b9b · outbound

This paper cites CSTR VCTK Corpus: En- glish multi-speaker corpus for cstr voice cloning toolkit (version 0.92).

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CSTR VCTK Corpus: En- glish multi-speaker corpus for cstr voice cloning toolkit (version 0.92)

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.532562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.357239Z digest=sha256:d2f6285802dfba838fc830a18601b50e4f089d273c9c299ea3dc458149260ae0

Observation 0b0598f3-d986-426f-9a83-05f99e54586e · outbound

This paper cites High Fidelity Neural Audio Compression.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement High Fidelity Neural Audio Compression

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.360299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.360299Z digest=sha256:e7d9c78fc07b53753cae773a7b594753f4ab9d9db0c3496a98309e54714a438b

Observation 050584c9-9979-4017-a8a1-00ce234971ad · outbound

This paper cites MLS: A large-scale multilingual dataset for speech research.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement MLS: A large-scale multilingual dataset for speech research

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.523449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.363655Z digest=sha256:d1486bcee30a96dbbbb1d250de7c1a19d4813a3bc73027bd9d26162528d2e4f4

Observation b8aac263-107c-4a80-9934-b90cdf3a6e72 · outbound

This paper cites Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:26:19.515151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.366424Z digest=sha256:b32847230118b1f90dae0eb222d9010291c7b694a60c0d01a9394f7ee3e0c807

Observation 06cef095-3980-4f5f-b5c1-4fef04f485f6 · outbound

This paper cites bit” vs. “bet.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement bit” vs. “bet

Reference 92

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T13:26:19.505856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:26:19.369265Z digest=sha256:cc1ea096380a1e3c1cd8d35569b75bf9a8550c8ad21e7ecf889b61496f60380c

Observation 7687b7be-def6-4817-b4ce-463f6fe28086 · outbound

This paper cites an unresolved cited work.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work

Reference 4186

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.293571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.293571Z digest=sha256:96849bef0ed69369c5d64591b97583e16dda39982f99980e948d513960c81dda

Pith citing papers

Observation abde2652-5844-45b5-bc1b-6b14c8aafb41 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.575812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.575812Z digest=sha256:946f118541a22a22726a8d092165d6264c2cd24d19e35b7263b93e8725c3ae6f

Observation 64546f18-8417-4c2a-b48b-0557f261eb3f · inbound

Entropy-based Coarse and Compressed Semantic Speech Representation Learning cites this paper.

Entropy-based Coarse and Compressed Semantic Speech Representation Learning Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:07.305867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:36:07.305867Z digest=sha256:3b9010613a8acfeb63eb653b727f22d251070ba1ce27d9730e200f04548cccf0

Observation 80a7b18a-b5da-4ab4-997e-a3a9b98d8558 · inbound

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck cites this paper.

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:50:51.780513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:51:38.030059Z digest=sha256:19c4d42217ce41375879899ab1f162067fd51b74e90afa1781230cee1f2c0da6

Observation 148ee6a2-99f8-4893-a50c-40d16a870387 · inbound

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora cites this paper.

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:01.383468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:57:07.894455Z digest=sha256:bd6c6dd36cee492a05703aa3b27256128470673985534e90aeaef46ad7dea7bf

Observation aef88493-6709-4379-9fe4-f2919aaf5837 · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:39:38.692633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:bb61e528e734526838bc8f7825bf08b7f4f263872a5546ca213c16f8874cc548

Observation 6c019909-a819-4aef-b337-348e77d88fbb · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.092101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:52158c45e792c5b1720c8513474a3632018fbdf95b335c9fddbcfd74d6f3ef4b

Observation cffc3372-92ee-43d3-99f7-c5a05a951d70 · inbound

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization cites this paper.

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T22:31:42.563915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:31:42.563915Z digest=sha256:94fc052ce69ebbaf007e422620cab9f2e70b55cb1a80912fde6fd23b5a5e0685