Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:26:19.369265Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 7 inbound Pith citation observations for arXiv:2502.07243.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:26:19.369265Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:21:59.575812Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:39:38.691071Z
93 of 93 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1234a951-ef09-4289-a4bb-218485f1202d · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural discrete representation learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9621dbf3-776a-4578-87e4-6d5cc7130f2d · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdcc9f2f-cdc6-46b8-8c46-646cb2c70740 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement An overview of voice conversion systems
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c384d50e-393a-4c38-9e76-05d4c1c59737 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement An overview of voice con- version and its challenges: From statistical modeling to deep learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef66d9b-3efe-4cc5-9fdb-32a19f32fb3b · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Foreign accent conversion in computer assisted pronunciation training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f2d5106-5bcc-4539-92f2-5c08a4174d82 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement L2-ARCTIC: A non-native english speech corpus
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84010b50-fc81-409d-b51c-3e124781da88 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Emotional voice conversion: Theory, databases and ESD
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f4141dd-7b92-4bc3-a3ed-6c8dade41c33 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural Text-to-Speech Synthesis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbf8051d-f030-4f39-8e3f-6990403e10e7 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Converting foreign accent speech without a reference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ddc4a1-f58f-452b-9dab-64b5718a6ba1 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Sahidullah, Aur ´elien Bellet, Marc Tom- masi, and Emmanuel Vincent
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a4d2f1-477d-49e4-9574-6aa2484f14a0 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab9f82f4-1ab7-4ad4-82d3-3aab0706d44f · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b705f3-918e-4255-a214-222525ab2f0e · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Maskgct: Zero-shot text-to- speech with masked generative codec transformer
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07ff5752-9eaa-4c8d-a501-508a918d48fe · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a2d694b-e80e-42a1-9d7c-b13db2eb9f4b · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Speech resynthesis from discrete disentan- gled self-supervised representations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e9e5b21-1953-45c9-9c85-9baf1e1642fb · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b580d5-50d5-4831-b8b8-aab0451b14bc · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48cc366c-9511-4c1d-b45c-122c4c30fd9f · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b16bd86c-730f-47cb-a719-d59fc2465f63 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b4e0abf5-deb8-4310-a6aa-de5bb23ebfa5 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e1cff98-0b03-451c-8ab3-a3b8d7261506 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Convert and speak: Zero-shot accent conversion with minimum supervision
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d63fbd8f-c1e4-4798-99a1-4a417dabffc5 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Au- tovc: Zero-shot voice style transfer with only autoencoder loss
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be8810d0-db54-481e-abb2-683a25f231af · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712c1ffd-5225-45d3-bd18-6aba84d44a47 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54dc202b-f995-44b5-9425-fc2300fe4d46 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Better speech synthesis through scaling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dd7fa07-0c42-4fbc-9577-7def92b7ad11 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bf9cba5-4724-401d-84a2-5dc3a569a6c9 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oicebox: Text- guided multilingual universal speech generation at scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fc2728d-5e2a-48df-ab02-34693677b7c2 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ff1696aa-260f-49fe-be8d-3e0e641aa098 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oice- preserving zero-shot multiple accent conversion
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d8b0e25-1f48-424f-a484-258480a55ff0 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Schuller, and Haizhou Li
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e4bf249b-c83e-41c0-bb00-de4227ca1539 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement PA VITS: exploring prosody-aware VITS for end-to-end emotional voice conversion
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb915b3c-08b9-4090-8a15-d8a50bf05a15 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Transfer the linguistic representa- tions from TTS to accent conversion with non-parallel data
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a52ec863-cd90-4e3b-97da-87c94d9798dc · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement U-style: Cascading u-nets with multi-level speaker and style modeling for zero-shot voice cloning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32c230e2-b170-4780-855b-e79bba4ccccb · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6baec9ef-0a8b-435f-909f-835a1966cab0 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement LLaMA: Open and Efficient Foundation Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb292e5-e1b2-4a92-b961-b30131fcd04a · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f59cd892-b39e-4fb8-b11f-f5882fcac97e · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Scalable diffusion models with transformers
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f9b274d1-daa9-4b5a-85f6-5411061b25d3 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02e7d8be-5b9d-4760-aba5-1cb3debb871c · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement One-shot voice conversion by vector quantization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fff07ec3-4590-4663-93c6-cc10cca13954 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unsupervised learning of disentangled speech content and style representation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b2c7c32-6e36-488c-843a-6e54cf7b2674 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Text- less speech-to-speech translation on real data
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f2dd42d1-ebc7-41b9-9bc5-b15b453909e6 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cce790f6-3be1-4ef6-8e45-4148352c4a91 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement A comparative study of self-supervised speech representation based voice conversion
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e0f126a-39a0-4447-90e0-55372264de06 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Leveraging diverse semantic-based audio pretrained models for singing voice conversion
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c6debe6e-8d3c-4c3c-9386-b7b172cb9bbb · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b21ccbbb-9ffe-4d95-b1a1-946d7527d093 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Stargan-vc: non- parallel many-to-many voice conversion using star generative adversarial networks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7ade7335-7583-48a0-bb39-d1564a094f0a · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e579a085-1702-4f69-8e7a-d3a41f7d7b54 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0117c55c-388b-4907-b30b-5a76b2feb2ed · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Tts-guided training for accent conversion without parallel data
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation efd90fa0-40a8-41eb-84b7-b3095e3c6b4f · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement End-to-end accent conversion without using native utterances
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2333a5bb-d67a-49cb-a334-c7bb3dee15cc · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e933a787-f410-4603-b343-601a48a25b3f · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement LM-VC: zero-shot voice conversion via speech generation based on language models.IEEE Signal Process
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 49e1554d-4460-4704-a07a-acb5915ffb54 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 616a3740-187f-4e40-9ae3-7ad81e2b7bca · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement A comparison of discrete and soft speech units for improved voice conversion
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f46fa70-3f63-4b79-80e6-1882ea7259c6 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement SEF-VC: speaker embedding free zero-shot voice conversion with cross attention
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 019df985-128f-4a50-9923-fc61b67370ae · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Neu- ral analysis and synthesis: Reconstructing speech from self-supervised representations
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d86c200-b042-4825-9fd0-b3f4f953501e · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement NANSY++: unified voice synthesis with neural analysis and synthesis
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation def82cc8-225b-4949-9917-0a981577767f · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottle- necks
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0e34687-921c-4eed-87e3-61b6123eb262 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Cox, Mark Hasegawa-Johnson, and Shiyu Chang
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0381617-1b1b-4451-9a2e-b8ba001f3301 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Multi-speaker expressive speech synthesis via multiple factors decoupling
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation efdb67c7-5795-43ff-878b-ab42b2040bc5 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CLUB: A contrastive log-ratio upper bound of mutual information
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44448173-2d8b-4098-ae9b-fd6b8fe82fd4 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Repcodec: A speech representation codec for speech tokenization
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ee6f654c-e8bf-43d3-92ba-67f5aa894cb3 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Soundstream: An end-to-end neural audio codec
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a7660a3-b4de-4b7a-852e-2f8cbe5b0437 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Wavlm: Large-scale self- supervised pre-training for full stack speech processing
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04a83850-4d53-4914-b52d-5c58321399a7 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb1b2cf2-77a4-4c63-8175-dbb8f8527d1a · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BERT: pre-training of deep bidirectional transformers for language understanding
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1caf02e0-b17b-4418-ab8a-847aa87da851 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Bigvgan: A universal neural vocoder with large-scale training
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97818603-d8b1-4009-b416-4ffbedb85cdf · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Tyers, and Gregor Weber
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a190bd8-a972-40ec-b49f-710944c8370f · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement The singing voice conversion challenge 2023
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 41e9396b-97de-4d23-847c-cb904661bfc9 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Robust speech recognition via large-scale weak supervision
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f129dbbf-ba2e-4d4d-b419-611e0c960a33 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d69f8ee8-2721-4124-a54a-686965e816ef · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement emotion2vec: Self-supervised pre-training for speech emotion representation
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 08f94af4-4fca-4428-8029-ea535e194cae · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d2b6dc8d-aa5c-4335-a1a2-5d2e740ad999 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Libri-light: A benchmark for ASR with limited or no supervision
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91b2c8cf-ee67-4ac5-a632-a16b75a59656 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Librispeech: An ASR corpus based on public domain audio books
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cc567952-405a-4834-803e-a1f8637b199a · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Amphion: An open-source audio, music and speech generation toolkit
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 35599e59-9f1a-431a-b55f-394a892cbc4b · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement V oicecraft: Zero-shot speech editing and text-to-speech in the wild
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d4d9660f-dee0-4fd3-836c-94376ab66820 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Overview of the Amphion Toolkit (v0.2)
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f007b9ae-b6b3-4469-a20f-6ff698ee87a3 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ea99c5d-0d73-46ab-a699-2c5b5fd2d259 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Decoupled weight decay regularization
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4437809-355c-4f4e-9839-d4877cffa94a · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Kingma and Jimmy Ba
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef3c1bcf-7fce-40a9-a5af-1a4ec8b76b51 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Classifier-Free Diffusion Guidance
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71139633-d1f4-49e7-9afe-ed8423788c6e · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Scaling speech technology to 1, 000+ languages
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fab7ebe3-0ec2-4843-ae1e-2de7f9ddbfb5 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Conditional variational autoencoder with adver- sarial learning for end-to-end text-to-speech
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0be682c-099a-45b8-9852-3d4430a7214b · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ed0f2eb6-4b91-4dec-9f18-5b4edbf627ae · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement wav2vec 2.0: A framework for self-supervised learning of speech representations
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2eb39b9-9991-480f-9f29-ec8fef4334f2 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c7454f28-2d70-4bc5-869a-936e0d8b3b9b · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement CSTR VCTK Corpus: En- glish multi-speaker corpus for cstr voice cloning toolkit (version 0.92)
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b0598f3-d986-426f-9a83-05f99e54586e · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement High Fidelity Neural Audio Compression
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050584c9-9979-4017-a8a1-00ce234971ad · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement MLS: A large-scale multilingual dataset for speech research
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b8aac263-107c-4a80-9934-b90cdf3a6e72 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06cef095-3980-4f5f-b5c1-4fef04f485f6 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement bit” vs. “bet
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7687b7be-def6-4817-b4ce-463f6fe28086 · outbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Unresolved cited work
Reference 4186
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abde2652-5844-45b5-bc1b-6b14c8aafb41 · inbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64546f18-8417-4c2a-b48b-0557f261eb3f · inbound
Entropy-based Coarse and Compressed Semantic Speech Representation Learning Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a7b18a-b5da-4ab4-997e-a3a9b98d8558 · inbound
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 148ee6a2-99f8-4893-a50c-40d16a870387 · inbound
MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aef88493-6709-4379-9fe4-f2919aaf5837 · inbound
An Evaluation Framework for Text-to-Speech Voice Reconstruction Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6c019909-a819-4aef-b337-348e77d88fbb · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cffc3372-92ee-43d3-99f7-c5a05a951d70 · inbound
NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.