Pith. sign in

Paper Citation Record · LEDGER

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion

As of 9 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2507.14534.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14534 v4

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:57:23.386903Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T21:17:48.886421Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:16:12.319512Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cedb4dbb-0fa5-4623-bdf8-6a18757a6d5d · outbound

This paper cites StreamVoice: Streamable context-aware language modeling for real-time zero-shot voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion StreamVoice: Streamable context-aware language modeling for real-time zero-shot voice conversion,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.317774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:20.058155Z digest=sha256:5795bcea49d0de70100cde3d26b25cfd46b006e7cdcbad0e72dc7c0bbf24765a

Observation 77e58ac2-15a7-4164-87f8-b751dd5abf80 · outbound

This paper cites Vqmivc: Vector quantization and mutual information-based unsuper- vised speech representation disentanglement for one-shot voice conver- sion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Vqmivc: Vector quantization and mutual information-based unsuper- vised speech representation disentanglement for one-shot voice conver- sion,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.301228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:20.124148Z digest=sha256:52742dd1721fed537bd1da50daccd7797d774de8452fe0e85793fceb06653987

Observation 6e9ee150-7279-4d55-a59c-157e9a33a042 · outbound

This paper cites Autovc: Zero-shot voice style transfer with only autoencoder loss,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Autovc: Zero-shot voice style transfer with only autoencoder loss,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.284013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:20.205359Z digest=sha256:c1e29e4059a81cbde742a235554341466ccb2cee0c7585309832dc5c56c50340

Observation 7dc0f064-a8a4-444a-98f9-1b116e5e5a28 · outbound

This paper cites Controlvc: Zero-shot voice conversion with time-varying controls on pitch and speed,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Controlvc: Zero-shot voice conversion with time-varying controls on pitch and speed,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.269197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:20.327182Z digest=sha256:54e2e2fd130bf25d572c91820fcf0fbbe8f8ebb9b931848f17a3db0b6a43fcb7

Observation 5b65b122-8ace-47f4-8fcd-ad16828b37ba · outbound

This paper cites Streamable Speech Representation Disentanglement and Multi-Level Prosody Modeling for Live One-Shot V oice Conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Streamable Speech Representation Disentanglement and Multi-Level Prosody Modeling for Live One-Shot V oice Conversion,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.253979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:20.387237Z digest=sha256:365b1689f125555efbcd439f7682f55c2299c2676a0d677a945c5cc0ce75ca94

Observation db8fc721-dc17-4452-8afd-bd97b4987118 · outbound

This paper cites Streamvc: Real-time low-latency voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Streamvc: Real-time low-latency voice conversion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.239188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:20.429437Z digest=sha256:97c4f341b0209209251468fea9e2977d91ccaef71c1c68f773a95790185aa2d4

Observation 5fd170c3-41f0-43cd-92d2-d365d32067cf · outbound

This paper cites TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:20.468550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:20.468550Z digest=sha256:fae07d9e53781919af0bb209bf033bb02f6afcc88d8e61d2442512fc655a58ed

Observation a770f3f7-d1b2-4e12-ab4a-000c0cdf4d44 · outbound

This paper cites Isdrama: Immersive spatial drama generation through multimodal prompting,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Isdrama: Immersive spatial drama generation through multimodal prompting,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:20.561847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:20.561847Z digest=sha256:111fe5171d4f63202642e24a4323b4e802e3ecf1ff50a082087aa7ec755fe5ba

Observation 2cb7ec70-99cb-4f02-9f9f-5c5fcae37b5e · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion A comparison of discrete and soft speech units for improved voice conversion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.224119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:20.642241Z digest=sha256:a87850d30f737a4b9921c668469e38e9a9f2c849438d1ee6a9e3b0bcae11820d

Observation d56cdefa-73c8-4ca9-8f2f-eabfa2678478 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:20.755757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:20.755757Z digest=sha256:df0112c91b93ab18fb9f4b24e8ec4bc79d133e9f971d40e742c910e3d07a654c

Observation 61d08d86-e121-47f0-8d73-727c8a3b5f26 · outbound

This paper cites Phonetic pos- teriorgrams for many-to-one voice conversion without parallel data training,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Phonetic pos- teriorgrams for many-to-one voice conversion without parallel data training,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.196974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:20.822747Z digest=sha256:7a6a2cffde9d64bb1580bcf47a25331848bb40e5224576e6938db69b93685d53

Observation 7a52f3ad-e015-45dc-8824-89840f32ea6a · outbound

This paper cites Starganv2-vc: A diverse, unsuper- vised, non-parallel framework for natural-sounding voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Starganv2-vc: A diverse, unsuper- vised, non-parallel framework for natural-sounding voice conversion,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.180932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:20.892516Z digest=sha256:5792550d1b898953b86fcddfa9c62ddef3aa2a8c99bb5cb3ff34f4e06cb24233

Observation 91124d21-72f8-47bc-90d3-cf7e4999109a · outbound

This paper cites End-to-end streaming model for low-latency speech anonymization,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion End-to-end streaming model for low-latency speech anonymization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.166104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:20.956605Z digest=sha256:ad5aca269ea6cb90cb3c6eebbaf6be2b6b73f79c5499e2904fdbdc99fc509d14

Observation 6a760ef3-0401-43bf-bd99-f3d878867ffd · outbound

This paper cites Contrastive predictive coding supported factorized variational autoen- coder for unsupervised learning of disentangled speech representations,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Contrastive predictive coding supported factorized variational autoen- coder for unsupervised learning of disentangled speech representations,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.151943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:21.020533Z digest=sha256:39e728931051d44cb4ea63b5b5e03c56db9a25dc4156f7d0d4221f07a374d647

Observation 96a890f9-a376-4685-8dd6-17cd857f0e3c · outbound

This paper cites Neural analysis and synthesis: Reconstructing speech from self-supervised representations,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Neural analysis and synthesis: Reconstructing speech from self-supervised representations,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.136970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:21.081301Z digest=sha256:9e46931deeefd6d46b50a4e8b9f00647367f048438f10b74804cb58ceca7eeec

Observation 1b9e8804-0a17-4676-90be-d75e25940d71 · outbound

This paper cites Lm-vc: Zero-shot voice conversion via speech generation based on language models,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Lm-vc: Zero-shot voice conversion via speech generation based on language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:21.209242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:21.209242Z digest=sha256:09533b60e51bcab51af9f3146906d998d5424419249dfbfe9384dfd12640dc0d

Observation 8c907b39-3523-476b-88cc-fb4502ebad7b · outbound

This paper cites An investigation of streaming non-autoregressive sequence-to-sequence voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion An investigation of streaming non-autoregressive sequence-to-sequence voice conversion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.111657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:21.349497Z digest=sha256:9af152cbd981a1061387d872c32b508ca2052f5b0b941acc1fe36f15ec597c2e

Observation 7b4fa9f6-cead-40b6-99e0-3546858cac79 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:29.078539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:21.451568Z digest=sha256:336ba0525dd21c7e3bedcabf389bf90cc1959e93ea5475a7c8521afec71a42fa

Observation 20d40fe8-b908-477f-8a3b-8040792bb92c · outbound

This paper cites Non- autoregressive sequence-to-sequence voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Non- autoregressive sequence-to-sequence voice conversion,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:28.987253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:21.521693Z digest=sha256:5bc5bc738eac41bfc1e8c9d551833a5ed29b606e3ffa9e1ccc7dbf0c6c6f878c

Observation 99af9cd1-284f-41ba-b65d-df1f8d8ce90e · outbound

This paper cites FastS2S-VC: Streaming Non-Autoregressive Sequence-to-Sequence Voice Conversion.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion FastS2S-VC: Streaming Non-Autoregressive Sequence-to-Sequence Voice Conversion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:21.646838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:21.646838Z digest=sha256:d8670f53e45cdfa688b6701cedb94d64bd36f74dea5421550b0aa6d2c0e9e6bb

Observation 413f6a08-89af-41ff-aa52-b939132cfaf1 · outbound

This paper cites Streaming voice conversion via intermediate bottleneck features and non-streaming teacher guidance,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Streaming voice conversion via intermediate bottleneck features and non-streaming teacher guidance,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:28.922925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:21.788288Z digest=sha256:4f122d09c2cd76b69aa9306a53dbc031b86a9a2c20e36af31a42f53edaa7e046

Observation d80a12b9-c892-4f3d-b002-b69ee35365d6 · outbound

This paper cites Dualvc 2: Dynamic masked convolution for unified streaming and non-streaming voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Dualvc 2: Dynamic masked convolution for unified streaming and non-streaming voice conversion,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:27.723679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:21.890854Z digest=sha256:88e1d632afc6d7a872b791c7f933cfd6a054c5c612b2b9c35a2c2e4b24d53516

Observation d29c1678-9cc9-459f-bf71-ea0f94c461cd · outbound

This paper cites Alo-vc: Any-to-any low-latency one-shot voice conversion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Alo-vc: Any-to-any low-latency one-shot voice conversion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:26.410264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.006855Z digest=sha256:383090152ac6583eab3ad082c81440a1297dc22527410dc6c9e61b961790bc71

Observation 29fa0b9f-0546-45cd-9807-a5059a1343d4 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:26.009254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.085174Z digest=sha256:7f6774b23a57ba16de939a1e80466b5f61ae3b2874d934e0b47ad0e83e3ffdd3

Observation 7f253c7c-e32b-4f9c-9cbe-6eff216ecd08 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:22.132763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:22.132763Z digest=sha256:f1d16fe86cfd1285a406f795bb09086d0a344d27149cdff32c75956793d13b62

Observation f02d83d8-498d-4482-b44d-365368ac1e63 · outbound

This paper cites Attentron: Few-shot text-to- speech utilizing attention-based variable-length embedding,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Attentron: Few-shot text-to- speech utilizing attention-based variable-length embedding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.846045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.184149Z digest=sha256:af1d6af129c1e16fac27cba5ec8b5b8c93d67f4b9147fc5feb7f402c937fb07c

Observation 6bf52a16-2edb-4b8d-8599-c30401a74128 · outbound

This paper cites Normalization driven zero- shot multi-speaker speech synthesis.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Normalization driven zero- shot multi-speaker speech synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.663916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.238284Z digest=sha256:449b2cb0c8257c6c4397e81158ec35d48f1471a59883d84edb3eeae3f0c6030a

Observation 6914e1ce-e1a9-4fd4-84a3-6b8419a2dfcc · outbound

This paper cites Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:57:23.740340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.279661Z digest=sha256:c4c7042825345b341dbf2d3e118b012f76aa6746221c17fbaec5c4364317ba09

Observation 59926bab-b637-4466-ac8a-85b31fa7c54d · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.482486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.352742Z digest=sha256:3aeb71981b31d5db179664a71040e4a03c868e53242365c564b770adb78efca9

Observation 14f439df-7073-4c64-87c6-958de311e3f5 · outbound

This paper cites Styler: Style factor modeling with rapidity and robustness via speech decomposition for expressive and controllable neural text to speech,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Styler: Style factor modeling with rapidity and robustness via speech decomposition for expressive and controllable neural text to speech,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.383341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.405331Z digest=sha256:b7ff5f0934130e95781c3b1a91f50ebedceacaa1e4d5fac8590e5135f93aa17c

Observation e10b327d-5ade-4bca-85bb-6fc4cceee77a · outbound

This paper cites Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.247102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.447883Z digest=sha256:86f3ea340afbb6c867f5a60d87192cf7c20730b3fd8c12a091edab1fa180280e

Observation b9c96fa5-2edf-40a3-b1a4-501bdbc90c3b · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:25.119522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.539901Z digest=sha256:fe24c8d264faea144a1f86eed02cbf02ede06c1a1b42d949de84a1e374da0606

Observation 0fabef11-8d8b-4b56-867a-9e69465e9f56 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:22.591196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:22.591196Z digest=sha256:4bdf1a7e66bd9df6acec44389da43d5ce86ee5221476b2959c642457b18a71e3

Observation df94fc8e-f56e-4d3c-b282-fb44fed0ee30 · outbound

This paper cites Fastpitch: Parallel text-to-speech with pitch prediction,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Fastpitch: Parallel text-to-speech with pitch prediction,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.969094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.641193Z digest=sha256:8083016513a0ea870ee481260d8ef23a5940e351255c40915db8e80628d6b042

Observation aa4e9205-55d7-43d1-98be-f6f2e64cde40 · outbound

This paper cites Tcsinger: Zero-shot singing voice synthesis with style transfer and multi-level style control,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Tcsinger: Zero-shot singing voice synthesis with style transfer and multi-level style control,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.822684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.706455Z digest=sha256:8f0e9303e1e1422ac694de36c1621ff5760755ef389ca628dbb7c2352398e6aa

Observation a45c2ba1-c090-4c1f-98f4-1dabd0e10053 · outbound

This paper cites Online clustered codebook,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Online clustered codebook,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.718540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.774621Z digest=sha256:bccf1834b99fa0f53e49124f0c6a099c5b08f441e0c30b6e4a41138ed1a552fb

Observation 413b1fd2-3a26-447e-b60c-260114e88022 · outbound

This paper cites Versatile framework for song generation with prompt-based control,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Versatile framework for song generation with prompt-based control,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:22.843384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:22.843384Z digest=sha256:1ab5ec720cf6e44b093c324e8f6c5273f73b3f654b651da00d2763cc780029e3

Observation d6d7515e-12de-4b4a-9708-c365d02980b0 · outbound

This paper cites Neural discrete representa- tion learning,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Neural discrete representa- tion learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.601983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.872499Z digest=sha256:b3a09bbaf38acb9217148315f57aaae4b7ec071afbb6a7e45bb9a401ec218544

Observation 192d2527-d4dd-4399-9079-c3b0aa3481c2 · outbound

This paper cites Stylesinger: Style transfer for out-of-domain singing voice synthesis,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Stylesinger: Style transfer for out-of-domain singing voice synthesis,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.472361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.907363Z digest=sha256:b7fc7d8f5d11fce866e53d331234b3fe7038180810ccd11cc18a734c2f6e6d7e

Observation 70e69050-c6de-4617-bf63-2846ad79f51a · outbound

This paper cites Attention is all you need,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Attention is all you need,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.366429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:22.946130Z digest=sha256:e2ac4b2930bb39c0405f02313431a305a169aec8b42b3a6d6c508b10c3264f83

Observation a3c0a13f-26ed-4cf6-8595-a7cfd7a99d5f · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:23.040525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:23.040525Z digest=sha256:b3182394e00ce0b9b627fd6168d66c611a3c34c211453b4c516fbe4aa5f079b8

Observation 2d7addaf-f07e-44a2-94c2-cd3f7f332b14 · outbound

This paper cites Deconvolution and checkerboard artifacts,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Deconvolution and checkerboard artifacts,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:23.075340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:23.075340Z digest=sha256:2aee1fb32d1d13e1ed42e7d41d910a3d0c127c4e63c93a992931411663d09853

Observation cc8f04d0-4f09-4b09-b8e2-1ba40b78caf3 · outbound

This paper cites Least squares generative adversarial networks,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Least squares generative adversarial networks,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:23.098616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:23.098616Z digest=sha256:058bb38a2fe60ca8839d75a9996a8cf55229310c3f7f51248eb3be34cc5d5327

Observation e00da223-e34f-4250-bffe-f10ee2f37467 · outbound

This paper cites Libritts: A corpus derived from librispeech for text-to-speech,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Libritts: A corpus derived from librispeech for text-to-speech,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.234647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:23.174327Z digest=sha256:210e9e9972b05965748274043fa7e3d104e8d2965fa02a8910b0ba378d2baff2

Observation 29c8e372-8a71-449e-84b6-078ce53fbe70 · outbound

This paper cites Gtsinger: A global multi-technique singing corpus with realistic music scores for all singing tasks,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Gtsinger: A global multi-technique singing corpus with realistic music scores for all singing tasks,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:24.096294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:23.233471Z digest=sha256:c3f92498142543219595510f2368666677da4ac983d65273d35005f72379877d

Observation 8969748c-db12-4245-bde3-787651920458 · outbound

This paper cites Any-to-any generation via composable diffusion,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Any-to-any generation via composable diffusion,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:23.957767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:23.261921Z digest=sha256:46a54ac169012b3a5369101ec72d00c6fbc7c8ee667c607844b7854fcb092976

Observation 4fa7bbd9-077c-481f-a3e6-0f43e14ac6e6 · outbound

This paper cites Any-to-many voice conversion with location-relative sequence-to-sequence modeling,.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion Any-to-many voice conversion with location-relative sequence-to-sequence modeling,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:23.295591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:23.295591Z digest=sha256:c5c41bb8b7839f4d3ea696c8e917a2226482b32f13500f45a42f762520e06df6

Observation 5a24e83c-6473-4627-860b-cd2028b88c1c · outbound

This paper cites QuickVC: Any-to-many Voice Conversion Using Inverse Short-time Fourier Transform for Faster Conversion.

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion QuickVC: Any-to-many Voice Conversion Using Inverse Short-time Fourier Transform for Faster Conversion

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:57:23.502029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:57:23.386903Z digest=sha256:a1daf4142f7ee55499a6290e88e009f35e4a056c2726e16dc9c3f1095da6420c

Pith citing papers

Observation 316d538f-1f07-4436-96e3-d160f469a179 · inbound

X-VC: Zero-shot Streaming Voice Conversion in Codec Space cites this paper.

X-VC: Zero-shot Streaming Voice Conversion in Codec Space Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:20:29.828049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:20:01.885544Z digest=sha256:a030921504f59c92fc2602f75d5af0582bc69a48e85eb0a6ccc5a1d53c83c868

Observation 67293642-b383-4b5c-8fda-7708339eed94 · inbound

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer cites this paper.

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:16:12.321170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T21:17:48.886421Z digest=sha256:3e47ba2a96094dcde95d63b9ed38e5b1650f2f6224a27b0699c245c674ce7cb2