Pith. sign in

Paper Citation Record · LEDGER

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion

As of 6 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2601.18094.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.18094 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T11:35:58.050305Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:00:27.852836Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19c87387-2e29-4079-8009-eae1873d6d4a · outbound

This paper cites Streaming voice con- version via intermediate bottleneck features and non- streaming teacher guidance.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Streaming voice con- version via intermediate bottleneck features and non- streaming teacher guidance

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.271534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:5f2b9a51f9b7951353cff82331ef1c129e186adf491c9cec615ee83db6c38580

Observation 91d816af-a723-4002-afea-6fabc1640502 · outbound

This paper cites Yingmusic-svc: Real- world robust zero-shot singing voice conversion with flow- grpo and singing-specific inductive biases.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Yingmusic-svc: Real- world robust zero-shot singing voice conversion with flow- grpo and singing-specific inductive biases.Arxiv

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.178893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:a5245102c6ee373160fb329a74c55216e9203508c1bce7f1b97515b1089f7212

Observation 9f190ef8-61d3-433d-a157-c20f0831b08c · outbound

This paper cites Neural analysis and synthesis: Reconstructing speech from self-supervised representations.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Neural analysis and synthesis: Reconstructing speech from self-supervised representations

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.290104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:e0927f2e3bb32713a3e19a6594d4e2164fa65ba004dc776e20d0b81f7a4ce1e2

Observation 8b30aaee-f30f-4754-8ea4-0e9033182326 · outbound

This paper cites an unresolved cited work.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-22T11:36:29.217619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:df89b513ab1dfa62ae0cc2f869a9d0263ea0d06d1165245717390732dbed98d6

Observation b498b656-2ec5-4a15-ad9b-f63a3b024ab6 · outbound

This paper cites The nus sung and spoken lyrics corpus: A quantitative comparison of singing and speech.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion The nus sung and spoken lyrics corpus: A quantitative comparison of singing and speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.202283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:f917c3b7ce99e2dc560cc952f7131e191aeb271b6b43caaedc3226ff3f3bd18a

Observation 88d8c067-597a-447c-85a9-736238fd5692 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Moshi: a speech-text foundation model for real-time dialogue

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.278134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:70d9d844448d9b10d04ec6153c516080db251bbae853b621a26748b7818078a0

Observation 5bb3d4b0-ccc1-4000-b6cb-903dd435bc0f · outbound

This paper cites Switch transformers: scaling to trillion parame- ter models with simple and efficient sparsity.Journal of Machine Learning Research, 23.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Switch transformers: scaling to trillion parame- ter models with simple and efficient sparsity.Journal of Machine Learning Research, 23

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.298051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:9fbdca1819d6f8cfedd9db0260a1f9f2a5cc70fb2b140ce365c89c247efa6463

Observation 3998f120-7296-481b-9081-a875bb7c3e99 · outbound

This paper cites Zico Kolter, and Kaiming He.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Zico Kolter, and Kaiming He

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.195657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:853b4504b01716fd440cea3b658a64ffa6175bfeda01b32ac77c958db7752b8e

Observation f88b34bf-2489-41e1-aa3f-91ed6ca779bc · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Bigvgan: A universal neural vocoder with large-scale training.Arxiv

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.188972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:16118a936c741119af436f51e8f9f54ba66ae72d30716f929aa5c700eeadc8e5

Observation 297ec85a-436d-4898-8281-49a3721bcb25 · outbound

This paper cites Emilia: A large-scale, extensive, multilingual, and diverse dataset for speech generation.Transactions on Audio, Speech and Language Processing, 33:4044–4054.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Emilia: A large-scale, extensive, multilingual, and diverse dataset for speech generation.Transactions on Audio, Speech and Language Processing, 33:4044–4054

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.185821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:2c74623b3cb7076840c0be7772bab5bfaacf02ee49dacb5f89c4c466e5c05e86

Observation 1c320591-e7cc-4d2c-b6c1-cc75f5a69cc1 · outbound

This paper cites HuBERT: Self- supervised speech representation learning by masked pre- diction of hidden units.Transactions on Audio, Speech, and Language Processing, 29:3451–3460.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion HuBERT: Self- supervised speech representation learning by masked pre- diction of hidden units.Transactions on Audio, Speech, and Language Processing, 29:3451–3460

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.205561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:2e9d49e15d16d7296556c81ac5f881c5487bbbbe4bfc07e3a933f5ffb4f777a2

Observation cfecea9d-92bc-453f-9db6-89d94ecbdf42 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion LoRA: Low-rank adaptation of large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.192479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:007276c965bf7019b084df9c8ba3155fcc4cb57797e38322a5e620e3dddf6e3a

Observation 29be1a8e-c65e-4da5-9f2c-aa131a215253 · outbound

This paper cites Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.232832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:e3357510d026fcb1df2d3cc8552b6d569eb9df73b89236581a1cac1b06635349

Observation 70f90b48-c7f0-4bf2-b541-b2e68a123495 · outbound

This paper cites The singing voice conversion challenge.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion The singing voice conversion challenge

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.268487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:644e038717bc37e11fd5149d155d7bb3f1eb23b8da48a91548d0433b974e0a2b

Observation 87851fa9-cddf-444d-b4b4-bf994d401481 · outbound

This paper cites DiTAR: Diffusion transformer autoregressive modeling for speech generation.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion DiTAR: Diffusion transformer autoregressive modeling for speech generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.199239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:28454101981fe2fb2da3bae53e19b0d716d08ee644a40147c9a3c5cc6361e490

Observation 227dce47-46fc-4752-9625-517637dab163 · outbound

This paper cites Ref-vc: Robust, expressive and fast zero-shot voice conversion with diffusion trans- formers.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Ref-vc: Robust, expressive and fast zero-shot voice conversion with diffusion trans- formers.Arxiv

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.214582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:402d3afc0e52cdf019ad3279b8fb0e732bef32d49f7c45fb8aa9b26e0689a97c

Observation 31ad7fda-1618-49d6-903c-3f66cefd5260 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion mod- els.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion mod- els

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.262253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:63c157bdbe374a4ca6833734649dde5045ba76a5cb467f93338279930b80d069

Observation 88103a3d-4d45-4173-840b-2f6a1c657f1b · outbound

This paper cites Efficient multilingual asr finetuning via lora language ex- perts.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Efficient multilingual asr finetuning via lora language ex- perts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.249541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:39f9f138bfb5eddcbb7313822a77e1e4960951b09d2ae91d25306332fcf2c392

Observation 48720ead-b5b8-4888-93f9-7e690c4a353e · outbound

This paper cites an unresolved cited work.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-22T11:36:29.211144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:af8e59a82965e7f44fce9fdc105f3abec52d61a24904e840252066c7c51a7da8

Observation 8db6212e-b3fc-4c1b-9f32-357467838f71 · outbound

This paper cites Transferring source style in non-parallel voice conversion.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Transferring source style in non-parallel voice conversion

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.208336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:9399486f63d82435ed02fd74bcdc9f0440e819423fd728f8479368dbad4308b3

Observation 481fda3a-022a-4a31-a607-7f7d42905c77 · outbound

This paper cites Learning the beauty in songs: Neural singing voice beautifier.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Learning the beauty in songs: Neural singing voice beautifier

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.182168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:d1895c24a6f6a14973232e168b22a7cefe6993e9d803f7a63e729c1aed1f3c45

Observation 45127fd4-855c-48d2-8891-cb7ac3ec736e · outbound

This paper cites Zero-shot voice conversion with diffusion transformers.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Zero-shot voice conversion with diffusion transformers.Arxiv

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.258872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:4bfa1fdde1b1933e241d7f4ce721fc190588af86c5d8d0af05e2dfc13718f319

Observation bd7e4782-d2d0-47c4-8dfd-529a7a296c17 · outbound

This paper cites Hdmole: Mixture of lora experts with hi- erarchical routing and dynamic thresholds for fine-tuning llm-based asr models.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Hdmole: Mixture of lora experts with hi- erarchical routing and dynamic thresholds for fine-tuning llm-based asr models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.255874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:be1e591c4703e9e5f3e4a050287ca9f68ab111ad992b2ce0a1ccebf4bf7adec1

Observation 5dc1f525-bcf4-4776-991a-63a43de8933a · outbound

This paper cites Scalable diffusion models with transformers.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Scalable diffusion models with transformers.Arxiv

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.242735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:0bf31189917f4c066d4e22cdd77d633c4176b59a27e3be23ddfde3b5953be210

Observation 2b7133db-176d-4531-8ea3-5d03cfea5c38 · outbound

This paper cites Vibevoice technical report.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Vibevoice technical report.Arxiv

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.229863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:4e9cbdce967e0014dd209e3138e6da5646c0191d16a0e7f82237b751e1f3ad0b

Observation b0e87392-298a-4c8b-b905-679510ad55c8 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.265309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:8c22b1977c14dcb0e660fe8638ecaa708992aa24ca4be33fa8e8628f18f464ff

Observation e6c60e03-6bee-4aef-86bb-5d839e6f9ab2 · outbound

This paper cites Singing voice data scaling-up: An intro- duction to ace-opencpop and ace-kising.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Singing voice data scaling-up: An intro- duction to ace-opencpop and ace-kising.Arxiv

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.294575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:89cdf8c592131897d3826b835cbb41dacd86b1aab21c702eb050a2280c9481d5

Observation 6d6c033f-3228-4f13-b64a-ff0bd6371b28 · outbound

This paper cites Li, Hao Wang, Shiyin Kang, and H.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Li, Hao Wang, Shiyin Kang, and H

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.235876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:8a5c1b083c1da8de761dd097af4c9be97f7190c2a85434d7144a4c36cd671456

Observation d3d40483-91aa-4c07-b1a5-855490497f68 · outbound

This paper cites Multimodal latent language modeling with next-token diffusion.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Multimodal latent language modeling with next-token diffusion.Arxiv

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.287788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:e2c14a2259738ccee9f426540ca3d6e15c80e42e11bf434b75bea13fa099803a

Observation 867c0439-611b-47dd-ae25-1ad9952d47dd · outbound

This paper cites Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis.Arxiv

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.239273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:26d0eaf70c3f156514911f9b7e20d2c88e2f4595e7b74420c68691b849c5e842

Observation 32192538-672b-41e9-b944-09775968a27b · outbound

This paper cites Metis: A foundation speech generation model with masked generative pre-training.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Metis: A foundation speech generation model with masked generative pre-training.Arxiv

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.246299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:fedcbcdc7dec782742b4567f63f0043bbb3c4985ccf8b83e3d6b179a3cf3f739

Observation 74b166b9-f8bb-4336-9097-cb386f322334 · outbound

This paper cites Moe-tts: Enhancing out-of-domain text understanding for description-based tts via mixture-of-experts.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Moe-tts: Enhancing out-of-domain text understanding for description-based tts via mixture-of-experts.Arxiv

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.252744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:f97e7d83a08ddf8dca24f59d068cd2a167082a6c4b904435db1d98ae9e031615

Observation d61cd101-c106-486f-b5f5-cbcdb990159a · outbound

This paper cites Uniaudio: An audio foundation model to- ward universal audio generation.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Uniaudio: An audio foundation model to- ward universal audio generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.284992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:061a84e9bb22c8739c4df5e76ecd977599534ce25a4fc62aa032a985df702cb8

Observation 48d0ee07-9e78-4e2c-88c5-b5cb466be1ab · outbound

This paper cites Llasa: Scal- ing train-time and inference-time compute for llama-based speech synthesis.Arxiv.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Llasa: Scal- ing train-time and inference-time compute for llama-based speech synthesis.Arxiv

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.281589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:941a84fc99756a47e4b08b053324c73925ba95ea9460da2a3a0eab6977399149

Observation 802ef9e6-633e-41de-9018-be6165b51762 · outbound

This paper cites Megabyte: Predicting million-byte sequences with multi- scale transformers.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Megabyte: Predicting million-byte sequences with multi- scale transformers

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.223750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:36ba756f420e59f979a50c60d4486a8d8f1a8b1ba3b333ea129f560df7f1c669

Observation 7fc96d1b-709d-462f-9f00-d01719115e12 · outbound

This paper cites Takin-VC: Expressive zero-shot voice conversion via adaptive hybrid content encoding and en- hanced timbre modeling.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Takin-VC: Expressive zero-shot voice conversion via adaptive hybrid content encoding and en- hanced timbre modeling

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.292338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:ba5e257f49d8673af703aa7e1141c2bb6b5e0054a37bdff8c22c96ea55f0a5a8

Observation 6c6efad3-e0fc-4932-8e3c-7e772ddf3c82 · outbound

This paper cites SoundStream: An end-to-end neural audio codec.Trans- actions on Audio, Speech, and Language Processing, 30:495–507.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion SoundStream: An end-to-end neural audio codec.Trans- actions on Audio, Speech, and Language Processing, 30:495–507

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.226774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:9b9f88cc73b4a8bb2d2f6d55b58993f69c24662c5d2e8558ce64ba7c51721eee

Observation 3f3a50f3-1873-4b26-8585-e27c36a65171 · outbound

This paper cites M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.220666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:60ae3ec6fabc997ff001ac2fc6583424d3874a80d94c4b6effc0231de9e5be90

Observation ece65de7-8748-4a7a-8970-0fc1230d80fc · outbound

This paper cites Transfusion: Predict the next token and diffuse images with one multi-modal model.

OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion Transfusion: Predict the next token and diffuse images with one multi-modal model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T11:36:29.274768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T11:35:58.050305Z digest=sha256:1e797bb1d83baba70543165805df920043d372753fcc21de811882f39aa2b919

Pith citing papers

Observation a47df495-4001-4464-829e-102a64b06aad · inbound

MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection cites this paper.

MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T18:00:27.852836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:00:27.852836Z digest=sha256:759d037ae5babd2fdcdd5009b3adb2d611f3a32e11e03df29382089e1dd46d52