Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:21.772183Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.11650.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:21.772183Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 522995db-b11a-42d5-a5e9-e411b986dd1d · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f0d584-4046-4020-99df-8a05eae7e989 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Better speech synthesis through scaling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 168de320-8cbc-4adb-9c74-5a0875765363 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08e61146-fd25-4f06-8e69-9f784b6de208 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 307bb490-e946-4a7c-92c3-5c23997d3d48 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder WavLM: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5d63568e-4495-4fc0-8521-de52398f9c33 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eaf9818-2954-4232-ad28-8a18644b280f · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d1e8297f-5517-4315-803e-1177a6098a76 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf21efb6-8216-4119-a169-78a80e1930b9 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2b3a91b3-eca5-4e8c-8f8c-d74484b897e8 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77a6d468-c0c5-4318-89ad-8d6560c533b1 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c9e8fee-bcc0-4b66-8f6a-1823e3024333 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12761c13-59d4-4838-8a85-fc06de9715be · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder ElevenLabs multilingual text-to-speech, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5abfd9d1-ad6e-4e63-9de9-ed7c959a217f · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0ac40676-6dfb-435f-aab8-561f0b362cc6 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Fish audio speech-2, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d1b00463-796a-4ac9-9366-52cceed22412 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder CV3-Eval: The cross-lingual evaluation benchmark of CosyV oice 3
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a0972a56-f80c-46fc-a82b-3175a160d78c · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition.Proceedings of Interspeech, 2022
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0714bb5c-6e1f-4cac-bf3f-280becf7cd01 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder MOSS-TTS technical report.arXiv preprint arXiv:2603.18090, 2026
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a065e00d-6607-4303-afb9-a04c549ece6d · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Classifier-Free Diffusion Guidance
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 254a2655-16a3-4a86-a49d-20a7c8408ab7 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Qwen3-TTS Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a488129c-6f09-40eb-8ab4-ef4ab1757731 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Mistral 7B
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457fcc10-450a-4318-bd55-271b957749f9 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0556a7c5-451b-4c06-94bb-3e5db3357557 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee8c75a-397e-4216-9bbc-02691f87df44 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder BigVGAN: A universal neural vocoder with large-scale training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e16a0760-7edd-4ee4-b87d-cdc3a09394ed · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder IndexTTS 2.5 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 334e7c7d-3041-48b5-a39b-e8664ebd08f0 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Flow matching for generative modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eb3de67e-0a51-43a5-b732-6d357ce07bf0 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Cross-lingual F5-TTS: Towards language-agnostic voice cloning and speech synthesis.arXiv preprint arXiv:2509.14579, 2025
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5dd4ecca-953a-4bad-8f56-b72790ca8ef1 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Zero-shot Voice Conversion with Diffusion Transformers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc50cef-a039-4e7c-ae10-aad928a72906 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Decoupled weight decay regularization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00901e98-f154-4607-9daa-0868ec04a7d7 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Matcha-TTS: A fast TTS architecture with conditional flow matching
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68779867-47e2-48f7-bca3-bd55533e7e13 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Attentive statistics pooling for deep speaker embedding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 74a6e69e-a1f5-4461-8d6a-68c182873978 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Scalable diffusion models with transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73795dc4-a4d7-496c-816f-3f3d6294f3ad · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Qwen-audio-3.0-tts: Freely controllable and highly robust speech synthesis with multi-stage training paradigm.arXiv preprint, 2026
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ee23f655-280d-4ca3-83a9-fd35fa2e5798 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Robust speech recognition via large-scale weak supervision
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417f6fb2-4093-4889-b312-c400a4e0a83d · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Language models are unsupervised multitask learners.OpenAI blog, 2019
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1844908d-efb1-4af3-9bc8-dc0d07380c8b · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Chatterbox-TTS
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4ebfec16-3416-4c58-9d42-04eb0f5a098d · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Natu- ralspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2b95ccd5-ab9b-4332-a760-f7c158ac93cf · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cfb784c-0ac9-4fd6-a210-2d260750e22f · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2eecb1a-ea0b-4bf8-93c3-ff37f5c41f39 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81cb5a0f-c28f-4f44-b65c-1dff527c1049 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a4408303-cfc5-49ef-bf47-55239e67808e · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 846ea7d4-6523-4177-a59c-7b1b2f59c24c · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder IndexTTS2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 242cc360-1cfa-4113-afcc-df98209f47ca · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder VoxCPM2 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92b450f3-d50b-4a05-9438-f792287449f3 · outbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.