Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2304.09116.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:56.312162Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T18:40:03.192889Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f47485fe-6dab-43e4-b8fe-365529144a4b · inbound
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 139
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c322811-7749-48e4-bc4b-4936f82b6af3 · inbound
Movie Gen: A Cast of Media Foundation Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5407115-8935-4156-98a6-2901d3393e33 · inbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74c40840-3b67-4339-a24d-74fc4cd50c58 · inbound
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84115302-5a8a-4079-8810-af0dbd16bb2d · inbound
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e224cf6-1374-4e0e-851a-8f60fdf62d64 · inbound
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec02f16b-1a6e-4d7c-89f3-9d90ad4f2018 · inbound
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96b3cfc-3d52-47d9-b621-0ef2835b38f1 · inbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b20f27b-795a-4b6b-b9ef-dba358d4ba82 · inbound
SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faba2588-9e08-487e-b264-81415b3455e9 · inbound
Next Tokens Denoising for Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dd2c9a9-4ed4-4bde-be05-e8b4c8dbfb64 · inbound
Marco-Voice Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8139e27-ff8e-4e4d-86e3-ba24bdd26ffb · inbound
Inference-time Scaling for Diffusion-based Audio Super-resolution NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6562b558-0a3e-403b-9f40-55481a3dceab · inbound
Audio-Guided Visual Editing with Complex Multi-Modal Prompts NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74f964eb-a931-415f-8f63-f15b96a40155 · inbound
FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ecbb115-ba05-42da-bc95-60a3ee269242 · inbound
TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for \"U-Tsang, Amdo and Kham Speech Dataset Generation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5aa886f0-7a3f-46af-a0b0-bdfb212737bb · inbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2953e945-6422-45e4-b6ab-0d92a97d0c64 · inbound
Qwen3-TTS Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa40bc84-52a8-4ae8-914d-109341bc5d0d · inbound
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4843fa1-2926-4e71-8322-db573e7e4dc3 · inbound
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e7439ab-4d74-46fe-b29a-8072b6e7f30e · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 46c4a353-49c8-4094-9e8c-fc49dce02c09 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c603281-0c77-48ca-a921-0538cb1e2ce2 · inbound
Scaling Properties of Continuous Diffusion Spoken Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 648a05cf-8f2f-41f5-a8bc-d95b2bd7ed38 · inbound
Voice "Cloning" is Style Transfer NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5a2da5f-67c4-4826-b9c7-27bf2f535ccc · inbound
SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f93442d7-5bf7-464a-b5f6-aadf112dbf63 · inbound
UniVoice: A Unified Model for Speech and Singing Voice Generation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b70f74c-f86a-40bd-aac6-a1530e6b36e5 · inbound
EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99b1b84c-9075-47ee-a0fb-c22b443a5885 · inbound
MeshFlow: Mesh Generation with Equivariant Flow Matching NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef844789-06ca-4d27-aa25-2caf8243ba5c · inbound
ZONOS2 Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 190
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f2f3709-6fd6-46c7-a37b-afc570b0240e · inbound
ZONOS2 Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 190
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 78488924-78b0-442e-9b36-f069452bdea2 · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f91965fd-1d8d-4831-aa05-a7edb4019bbb · inbound
SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb60fb7-d20c-49c9-93c2-cc2bafa5bac3 · inbound
StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43fc6afc-8c13-4382-8c7d-7ab8c9e49b12 · inbound
Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21be1899-e25c-4269-8f96-a72d1861f903 · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 265648e2-a0c1-4985-8af3-e4ba1b4ce885 · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.