Pith. sign in

Paper Citation Record · LEDGER

NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2304.09116.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.09116 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:56.312162Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:03.192889Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f47485fe-6dab-43e4-b8fe-365529144a4b · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.517814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:0a87c5c08b99807b4dcb632ed344065d6d584784f21a1f92b7bad9044434d9b6

Observation 0c322811-7749-48e4-bc4b-4936f82b6af3 · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.914341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:a4df50479d594001c094001e0945724258ed9d0fee61a6c0181cc1ff6a375d85

Observation d5407115-8935-4156-98a6-2901d3393e33 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.312162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.312162Z digest=sha256:5e9ee2f81358fd1f1bc911843295f73ceb90bda176bdb0d915441d4bd32e2f11

Observation 74c40840-3b67-4339-a24d-74fc4cd50c58 · inbound

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model cites this paper.

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:56.445581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:56.445581Z digest=sha256:5d4030ea8853eeabe086b3fd5f9d47dd05c765acdc2154406340487ab9a6e468

Observation 84115302-5a8a-4079-8810-af0dbd16bb2d · inbound

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations cites this paper.

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.184518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.184518Z digest=sha256:f121c9ea087e28f7346cdcde5c0d0bc0142e13d16ccb9ebd4a328cefaef54259

Observation 4e224cf6-1374-4e0e-851a-8f60fdf62d64 · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.041243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:4a9fe4270ee6824607fa27b04b0bda8ef1ae9a75c22594e36235cf2347d3e98f

Observation ec02f16b-1a6e-4d7c-89f3-9d90ad4f2018 · inbound

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech cites this paper.

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:00:07.579047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:00:07.579047Z digest=sha256:e1ff8002bc695e923747d51a099fd1e8885677c0ea20f6b67ccd3eab9930bce6

Observation a96b3cfc-3d52-47d9-b621-0ef2835b38f1 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.184511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.184511Z digest=sha256:25cacd3e6f51935ea0a56664efdf9418ffa8bbc25aea3645f39f49e85eff1c02

Observation 3b20f27b-795a-4b6b-b9ef-dba358d4ba82 · inbound

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations cites this paper.

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:00.825953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:00.825953Z digest=sha256:9a8f234a7d077b4840b37bd314910a57f3279cd470fe0e757157c82f357fab31

Observation faba2588-9e08-487e-b264-81415b3455e9 · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.121148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.121148Z digest=sha256:1426a2893eae456d446c196660610bcafc63076f37ab2e36c295f2f187dd2853

Observation 3dd2c9a9-4ed4-4bde-be05-e8b4c8dbfb64 · inbound

Marco-Voice Technical Report cites this paper.

Marco-Voice Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.401559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.401559Z digest=sha256:6fd7052cdc93378ec0d79b5f8ec2bad4ac9cc4d152520621e76e5438dfe204de

Observation e8139e27-ff8e-4e4d-86e3-ba24bdd26ffb · inbound

Inference-time Scaling for Diffusion-based Audio Super-resolution cites this paper.

Inference-time Scaling for Diffusion-based Audio Super-resolution NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.156213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.156213Z digest=sha256:e69cdc82aaeb1322c578cc36157668c189ef776592297b9ef719f6c1afd9ca64

Observation 6562b558-0a3e-403b-9f40-55481a3dceab · inbound

Audio-Guided Visual Editing with Complex Multi-Modal Prompts cites this paper.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.804516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.804516Z digest=sha256:e9ed09f979954ff9a0d6e743532a58e36968f564417e095a02a4b342479f9f4e

Observation 74f964eb-a931-415f-8f63-f15b96a40155 · inbound

FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks cites this paper.

FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:35:01.874077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:35:01.874077Z digest=sha256:45da6947aa885961638f5807a14f945db25b91d359ce3d281f8b0496ad2d22f6

Observation 2ecbb115-ba05-42da-bc95-60a3ee269242 · inbound

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for \"U-Tsang, Amdo and Kham Speech Dataset Generation cites this paper.

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for \"U-Tsang, Amdo and Kham Speech Dataset Generation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:27.981615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T14:25:40.217180Z digest=sha256:44c1c2c3a3f13e72fa535f930493a0a71e989268cd2c1bf8957aadebbaf0b25c

Observation 5aa886f0-7a3f-46af-a0b0-bdfb212737bb · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:35.214098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:35.214098Z digest=sha256:7d8dcc229c281b9d8a2c942fae707150c541c22b5e290ab0782929b9040f18cc

Observation 2953e945-6422-45e4-b6ab-0d92a97d0c64 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.145264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:ad522c2a1d344fb74e0227db8a3b60b9a750c108a345aa8a53950b3d42c5a292

Observation aa40bc84-52a8-4ae8-914d-109341bc5d0d · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:16.339880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:16.339880Z digest=sha256:fb2ba1f967cd40ed9c0c9f8bcd70e4cb2004da59cbcf063a3e4c52d21fe9c3f5

Observation b4843fa1-2926-4e71-8322-db573e7e4dc3 · inbound

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech cites this paper.

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:00.491000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:24:54.627215Z digest=sha256:0541f91f833a6d417be28c97c67910ae1fd22085ae8fea01257ea4011ea3817c

Observation 5e7439ab-4d74-46fe-b29a-8072b6e7f30e · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.199763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:dfd2102308bc1cf0ba20aabe0b53c72c8148f688994820df8a3b1990024adb52

Observation 46c4a353-49c8-4094-9e8c-fc49dce02c09 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:a54845d29fb6d316646197faf0b5d54c52fa98584b2177e449b920cabcf1e999

Observation 2c603281-0c77-48ca-a921-0538cb1e2ce2 · inbound

Scaling Properties of Continuous Diffusion Spoken Language Models cites this paper.

Scaling Properties of Continuous Diffusion Spoken Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:28.356045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:43:10.132447Z digest=sha256:da110086c7f1dfbd5718c3a1d523b783da2b42b420fa60f7039fe15016ecee6d

Observation 648a05cf-8f2f-41f5-a8bc-d95b2bd7ed38 · inbound

Voice "Cloning" is Style Transfer cites this paper.

Voice "Cloning" is Style Transfer NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:01.038014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:58:45.582395Z digest=sha256:55c04d7799d6c1823ad3af6ab420eaed859819982ca41f1e8918897441257b07

Observation b5a2da5f-67c4-4826-b9c7-27bf2f535ccc · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.563073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:74c41bbfb532e7944d1b0eee3b71798e427aa3e9b0f47876fbcc45cd285a8db0

Observation f93442d7-5bf7-464a-b5f6-aadf112dbf63 · inbound

UniVoice: A Unified Model for Speech and Singing Voice Generation cites this paper.

UniVoice: A Unified Model for Speech and Singing Voice Generation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:08.298843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T23:56:14.198308Z digest=sha256:3a6655f62cefd8cc996bad89537318e200d8a2bd7e3da1e26917ca102e451a99

Observation 0b70f74c-f86a-40bd-aac6-a1530e6b36e5 · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.531681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:c192cbdd15fcb5cf7a1f0af91eab3462b427995b600b4fffe6ff9af478807814

Observation 99b1b84c-9075-47ee-a0fb-c22b443a5885 · inbound

MeshFlow: Mesh Generation with Equivariant Flow Matching cites this paper.

MeshFlow: Mesh Generation with Equivariant Flow Matching NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:49:52.683546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T05:52:47.536542Z digest=sha256:44fe0b8f2e292c4b511f97316cfd6a236b902106c45c5b79978362e3abead0ca

Observation ef844789-06ca-4d27-aa25-2caf8243ba5c · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 190

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:40:03.194270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:321686188a35491f75040f2e7ebfc20c843e094983ff731069591335ffa318e2

Observation 6f2f3709-6fd6-46c7-a37b-afc570b0240e · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 190

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:15:59.065739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:3ab7cab620118fad4dab606d7572308da898b82d49354b8e45621ab61e67e03a

Observation 78488924-78b0-442e-9b36-f069452bdea2 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.268582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:c5c739d780a8711d78a4364bb365c3c8f34636af32ab33937393ffd66c4e0ff1

Observation f91965fd-1d8d-4831-aa05-a7edb4019bbb · inbound

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision cites this paper.

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T11:43:06.058605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:43:06.058605Z digest=sha256:96a51d94e34c7e9e1f474454bac37b5639d4985e99158b028f01ee48b07afbd4

Observation feb60fb7-d20c-49c9-93c2-cc2bafa5bac3 · inbound

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis cites this paper.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.827054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.827054Z digest=sha256:b6ff3b1722859a47569ffe53b2319500088c6b97c7bae67d83d0d8f6abc7596e

Observation 43fc6afc-8c13-4382-8c7d-7ab8c9e49b12 · inbound

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces cites this paper.

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:31.445968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:40:31.445968Z digest=sha256:651a644f5cc1c1b761ab9223b47f0eb38ac7a68f1ee70c75092fc147f44c9fc0

Observation 21be1899-e25c-4269-8f96-a72d1861f903 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:27.768440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:27.768440Z digest=sha256:736e823feb80a40587839abbb277b5d7d28efdb220aefee8083b89bad7871445

Observation 265648e2-a0c1-4985-8af3-e4ba1b4ce885 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.793573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.793573Z digest=sha256:0333efb9676cd4829d71a21e4b0397004e6a5cafe0f9328d026087f2c30d5fd4