Pith. sign in

Paper Citation Record · LEDGER

NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2304.09116.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.09116 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:54:18.726951Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:03.192889Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f47485fe-6dab-43e4-b8fe-365529144a4b · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.517814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:a3cc7dd46e92f8ca5f980bf0eb8a42231123ae50a3c29646e82c0b8219e97857

Observation 0c322811-7749-48e4-bc4b-4936f82b6af3 · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.914341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:7a5a6bb616dd20efcf09af9a968bff94546b3d8d89d7d2c431a3489070ae0539

Observation 6315151a-bfd4-449a-b68c-b0427656b917 · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.726951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.726951Z digest=sha256:f9ce180826aa6766926a62b3a5a78136adf0d324bb21beb8028b98472e9a2bc7

Observation d5407115-8935-4156-98a6-2901d3393e33 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.312162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.312162Z digest=sha256:603993f08253db9969cb70f7be655d361386f099ad9c60964c56fc3a54042455

Observation 74c40840-3b67-4339-a24d-74fc4cd50c58 · inbound

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model cites this paper.

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:56.445581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:56.445581Z digest=sha256:5d4030ea8853eeabe086b3fd5f9d47dd05c765acdc2154406340487ab9a6e468

Observation 84115302-5a8a-4079-8810-af0dbd16bb2d · inbound

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations cites this paper.

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.184518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.184518Z digest=sha256:a09ce4464a3589544e04694611491dfcbf71a65153b6f84ac439fa54f1c2b9a1

Observation 4e224cf6-1374-4e0e-851a-8f60fdf62d64 · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.041243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:53499780008a47dc9ce920ed37fb70cb7857d83e927d33aaf186858e8d585aa3

Observation ec02f16b-1a6e-4d7c-89f3-9d90ad4f2018 · inbound

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech cites this paper.

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:00:07.579047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:00:07.579047Z digest=sha256:e1ff8002bc695e923747d51a099fd1e8885677c0ea20f6b67ccd3eab9930bce6

Observation a96b3cfc-3d52-47d9-b621-0ef2835b38f1 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.184511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.184511Z digest=sha256:25cacd3e6f51935ea0a56664efdf9418ffa8bbc25aea3645f39f49e85eff1c02

Observation 3b20f27b-795a-4b6b-b9ef-dba358d4ba82 · inbound

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations cites this paper.

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:00.825953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:00.825953Z digest=sha256:9a8f234a7d077b4840b37bd314910a57f3279cd470fe0e757157c82f357fab31

Observation faba2588-9e08-487e-b264-81415b3455e9 · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.121148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.121148Z digest=sha256:1426a2893eae456d446c196660610bcafc63076f37ab2e36c295f2f187dd2853

Observation 3dd2c9a9-4ed4-4bde-be05-e8b4c8dbfb64 · inbound

Marco-Voice Technical Report cites this paper.

Marco-Voice Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.401559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.401559Z digest=sha256:6fd7052cdc93378ec0d79b5f8ec2bad4ac9cc4d152520621e76e5438dfe204de

Observation e8139e27-ff8e-4e4d-86e3-ba24bdd26ffb · inbound

Inference-time Scaling for Diffusion-based Audio Super-resolution cites this paper.

Inference-time Scaling for Diffusion-based Audio Super-resolution NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:28.156213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:28.156213Z digest=sha256:e69cdc82aaeb1322c578cc36157668c189ef776592297b9ef719f6c1afd9ca64

Observation 6562b558-0a3e-403b-9f40-55481a3dceab · inbound

Audio-Guided Visual Editing with Complex Multi-Modal Prompts cites this paper.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.804516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.804516Z digest=sha256:e9ed09f979954ff9a0d6e743532a58e36968f564417e095a02a4b342479f9f4e

Observation 74f964eb-a931-415f-8f63-f15b96a40155 · inbound

FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks cites this paper.

FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:35:01.874077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:35:01.874077Z digest=sha256:45da6947aa885961638f5807a14f945db25b91d359ce3d281f8b0496ad2d22f6

Observation 2ecbb115-ba05-42da-bc95-60a3ee269242 · inbound

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for \"U-Tsang, Amdo and Kham Speech Dataset Generation cites this paper.

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for \"U-Tsang, Amdo and Kham Speech Dataset Generation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:27.981615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:25:40.217180Z digest=sha256:33e3257c54aeaac4c8cdbd5d581fb4893ba5e55a81783ed385cd0f5238d8933d

Observation 5aa886f0-7a3f-46af-a0b0-bdfb212737bb · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:35.214098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:35.214098Z digest=sha256:7d8dcc229c281b9d8a2c942fae707150c541c22b5e290ab0782929b9040f18cc

Observation 2953e945-6422-45e4-b6ab-0d92a97d0c64 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.145264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:e94f316b8631d5f91036727f100e930882e7b9aba0a16ed55a85de335d24e770

Observation aa40bc84-52a8-4ae8-914d-109341bc5d0d · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:16.339880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:16.339880Z digest=sha256:fb2ba1f967cd40ed9c0c9f8bcd70e4cb2004da59cbcf063a3e4c52d21fe9c3f5

Observation b4843fa1-2926-4e71-8322-db573e7e4dc3 · inbound

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech cites this paper.

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:00.491000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:24:54.627215Z digest=sha256:f2f84f62c6b9d2bf1411d34afee1106c7b44a544ac4a0bb5c5bd6762cba95f5a

Observation 5e7439ab-4d74-46fe-b29a-8072b6e7f30e · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.199763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:62f04e2220e64f495034ef8db3b16214c6a09863302d0efb96df84a2e390b159

Observation 46c4a353-49c8-4094-9e8c-fc49dce02c09 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:a54845d29fb6d316646197faf0b5d54c52fa98584b2177e449b920cabcf1e999

Observation 2c603281-0c77-48ca-a921-0538cb1e2ce2 · inbound

Scaling Properties of Continuous Diffusion Spoken Language Models cites this paper.

Scaling Properties of Continuous Diffusion Spoken Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:28.356045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:43:10.132447Z digest=sha256:869dcd9cf9a58095a271fe478265d20ebbb5240d0332af0a33c2ec6e0474e94d

Observation 648a05cf-8f2f-41f5-a8bc-d95b2bd7ed38 · inbound

Voice "Cloning" is Style Transfer cites this paper.

Voice "Cloning" is Style Transfer NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:01.038014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:58:45.582395Z digest=sha256:672a13e15527a1bae42e49ae6409d0c6b598fc528c4848909d0bb6ef7935286b

Observation b5a2da5f-67c4-4826-b9c7-27bf2f535ccc · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.563073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:fb0d425d531253991e5ae815f783d2ac26211be029db8ee186126b5ce162ba08

Observation f93442d7-5bf7-464a-b5f6-aadf112dbf63 · inbound

UniVoice: A Unified Model for Speech and Singing Voice Generation cites this paper.

UniVoice: A Unified Model for Speech and Singing Voice Generation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:08.298843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T23:56:14.198308Z digest=sha256:96d0687de88abeb73991218614478b03ea65430adf7770fa306033448c101bf3

Observation 0b70f74c-f86a-40bd-aac6-a1530e6b36e5 · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.531681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:6805782fadea70514002bc86967007601490e272789566fff4a24365babadc96

Observation 99b1b84c-9075-47ee-a0fb-c22b443a5885 · inbound

MeshFlow: Mesh Generation with Equivariant Flow Matching cites this paper.

MeshFlow: Mesh Generation with Equivariant Flow Matching NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:49:52.683546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T05:52:47.536542Z digest=sha256:318f9d594f9f1d4ff941cec531a565077bd251eb5eb01750e208bbe85136a1a9

Observation ef844789-06ca-4d27-aa25-2caf8243ba5c · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 190

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:40:03.194270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:b5f87c12fc4ee3a9a3b96ed32020aa24dcbf7ceb48624b092d82df0dcbc3ef44

Observation 6f2f3709-6fd6-46c7-a37b-afc570b0240e · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 190

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:15:59.065739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:23de52507492a7da9dbc4c8413b88e1d89e8dfcf6fe24ba80b5a0bd40d418fa0

Observation 78488924-78b0-442e-9b36-f069452bdea2 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.268582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:b75eda01da69e5e769432698202d87bd34f20c0e68e40ef0a1d295157a972ff0

Observation f91965fd-1d8d-4831-aa05-a7edb4019bbb · inbound

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision cites this paper.

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T11:43:06.058605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:43:06.058605Z digest=sha256:96a51d94e34c7e9e1f474454bac37b5639d4985e99158b028f01ee48b07afbd4

Observation feb60fb7-d20c-49c9-93c2-cc2bafa5bac3 · inbound

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis cites this paper.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.827054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.827054Z digest=sha256:b6ff3b1722859a47569ffe53b2319500088c6b97c7bae67d83d0d8f6abc7596e

Observation 43fc6afc-8c13-4382-8c7d-7ab8c9e49b12 · inbound

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces cites this paper.

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:31.445968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:40:31.445968Z digest=sha256:696a8eb3f55750eda193764a89fe9060d1684d4f7db1a9d94cacd2f651b664a5

Observation 21be1899-e25c-4269-8f96-a72d1861f903 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:27.768440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:27.768440Z digest=sha256:12bb7f2c442cf8c1b11371ba8cdd48082f0d873d428c348d50b780cb1ba84c14

Observation 265648e2-a0c1-4985-8af3-e4ba1b4ce885 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.793573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.793573Z digest=sha256:a701f7ca75f162301a926dd9e2d7088836b9d412946b288bcf73b21448b37173