Pith. sign in

Paper Citation Record · LEDGER

A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2303.13336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.13336 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:17:12.793801Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.950402Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 785047e2-ff42-4ee9-a1b2-6dceb08f0623 · inbound

Boost-and-Skip: A Simple Guidance-Free Diffusion for Minority Generation cites this paper.

Boost-and-Skip: A Simple Guidance-Free Diffusion for Minority Generation A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:17:12.793801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:17:12.793801Z digest=sha256:fe39cf0c9410cefcf4813f16400458e258e0c740f3c4141e29f43eca43f47430

Observation c57031df-7dd5-4e0b-8e73-f699b9d5cf01 · inbound

SyntheticPop: Attacking Speaker Verification Systems With Synthetic VoicePops cites this paper.

SyntheticPop: Attacking Speaker Verification Systems With Synthetic VoicePops A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T21:10:02.504068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:10:02.504068Z digest=sha256:ed92d50911ff1da206bdf9a8d34693e12e54ba5721c488eda1ba1e74677d5dd0

Observation decb9ae9-765b-439a-8be8-2b9f73f8a154 · inbound

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences cites this paper.

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:58.845200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:58.845200Z digest=sha256:0b8c905d7643324d3356af4b7f77bdc9ba52e4dba11b4bc907bfeb668b513198

Observation ba40f52d-e387-48a7-9d1b-dcec9d791bdd · inbound

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching cites this paper.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.307090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.307090Z digest=sha256:39ed86cdb57479cba38c5f6ddcc969fad8e35cfd7a3e2ab95da5495f942da446

Observation ed6b3a6c-29f5-478f-83a8-d6b4cdc25168 · inbound

The Effect of Stochasticity in Score-Based Diffusion Sampling: a KL Divergence Analysis cites this paper.

The Effect of Stochasticity in Score-Based Diffusion Sampling: a KL Divergence Analysis A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:40.247061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:40.247061Z digest=sha256:68e1093f26320de414f373f2a26e8419c1a6078a50adbd132b05c08d3a163954

Observation bdb366e9-50f0-4711-99f6-962014b3a2e2 · inbound

MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation cites this paper.

MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:18.263678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:57:18.263678Z digest=sha256:5df4e5ed2d76a36902c680db852f631719d146927cba0bcc229c468889e4ce10

Observation 605f7043-c24c-42d9-82e5-de471472d079 · inbound

Marco-Voice Technical Report cites this paper.

Marco-Voice Technical Report A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.655722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.655722Z digest=sha256:e1aeb91f46c98db0537d8e29a53779a78af93b85c25769a9c78e4155c7379915

Observation c4712f20-4b91-46b7-81db-1de49dcf08e2 · inbound

Permutation-Invariant Spectral Learning via Dyson Diffusion cites this paper.

Permutation-Invariant Spectral Learning via Dyson Diffusion A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:26.632802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:26.632802Z digest=sha256:89b488c9f323592c2f8878fa923f2aa141c119a09de2cc817800a2356eb4dca7

Observation f121603c-f82c-4356-bc00-ee365d68637c · inbound

Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs cites this paper.

Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:02.720588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:23:50.751356Z digest=sha256:0a38b8b9e1558a841a71867c18e6e065de4481ce59f2b1828f06d89293250e0e

Observation e561caa1-c2cb-4966-adf2-af054f964a90 · inbound

Grokking of Diffusion Models: Case Study on Modular Addition cites this paper.

Grokking of Diffusion Models: Case Study on Modular Addition A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.869487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:28:08.221886Z digest=sha256:fca8b37e9d55233c985ea6c829224ed4a2a513a62ee37d6f79757cd3d887e4e0

Observation b9d300d0-be12-4912-819b-708ed15ab9dc · inbound

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges cites this paper.

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:01.259983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T15:21:02.100545Z digest=sha256:c331bbed64d0011be2e1435bfb438faed22d893e0678cd4560de4e7a31d4d988

Observation 94a39d1e-6cab-447d-8a3e-e1ea1d83cc1b · inbound

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges cites this paper.

Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:12:28.683468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T07:08:43.088643Z digest=sha256:63a176d218c286d901a4991eafe7ae26e6877aaca657896a63676dbd3e02d94e

Observation ea88926b-e53f-435c-8d21-638ff2aa0b9c · inbound

Inverse Design for Conditional Distribution Matching cites this paper.

Inverse Design for Conditional Distribution Matching A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:56:29.547249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:46:08.232616Z digest=sha256:36324456e625f02f744958fd3d1769775ca4480744fdffc05978c96ce593830e

Observation 3af78a66-4dca-439b-a3ff-dd229d8d4773 · inbound

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining cites this paper.

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:52:35.108279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T18:53:10.978631Z digest=sha256:2ea81208d0629a0b12f0cd23db78b1475888c8bcf7e9d7ab08c0f9b4b54b88d9

Observation 62e40a7a-ec40-46d4-98a8-9300c21c4415 · inbound

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS cites this paper.

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:20:07.953940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T20:11:18.045777Z digest=sha256:9e7c018beb10d2f2469a63698112cdf4de5f9cbddc28df2edf3cdb78538e8341