Pith. sign in

Paper Citation Record · LEDGER

Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2406.05551.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05551 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:09:20.716820Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:35.627140Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a37e1d39-7169-43b7-b78e-807172d6bac5 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.440978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:390172d7038e4974561ddc990f37846520c6fe1e32ba1a8ff488031b751e3b90

Observation 409fb9d5-cb23-48db-a668-1044ea7a94e1 · inbound

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling cites this paper.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:20.716820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:20.716820Z digest=sha256:630a98773f2c5735106f5df93e5faa66d548d87f7755486350ea3073b2b8af63

Observation cda5b317-bf10-4a5b-9491-bd3be58499d9 · inbound

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion cites this paper.

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:36:53.326258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:36:53.029590Z digest=sha256:7464f6ff7fa65249dd72be9e469d950b494f98efe6e0e8a3ebf3eb6e0634427c

Observation 9a8e6f38-969d-4cdd-bd82-df9d885b2315 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.922139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.922139Z digest=sha256:06f185d6ed8eef24b9f8346a7503fd0f6cb15424847a720d192cbf1b8eb471c8

Observation 9fbfafa0-f539-453a-95f2-15b3161ffe4c · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.701539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.701539Z digest=sha256:62915ea48f319c19dfddd73eac1ba73ded288dd1e4facc7b4836d211ad30f3d3

Observation ad2d839b-f8e5-48f7-82d1-2b7dae1305fa · inbound

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling cites this paper.

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:53.301154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:08:56.588282Z digest=sha256:02b21a1113d6dfb75a955db05944082bb9348bd66a570415807ef7b0bafc815f

Observation 49e4c55a-ae76-4246-988d-bf6b3d735819 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.216098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:314a3c67f491808b0b45bb87a5005c899d7ce7773ee03c5fff0f783a624123aa

Observation 5cc8b4ba-b737-4398-9219-3a8d3acb7bc2 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 138

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:bfc6558ad31d1456f0a54c52a203374c97c8093067a449c8cef34080b079c6fe

Observation 742112f1-e411-4b25-99dd-9cc91516fdfb · inbound

V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data cites this paper.

V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:14.507047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T07:46:23.931059Z digest=sha256:d9d87be58ff5c260d9af312d8a6a30312a6ad5c56e1d954a3f5775dcbbb5d6d5

Observation 054760b5-1092-498f-a46d-067d866ca8d2 · inbound

Taming Audio VAEs via Target-KL Regularization cites this paper.

Taming Audio VAEs via Target-KL Regularization Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.224858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:1cbcfc54bbc654ce2ae313cd935622244f8a3d5bc71935bf7eacab9357665480

Observation efba0c9f-6502-4976-9253-af1178854f41 · inbound

dots.tts Technical Report cites this paper.

dots.tts Technical Report Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.039783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:67e078a4254415d7f952b2866063f74aa4c2e4ee90d94d700a2239961fc90919

Observation 6a57ac6d-c5ad-41e0-949a-4c268b81d625 · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.628751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:c0536a58fc75708f277a42a1558450cee5ad10075e38d9c828b34eb90797843d

Observation 888525bb-4f74-4615-9b35-cae730747879 · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:e8f67d2f83e42d044e5b2032edf0cd8ec8cad46116a3213f34844cb3281a6e46

Observation 9da3a33a-10a0-4a4d-93d7-2e50bb0c1f16 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:52.386314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:52.386314Z digest=sha256:b71f6796545edc41226c1c863c99525dbd1b512737b6d810197a5b09aa52874e