Pith. sign in

Paper Citation Record · LEDGER

Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2406.05551.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05551 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:09:20.716820Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:35.627140Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a37e1d39-7169-43b7-b78e-807172d6bac5 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.440978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:032f0b8063b731a4808837080f4d92eb7b2474c38db533deec50d1b0fd38975f

Observation 409fb9d5-cb23-48db-a668-1044ea7a94e1 · inbound

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling cites this paper.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:20.716820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:20.716820Z digest=sha256:630a98773f2c5735106f5df93e5faa66d548d87f7755486350ea3073b2b8af63

Observation cda5b317-bf10-4a5b-9491-bd3be58499d9 · inbound

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion cites this paper.

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:36:53.326258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:36:53.029590Z digest=sha256:e7c6e1664b33316b672a06b6acdadedf8143b6d115665afa4f9978f9f164c4c9

Observation 9a8e6f38-969d-4cdd-bd82-df9d885b2315 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.922139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.922139Z digest=sha256:06f185d6ed8eef24b9f8346a7503fd0f6cb15424847a720d192cbf1b8eb471c8

Observation 9fbfafa0-f539-453a-95f2-15b3161ffe4c · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.701539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.701539Z digest=sha256:62915ea48f319c19dfddd73eac1ba73ded288dd1e4facc7b4836d211ad30f3d3

Observation ad2d839b-f8e5-48f7-82d1-2b7dae1305fa · inbound

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling cites this paper.

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:53.301154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:08:56.588282Z digest=sha256:3a42ba12db62f5a15a13d636a659952184b764d1c1935899ff59616ad048d29f

Observation 49e4c55a-ae76-4246-988d-bf6b3d735819 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.216098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:8067c731998d8bd56895105a68181c5a88ba8825afeb774bfdfd086813be60a5

Observation 5cc8b4ba-b737-4398-9219-3a8d3acb7bc2 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 138

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:bfc6558ad31d1456f0a54c52a203374c97c8093067a449c8cef34080b079c6fe

Observation 742112f1-e411-4b25-99dd-9cc91516fdfb · inbound

V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data cites this paper.

V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:14.507047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T07:46:23.931059Z digest=sha256:3a55206105884b30c991bce5cc5652e7cd3ed1b29fdc0172ee5805373025760a

Observation 054760b5-1092-498f-a46d-067d866ca8d2 · inbound

Taming Audio VAEs via Target-KL Regularization cites this paper.

Taming Audio VAEs via Target-KL Regularization Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.224858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:b9873898ee0b4fa14ecd83366d640c967d7cae688499533e2ddc55fd5d417b05

Observation efba0c9f-6502-4976-9253-af1178854f41 · inbound

dots.tts Technical Report cites this paper.

dots.tts Technical Report Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.039783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:b034c7c7fe147a0485f9e17773a763563bc422333cb8a438b5ddef0b0ac4cc80

Observation 6a57ac6d-c5ad-41e0-949a-4c268b81d625 · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.628751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:9dbed0321ac15727ca1e9e96e0abd92809b970bbe3f69593dcd91fd8d9951524

Observation 888525bb-4f74-4615-9b35-cae730747879 · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:e8f67d2f83e42d044e5b2032edf0cd8ec8cad46116a3213f34844cb3281a6e46

Observation 9da3a33a-10a0-4a4d-93d7-2e50bb0c1f16 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:52.386314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:52.386314Z digest=sha256:b71f6796545edc41226c1c863c99525dbd1b512737b6d810197a5b09aa52874e