Pith. sign in

Paper Citation Record · LEDGER

DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2406.11427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11427 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:21:55.843335Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:58:08.367612Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation de99c861-4fa9-4af9-b871-0aa3d7d9c1e7 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.413952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:ff0ea326f8f7761602155f46983b8aa9712ac5d9d5e04d9deea8ac152cb78fa9

Observation 5fd9797a-1517-42f5-a34d-5fe417ab5401 · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.559170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:ba59d1fce8874c7e5c3e63cac9d279dcfe27aa61d572f41749f301ce7919b898

Observation ddbce4e2-199e-428a-8189-e8123457ed81 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.522200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:82bf59e5caf8a3fedd59dfb9e20e7215c10d3ecb8307d5304f1f3de666117cdc

Observation 76fed29e-242d-4dfa-b6fc-41b2f8fc9140 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.843335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.843335Z digest=sha256:eb3136ead2f4a9e76004bd982c23558f39f4c6b0b593b3c01d3587cb470b5192

Observation 23782b60-0a99-4477-88f2-c59be7d64832 · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.024503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:4cc6766c22b920c720b137628de70028f2c92343afc40f54ab8d84437d3dc128

Observation b3eef67e-c8b2-4d44-a2e5-8e5a1da552c8 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.419026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.419026Z digest=sha256:0496e19a62e758e63966a6bfcba8546aa519167da4b3336b4ac20bd07239fc4f

Observation 2c943391-9c5c-40ac-a1ca-27fd64d2e275 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.542297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.542297Z digest=sha256:85a7900e1d815e46d49988c49efbca8e833eca40dbb62f78d5d11537c3a74236

Observation 89c0ecd3-ed9d-45c2-9da5-81d777db0f0b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.183773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:927f5e6a4360d64e03121b01e4a1ecc025a2697d436aa3aed68c8ec8242872f1

Observation 0e3aaad8-48ce-4b1e-a5c5-1a1eb2beeeec · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:4e67e931876669d5825a9da472afa6f282514fad0afae59e7c52941cfc313c8b

Observation 51c414b0-ffcf-41d2-806b-9a5443fc2255 · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.447788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:3b0f6b8f4fdaa81d3e562898695f4b0b631238e741ef9ebadf8d6c5d22286529

Observation b5d83bd4-65ad-4c8b-ac08-24fed19301e4 · inbound

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations cites this paper.

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.369088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:38:51.371500Z digest=sha256:e3d571668e8b705014488f09c3dd6ef2122651379dcb58271ac661212199e673

Observation 4af0ea1f-98f3-4fe0-9baf-4e6e14e5981a · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 181

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.165963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:eae84b27ec230e0ce9a45534492bcf733139d3f16df7a5e76d8389f51c6c25be

Observation 42a7f1bd-868f-46a0-872d-4e05a7e68c89 · inbound

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech cites this paper.

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T21:24:36.925360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:24:36.925360Z digest=sha256:2bf0130ab12f05ba5427f802cb4736745b96301c672f14c6c9837cade775ce32