Pith. sign in

Paper Citation Record · LEDGER

DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2406.11427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11427 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:21:55.843335Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:58:08.367612Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation de99c861-4fa9-4af9-b871-0aa3d7d9c1e7 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.413952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:69ad86b83ff2f7b14c25602e912c420f2f4d315135df5883bf70ee04c43a7e5c

Observation 5fd9797a-1517-42f5-a34d-5fe417ab5401 · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.559170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:94af37eb0adac2c498c621d01fb7ac889d32303ed24c97e106c004db0092c663

Observation ddbce4e2-199e-428a-8189-e8123457ed81 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.522200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:72c7cef7e2198da93a9b80181a6af32e30912dfd20c10ec8a6100ca0db5681fd

Observation 76fed29e-242d-4dfa-b6fc-41b2f8fc9140 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.843335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.843335Z digest=sha256:eb3136ead2f4a9e76004bd982c23558f39f4c6b0b593b3c01d3587cb470b5192

Observation 23782b60-0a99-4477-88f2-c59be7d64832 · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.024503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:f3ce1598749c59f1021d405cecd88cf7fb6aa015bf1c4c799dd0e7dd3522bc2f

Observation b3eef67e-c8b2-4d44-a2e5-8e5a1da552c8 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.419026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.419026Z digest=sha256:0496e19a62e758e63966a6bfcba8546aa519167da4b3336b4ac20bd07239fc4f

Observation 2c943391-9c5c-40ac-a1ca-27fd64d2e275 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.542297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.542297Z digest=sha256:85a7900e1d815e46d49988c49efbca8e833eca40dbb62f78d5d11537c3a74236

Observation 89c0ecd3-ed9d-45c2-9da5-81d777db0f0b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.183773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:591668f918bb49761ab3c80d8caf66563ef6f9c58be4e6c3acc8e993c8e825e4

Observation 0e3aaad8-48ce-4b1e-a5c5-1a1eb2beeeec · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:4e67e931876669d5825a9da472afa6f282514fad0afae59e7c52941cfc313c8b

Observation 51c414b0-ffcf-41d2-806b-9a5443fc2255 · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.447788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:2a1331df6edad3923d7b81aa3f85047b274db0f7bac10c00087be86e59d8cc3a

Observation b5d83bd4-65ad-4c8b-ac08-24fed19301e4 · inbound

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations cites this paper.

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.369088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T08:38:51.371500Z digest=sha256:c5f612aadb19e85ecdba7565a0708dc5c2c3a232eb9b5f9cb2242003f1ec762b

Observation 4af0ea1f-98f3-4fe0-9baf-4e6e14e5981a · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 181

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.165963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:d793ab6e5e49e35336941e1e519259e65ac989a1153eca2f0236a83663b5fc36

Observation 42a7f1bd-868f-46a0-872d-4e05a7e68c89 · inbound

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech cites this paper.

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T21:24:36.925360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:24:36.925360Z digest=sha256:2bf0130ab12f05ba5427f802cb4736745b96301c672f14c6c9837cade775ce32