Pith. sign in

Paper Citation Record · LEDGER

VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2401.14321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.14321 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:31.416374Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T06:06:41.477473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e2b1200-720a-4f80-9466-db9ffa8ffb45 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.479977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:63e876df71559fda43d64df3586f058318a2569089114a5578a0b76f94fce5c0

Observation 3498705f-a958-4917-b168-e003ba4c70ef · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.541520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:3d81df7c8dbec90d3ac4fe58afbeea63f51ad11b967e605924fa52d807327c66

Observation f3c4b565-7901-46b2-9a28-bc5704648bbd · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.479961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:7f0a139a170621129328ea6e3671e8b7377ffb6c732427ed4d8c1cf9a3635009

Observation e261da5d-e116-4fe1-b3ab-5a9103f733a5 · inbound

SpeakStream: Streaming Text-to-Speech with Interleaved Data cites this paper.

SpeakStream: Streaming Text-to-Speech with Interleaved Data VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.416374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.416374Z digest=sha256:de58b9ee4df911080f5e25bbe60d06ce619e248e847e405009ccf6c734969aba

Observation a3208703-8afc-40fe-955b-7e983c385967 · inbound

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding cites this paper.

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:01.847943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:01.847943Z digest=sha256:95831bf882c2e096d220d48b563843f3e873f809419387a97d9dfe127a29af6e

Observation ba1f6b20-a679-4c40-94ab-177822983e7f · inbound

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding cites this paper.

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:01.968041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:01.968041Z digest=sha256:60016871a69fbb130b939e678e951f3d46d6e21ab6431e8b27455d9263f6ea6e