Pith. sign in

Paper Citation Record · LEDGER

VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2401.14321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.14321 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:31.416374Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T06:06:41.477473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e2b1200-720a-4f80-9466-db9ffa8ffb45 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.479977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:ac8d923457e98760a686b011ff366c8d88eb83135cdd184aaaf02f3a149d6000

Observation 3498705f-a958-4917-b168-e003ba4c70ef · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.541520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:7757ba185ed2ffb226f578e16bfd1e6e23beb38756552b90430f17a5e903e752

Observation f3c4b565-7901-46b2-9a28-bc5704648bbd · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.479961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:ed508ca64e6bc7615268a5e9e5af61939e6bb237bcf780ab1b61624436a29d2c

Observation e261da5d-e116-4fe1-b3ab-5a9103f733a5 · inbound

SpeakStream: Streaming Text-to-Speech with Interleaved Data cites this paper.

SpeakStream: Streaming Text-to-Speech with Interleaved Data VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.416374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.416374Z digest=sha256:65f761c2386c32308f1e0ce90eb2cd81884d15e7dbe1fd83adfd754c1dcf5aaa

Observation a3208703-8afc-40fe-955b-7e983c385967 · inbound

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding cites this paper.

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:01.847943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:01.847943Z digest=sha256:95831bf882c2e096d220d48b563843f3e873f809419387a97d9dfe127a29af6e

Observation ba1f6b20-a679-4c40-94ab-177822983e7f · inbound

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding cites this paper.

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:01.968041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:01.968041Z digest=sha256:60016871a69fbb130b939e678e951f3d46d6e21ab6431e8b27455d9263f6ea6e