Pith. sign in

Paper Citation Record · LEDGER

InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2301.13662.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.13662 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:41:36.305336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T08:30:57.168348Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cf387bbf-4908-448c-aa4d-445cbbf9fb8f · inbound

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles cites this paper.

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:36.305336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:41:36.305336Z digest=sha256:0dd9e8e181c5894703f8d6cd077d6b9be8de7ae67d9ab7d88354f7cac28181a4

Observation 247c7768-c071-4474-85a1-1ce7ff51ca9c · inbound

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control cites this paper.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:35.018333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:35.018333Z digest=sha256:ce5c353fd9a862a69226cbdcc1f9a03906bd30c65c2e19762570b8c3fe3b592b

Observation fc53d6c9-e5c1-4e46-8f80-0122508deb35 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:54.567317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:54.567317Z digest=sha256:26aa62257b6509236ad49bad3b5ca3ba02754cc1b458a1d4952ce869efe6c201

Observation fde8ad37-6ea3-473a-9080-36dd28b639c0 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.287526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:6d8e7af8341075052256b955c0a5140c402a46e423f016887b123b9c40db0c35

Observation 13056a05-d758-4bbc-982f-776dd6317ce2 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.169910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:1324da5cfa570e76ed1d23af225623493fce0857baa4d017539c30ae6d429356

Observation d7dca130-2869-4c72-9d71-5b43ab92708f · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:a66e8aabcc6071c036f7ccd65429129e9c98ced8468a5acb2e79ee77a2cdc052