Pith. sign in

Paper Citation Record · LEDGER

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.19859.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19859 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:33:14.254397Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 762d9abe-c9e1-45d6-b5c6-8b016af7e3d8 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:11.240813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:11.240813Z digest=sha256:528ca1aa2f36f84c5cb74ab14c8ebdab0da170a47a3f19b1d6fa289c332dbdb4

Observation e95087c6-65a3-4578-b689-5d7c7a80761a · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:11.360312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:11.360312Z digest=sha256:6192d3321aac9b9189986fb13af46b778549c90859895b161980544ac34899b5

Observation ba334b15-2f8f-4727-adbf-68c4606b6ace · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:11.513298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:11.513298Z digest=sha256:b46bc7d4bc09b85fbdaf1fc38cec1404cad8583538f2c4f2fffeffcaaaff44e9

Observation 032aedcf-b139-4bfc-bec3-cb8564c8078b · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis SoundStorm: Efficient Parallel Audio Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:11.673952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:11.673952Z digest=sha256:9caf689bafaecb2e678c22aaf50d9cfcbfaa66a5d06291b1850355612b97daa4

Observation 60ab2a50-c696-4b5d-905b-9e54efd22029 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:11.847608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:11.847608Z digest=sha256:8aa806418e7d9533dc64ede26546bce529ed7b60793e38669d17bc3df69afe73

Observation eb22d84c-56d5-40f1-a14d-1a95d0c6e2be · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.021514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.021514Z digest=sha256:c76b03c3173874f8b3e98e5efa0ba6b18c80f9ea0474131b72b2899c2123a50d

Observation 559dbbb9-104c-4d86-b497-679ab7fe43a6 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.225012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.225012Z digest=sha256:c54f52b7a7be7367ac49c4f2231f7310051bdf22c6c4cb7d05547620096c05e5

Observation c255864a-fb98-409a-ad4e-539c5adf2442 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.331909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.331909Z digest=sha256:151dfb5009975f1dc140f4cb44fe2eddbae1cd3ad93c1edf37aeb5e58c25cb70

Observation 4ad4210d-96e1-4a88-bc55-07a6f53cb46b · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis Fastspeech: Fast, robust and controllable text to speech,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.392017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.392017Z digest=sha256:bd6de275601e957f3a1dbbe5947f0380cd2c0e567c649b1052a6c79be9553ddf

Observation a0d3e1bf-75f0-4f90-b74b-e46458c72f28 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.513735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.513735Z digest=sha256:40481bc6408a8e82bd187798bf37814da53259121bbdb85fc286e93027ece8ee

Observation 927b2462-ee40-49c2-987a-e358666ffbf7 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.652904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.652904Z digest=sha256:38d14b3224fa9a1e6f030a8a6348097b46873853051a906417bd6a3c50d7dd96

Observation feb60fb7-d20c-49c9-93c2-cc2bafa5bac3 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.827054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.827054Z digest=sha256:b6ff3b1722859a47569ffe53b2319500088c6b97c7bae67d83d0d8f6abc7596e

Observation b0a25f98-c10c-4f9c-8a48-09af0406945a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis LLaMA: Open and Efficient Foundation Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.964112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.964112Z digest=sha256:25275cf3a231b1222a6a045b669ef3896cb20442a17149e9c7f04f0ca2e7b420

Observation b382ecb4-1ab2-4503-b3d6-547b06d9e204 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:13.102191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:13.102191Z digest=sha256:f808aa7412061817a470f1c9d5b4fea0ae45794714e5b72a3e525ca548a4b158

Observation 2229f2bd-e850-4bf5-b9e1-0b068a98df86 · outbound

This paper cites Neural discrete representation learning,.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis Neural discrete representation learning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:13.210604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:13.210604Z digest=sha256:6f812705cb8a583fb450c66e18ff7403c8f0ac55e17c9246cf23a73961f4d748

Observation 4fc1f77a-fc24-40d7-9683-4155eee4a389 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis Soundstream: An end-to-end neural audio codec,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:13.347635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:13.347635Z digest=sha256:4e42be7dadc722d83722d20c0064d309ff305b8262f3bc6d7b2ee3568eec84bd

Observation 9dabd932-87a8-4315-9da8-06b623d3b42a · outbound

This paper cites High Fidelity Neural Audio Compression.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis High Fidelity Neural Audio Compression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:13.485060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:13.485060Z digest=sha256:b37b380e61938f41afb87aa9e3facc917472de7cf1998e74a32aeb0596d9500e

Observation 4c54927a-3390-40db-af24-1c7d8fa7252b · outbound

This paper cites Hubert: Self-supervised speech repre- sentation learning by masked prediction of hidden units,.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis Hubert: Self-supervised speech repre- sentation learning by masked prediction of hidden units,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:13.624740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:13.624740Z digest=sha256:e7bdb7f36348cccefaca64b8af2e5e2de050b2521d4972d12c854a39e63db618

Observation ddee6882-f116-4dbf-ad7e-24097741fbee · outbound

This paper cites Wavlm: Large-scale self- supervised pretraining for full stack speech processing,.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis Wavlm: Large-scale self- supervised pretraining for full stack speech processing,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:13.767556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:13.767556Z digest=sha256:599b9b68ddc6a692abd46d9a245aaf847f492f88fbb462773211c2e603cadd76

Observation 98219372-94ca-4c97-b7cd-f74e065db327 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:13.914349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:13.914349Z digest=sha256:42053cc1edad72b668ce45d0b67e4f7673adfbe5e350ab2a668b7534100b8042

Observation 505248d2-3a49-4223-b368-65e25321864d · outbound

This paper cites 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:14.026639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:14.026639Z digest=sha256:76d82c22e0315c4192536b6b8d0c639baadd21b630fc10ffa3ac8dd7d2399c16

Observation dcfe1bb0-7916-43a6-ad3b-06d4a79bf3d3 · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:14.162197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:14.162197Z digest=sha256:5ccc5b64cd39d57c1bdfbe8021cb5c1e78e05fb0ca88079e17b3335132565187

Observation 1c9468fe-a430-472d-8256-dce62ab51a78 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:14.254397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:14.254397Z digest=sha256:af28804bcf143f47114395a048c41c02f9e14a57c90bb1a5f18229785e2e7ac9

Pith citing papers

No inbound Pith citation observations are available.