Pith. sign in

Paper Citation Record · LEDGER

SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2406.02328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.02328 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:16:17.331703Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:08:37.504398Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e0a27129-931a-4f22-92d5-d9946d014e29 · inbound

Speech Separation using Neural Audio Codecs with Embedding Loss cites this paper.

Speech Separation using Neural Audio Codecs with Embedding Loss SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.432669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.432669Z digest=sha256:3683be16440b086cc92fa1063f1cf0950c377e3d82af3405a7747b7ad3fa1d60

Observation 07ee884e-10a3-469e-a0a5-68388c88c031 · inbound

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder cites this paper.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.331703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.331703Z digest=sha256:14a1a3e817cc84e9fd2575415b067d374a0f53e0a1dedab6373fed38b8f9c93e

Observation cb511548-a040-4418-975d-3dd8512cc73a · inbound

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model cites this paper.

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:56.625844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:56.625844Z digest=sha256:148b2d5ead5bbd6f8bea31bd6d66b06622214b1b0bf31bd1555d3fedad81a3e4

Observation 3532d1ea-3948-4fe4-a0b4-111f73970903 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.444748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.444748Z digest=sha256:26b87fa3309fc1c3044bb84968453cac973e766b2704e9577277254690576956

Observation 9cb08e55-b39c-4f59-ab51-4d28892893d5 · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.476364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.476364Z digest=sha256:c04c3e8fd3e476a630046513d2615c2e457db222ec289296df7b5932bc8d57ad

Observation b78d2d07-028f-4631-9aff-1f837a23dda8 · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:51.164939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:51.164939Z digest=sha256:b4feb4c570aa930e90797837a89d9a04e2057605302fd107cc26c76dbd2ef76c

Observation ea9e05b2-2663-4396-bd66-7eaa480045e2 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.175358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:2d91751c39458a637e5a7a7f3730cf6f07b23fe192ac3b27b99b9df2edf30fe9

Observation 5645eff3-53d7-43e8-8aae-52110c772c1a · inbound

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment cites this paper.

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:08:37.506013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T06:05:26.735340Z digest=sha256:3d48a40b37116de9b3ebea267354ca4a3b3f6d7d3011757a0255c1f1039111c8

Observation 6ca38534-66b4-4f39-af00-cc40940c04a9 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 180

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.163389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:03f3e7ea76b4b5b66d5a74936c7e77511912d0f248dd6907848e6808be1eb7c9