Pith. sign in

Paper Citation Record · LEDGER

Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2306.03509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.03509 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:15:52.145073Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:19:44.612291Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 53c5908d-2c1c-4b39-8fe2-588f9940ae35 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.342857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:87ce27fffb07d4a0a212d3dffdf5d73b6a2cd176f9d56836559c81e08674935f

Observation eaa02832-2131-4664-832e-6c0440586611 · inbound

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model cites this paper.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.145073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.145073Z digest=sha256:8146131a7c90180fcce546b727ed9e5a53f548e26fa220b30301f7e8c9247ad1

Observation 6e9e5b21-1953-45c9-9c85-9baf1e1642fb · inbound

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement cites this paper.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.136859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.136859Z digest=sha256:65a417afdf6fd71ecf67d9e488eb1e16b44e14f6100074d5d3216058a49c3560

Observation 978f7a07-4e93-4da1-9396-af63e0416aa4 · inbound

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt cites this paper.

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:38.401645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:38.401645Z digest=sha256:1318adf50b1df15356075510bf8ee529de3eb89afadcb739c7f9daa634c58660

Observation 3edb39cc-aa46-4f93-b2bc-0f6f5ec4b43f · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:36.096538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:36.096538Z digest=sha256:3cbdaf709c97f0791cfc55eac31e761b746826d36ec6a8d99aa4b5aa5b4792da

Observation b78bc22a-fcd7-4a2c-84c2-66b3954ae4d2 · inbound

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning cites this paper.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.469169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.469169Z digest=sha256:bb04e457163f180b71463001767547f4fc83ccd21da70611857448143d2fac73

Observation 2146cef1-0d9c-445e-8e79-3a3c1cf73e2a · inbound

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs cites this paper.

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:07:03.766330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:07:03.766330Z digest=sha256:72120f8f87d4b5bf0fabff6c19296d893e5819259ec011f468ef679e3b2beea0

Observation 78e4abf9-eb85-44cf-9e23-1d155d9b1699 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.186280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.186280Z digest=sha256:c736738f7a718efb4c35f5891c7d815edbd06383306eb095a8b7ff0b6de604ef

Observation 8c88c9e2-06e9-4bc7-9437-475bd62904d4 · inbound

Two-Dimensional Quantization for Geometry-Aware Audio Coding cites this paper.

Two-Dimensional Quantization for Geometry-Aware Audio Coding Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:20:29.207437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T18:16:51.486807Z digest=sha256:15c1037a2d7f0bf514dfc0e2794cc64b54a50cc2a158e4974d10fb18bc5ef8f0

Observation 18b57f1e-c3ff-45f2-985d-6beb4f49fef2 · inbound

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion cites this paper.

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:44.613832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T11:49:04.308326Z digest=sha256:9df276952173270b7e0363cb460f7499fb0890a8eae1bf832e3ffe4914bf7f12