Pith. sign in

Paper Citation Record · LEDGER

Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2306.03509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.03509 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:13:57.482560Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:19:44.612291Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 53c5908d-2c1c-4b39-8fe2-588f9940ae35 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.342857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:3ca8da43dbaf26a4a85b9804a93d4336e8d85e8f5c5a67cbfc16b572fbb18788

Observation f123214f-8997-456f-9770-df9e795a12df · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.482560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.482560Z digest=sha256:3b38f025a21a39b8c472387006560c1831193aa0284a7bc363b36002dc32c332

Observation e3b3b352-6496-40a1-9fc5-5e9db1de9ab3 · inbound

Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook cites this paper.

Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 162

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:48.323828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:08:48.323828Z digest=sha256:5ba20c3a14cfeb1351729630b47c85c1343c2cdaef21da757d220c5d14b99116

Observation 8dd98ab2-a986-426a-b606-ca356a5e976b · inbound

FreeCodec: A disentangled neural speech codec with fewer tokens cites this paper.

FreeCodec: A disentangled neural speech codec with fewer tokens Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:47:55.423144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:47:55.423144Z digest=sha256:85ee978ee01eee8c3a9ccbd18991b3f6f2a749b01c384286dcf83db3bdf99548

Observation fcd5234d-89a1-4dd1-a08e-24d446da1fa0 · inbound

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor cites this paper.

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:52:03.796299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:52:03.796299Z digest=sha256:7d7b0cac0a48c70e09b96adaa64c5dfa07d1dd0f9874c6e60d14377f7dddec19

Observation c43b41df-fc1b-4f15-99f1-05aa0107e086 · inbound

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting cites this paper.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:33:55.993734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:33:55.993734Z digest=sha256:22c6831716310617249aaf84ca8ecd2a1875d10a96fa7fce2c6b46de833eec82

Observation 2e561cb2-75f7-4f63-8ced-106c91d58529 · inbound

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles cites this paper.

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:36.186653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:41:36.186653Z digest=sha256:d9d6c428ee4ba1a6f74ed58eb8b479980be07ac58786df16b217bddc24e063ce

Observation eaa02832-2131-4664-832e-6c0440586611 · inbound

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model cites this paper.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.145073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.145073Z digest=sha256:8d0dd4db98f12c04fc3d286b7d61dc69499f0906de78fd2914432edda7ed057e

Observation 6e9e5b21-1953-45c9-9c85-9baf1e1642fb · inbound

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement cites this paper.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.136859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.136859Z digest=sha256:4cb8574a244dc583a17dec3a87d6cba6f6678c1f9607fcfe3fa8378d8797c70f

Observation 978f7a07-4e93-4da1-9396-af63e0416aa4 · inbound

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt cites this paper.

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:38.401645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:38.401645Z digest=sha256:9bd48d50bb0244d33f550f5f1b31ebe286e55d91606f86ca955f9ffd62b29318

Observation 3edb39cc-aa46-4f93-b2bc-0f6f5ec4b43f · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:36.096538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:36.096538Z digest=sha256:f9cf7bc76e96cfddfc7f2ba932f2765f5f03781764e11fc422cd0721cbbbcc1a

Observation b78bc22a-fcd7-4a2c-84c2-66b3954ae4d2 · inbound

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning cites this paper.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.469169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.469169Z digest=sha256:eae912f350410d50246be3a9532271863f0ddf295f3629fb0170d6c35807b437

Observation 2146cef1-0d9c-445e-8e79-3a3c1cf73e2a · inbound

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs cites this paper.

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:07:03.766330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:07:03.766330Z digest=sha256:f706736bc5b9a81343c35f8cd44a25e3f288eb965c68e8e0d49b3a59c51cde6c

Observation 78e4abf9-eb85-44cf-9e23-1d155d9b1699 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.186280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.186280Z digest=sha256:ada5ae09af296a17c35652f150171647e2c50ee0ba2a55a6a410923349acc43e

Observation 8c88c9e2-06e9-4bc7-9437-475bd62904d4 · inbound

Two-Dimensional Quantization for Geometry-Aware Audio Coding cites this paper.

Two-Dimensional Quantization for Geometry-Aware Audio Coding Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:20:29.207437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-21T18:16:51.486807Z digest=sha256:a79e9ba78d79f33b00a6769326be65a24292e9888e9a33ad248d396941d3ef59

Observation 18b57f1e-c3ff-45f2-985d-6beb4f49fef2 · inbound

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion cites this paper.

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:44.613832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T11:49:04.308326Z digest=sha256:31ff57e5c04e8988531f75f9c8e6869b1c46fd58a5471c9244d592b7d9a80aef