Pith. sign in

Paper Citation Record · LEDGER

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting

As of 13 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2412.20155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20155 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:33:56.087718Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57b2daa0-28dc-43fe-98e5-4938d53c2b1a · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.474405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:55.978812Z digest=sha256:fc1158041d525bd9e178f26c514b7a1351629f9793c18e1ca2ad0de39a6224d4

Observation 3637ec4e-14db-48f5-85d7-b43f73f21954 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:33:55.983782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:33:55.983782Z digest=sha256:d5948b6b40f5826cc1241cf75487096dd6305d2c8437f5514a117910111767b9

Observation ff02bd9e-0d6f-41f9-b8fb-d736d2c1e42c · outbound

This paper cites Grad-StyleSpeech: Any-Speaker Adaptive Text-to-Speech Synthesis with Diffusion Models,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Grad-StyleSpeech: Any-Speaker Adaptive Text-to-Speech Synthesis with Diffusion Models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.460667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:55.989046Z digest=sha256:1ee541bfb936811879a392a3a5479ec9378db07d7b7d2035371ea361e5ba3272

Observation c43b41df-fc1b-4f15-99f1-05aa0107e086 · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:33:55.993734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:33:55.993734Z digest=sha256:22c6831716310617249aaf84ca8ecd2a1875d10a96fa7fce2c6b46de833eec82

Observation cd36bbc1-d4a0-4f91-acc1-c84c63eb359c · outbound

This paper cites YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.447310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:55.998428Z digest=sha256:e20a02e764424c1cfb657083de2e5c331745756cff958225f81f70aa5d1eabc0

Observation 9a221940-a3e4-46a4-8058-23e060163828 · outbound

This paper cites UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.433614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.002933Z digest=sha256:3e456b7892635338f9568f3f298e2f85b5d6e5867363c085d908045d39eb2d11

Observation 0a1b30dc-906f-4ccd-afbe-7aaf74273569 · outbound

This paper cites AdaSpeech: Adaptive Text to Speech for Custom V oice,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting AdaSpeech: Adaptive Text to Speech for Custom V oice,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.417651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.008124Z digest=sha256:39ce7ccc3cc0e23ae50719706d7d45ab28f72f94545445e0aa852f3ce1a5433c

Observation f9976d44-a8be-4a21-b6df-f90fd25a76c5 · outbound

This paper cites Adapting TTS models For New Speakers using Transfer Learning.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Adapting TTS models For New Speakers using Transfer Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:33:56.012265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:33:56.012265Z digest=sha256:b27e7df71fec8ce660c868a5bf42a6c73f5d942447bf6f40cb161ca23e9ba1a9

Observation a661e80a-66ec-4362-a7c0-a0b0b701b009 · outbound

This paper cites V oxCeleb: A Large- Scale Speaker Identification Dataset,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting V oxCeleb: A Large- Scale Speaker Identification Dataset,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.403262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.016452Z digest=sha256:88000381709882fc505b65194d1025b346028d75f23618e82890c09c8ca4d347

Observation 38d93117-beb9-485e-989d-916afa95857c · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.388986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.020445Z digest=sha256:6c3c13749999938a1fd4f25fb67d350f264e32aae2d9b8dbfd08628ca6f6e0ce

Observation f1afb0c0-87dc-453c-938c-e8943250de66 · outbound

This paper cites DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject- Driven Generation,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject- Driven Generation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.374096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.024609Z digest=sha256:30bc0b37c8aa888c0a4618bf72566469dfbe614ba761bc2fc2080cd49273402e

Observation fe16b6a9-17dc-432b-b4b2-35952806d569 · outbound

This paper cites Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.358891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.028861Z digest=sha256:d062d521fd6b704e712f740c93a8cc8d6e9ad604187e10f0c23cd0290eac16e3

Observation 9382b176-b8b7-4f09-88c6-f410b7890aaf · outbound

This paper cites Fast- Speech: Fast, Robust and Controllable Text to Speech,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Fast- Speech: Fast, Robust and Controllable Text to Speech,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.342524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.033429Z digest=sha256:8c308fdc180848aa12dbfc9314f01cbb76b27d284a765ff9d7108a556556b7a9

Observation f0ef5ba8-ea54-4ef7-a2e0-17783fb8252f · outbound

This paper cites One TTS alignment to rule them all,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting One TTS alignment to rule them all,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.326971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.038207Z digest=sha256:1ea2d058d73e993b2fc3ca59b7622c1e96b68d0285e88557949dfab855c7c554

Observation d9078847-cc98-4bba-a151-30d4a22038c3 · outbound

This paper cites Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.312197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.042644Z digest=sha256:7399adbf0b250c52204556bef447778733fde883f9b19d31dc2a6df5d174ae24

Observation 4eba0b97-1f60-431f-bbfd-390c5b60ad1c · outbound

This paper cites Neural Discrete Representation Learning,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Neural Discrete Representation Learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.297286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.047331Z digest=sha256:708af9515803628a4464c5efd1ad69c4452507bab2695249312201c672a67f61

Observation 29d51dbf-2c53-450c-8a62-205117f9797d · outbound

This paper cites Denoising Diffusion Probabilistic Models,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Denoising Diffusion Probabilistic Models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.283889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.051721Z digest=sha256:1828e3d05382d3f996c6e9b67d9bd9be6d07b217e41e4a56d7ce6d925c54d0c1

Observation 1ec483cc-4d90-4046-adfc-852af9108672 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Score-Based Generative Modeling through Stochastic Differential Equations,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:33:56.056337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:33:56.056337Z digest=sha256:f1b9866e5a1dfe3852d791a20ac491f1ed6fdfc826dc4df4865ca4c572b11a25

Observation 0e7e9042-33f4-4d02-baf8-6208fb0c8a6d · outbound

This paper cites ProsoSpeech: Enhancing Prosody With Quantized Vector Pre-training in Text-to-Speech,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting ProsoSpeech: Enhancing Prosody With Quantized Vector Pre-training in Text-to-Speech,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.260513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.060818Z digest=sha256:defe81a7a2156db9404425aeea755bfb3412d92c6fd0c68b584a5ecd79c80e49

Observation 23d7844a-e2f6-4d83-9fef-2ed788c1df22 · outbound

This paper cites Language Models are Unsupervised Multitask Learners,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Language Models are Unsupervised Multitask Learners,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.245730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.065141Z digest=sha256:3cc0484efeceef2e99155ab0529267c04718e28c9da8ce75d260b50e3284577a

Observation d3453a05-c41e-48a2-9416-c4e4e36264bd · outbound

This paper cites CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit (version 0.92),.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit (version 0.92),

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.229694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.069431Z digest=sha256:4d9028714e25fd62326329c0ea2587cf0b261b48392b9c3c6382d866d0b2dc2b

Observation 18f8a39d-d911-4600-8832-ba1880980335 · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.214818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.074175Z digest=sha256:1b313002fbf5c3da1ca5fede638db366ee940e2c7c507935877bc95ea9540423

Observation 5eb6ee85-955e-4622-9db4-a63d68472191 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Su- pervision,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Robust Speech Recognition via Large-Scale Weak Su- pervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.199648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.078711Z digest=sha256:1cbe83a7c466bf8a2c809af092b223a7df13781e7b5dfea5dbfd2441192c4646

Observation d7fb2b4b-2b90-4e08-b1f8-223f6ea86706 · outbound

This paper cites Resemblyzer,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting Resemblyzer,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.184094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.083238Z digest=sha256:95e36f2fe070763b1dc95118e3f11ce81413c71df807610012ad71e3dcc5770a

Observation 865c6e60-bb0b-43e5-85db-273d0e64fde6 · outbound

This paper cites HiFi-GAN: Generative Adversarial Net- works for Efficient and High Fidelity Speech Synthesis,.

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting HiFi-GAN: Generative Adversarial Net- works for Efficient and High Fidelity Speech Synthesis,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:33:56.168979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:33:56.087718Z digest=sha256:dfd14c21cb95ca82ccd944a16bcddde81f1a96a8072b2f995a09b0c26a666a36

Pith citing papers

No inbound Pith citation observations are available.