Pith. sign in

Paper Citation Record · LEDGER

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation

As of 17 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2506.02997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02997 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:32.475093Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98ee0846-7ee5-4696-a363-c8518914cb4d · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Prompttts: Controllable text-to-speech with text descriptions,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.432358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:16:31.033516Z digest=sha256:99195a50479801462072a76ac781dc68e600c9169d9a1eac7b4b266738a05420

Observation 43e8d4ab-718e-47cb-bf3e-fabe5fd92d5d · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.114971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.114971Z digest=sha256:3e372ce0a90cdf2e645e395cf74c0966d1ff9406fc2662fccb1d355748bdce06

Observation 01f03004-b063-4a86-81c5-1c4eb5d14d2c · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Textrolspeech: A text style control speech corpus with codec language text-to-speech models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.257590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:16:31.233871Z digest=sha256:890fe361f61894da3317f1df7cd16a9dfdba9aaac7c85faba1cde82a96e34c2f

Observation 9d2bd77f-ec90-4f34-8bdb-9bbeb17e9ad7 · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.065052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:16:31.305771Z digest=sha256:2b99b4ee9e0624855a66100535dbeb23e226547135da6e08983269971635b61a

Observation 4274f1ba-67a4-4902-ae9d-8cb9510965e3 · outbound

This paper cites VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:16:32.936942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:16:31.421450Z digest=sha256:191e0acd87e7c5e8c93729c6ea825d338a97eaf710c3359dda484d1dcd74c9ba

Observation 36e3e659-0c81-48b7-8ddc-5fbf05c4dd55 · outbound

This paper cites Prosody-tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Prosody-tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.848661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:16:31.520970Z digest=sha256:9b06f99fdc712bf95799c74fc0e2a2ebdbba941a82a134297cd0df0a5b638b4d

Observation 7fc73b80-3143-46a8-8b40-71a6b50a205b · outbound

This paper cites Uniaudio: Towards universal audio generation with large language models,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Uniaudio: Towards universal audio generation with large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.625603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:16:31.600517Z digest=sha256:5f5f73eb69b9ef0f6fd76f46119ad87c249d2b55750ee9fdfa3d1e1248d96360

Observation 71595cf2-7d84-4261-bb37-a801c98d8cf4 · outbound

This paper cites Classifier-free diffusion guidance,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Classifier-free diffusion guidance,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.676312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.676312Z digest=sha256:ddcb1d02116b4828c0314b08795151fedf3b3aec7df2afbfa7f56caf8af24d00

Observation 2c4cdf49-992a-4d22-a6ff-09981ad38bb5 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.762687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.762687Z digest=sha256:1886c6edae386a72f876cd91c92c3c326149d399423cb863b6cd104b858adce2

Observation 827789aa-6c87-4b25-b9ce-1b1bc304e76f · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Librispeech: an asr corpus based on public domain audio books,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.849446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.849446Z digest=sha256:08397097c81a3ef7a0bffa06b2ecb65196c6e18ab387084f07c7093657f186c2

Observation fb796082-112f-4c2d-b715-ee251fdd9099 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.921256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.921256Z digest=sha256:559aa962b4f7dc8fa9273e9d44b5739e13bcce16bfd94bb83c3944b39140b0b8

Observation 6af88a1a-5f9a-498a-b07e-3a8ffabb8bc3 · outbound

This paper cites Dailytalk: Spoken dialogue dataset for conversational text-to-speech,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Dailytalk: Spoken dialogue dataset for conversational text-to-speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.382842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:16:31.993650Z digest=sha256:a9df271b249597970a9d019169a240c41f87bbb663fd249e15f162d8c11a7d18

Observation f3069eef-0cf1-4c51-be97-96203ce4197e · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.057240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.057240Z digest=sha256:bdbf0ae70af9cc761a39cccd3a08261fb1c62d40cc5dc0ddafe276b09df08a21

Observation c89fd632-0615-417d-824b-23b60a90d21b · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.166827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.166827Z digest=sha256:447ecfb9a7a82e029aa523eb809ac6bf26a4df35337253f0ef2bf37dce4177cd

Observation 1d45114f-5714-4f3e-a082-87767117d8d6 · outbound

This paper cites High Fidelity Neural Audio Compression.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation High Fidelity Neural Audio Compression

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.279195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.279195Z digest=sha256:73573d9bd4d6c7aa9ad38ae10cacb05264df516409f9b624aac6cbaf8b0aa24a

Observation 27a2671c-87a4-4c6d-a93e-23daa4105847 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.190288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:16:32.337391Z digest=sha256:2d71efbcd983a6259b5f24ee2b666d5b17b287dfe8d9230ff7f34f6f7694ddb0

Observation 975c135b-5556-4155-bad1-0e7301bc06e9 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.422722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.422722Z digest=sha256:dc952986a5c85c34ab2e1999774815998baefc726f1a83e65c86c2095f029d1a

Observation b4d329d4-8d54-4a05-b4da-311e7867885a · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.475093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.475093Z digest=sha256:012af1a969e8d32bbf7d415a4dc18447b591f1bae47b5190f0f18265455638cb

Pith citing papers

No inbound Pith citation observations are available.