Pith. sign in

Paper Citation Record · LEDGER

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2506.02997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02997 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:32.475093Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98ee0846-7ee5-4696-a363-c8518914cb4d · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Prompttts: Controllable text-to-speech with text descriptions,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.432358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:16:31.033516Z digest=sha256:c5e3dd8b7e1fc18ca4a4dc7ad5785a4f9a261b16aa7e31ba1ddce4fd7c5ea1b7

Observation 43e8d4ab-718e-47cb-bf3e-fabe5fd92d5d · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.114971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.114971Z digest=sha256:da1dc9194087a91eff851a616c55b7fe3b8351a3c55c1322d97eaa07195e526a

Observation 01f03004-b063-4a86-81c5-1c4eb5d14d2c · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Textrolspeech: A text style control speech corpus with codec language text-to-speech models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.257590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:16:31.233871Z digest=sha256:0d8b69d6b9120d9adf4b015c21aa6115d7f120eda74e6fc46f9670f3bd7e9067

Observation 9d2bd77f-ec90-4f34-8bdb-9bbeb17e9ad7 · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.065052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:16:31.305771Z digest=sha256:f21e34eb6aa552c2cc63ffcce24e979bb4442df4069e4fcd6dace74eedf31045

Observation 4274f1ba-67a4-4902-ae9d-8cb9510965e3 · outbound

This paper cites VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:16:32.936942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:16:31.421450Z digest=sha256:1a41c3e655e7fc148510c96decdc4c37edc71c3e48f2b5447e40794df5808829

Observation 36e3e659-0c81-48b7-8ddc-5fbf05c4dd55 · outbound

This paper cites Prosody-tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Prosody-tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.848661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:16:31.520970Z digest=sha256:a8514130b1ca31d1dda7455860086b1fb7252309bd1dafa8265c06b69b3f6222

Observation 7fc73b80-3143-46a8-8b40-71a6b50a205b · outbound

This paper cites Uniaudio: Towards universal audio generation with large language models,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Uniaudio: Towards universal audio generation with large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.625603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:16:31.600517Z digest=sha256:9a4c6bc0b889c1b9c62a24edd8d7a4c006e4bd15d3fbac3b94a885ec1f1ead0f

Observation 71595cf2-7d84-4261-bb37-a801c98d8cf4 · outbound

This paper cites Classifier-free diffusion guidance,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Classifier-free diffusion guidance,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.676312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.676312Z digest=sha256:54581d22259237cb56a98893cca571ec4164f5e0b7f2216aa16b07c53d1cfffa

Observation 2c4cdf49-992a-4d22-a6ff-09981ad38bb5 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.762687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.762687Z digest=sha256:e51ac229c20ec0b340d2e337f09c7a2db83efc0d8feccea04885e075541e1160

Observation 827789aa-6c87-4b25-b9ce-1b1bc304e76f · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Librispeech: an asr corpus based on public domain audio books,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.849446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.849446Z digest=sha256:fd97d498e85f556d5dc80b1e85a2ecd74e8638094522cab303c67fcdb0975421

Observation fb796082-112f-4c2d-b715-ee251fdd9099 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.921256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.921256Z digest=sha256:a0d7057522537f00ecde26607e64738a7fe83c437fd1354b7362eb5b9e2ab5a4

Observation 6af88a1a-5f9a-498a-b07e-3a8ffabb8bc3 · outbound

This paper cites Dailytalk: Spoken dialogue dataset for conversational text-to-speech,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Dailytalk: Spoken dialogue dataset for conversational text-to-speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.382842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:16:31.993650Z digest=sha256:1dbcab340fd8c567965727b87cc08b7d0242b13ad77b5021b80d79282322d2d9

Observation f3069eef-0cf1-4c51-be97-96203ce4197e · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.057240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.057240Z digest=sha256:56573d0e1e19d7c89fdbf419e7fcd67dad31de5c834c5ffac93c2521feb99f30

Observation c89fd632-0615-417d-824b-23b60a90d21b · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.166827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.166827Z digest=sha256:dca65be6a4bd63f6118e5f1eb88d4106ed8f8fe0242eafdbb79918814701d814

Observation 1d45114f-5714-4f3e-a082-87767117d8d6 · outbound

This paper cites High Fidelity Neural Audio Compression.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation High Fidelity Neural Audio Compression

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.279195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.279195Z digest=sha256:9aee1df67339bbaab1d3a988fd44f0764f5dd5a2b5a09dd9bdf4787e736241f9

Observation 27a2671c-87a4-4c6d-a93e-23daa4105847 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.190288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:16:32.337391Z digest=sha256:415804a24da74fb3b66d37a7bb16b1fc705b9476b9feb20f54127b0bc746dc55

Observation 975c135b-5556-4155-bad1-0e7301bc06e9 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.422722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.422722Z digest=sha256:8c34a48b76f7158765aded3c53e23f826740b3d8e7c7f72c8235d54ebc3fdd5e

Observation b4d329d4-8d54-4a05-b4da-311e7867885a · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.475093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.475093Z digest=sha256:f8942ae803bce5528cae8705dcf67c67c082df4fe01539917951b44662e93473

Pith citing papers

No inbound Pith citation observations are available.