Pith. sign in

Paper Citation Record · LEDGER

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2506.20945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20945 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:43:52.529754Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27adb6b3-b24d-4ae4-8078-4cbcf618dc2c · outbound

This paper cites Mega- tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Mega- tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.933915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:50.228144Z digest=sha256:b9bc0448942c419305d04834b99eeafe4e669d786497be7822d2a6bb416b9323

Observation 49a3a70a-a4b9-4624-8064-993c0af597aa · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.265834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.265834Z digest=sha256:0eaa71e4c1e169c8cfdab3cbe55f5f4f75b21071ee83c301462a7c00e77e3364

Observation 663c1c04-c3e4-47e7-9303-e95d307ed5f0 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.332197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.332197Z digest=sha256:eef7186a3013a40039341b10dc054ddcf86896be8b916bad8dab164c99da80db

Observation 80bebff5-0883-408a-bb4e-bb46906172b1 · outbound

This paper cites Imaginary voice: Face-styled diffusion model for text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Imaginary voice: Face-styled diffusion model for text-to-speech,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.768449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:50.379308Z digest=sha256:7d7f174869bd56db76cff87e34bf3c0a33955c9ff90d25e729d9e01b1feb8fb4

Observation 3bae3fb5-4477-40bf-a466-d1206948ef89 · outbound

This paper cites SYNTHE-SEES: Face based text-to-speech for virtual speaker,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis SYNTHE-SEES: Face based text-to-speech for virtual speaker,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.560254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:50.454168Z digest=sha256:df57f3d7b0b8e176db0473af8352f0b52837c3bce476909de1bb5ddbff79c4d4

Observation 921cc365-a7bf-4ae9-aafd-b477a96c2b58 · outbound

This paper cites Face2Speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Face2Speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.395195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:50.528613Z digest=sha256:a21c4000af8d2cce50ab3c2fbe70543a4cad0381ae21588d2b84763fe6b6383b

Observation 3cf14a4c-5247-4d3f-a3fd-edced867f33d · outbound

This paper cites FVTTS : Face based voice synthesis for text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis FVTTS : Face based voice synthesis for text-to-speech,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.239168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:50.580539Z digest=sha256:8550fa8e763dc9709cb4478865b5f48db46ad38f33924b1ffcdab2714949d443

Observation a53c088a-a93b-47ee-9a23-c88649a88f76 · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.088342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:50.633123Z digest=sha256:568c3ddc8c529ab381a903a1505049a79777b4bd8dc75364d4602b7792f17f41

Observation ffd35756-5c7b-4282-9d92-0eb646be2e2d · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Prompttts: Controllable text-to-speech with text descriptions,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.942659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:50.701372Z digest=sha256:81e45018094b30d948d39fbdaee8cd5fc00e604d5f69dfa07b065df66ad8d0f1

Observation 3556fb19-c8c4-4690-a30d-2355895035fb · outbound

This paper cites PromptTTS++: Controlling speaker identity in prompt-based text-to- speech using natural language descriptions,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis PromptTTS++: Controlling speaker identity in prompt-based text-to- speech using natural language descriptions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.781073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:50.766499Z digest=sha256:856c263971f434e836602eb8825fba2d948674f4069f893fdb35da4856aae4b2

Observation 96b659af-c3d2-464f-83f5-8ed7af6292f7 · outbound

This paper cites UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.830879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.830879Z digest=sha256:a7816575da7176c1cfb052d1c8f624c86aa606adeb1e0b94140792cd81c9d947

Observation 6de8182b-399d-45b8-bde7-dc17676d2a97 · outbound

This paper cites MM-TTS: Multi- modal prompt based style transfer for expressive text-to-speech synthe- sis,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MM-TTS: Multi- modal prompt based style transfer for expressive text-to-speech synthe- sis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.631249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:50.902999Z digest=sha256:30abe443de3cebb61bc3ce799ebbd9c84ca87efe0d24f80f916725e85c984f5c

Observation 9b3f1fe2-a79f-49e2-ac7e-39310830e03f · outbound

This paper cites Gen- eralized end-to-end loss for speaker verification,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Gen- eralized end-to-end loss for speaker verification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.479581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:50.984883Z digest=sha256:78d8eedbb8e94938379e6ce475e00aa41cb944883fad263a687c10ffb3c67f05

Observation 42cd64a7-6d95-4511-bf4f-adf0c4a098f7 · outbound

This paper cites Additive margin softmax for face verification,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Additive margin softmax for face verification,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.350835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:51.073749Z digest=sha256:0951add9b9ae5149e59f37b7045eb2d7ea13388f87fc93c0cf322441016492a7

Observation 513fb0e2-5d0e-4c9d-980a-02ad682a3cba · outbound

This paper cites Bridging the gap between object and image-level representations for open-vocabulary detection,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Bridging the gap between object and image-level representations for open-vocabulary detection,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.148216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:51.143577Z digest=sha256:957a5114186e0eec6f33f48fe012905b48197eafb364e81e883bf7c7e8f1d828

Observation 1a202008-871a-415c-9415-9faaffdcd3d8 · outbound

This paper cites Joint- teaching: Learning to refine knowledge for resource-constrained un- supervised cross-modal retrieval,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Joint- teaching: Learning to refine knowledge for resource-constrained un- supervised cross-modal retrieval,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.002813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:51.255727Z digest=sha256:697814013a2f8eb6196c32299215f208d387e7ca25abd74c43813b3d36604cf4

Observation 8de811c5-5e1f-49dc-865e-18f221ead62d · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Representation Learning with Contrastive Predictive Coding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.359236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.359236Z digest=sha256:f26433b3829dab317bc308d624ca1b3587951760a791a6908f812b775df09131

Observation 283108dd-26b0-4501-ae07-6e229d58e159 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.818807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:51.427727Z digest=sha256:9cafaf20b1056b3e02c3f88dca0a39291b4c88ba8533cc239c9b970a6d23dbce

Observation a29d89ed-059f-4cef-a988-bf5424033794 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LRS3-TED: a large-scale dataset for visual speech recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.542194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.542194Z digest=sha256:c7f270db52bc63b7e183ab58261163e5cd59abf9e0ee475f7992556b5026e31e

Observation 4f76282e-dc9d-4e21-a18a-023dd6db6970 · outbound

This paper cites Multi- caption text-to-face synthesis: Dataset and algorithm,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Multi- caption text-to-face synthesis: Dataset and algorithm,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.588907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:51.614760Z digest=sha256:bf9754734c7ff149c054e895ba84eb74ff3a8be7d005ed404bbf2c77985a7c07

Observation 5eb4f155-d80c-43eb-b966-9739821a71fa · outbound

This paper cites LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.707283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.707283Z digest=sha256:86ef9989070daebb7bf7fddb4ed0a210e257b481fa440771fb07919187d6c701

Observation d4b436d8-9b4e-4335-b459-f4b37136a926 · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.837350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.837350Z digest=sha256:8e55595d3557ed11e2a8bd2be8aacd65d4ce0bd24158e286d39428a3f54861fb

Observation 1eb94adb-cfed-44c4-8ab4-2efea73e8fff · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.942109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.942109Z digest=sha256:b1990dd57e3cd0b3f04186fdd5d997bd9ee8e03d7e230e7d18e0bb0c20141ccf

Observation 30bb92a6-34a9-4db0-8670-fc1c27a258f0 · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Facenet: A unified embedding for face recognition and clustering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.414214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:52.028232Z digest=sha256:8dee3e3e30baf428cfa667c11dd3e991580773b4f13129729d157aa622b7bac4

Observation a1b04ad8-8725-496d-a7bc-e9b50f752036 · outbound

This paper cites Vggface2: A dataset for recognising faces across pose and age,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Vggface2: A dataset for recognising faces across pose and age,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.221017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:52.103593Z digest=sha256:60e470ab46799db9426d45d2eccff5030beeb3118bb0abd2eb88c0caf78f7861

Observation 0398b027-d2a8-43c4-b82e-21d5e5777b9b · outbound

This paper cites Joint face detection and alignment using multitask cascaded convolutional networks,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Joint face detection and alignment using multitask cascaded convolutional networks,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.040774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:52.210637Z digest=sha256:cc535e64620d40039707dbd6a5856b393e1734d3fc0851ce2313be6769fe0db1

Observation 0859996d-b0fe-41e4-8386-aaebfc5501b1 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:52.286082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:52.286082Z digest=sha256:77d17ab758442065338fedbb4e6a25410acfeea82ee5efd2d45990d5c0edd490

Observation 5e0dcf08-eb5f-4c7c-b3d9-6e454ef07ed8 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis VoxCeleb2: Deep Speaker Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:52.418324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:52.418324Z digest=sha256:23346250e690a1c974c6e13b38523ebc513d0da042ce488ce0b0d1eb0ebfe813

Observation 385dc97e-106a-4cb1-b335-b4c575076666 · outbound

This paper cites Explor- ing the limits of transfer learning with a unified text-to-text transformer,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Explor- ing the limits of transfer learning with a unified text-to-text transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:52.825995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:43:52.529754Z digest=sha256:f271e17928b0be1ef07f775064827ce91b5f9f11ab05dc79305558820f28671a

Pith citing papers

No inbound Pith citation observations are available.