Pith. sign in

Paper Citation Record · LEDGER

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis

As of 11 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2506.20945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20945 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:43:52.529754Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27adb6b3-b24d-4ae4-8078-4cbcf618dc2c · outbound

This paper cites Mega- tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Mega- tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.933915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:50.228144Z digest=sha256:459a495d2bfb1c0d91ac8e8994a388492b3c031df157ae8b2aa7c6149e00953f

Observation 49a3a70a-a4b9-4624-8064-993c0af597aa · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.265834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.265834Z digest=sha256:2a482357ae030497ed8f1a8c3b9b2593ee9d1349140082de4ead7b656a79ddf8

Observation 663c1c04-c3e4-47e7-9303-e95d307ed5f0 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.332197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.332197Z digest=sha256:f6c5b0de6bf71a21d9e918692b3ccb8f7374d9351eef2aa9ed589235246320da

Observation 80bebff5-0883-408a-bb4e-bb46906172b1 · outbound

This paper cites Imaginary voice: Face-styled diffusion model for text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Imaginary voice: Face-styled diffusion model for text-to-speech,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.768449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:50.379308Z digest=sha256:f47e3f9bc8321d2e0947fd52582398198553dd7014666b7ec32227af1e1c48a6

Observation 3bae3fb5-4477-40bf-a466-d1206948ef89 · outbound

This paper cites SYNTHE-SEES: Face based text-to-speech for virtual speaker,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis SYNTHE-SEES: Face based text-to-speech for virtual speaker,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.560254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:50.454168Z digest=sha256:517dbb00578e00cb4fa4bf2e183db41cd66bbfdc8a461e4fe889e54060911b43

Observation 921cc365-a7bf-4ae9-aafd-b477a96c2b58 · outbound

This paper cites Face2Speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Face2Speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.395195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:50.528613Z digest=sha256:23ad8aed5238ac0a16899fad181930406592f57630eb625011a63f5ab9a7e470

Observation 3cf14a4c-5247-4d3f-a3fd-edced867f33d · outbound

This paper cites FVTTS : Face based voice synthesis for text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis FVTTS : Face based voice synthesis for text-to-speech,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.239168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:50.580539Z digest=sha256:50a8c47a004d1f7873877cb0c8d971a18c0bfec1e9bbda27810cb3068ec66313

Observation a53c088a-a93b-47ee-9a23-c88649a88f76 · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:55.088342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:50.633123Z digest=sha256:e085ce68229f10e8a9ee95dd7fd54fc6f6d78c9c466f6f34085b0c032ed0c882

Observation ffd35756-5c7b-4282-9d92-0eb646be2e2d · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Prompttts: Controllable text-to-speech with text descriptions,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.942659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:50.701372Z digest=sha256:ab808f8eac117575bd5ad3b4ae80ad9930933d8566263da0a47339658d9e3a7f

Observation 3556fb19-c8c4-4690-a30d-2355895035fb · outbound

This paper cites PromptTTS++: Controlling speaker identity in prompt-based text-to- speech using natural language descriptions,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis PromptTTS++: Controlling speaker identity in prompt-based text-to- speech using natural language descriptions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.781073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:50.766499Z digest=sha256:c63b9d066fc41190a004623f15ec0c9e18372be03048ebe09daa0695c25fd7b4

Observation 96b659af-c3d2-464f-83f5-8ed7af6292f7 · outbound

This paper cites UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.830879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.830879Z digest=sha256:a7816575da7176c1cfb052d1c8f624c86aa606adeb1e0b94140792cd81c9d947

Observation 6de8182b-399d-45b8-bde7-dc17676d2a97 · outbound

This paper cites MM-TTS: Multi- modal prompt based style transfer for expressive text-to-speech synthe- sis,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MM-TTS: Multi- modal prompt based style transfer for expressive text-to-speech synthe- sis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.631249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:50.902999Z digest=sha256:98b1cb612125cdad7421537a77a1a43edee8ec2b854897a1cfe83e9ea2133e3c

Observation 9b3f1fe2-a79f-49e2-ac7e-39310830e03f · outbound

This paper cites Gen- eralized end-to-end loss for speaker verification,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Gen- eralized end-to-end loss for speaker verification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.479581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:50.984883Z digest=sha256:11795d924885544ed4470de1dabf3d4f4084b75a4867836678d8f3bb18198ddf

Observation 42cd64a7-6d95-4511-bf4f-adf0c4a098f7 · outbound

This paper cites Additive margin softmax for face verification,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Additive margin softmax for face verification,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.350835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:51.073749Z digest=sha256:02f0fc8de779714066290ec61819e1cbea80b74db9d01bf7c0480e2514f18c41

Observation 513fb0e2-5d0e-4c9d-980a-02ad682a3cba · outbound

This paper cites Bridging the gap between object and image-level representations for open-vocabulary detection,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Bridging the gap between object and image-level representations for open-vocabulary detection,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.148216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:51.143577Z digest=sha256:b9668ac5ee6e3cb607da1971a39a20f34b70b36028ce2994cf18ac56a1e33a83

Observation 1a202008-871a-415c-9415-9faaffdcd3d8 · outbound

This paper cites Joint- teaching: Learning to refine knowledge for resource-constrained un- supervised cross-modal retrieval,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Joint- teaching: Learning to refine knowledge for resource-constrained un- supervised cross-modal retrieval,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:54.002813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:51.255727Z digest=sha256:57ac0d8c2c49df42a8246e3428459aac2feb9785dafe6002b22e29675ffa5b09

Observation 8de811c5-5e1f-49dc-865e-18f221ead62d · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Representation Learning with Contrastive Predictive Coding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.359236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.359236Z digest=sha256:f26433b3829dab317bc308d624ca1b3587951760a791a6908f812b775df09131

Observation 283108dd-26b0-4501-ae07-6e229d58e159 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.818807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:51.427727Z digest=sha256:8f3b3b74101e41a0301b97514f6afe038c818f630f452dab1869b18bead2b6fa

Observation a29d89ed-059f-4cef-a988-bf5424033794 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LRS3-TED: a large-scale dataset for visual speech recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.542194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.542194Z digest=sha256:c7f270db52bc63b7e183ab58261163e5cd59abf9e0ee475f7992556b5026e31e

Observation 4f76282e-dc9d-4e21-a18a-023dd6db6970 · outbound

This paper cites Multi- caption text-to-face synthesis: Dataset and algorithm,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Multi- caption text-to-face synthesis: Dataset and algorithm,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.588907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:51.614760Z digest=sha256:dd1e4e3a84963c0451b52e77224f23fd6c6a6ab65d060e1eb7fa705f858ca119

Observation 5eb4f155-d80c-43eb-b966-9739821a71fa · outbound

This paper cites LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.707283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.707283Z digest=sha256:86ef9989070daebb7bf7fddb4ed0a210e257b481fa440771fb07919187d6c701

Observation d4b436d8-9b4e-4335-b459-f4b37136a926 · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.837350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.837350Z digest=sha256:88785ebff2097d0a29acddfb5630b32ecfa2fdc75d73606d4223b0ddd4b8132a

Observation 1eb94adb-cfed-44c4-8ab4-2efea73e8fff · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:51.942109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:51.942109Z digest=sha256:ab4632833f5276b779cb18a5605916b90c4fe5e7c977113b02fcc41f354f49c3

Observation 30bb92a6-34a9-4db0-8670-fc1c27a258f0 · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Facenet: A unified embedding for face recognition and clustering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.414214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:52.028232Z digest=sha256:06acfbff7bcabc9072cc68035dfcce9b9b18412660b419e2ace984b913bf6229

Observation a1b04ad8-8725-496d-a7bc-e9b50f752036 · outbound

This paper cites Vggface2: A dataset for recognising faces across pose and age,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Vggface2: A dataset for recognising faces across pose and age,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.221017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:52.103593Z digest=sha256:00d1f211a7ba3676f0b44969dc006d3d5e45cc337e2f374f327d3eaa9769213b

Observation 0398b027-d2a8-43c4-b82e-21d5e5777b9b · outbound

This paper cites Joint face detection and alignment using multitask cascaded convolutional networks,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Joint face detection and alignment using multitask cascaded convolutional networks,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:53.040774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:52.210637Z digest=sha256:d53c3e8d57a09d79b244ec81fd28ed544ab507e2f0b3ee0a55cf84acf525f635

Observation 0859996d-b0fe-41e4-8386-aaebfc5501b1 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:52.286082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:52.286082Z digest=sha256:77d17ab758442065338fedbb4e6a25410acfeea82ee5efd2d45990d5c0edd490

Observation 5e0dcf08-eb5f-4c7c-b3d9-6e454ef07ed8 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis VoxCeleb2: Deep Speaker Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:52.418324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:52.418324Z digest=sha256:23346250e690a1c974c6e13b38523ebc513d0da042ce488ce0b0d1eb0ebfe813

Observation 385dc97e-106a-4cb1-b335-b4c575076666 · outbound

This paper cites Explor- ing the limits of transfer learning with a unified text-to-text transformer,.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis Explor- ing the limits of transfer learning with a unified text-to-text transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:43:52.825995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:43:52.529754Z digest=sha256:e35b95b43c928517a60e8292852910fe3125debf0ba93bd9a04074ca5881cd38

Pith citing papers

No inbound Pith citation observations are available.