Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T20:50:17.390376Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2501.19258.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T20:50:17.390376Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T20:50:17.213254Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-09T20:50:17.551912Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5a1cde93-cc78-44a7-a999-ef99aa50718d · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Despite the high-quality output, the prosody of the generated speech sometimes be- comes inappropriate for the context
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 46264d7c-d2c7-418f-8164-aa4dcda14f9b · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video VisualSpeech: Enhancing Prosody Modeling in TTS Using Video
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 614772b2-901a-40fa-901d-0bc9060921b3 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Experimental Setup To assess the impact of visual information on speech generation performance, a dataset containing both diverse prosodic and vi- sual data is essential
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b65af53c-99b6-4490-8153-84449ea9e73f · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Using two dis- tinct video feature extractors, we demonstrate that these visual features encapsulate prosodic information
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5116b954-f3ef-435a-bc36-ba1e01d24db1 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9638055a-4317-40f2-8b01-ae94bdd9e1ea · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Deep mixture density networks for acous- tic modeling in statistical parametric speech synthesis,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6ef20369-d217-483d-b630-cd32c8fcf63c · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88fe2c5f-9cc4-4a1b-b191-bdac925e8e76 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Unidirectional long short-term memory re- current neural network with recurrent output layer for low-latency speech synthesis,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7bc7dcca-935c-4562-86a2-3d334cd4375a · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video NaturalSpeech: End-to-end text-to- speech synthesis with human-level quality,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d9b9273-93e7-4dd2-bef1-3602fcccba1b · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastSpeech: Fast, robust and controllable text to speech,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6bf8f256-6251-4a71-921f-55f767396b09 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video On granularity of prosodic representations in expressive text-to-speech,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b3401b66-54ee-4800-a4b7-f3fed282a958 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c0bffc1-dbce-4135-85f1-968810ebaf38 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Chive: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0eea874-c8af-416c-b8dc-55541afae4b7 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Mellotron: Mul- tispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94ffe127-0a7d-4bff-af9a-4a3d08338dee · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b094d9-a8b9-4d94-819b-d3989ab742f9 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video The LJ Speech Dataset,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed07cf1c-e7a6-4559-b713-9ea96e7bb626 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a95dddd-a602-4907-a7dd-a1b6023ca39c · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Ego4d: Around the world in 3,000 hours of egocentric video,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3fab26cc-05a9-4223-80fa-deed08cd241a · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Condensed movies: Story based retrieval with contextual embeddings,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c26d7058-5e60-4d59-8f14-ba95f6a3a9f1 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Robust speech recognition via large-scale weak supervision,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f03ce37d-88da-4aec-a4cd-b3e2861d5cc6 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video CMD+: A D.I.Y . Audiovisual Dataset for Multi- Speaker TTS,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ec659a45-2c03-4b10-bd2e-179f4d146a1d · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video The Sound Demixing Challenge 2023 $\unicode{x2013}$ Music Demixing Track
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c18112-6f8c-482e-b89b-e2c661ee6d81 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Resemblyzer,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ba086c40-95b3-425b-8253-8c5412fe4d0d · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d23f2617-43ca-43af-9b6f-d7c6a5031199 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video A short- time objective intelligibility measure for time-frequency weighted noisy speech,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fc38aa8-fb53-4dd2-9a6f-7a97c3cbc26a · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8ac8f403-5036-45c5-b5a3-c09bb07886bd · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video SDR– half-baked or well done?
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0f45bc67-4145-4ae8-ac5e-f4f90fd1eec3 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Torchaudio-Squim: Reference-less speech quality and intelligibility measures in torchaudio,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d918af71-c765-4722-8453-a3ce1a8793bc · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Omnivore: A single model for many visual modalities,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a12931ea-8921-4708-afa1-3bc265c6a23e · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Deep residual learning for image recognition,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de6c5db6-6bc8-4274-87e0-efa7ea411131 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Adam: A Method for Stochastic Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe24b21-8c07-4d94-97e1-6b4b4d3a5c46 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastPitch: Parallel text-to-speech with pitch pre- diction,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e3e3b893-f7fc-4721-9143-ac3f359b7a27 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Sonicvisionlm: Playing sound with vision language models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4eeae4f1-ef67-4de6-8687-25c0e93965a3 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video End-to-end video-to-speech synthesis using gener- ative adversarial networks,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dce0b177-93b1-4bad-ac6b-1a923b5ebdff · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Intelligible Lip-to-Speech Synthesis with Speech Units
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f82f58ef-f399-4ae2-8df7-c33c5e669b64 · outbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Camp: a two-stage approach to modelling prosody in context,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 46264d7c-d2c7-418f-8164-aa4dcda14f9b · inbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video VisualSpeech: Enhancing Prosody Modeling in TTS Using Video
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.