Pith. sign in

Paper Citation Record · LEDGER

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video

As of 10 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2501.19258.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.19258 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:50:17.390376Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:50:17.213254Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-09T20:50:17.551912Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a1cde93-cc78-44a7-a999-ef99aa50718d · outbound

This paper cites Despite the high-quality output, the prosody of the generated speech sometimes be- comes inappropriate for the context.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Despite the high-quality output, the prosody of the generated speech sometimes be- comes inappropriate for the context

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.994955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.207333Z digest=sha256:4a2352efab8daeeb60e764053ec19608be455f6daf33bc8b6cb65a3400584968

Observation 46264d7c-d2c7-418f-8164-aa4dcda14f9b · outbound

This paper cites VisualSpeech: Enhancing Prosody Modeling in TTS Using Video.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video VisualSpeech: Enhancing Prosody Modeling in TTS Using Video

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T20:50:17.559362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.213254Z digest=sha256:e93ba54edc421897ead3953a425c921e1eb3f7529bb2bfb1a0fcd755b757c20b

Observation 614772b2-901a-40fa-901d-0bc9060921b3 · outbound

This paper cites Experimental Setup To assess the impact of visual information on speech generation performance, a dataset containing both diverse prosodic and vi- sual data is essential.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Experimental Setup To assess the impact of visual information on speech generation performance, a dataset containing both diverse prosodic and vi- sual data is essential

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.978937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.218991Z digest=sha256:0d19f1a2b664ebf31872a48f8bab43cfe33407d1f33de9adb1245b96eaae971f

Observation b65af53c-99b6-4490-8153-84449ea9e73f · outbound

This paper cites Using two dis- tinct video feature extractors, we demonstrate that these visual features encapsulate prosodic information.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Using two dis- tinct video feature extractors, we demonstrate that these visual features encapsulate prosodic information

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.944262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.229624Z digest=sha256:20518adf93a100436c15ee2f9643ed4c0822a12b6e886a0e8168edd2e8133f28

Observation 5116b954-f3ef-435a-bc36-ba1e01d24db1 · outbound

This paper cites FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.254499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.254499Z digest=sha256:a8e5a360d301c9b30d73d07624a7066fb7c9d04015756e5210918c59da600f7a

Observation 9638055a-4317-40f2-8b01-ae94bdd9e1ea · outbound

This paper cites Deep mixture density networks for acous- tic modeling in statistical parametric speech synthesis,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Deep mixture density networks for acous- tic modeling in statistical parametric speech synthesis,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.925404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.234534Z digest=sha256:ba926481f090d2d7fc07b698f3eb9f8a2f353d940a82fabd75ba8833840e8263

Observation 6ef20369-d217-483d-b630-cd32c8fcf63c · outbound

This paper cites an unresolved cited work.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:50:17.961376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.224287Z digest=sha256:9f9263704ec16a313250ce67d127000e8ecbad1ef0cafa45440f8a4a90a9e660

Observation 88fe2c5f-9cc4-4a1b-b191-bdac925e8e76 · outbound

This paper cites Unidirectional long short-term memory re- current neural network with recurrent output layer for low-latency speech synthesis,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Unidirectional long short-term memory re- current neural network with recurrent output layer for low-latency speech synthesis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.906544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.239311Z digest=sha256:66027f9051efb5d6e487b7a14159771bdf7068b5838239d49f11e72fd44c644d

Observation 7bc7dcca-935c-4562-86a2-3d334cd4375a · outbound

This paper cites NaturalSpeech: End-to-end text-to- speech synthesis with human-level quality,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video NaturalSpeech: End-to-end text-to- speech synthesis with human-level quality,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.888796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.244348Z digest=sha256:1ee14d4ec537863c77df981865ed3b310480a6adde42b363ef233c4a829ecbf0

Observation 9d9b9273-93e7-4dd2-bef1-3602fcccba1b · outbound

This paper cites FastSpeech: Fast, robust and controllable text to speech,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastSpeech: Fast, robust and controllable text to speech,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.871652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.249595Z digest=sha256:d4acfe6de007612d9cedd1268840984868c015e1c2b7f4a3655be5fcbb086f72

Observation 6bf8f256-6251-4a71-921f-55f767396b09 · outbound

This paper cites On granularity of prosodic representations in expressive text-to-speech,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video On granularity of prosodic representations in expressive text-to-speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.854616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.260279Z digest=sha256:30dd11d289e533ea9edbb323c8b872d7d6956619b30c953296a58a7cfc53985c

Observation b3401b66-54ee-4800-a4b7-f3fed282a958 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.265052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.265052Z digest=sha256:429713b7ec326278daefce352917e5fd1883c80f6ed8750235fdfec9cf5274c7

Observation 0c0bffc1-dbce-4135-85f1-968810ebaf38 · outbound

This paper cites Chive: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Chive: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.838085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.275553Z digest=sha256:d37e04623f6d812248fd0e70b720409b72c8183e27fb0ee9111b2f9ee325b5b0

Observation d0eea874-c8af-416c-b8dc-55541afae4b7 · outbound

This paper cites Mellotron: Mul- tispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Mellotron: Mul- tispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.280420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.280420Z digest=sha256:f192147bc5ecd0b3ac2673b28694fb76f4f6aaef9657c0c09288bf9d23e0067f

Observation 94ffe127-0a7d-4bff-af9a-4a3d08338dee · outbound

This paper cites ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.284979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.284979Z digest=sha256:52e5573d00008c7f28ea4b2f4f3c3cebc04b7365584c5fcde604ec78a133f13e

Observation 46b094d9-a8b9-4d94-819b-d3989ab742f9 · outbound

This paper cites The LJ Speech Dataset,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video The LJ Speech Dataset,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.289999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.289999Z digest=sha256:b3ea75561c921a4ab168ed42bf52986120ab6691471dac210ec0cff0187e90b8

Observation ed07cf1c-e7a6-4559-b713-9ea96e7bb626 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.294686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.294686Z digest=sha256:ff54405b8ce35eb9774f6942bdcc62868014a80233078e260a2120fb3e3b4043

Observation 9a95dddd-a602-4907-a7dd-a1b6023ca39c · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.801073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.300516Z digest=sha256:0b958d3b84a04ce86cea3e382fa818721d75df5f302f56db0cd459fde77d60b8

Observation 3fab26cc-05a9-4223-80fa-deed08cd241a · outbound

This paper cites Condensed movies: Story based retrieval with contextual embeddings,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Condensed movies: Story based retrieval with contextual embeddings,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.784934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.305703Z digest=sha256:bf734fef8bed4ee293daa6377b42b97c95eef722121a4add9277e183b54f070f

Observation c26d7058-5e60-4d59-8f14-ba95f6a3a9f1 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Robust speech recognition via large-scale weak supervision,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.769915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.310629Z digest=sha256:a47b6c17b9824d965350c2512d99e28ddc5876f2d19e0ca1a137527d57decd9c

Observation f03ce37d-88da-4aec-a4cd-b3e2861d5cc6 · outbound

This paper cites CMD+: A D.I.Y . Audiovisual Dataset for Multi- Speaker TTS,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video CMD+: A D.I.Y . Audiovisual Dataset for Multi- Speaker TTS,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.754642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.316038Z digest=sha256:1137f5ef1f050855e67df0876c79a117f5abee94e7e1f938174e21d3a8d80eb4

Observation ec659a45-2c03-4b10-bd2e-179f4d146a1d · outbound

This paper cites The Sound Demixing Challenge 2023 $\unicode{x2013}$ Music Demixing Track.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video The Sound Demixing Challenge 2023 $\unicode{x2013}$ Music Demixing Track

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.320971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.320971Z digest=sha256:4958a46526fc376b05243a7c36038524ab4c61b9441eb15a5bc86b0604100ea4

Observation 77c18112-6f8c-482e-b89b-e2c661ee6d81 · outbound

This paper cites Resemblyzer,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Resemblyzer,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.738232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.326282Z digest=sha256:bd8c72a9a47680cc6b99e4ab87b028644a199ef38298f68a52a98c514b064a15

Observation ba086c40-95b3-425b-8253-8c5412fe4d0d · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.331087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.331087Z digest=sha256:5e99b97607ba5961d8cbe417bea41c3bfe29f071a17ae694e09be1af0fcf2155

Observation d23f2617-43ca-43af-9b6f-d7c6a5031199 · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.721365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.336097Z digest=sha256:8c6189a8f120cddb2e263d4eba4debe9bcafe0c4c709b4e203e6b6eb834d80c3

Observation 1fc38aa8-fb53-4dd2-9a6f-7a97c3cbc26a · outbound

This paper cites Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.705037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.341247Z digest=sha256:0062649cb4c56d0132fbd95d22e432178f6ace7f14eb73cd4520870f1fe5486d

Observation 8ac8f403-5036-45c5-b5a3-c09bb07886bd · outbound

This paper cites SDR– half-baked or well done?.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video SDR– half-baked or well done?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.686765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.346185Z digest=sha256:34ffc74f3cf838ff460598455264aac04a0d9990b10d7d46a11e2811d1595aea

Observation 0f45bc67-4145-4ae8-ac5e-f4f90fd1eec3 · outbound

This paper cites Torchaudio-Squim: Reference-less speech quality and intelligibility measures in torchaudio,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Torchaudio-Squim: Reference-less speech quality and intelligibility measures in torchaudio,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.670742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.350971Z digest=sha256:b9f9e29a24f25065737cd3b40abf75a3bfa3de9201e1ddc938360ce8f9927f60

Observation d918af71-c765-4722-8453-a3ce1a8793bc · outbound

This paper cites Omnivore: A single model for many visual modalities,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Omnivore: A single model for many visual modalities,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.651564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.356165Z digest=sha256:d9065355bd65c1f822688e94e4005bd93d7a978263c2561eb29d48e1a9092b6a

Observation a12931ea-8921-4708-afa1-3bc265c6a23e · outbound

This paper cites Deep residual learning for image recognition,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Deep residual learning for image recognition,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.360860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.360860Z digest=sha256:61607d09aa485d8a22ae5a297dbc6de29dab7e873fdd1832428d36ef57c6fb4f

Observation de6c5db6-6bc8-4274-87e0-efa7ea411131 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Adam: A Method for Stochastic Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.365484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.365484Z digest=sha256:58d962cb013df4556d1497931194b7df6ee6e4066f8cb3079f30ec2880808866

Observation dfe24b21-8c07-4d94-97e1-6b4b4d3a5c46 · outbound

This paper cites FastPitch: Parallel text-to-speech with pitch pre- diction,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastPitch: Parallel text-to-speech with pitch pre- diction,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.624813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.370920Z digest=sha256:e2f7acb578758fdb6fcdcd6e212287b1225aa4219ab87c7afb3a5706bcdb7f88

Observation e3e3b893-f7fc-4721-9143-ac3f359b7a27 · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Sonicvisionlm: Playing sound with vision language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.608653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.375699Z digest=sha256:b98488d8aabf6c6a9e70a9edf79becb70c41ec7a9681bf03d49773fa432b37b3

Observation 4eeae4f1-ef67-4de6-8687-25c0e93965a3 · outbound

This paper cites End-to-end video-to-speech synthesis using gener- ative adversarial networks,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video End-to-end video-to-speech synthesis using gener- ative adversarial networks,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.593035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.380409Z digest=sha256:049604df7ab1f21cd52500d044c8cc810cdb7f60d4b147b3b017303d0e47547d

Observation dce0b177-93b1-4bad-ac6b-1a923b5ebdff · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Intelligible Lip-to-Speech Synthesis with Speech Units

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.385174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.385174Z digest=sha256:342c50f039e71e2dfce73cb002a0d1c58ebfd7ac5d006cfcdb8dc8689eb6534c

Observation f82f58ef-f399-4ae2-8df7-c33c5e669b64 · outbound

This paper cites Camp: a two-stage approach to modelling prosody in context,.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Camp: a two-stage approach to modelling prosody in context,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:50:17.577819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.390376Z digest=sha256:cd9918fe0c28d8b2d355aa77b0a951de3d6c002541e099f57dae00d97fbcf00a

Pith citing papers

Observation 46264d7c-d2c7-418f-8164-aa4dcda14f9b · inbound

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video cites this paper.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video VisualSpeech: Enhancing Prosody Modeling in TTS Using Video

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T20:50:17.559362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T20:50:17.213254Z digest=sha256:e93ba54edc421897ead3953a425c921e1eb3f7529bb2bfb1a0fcd755b757c20b