Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:15.460678Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2505.22024.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:15.460678Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:13.022735Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T13:21:15.811155Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ed1c5e6b-449f-4e0f-becd-290f036641e4 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling L2S holds great potential across various domains, from enhancing communication in noisy environments to providing assistive technologies for individuals with aphonia
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9a1b0569-55db-4704-b26b-f58a996c0814 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 112078a0-0eb7-4c59-9ab8-6ed09a77ef0f · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Datasets LRS2-BBC[25] is an English audio-visual dataset from BBC programs, comprising over 220 hours of video
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94b33b3c-1bc4-4513-9272-94e998abd5bb · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ac18ebce-cb78-4f5a-9a3f-5ec3710a5796 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Grounded in source-filter theory, RESOUND separates speech generation into acoustic and semantic branches, capturing prosody and linguistic con- tent
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9fcaeeb4-fdd8-4bec-9b4f-409ef1f3fd91 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling An audio- visual corpus for speech perception and automatic speech recog- nition,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 939e9ab7-2f22-416b-8828-58a7b302dde1 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Flow- based unconstrained lip to speech generation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a14423a1-dcf1-45af-8248-7df01ac594da · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Let there be sound: Recon- structing high quality speech from silent videos,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 12590acf-fc47-4f56-bde0-69b8355a6598 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Lip to speech synthesis with visual context attentional gan,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e5d87d50-5322-49b3-8b42-01c5a04cb33c · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling End-to-end video-to-speech synthesis using gener- ative adversarial networks,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5a7e119c-0cf9-4212-881a-cdc5ec7dc7e1 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Tcd-timit: An audio-visual corpus of continuous speech,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7aa21951-6345-4de6-8623-56c6f6075bdc · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Lipvoicer: Generating speech from silent videos guided by lip reading,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8cb68e3a-1fd5-4981-a86b-4bddec5dea5e · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Learning individual speaking styles for accurate lip to speech synthesis,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ba66d99b-65b1-430c-89d0-c03af3704cfd · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Revise: Self-supervised speech resynthesis with visual input for univer- sal and generalized speech regeneration,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 69278784-47c7-453e-a891-8d07d690fb45 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Intelligible lip-to-speech synthe- sis with speech units,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1b1b669b-876a-4acf-b18f-b553bbbf7346 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Uni-dubbing: Zero-shot speech synthe- sis from visual articulation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5ae47607-b2d3-45b9-bbdd-cb089a3c1398 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Lip-to-speech synthesis in the wild with multi-task learning,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0a94b2e8-b246-455e-be65-cd0701c47e14 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling MultiVerse: Efficient and expressive zero-shot multi-task text-to-speech,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation aeb6d7a5-5311-478a-b337-e08c1dd8598e · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Towards accurate lip-to-speech synthesis in-the-wild,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e6f5ef46-6659-41b6-ad93-fbd07d562055 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embed- ding ,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16226863-87a8-48ca-91d2-df9c895c5109 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling The source–filter theory of speech,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4bf9742b-496b-4a06-967f-7eae7023f0a0 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Fant,Acoustic theory of speech production,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e3f49fac-acb4-4224-a430-1228f71dab60 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Learning pronunciation from a foreign language in speech synthesis networks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08588f09-25f7-4c70-839c-e21b84f64bfc · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Fastpitchfor- mant: Source-filter based decomposed modeling for speech syn- thesis,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3a36e45a-99bf-457a-a4dd-989f8aecfb3a · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Deep audio-visual speech recognition,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d377d047-9ca5-4eb2-94c8-7c1839a2cd8a · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Clova baseline system for the voxceleb speaker recognition challenge 2020,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 580f29b4-8fee-4de9-9cf5-e68cb316e3af · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Svts: Scalable video-to-speech synthe- sis,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2f5f3524-fde5-4c11-aa65-0ebf158c048b · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling SECS scores are computed viaResemblyzer2, while WER is derived from Auto-A VSR [22]
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0b73cfd3-47bd-4fe8-87e2-5abed2767698 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Fastspeech 2: Fast and high-quality end-to-end text to speech,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 70147acb-db85-4742-9d2f-38abdbbae88f · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Learning audio-visual speech representation by masked multimodal cluster prediction,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ae494202-97d9-4c93-918f-92eef79ed1af · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Auto-avsr: Audio-visual speech recognition with automatic labels,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 55211103-8a43-46bb-b29d-a37c082b9928 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Conformer: Convolution-augmented transformer for speech recognition,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac304ed-52b1-4e77-89f7-84f1048d182f · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling LRS3-TED: a large-scale dataset for visual speech recognition
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a88cb763-94ef-4bbd-8982-eae46c053ccc · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Utmos: Utokyo-sarulab system for voicemos challenge 2022,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f2858f7-cf06-4743-b941-09eec760ae53 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c69e97d-8e43-45b7-a7b7-dec921497fd8 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling V2c: Visual voice cloning,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a890cb67-21a3-48f4-8b45-c3a739fb3ee1 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling MediaPipe: A Framework for Building Perception Pipelines
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34789cbc-584f-4bd5-ba4c-82fde38d268d · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Available: https://www.amazon.com/ Acoustic-Production-Description-Analysis-Contemporary/dp/ 9027916004
Reference 1960
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4701dbe3-f101-4bd3-aeb1-a733f740a460 · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf85582c-7e86-4658-a5b5-f5e2cdf7c42d · outbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a1b0569-55db-4704-b26b-f58a996c0814 · inbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.