Pith. sign in

Paper Citation Record · LEDGER

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2505.22024.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22024 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:15.460678Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:13.022735Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:21:15.811155Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed1c5e6b-449f-4e0f-becd-290f036641e4 · outbound

This paper cites L2S holds great potential across various domains, from enhancing communication in noisy environments to providing assistive technologies for individuals with aphonia.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling L2S holds great potential across various domains, from enhancing communication in noisy environments to providing assistive technologies for individuals with aphonia

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:21.782699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:12.974179Z digest=sha256:25b9faed67b1414a3415ef634035515898943efb977266be1b4a296825e52249

Observation 9a1b0569-55db-4704-b26b-f58a996c0814 · outbound

This paper cites RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:21:15.863647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.022735Z digest=sha256:1bab36007ed92d3d97a9af5883cedb12137916f733ce0c5d6d89bfc691252000

Observation 112078a0-0eb7-4c59-9ab8-6ed09a77ef0f · outbound

This paper cites Datasets LRS2-BBC[25] is an English audio-visual dataset from BBC programs, comprising over 220 hours of video.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Datasets LRS2-BBC[25] is an English audio-visual dataset from BBC programs, comprising over 220 hours of video

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:21.506731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.105023Z digest=sha256:18d7b75f00718de7b66c8a7c04685829151ac2223e477251df6c0c2df2677c72

Observation 94b33b3c-1bc4-4513-9272-94e998abd5bb · outbound

This paper cites an unresolved cited work.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Unresolved cited work

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:21:20.982518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.304657Z digest=sha256:4d9d4e5216fc5b1a7847bebd8132a021c16036df572998aaf9981995cbace26e

Observation ac18ebce-cb78-4f5a-9a3f-5ec3710a5796 · outbound

This paper cites Grounded in source-filter theory, RESOUND separates speech generation into acoustic and semantic branches, capturing prosody and linguistic con- tent.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Grounded in source-filter theory, RESOUND separates speech generation into acoustic and semantic branches, capturing prosody and linguistic con- tent

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:20.884251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.354854Z digest=sha256:bccd0a6eae9b99332185ceb91b9461d36847ed7ec397f1a7098eaf56389dcb5b

Observation 9fcaeeb4-fdd8-4bec-9b4f-409ef1f3fd91 · outbound

This paper cites An audio- visual corpus for speech perception and automatic speech recog- nition,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling An audio- visual corpus for speech perception and automatic speech recog- nition,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:13.745946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:13.745946Z digest=sha256:21c02f2168c080e563c23b884ee8747edde833b729d88367d4a5f21992fdf4ba

Observation 939e9ab7-2f22-416b-8828-58a7b302dde1 · outbound

This paper cites Flow- based unconstrained lip to speech generation,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Flow- based unconstrained lip to speech generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:20.690179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.439345Z digest=sha256:b5934565d486535638c9cc654a911fe421f0e37627fb1675e7e06ffcb511e278

Observation a14423a1-dcf1-45af-8248-7df01ac594da · outbound

This paper cites Let there be sound: Recon- structing high quality speech from silent videos,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Let there be sound: Recon- structing high quality speech from silent videos,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:20.485338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.530458Z digest=sha256:7f0ba5dada1b83c97fa1ad162920718c009ac5837f1137156d6df548bec217f7

Observation 12590acf-fc47-4f56-bde0-69b8355a6598 · outbound

This paper cites Lip to speech synthesis with visual context attentional gan,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Lip to speech synthesis with visual context attentional gan,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:20.160550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.594360Z digest=sha256:ca446436909598ed1c26b09ec2e313830e89e309a010b41255ac0ccf9a1c1743

Observation e5d87d50-5322-49b3-8b42-01c5a04cb33c · outbound

This paper cites End-to-end video-to-speech synthesis using gener- ative adversarial networks,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling End-to-end video-to-speech synthesis using gener- ative adversarial networks,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.969415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.643137Z digest=sha256:49475b43ece3aa1a1d8d367e5ae155063eff4c81fa67c02520c4977a043d3db5

Observation 5a7e119c-0cf9-4212-881a-cdc5ec7dc7e1 · outbound

This paper cites Tcd-timit: An audio-visual corpus of continuous speech,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Tcd-timit: An audio-visual corpus of continuous speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.779378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.687135Z digest=sha256:586941f2b6b80e1de5ec78906cc49348d35d369ef21f8b1561696f746c52ae62

Observation 7aa21951-6345-4de6-8623-56c6f6075bdc · outbound

This paper cites Lipvoicer: Generating speech from silent videos guided by lip reading,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Lipvoicer: Generating speech from silent videos guided by lip reading,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:18.541878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.114697Z digest=sha256:bc56b2709e43bd145c0786aa064c1ae5e49696bbf9d5ac98b185f89f26ecebce

Observation 8cb68e3a-1fd5-4981-a86b-4bddec5dea5e · outbound

This paper cites Learning individual speaking styles for accurate lip to speech synthesis,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Learning individual speaking styles for accurate lip to speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.586523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.827137Z digest=sha256:1856a9b79f6fe64c7f7d38574443730dc5861099f7ef9952f1f85c7b3b104ac2

Observation ba66d99b-65b1-430c-89d0-c03af3704cfd · outbound

This paper cites Revise: Self-supervised speech resynthesis with visual input for univer- sal and generalized speech regeneration,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Revise: Self-supervised speech resynthesis with visual input for univer- sal and generalized speech regeneration,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.425846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.913971Z digest=sha256:e75ed4b77223dbf8eae0c9dd67524270e1fbc209be40aac6d4b5eeb104875dd4

Observation 69278784-47c7-453e-a891-8d07d690fb45 · outbound

This paper cites Intelligible lip-to-speech synthe- sis with speech units,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Intelligible lip-to-speech synthe- sis with speech units,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.256905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.945608Z digest=sha256:c4cff7c7364a540ca98127c8cb673afb05ac1e7252841ca2099687b5af59d583

Observation 1b1b669b-876a-4acf-b18f-b553bbbf7346 · outbound

This paper cites Uni-dubbing: Zero-shot speech synthe- sis from visual articulation,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Uni-dubbing: Zero-shot speech synthe- sis from visual articulation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.082692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.981558Z digest=sha256:dd967d7cab16e5e90d424dcdded9a84c93989929333373e836692a1fb0e78b2c

Observation 5ae47607-b2d3-45b9-bbdd-cb089a3c1398 · outbound

This paper cites Lip-to-speech synthesis in the wild with multi-task learning,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Lip-to-speech synthesis in the wild with multi-task learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:18.875851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.048257Z digest=sha256:34c7a4ad79874596b025ba9efe83af8babb01067af8d1a17b1a80972586ee56b

Observation 0a94b2e8-b246-455e-be65-cd0701c47e14 · outbound

This paper cites MultiVerse: Efficient and expressive zero-shot multi-task text-to-speech,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling MultiVerse: Efficient and expressive zero-shot multi-task text-to-speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:17.532690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.535892Z digest=sha256:d689975bf01d37cb26c03f36adbe6ad1cef75d5d1176c66afab76cb66a4a921f

Observation aeb6d7a5-5311-478a-b337-e08c1dd8598e · outbound

This paper cites Towards accurate lip-to-speech synthesis in-the-wild,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Towards accurate lip-to-speech synthesis in-the-wild,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:18.365679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.221956Z digest=sha256:c5cc9524539a174c2394cf73ba53ef89831d500de4040301f421b2f52915792d

Observation e6f5ef46-6659-41b6-ad93-fbd07d562055 · outbound

This paper cites DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embed- ding ,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embed- ding ,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:14.257768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:14.257768Z digest=sha256:3cef334e7d97a2c50876e17f47d82bca1315b7173c877f7f29b0af26424da7e9

Observation 16226863-87a8-48ca-91d2-df9c895c5109 · outbound

This paper cites The source–filter theory of speech,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling The source–filter theory of speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:18.211795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.317395Z digest=sha256:65ca2f0ed8139c5ff7fba1acce5b00b15664c29517c111237699dfe263204cdc

Observation 4bf9742b-496b-4a06-967f-7eae7023f0a0 · outbound

This paper cites Fant,Acoustic theory of speech production,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Fant,Acoustic theory of speech production,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:18.037421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.377177Z digest=sha256:c0929c41e56980580befb0155cce8ac0a3632e6a719310bf53c1108ffb6fcb7f

Observation e3f49fac-acb4-4224-a430-1228f71dab60 · outbound

This paper cites Learning pronunciation from a foreign language in speech synthesis networks.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Learning pronunciation from a foreign language in speech synthesis networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:14.918183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:14.918183Z digest=sha256:368e339cd3989fac38670232160c222a774b3e562f1832e57f5831d05dbc1715

Observation 08588f09-25f7-4c70-839c-e21b84f64bfc · outbound

This paper cites Fastpitchfor- mant: Source-filter based decomposed modeling for speech syn- thesis,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Fastpitchfor- mant: Source-filter based decomposed modeling for speech syn- thesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:17.675755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.503233Z digest=sha256:1d53e81f2e37545e4305189976e9045309c5e4578347c67fb5057bb914c6616d

Observation 3a36e45a-99bf-457a-a4dd-989f8aecfb3a · outbound

This paper cites Deep audio-visual speech recognition,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Deep audio-visual speech recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.079405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.079405Z digest=sha256:97ac8275b10e9b2e45c062089c40ced4c3992e700ee7c7a440f6889e04ffbc79

Observation d377d047-9ca5-4eb2-94c8-7c1839a2cd8a · outbound

This paper cites Clova baseline system for the voxceleb speaker recognition challenge 2020,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Clova baseline system for the voxceleb speaker recognition challenge 2020,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:17.309864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.567438Z digest=sha256:6dc623e5936230f7d7ed9db8e5f6a00bf4ea753a519332e5dedc31a6d5259d3f

Observation 580f29b4-8fee-4de9-9cf5-e68cb316e3af · outbound

This paper cites Svts: Scalable video-to-speech synthe- sis,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Svts: Scalable video-to-speech synthe- sis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:16.263092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:15.194847Z digest=sha256:4dee8219dcf9286ed258ab0b3535ddcefb2403c8e4884343aa97a3583b88b051

Observation 2f5f3524-fde5-4c11-aa65-0ebf158c048b · outbound

This paper cites SECS scores are computed viaResemblyzer2, while WER is derived from Auto-A VSR [22].

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling SECS scores are computed viaResemblyzer2, while WER is derived from Auto-A VSR [22]

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:21.255056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.194903Z digest=sha256:257796299994e66a5bc5241a48c2cc92d83b1d29c6a929121d882d81f2c991af

Observation 0b73cfd3-47bd-4fe8-87e2-5abed2767698 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:17.012995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.710994Z digest=sha256:567823d5a7741a655a24d2d7c942659ae19aa2e77cd0443cd30274a2e203ee7c

Observation 70147acb-db85-4742-9d2f-38abdbbae88f · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Learning audio-visual speech representation by masked multimodal cluster prediction,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:16.763456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.826333Z digest=sha256:284802798c954bf5f4c34a722f880f738f054d355ef2a513fbfc8757f6ac43d0

Observation ae494202-97d9-4c93-918f-92eef79ed1af · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Auto-avsr: Audio-visual speech recognition with automatic labels,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:16.525923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.856745Z digest=sha256:919da206b188db310753a0fa61737420e993bccbf26d15273205aa4b730c514a

Observation 55211103-8a43-46bb-b29d-a37c082b9928 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Conformer: Convolution-augmented transformer for speech recognition,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.011648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.011648Z digest=sha256:2e0fb6478923902d587c1c270f3247104cdd8c9a2833aab761cfdb7b18d37b82

Observation 7ac304ed-52b1-4e77-89f7-84f1048d182f · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling LRS3-TED: a large-scale dataset for visual speech recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.124650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.124650Z digest=sha256:1fd863ea7462f2940c2206d8a8f3a7a75168cf103c1c519f130dd69cf7acb094

Observation a88cb763-94ef-4bbd-8982-eae46c053ccc · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Utmos: Utokyo-sarulab system for voicemos challenge 2022,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.289377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.289377Z digest=sha256:6191b757ab3e8e7a416f84db46f3a5d894d60da1f3b75946d22a61ee31c8146d

Observation 7f2858f7-cf06-4743-b941-09eec760ae53 · outbound

This paper cites An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.365940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.365940Z digest=sha256:13ce1ea7a00a27695bed409a933a3275ea0ab6e92ed6306d92f4a8be18bf8033

Observation 0c69e97d-8e43-45b7-a7b7-dec921497fd8 · outbound

This paper cites V2c: Visual voice cloning,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling V2c: Visual voice cloning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:16.083512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:15.419786Z digest=sha256:6c7cedf9198e0a485d035389a8e2f0809debfd6033cf1462ad6a17078b54bb62

Observation a890cb67-21a3-48f4-8b45-c3a739fb3ee1 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling MediaPipe: A Framework for Building Perception Pipelines

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.460678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.460678Z digest=sha256:fecc9d9de1373674aa5464c8dcbe23ca244f6e995031173c799ea4d910104119

Observation 34789cbc-584f-4bd5-ba4c-82fde38d268d · outbound

This paper cites Available: https://www.amazon.com/ Acoustic-Production-Description-Analysis-Contemporary/dp/ 9027916004.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Available: https://www.amazon.com/ Acoustic-Production-Description-Analysis-Contemporary/dp/ 9027916004

Reference 1960

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:17.826042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:14.451358Z digest=sha256:e4df958612e4cfc86bc2c7f6469c47de0c189d6b3d7bb2a5e0eddf0e933be667

Observation 4701dbe3-f101-4bd3-aeb1-a733f740a460 · outbound

This paper cites Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:14.602553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:14.602553Z digest=sha256:049fe9195f84a255f54971d6a0dcf6d0c517a914f551a3fd835d9c649b051e7c

Observation bf85582c-7e86-4658-a5b5-f5e2cdf7c42d · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:14.783720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:14.783720Z digest=sha256:a357afe3a07d9a773ef33299103faf72c529fa168418fe5522818108b9ff2ff7

Pith citing papers

Observation 9a1b0569-55db-4704-b26b-f58a996c0814 · inbound

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling cites this paper.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:21:15.863647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:21:13.022735Z digest=sha256:1bab36007ed92d3d97a9af5883cedb12137916f733ce0c5d6d89bfc691252000