Pith. sign in

Paper Citation Record · LEDGER

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2505.22024.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22024 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:15.460678Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:13.022735Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:21:15.811155Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed1c5e6b-449f-4e0f-becd-290f036641e4 · outbound

This paper cites L2S holds great potential across various domains, from enhancing communication in noisy environments to providing assistive technologies for individuals with aphonia.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling L2S holds great potential across various domains, from enhancing communication in noisy environments to providing assistive technologies for individuals with aphonia

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:21.782699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:12.974179Z digest=sha256:0b54564df3f9dcf077629dd0f4c400ad259843f5c6d5a5872024f96045045367

Observation 9a1b0569-55db-4704-b26b-f58a996c0814 · outbound

This paper cites RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:21:15.863647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.022735Z digest=sha256:fed99d9704abd753246c707b18fec84a75774da838b26b5ab82278da4e1dd06a

Observation 112078a0-0eb7-4c59-9ab8-6ed09a77ef0f · outbound

This paper cites Datasets LRS2-BBC[25] is an English audio-visual dataset from BBC programs, comprising over 220 hours of video.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Datasets LRS2-BBC[25] is an English audio-visual dataset from BBC programs, comprising over 220 hours of video

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:21.506731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.105023Z digest=sha256:87f4f35fa483f1425e192ef9d90d669cdd9b17ff7929ef3a8bd9d56ce8ddf0b6

Observation 94b33b3c-1bc4-4513-9272-94e998abd5bb · outbound

This paper cites an unresolved cited work.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Unresolved cited work

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:21:20.982518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.304657Z digest=sha256:9a238b74d68e7551e40a21b663c047b1ec9927b0080da56044fdec537e72bbf9

Observation ac18ebce-cb78-4f5a-9a3f-5ec3710a5796 · outbound

This paper cites Grounded in source-filter theory, RESOUND separates speech generation into acoustic and semantic branches, capturing prosody and linguistic con- tent.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Grounded in source-filter theory, RESOUND separates speech generation into acoustic and semantic branches, capturing prosody and linguistic con- tent

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:20.884251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.354854Z digest=sha256:110973d5b4301cc02f688303bb4855bb70d9546e76d91fe5478f8335103c064b

Observation 9fcaeeb4-fdd8-4bec-9b4f-409ef1f3fd91 · outbound

This paper cites An audio- visual corpus for speech perception and automatic speech recog- nition,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling An audio- visual corpus for speech perception and automatic speech recog- nition,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:13.745946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:13.745946Z digest=sha256:cf4faf946e6566e5e9936dd417e8d4b7658f6e465fffd7d14f5bfe49b2a5cc83

Observation 939e9ab7-2f22-416b-8828-58a7b302dde1 · outbound

This paper cites Flow- based unconstrained lip to speech generation,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Flow- based unconstrained lip to speech generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:20.690179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.439345Z digest=sha256:7ea6534885bf350dfed2208f8f83ed37a717eebb13148219a9bacee0ab71673c

Observation a14423a1-dcf1-45af-8248-7df01ac594da · outbound

This paper cites Let there be sound: Recon- structing high quality speech from silent videos,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Let there be sound: Recon- structing high quality speech from silent videos,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:20.485338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.530458Z digest=sha256:f955b15445e5805be9ccaeb1e2631fc3cec274f22706a4b173957f922c694b41

Observation 12590acf-fc47-4f56-bde0-69b8355a6598 · outbound

This paper cites Lip to speech synthesis with visual context attentional gan,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Lip to speech synthesis with visual context attentional gan,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:20.160550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.594360Z digest=sha256:9dc170283dfe27c307158b1134a2156b94c2b312821a7f3c6b8578bdc52d9206

Observation e5d87d50-5322-49b3-8b42-01c5a04cb33c · outbound

This paper cites End-to-end video-to-speech synthesis using gener- ative adversarial networks,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling End-to-end video-to-speech synthesis using gener- ative adversarial networks,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.969415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.643137Z digest=sha256:a00c1d6fd38cf457c3a745fca6b0d6566cb34df8982acfb0f0ad7eadb1d7eb9b

Observation 5a7e119c-0cf9-4212-881a-cdc5ec7dc7e1 · outbound

This paper cites Tcd-timit: An audio-visual corpus of continuous speech,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Tcd-timit: An audio-visual corpus of continuous speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.779378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.687135Z digest=sha256:ecb9363ad4cfa6604ad57cdf6df6548f438fb8d650f52f5abe74912e9945b50a

Observation 7aa21951-6345-4de6-8623-56c6f6075bdc · outbound

This paper cites Lipvoicer: Generating speech from silent videos guided by lip reading,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Lipvoicer: Generating speech from silent videos guided by lip reading,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:18.541878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.114697Z digest=sha256:05fa845963e31036c332e4631dadc4fb24b3c3b6c8f335d992f07e0ed28b7014

Observation 8cb68e3a-1fd5-4981-a86b-4bddec5dea5e · outbound

This paper cites Learning individual speaking styles for accurate lip to speech synthesis,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Learning individual speaking styles for accurate lip to speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.586523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.827137Z digest=sha256:2909a873836db194c9c269ff5d742aa1b12df7d45a9030444b4480c98d1e63e1

Observation ba66d99b-65b1-430c-89d0-c03af3704cfd · outbound

This paper cites Revise: Self-supervised speech resynthesis with visual input for univer- sal and generalized speech regeneration,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Revise: Self-supervised speech resynthesis with visual input for univer- sal and generalized speech regeneration,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.425846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.913971Z digest=sha256:72862a2a575df1e3cd902b78afb4a5c1e7559cedbc5017af51fb1fdca413ff52

Observation 69278784-47c7-453e-a891-8d07d690fb45 · outbound

This paper cites Intelligible lip-to-speech synthe- sis with speech units,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Intelligible lip-to-speech synthe- sis with speech units,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.256905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.945608Z digest=sha256:2441b2f993050da46e95911a225fcd95b455ccf1d5db53b8c39f0c559ba24fe2

Observation 1b1b669b-876a-4acf-b18f-b553bbbf7346 · outbound

This paper cites Uni-dubbing: Zero-shot speech synthe- sis from visual articulation,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Uni-dubbing: Zero-shot speech synthe- sis from visual articulation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:19.082692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.981558Z digest=sha256:bde267b87eba469782a6e000479b7624d627e058ef7696ceb53fca785cb84688

Observation 5ae47607-b2d3-45b9-bbdd-cb089a3c1398 · outbound

This paper cites Lip-to-speech synthesis in the wild with multi-task learning,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Lip-to-speech synthesis in the wild with multi-task learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:18.875851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.048257Z digest=sha256:548f416367b209a12ad9630e0adbfe26215c665a0cc7c2f7f4a09ef008aecfa7

Observation 0a94b2e8-b246-455e-be65-cd0701c47e14 · outbound

This paper cites MultiVerse: Efficient and expressive zero-shot multi-task text-to-speech,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling MultiVerse: Efficient and expressive zero-shot multi-task text-to-speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:17.532690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.535892Z digest=sha256:7b2f8756d80f98bfada025e49449a51fb1b6b1d1fcd9c0de04984bbcdb93734d

Observation aeb6d7a5-5311-478a-b337-e08c1dd8598e · outbound

This paper cites Towards accurate lip-to-speech synthesis in-the-wild,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Towards accurate lip-to-speech synthesis in-the-wild,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:18.365679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.221956Z digest=sha256:5f4a72b1b8f856d0ae5cfcd73aa6e9da82432ddc692975769c90d3d355e3314d

Observation e6f5ef46-6659-41b6-ad93-fbd07d562055 · outbound

This paper cites DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embed- ding ,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embed- ding ,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:14.257768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:14.257768Z digest=sha256:1557f9e861abc1088facb0e102c99c7f76a93be40be7db5993eb206a1a3bcf0f

Observation 16226863-87a8-48ca-91d2-df9c895c5109 · outbound

This paper cites The source–filter theory of speech,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling The source–filter theory of speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:18.211795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.317395Z digest=sha256:a317015357800c3017ed02f56f5139d8f01d388aeccd9d002f87777c0ad8cebe

Observation 4bf9742b-496b-4a06-967f-7eae7023f0a0 · outbound

This paper cites Fant,Acoustic theory of speech production,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Fant,Acoustic theory of speech production,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:18.037421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.377177Z digest=sha256:06dae725ab7509e0e3e00be7be0384d3a32a9e0fe500a329f80e833ba3493e78

Observation e3f49fac-acb4-4224-a430-1228f71dab60 · outbound

This paper cites Learning pronunciation from a foreign language in speech synthesis networks.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Learning pronunciation from a foreign language in speech synthesis networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:14.918183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:14.918183Z digest=sha256:3712c3299d2724f9fe48e9e133db5a944439712387e59d2e845278b89d31633c

Observation 08588f09-25f7-4c70-839c-e21b84f64bfc · outbound

This paper cites Fastpitchfor- mant: Source-filter based decomposed modeling for speech syn- thesis,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Fastpitchfor- mant: Source-filter based decomposed modeling for speech syn- thesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:17.675755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.503233Z digest=sha256:a911d00055fffbf230563948f205380367e745c444cc02d9f0c0a80796038204

Observation 3a36e45a-99bf-457a-a4dd-989f8aecfb3a · outbound

This paper cites Deep audio-visual speech recognition,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Deep audio-visual speech recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.079405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.079405Z digest=sha256:9163341d5d9a12224d1a0384dd47dd754f47fd6c38d609754779fca4e0e0de8b

Observation d377d047-9ca5-4eb2-94c8-7c1839a2cd8a · outbound

This paper cites Clova baseline system for the voxceleb speaker recognition challenge 2020,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Clova baseline system for the voxceleb speaker recognition challenge 2020,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:17.309864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.567438Z digest=sha256:6ebeb20d72f7c9fcc34ed37caa3b4d993c3a0311a891ff503b5daafafb2d818c

Observation 580f29b4-8fee-4de9-9cf5-e68cb316e3af · outbound

This paper cites Svts: Scalable video-to-speech synthe- sis,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Svts: Scalable video-to-speech synthe- sis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:16.263092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:15.194847Z digest=sha256:cba42846e4b58c9f3f86c87ffd16259c1dda0ba6c6f0cfe2ff93cf1f262fa8cf

Observation 2f5f3524-fde5-4c11-aa65-0ebf158c048b · outbound

This paper cites SECS scores are computed viaResemblyzer2, while WER is derived from Auto-A VSR [22].

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling SECS scores are computed viaResemblyzer2, while WER is derived from Auto-A VSR [22]

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:21.255056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.194903Z digest=sha256:39adfffb5551cb1b24569e242648fcf6744b95182aa1bae0ad596de16aa718a8

Observation 0b73cfd3-47bd-4fe8-87e2-5abed2767698 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:17.012995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.710994Z digest=sha256:f9c6b2a55a69b16b82a25f6a5282519cfbda30f42d3e9cf2df57c621ff81a70a

Observation 70147acb-db85-4742-9d2f-38abdbbae88f · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Learning audio-visual speech representation by masked multimodal cluster prediction,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:16.763456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.826333Z digest=sha256:575c57a787fc4a654d4fbfbe5987a7c737f36983b7832441468b233fc73e04ea

Observation ae494202-97d9-4c93-918f-92eef79ed1af · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Auto-avsr: Audio-visual speech recognition with automatic labels,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:16.525923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.856745Z digest=sha256:ec825cabf1cc3c9f21d8f7ac421ec9d29e70ccd1cfaf617e457d63f9acfe9182

Observation 55211103-8a43-46bb-b29d-a37c082b9928 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Conformer: Convolution-augmented transformer for speech recognition,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.011648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.011648Z digest=sha256:29768731e16b0fa94a268535c31937e5bcef0a56134d1bb18b680ddfb7bcf038

Observation 7ac304ed-52b1-4e77-89f7-84f1048d182f · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling LRS3-TED: a large-scale dataset for visual speech recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.124650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.124650Z digest=sha256:6e60458467e35c3f91fc9be1733520fff7d54a21cd5cf920687f8c2a7eed6c76

Observation a88cb763-94ef-4bbd-8982-eae46c053ccc · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Utmos: Utokyo-sarulab system for voicemos challenge 2022,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.289377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.289377Z digest=sha256:73a9398abd58866eddbbe0a5324761d6a4e581dcbba89f8b4c66c96c29667706

Observation 7f2858f7-cf06-4743-b941-09eec760ae53 · outbound

This paper cites An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.365940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.365940Z digest=sha256:3e2425dc811421652e3f026f7e17838012b5f631d1871a258de6f00008c3e487

Observation 0c69e97d-8e43-45b7-a7b7-dec921497fd8 · outbound

This paper cites V2c: Visual voice cloning,.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling V2c: Visual voice cloning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:16.083512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:15.419786Z digest=sha256:397816e9f1f652db2263c659d77285e0d1826b72494bc9700aa47d82c4eb9712

Observation a890cb67-21a3-48f4-8b45-c3a739fb3ee1 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling MediaPipe: A Framework for Building Perception Pipelines

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:15.460678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:15.460678Z digest=sha256:e7455e332a3e4bbccda347374942d13a84825e9779ea78d1789393db31f6783c

Observation 34789cbc-584f-4bd5-ba4c-82fde38d268d · outbound

This paper cites Available: https://www.amazon.com/ Acoustic-Production-Description-Analysis-Contemporary/dp/ 9027916004.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Available: https://www.amazon.com/ Acoustic-Production-Description-Analysis-Contemporary/dp/ 9027916004

Reference 1960

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:21:17.826042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:14.451358Z digest=sha256:68291ccbf5cd19f818c6bac03d9e7fdf5d21f97506517bfc7dccfc3641cefdbe

Observation 4701dbe3-f101-4bd3-aeb1-a733f740a460 · outbound

This paper cites Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:14.602553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:14.602553Z digest=sha256:fb58794d80f63ea4f45954b38d23ed23c12c55ccb4bae94607ab7578f1e85f3b

Observation bf85582c-7e86-4658-a5b5-f5e2cdf7c42d · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:14.783720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:14.783720Z digest=sha256:5de02e1c98e3fc50f59889e09de234638420f61bf2cefbeed056b4494a0abb5d

Pith citing papers

Observation 9a1b0569-55db-4704-b26b-f58a996c0814 · inbound

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling cites this paper.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:21:15.863647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:21:13.022735Z digest=sha256:fed99d9704abd753246c707b18fec84a75774da838b26b5ab82278da4e1dd06a