Pith. sign in

Paper Citation Record · LEDGER

MuteSwap: Visual-informed Silent Video Identity Conversion

As of 15 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2507.00498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00498 v3

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:53.157881Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact6
  • verified fuzzy3
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9291ee37-5d5e-4583-9f71-c7208b032c31 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

MuteSwap: Visual-informed Silent Video Identity Conversion , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.147531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.147531Z digest=sha256:ce5a2657c18483a5f1ce0bcaa158eaa8b240e18720cfd0ddc4884cc3eb4a5cce

Observation 5706bf15-168a-4b5c-b8d3-606e51028c32 · outbound

This paper cites write newline.

MuteSwap: Visual-informed Silent Video Identity Conversion write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.171962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.171962Z digest=sha256:dc634580a06a4be5cbefd37a4fb3dd9e898890775d5dbb71fe55be2e6a0d7b3e

Observation 9d74eef3-79db-482a-95fc-45a46764d337 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

MuteSwap: Visual-informed Silent Video Identity Conversion LRS3-TED: a large-scale dataset for visual speech recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.211857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.211857Z digest=sha256:4378a9bdee25b138042a02c9de88651578541563b2be0092d0229b66bf26d0bf

Observation fe53fff1-4baf-44a5-878b-030fb9ff06e0 · outbound

This paper cites CLUB: A Contrastive Log-ratio Upper Bound of Mutual Information.

MuteSwap: Visual-informed Silent Video Identity Conversion CLUB: A Contrastive Log-ratio Upper Bound of Mutual Information

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.154005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.301107Z digest=sha256:f372443e319137be9f0aff70b2b405eb45031b224cdcf7c6fa8fb152d22ceb50

Observation f1bd61be-3db6-4f8d-9f3e-756e67d0be9a · outbound

This paper cites DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embedding.

MuteSwap: Visual-informed Silent Video Identity Conversion DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embedding

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.028479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.343362Z digest=sha256:d3ed6becea9c2c88ec05bf4bc3e7027e1a3f15b8dcd04590e7986065dd8e07f5

Observation d0dcf465-cbdc-4f90-8470-770e0df21ade · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units.

MuteSwap: Visual-informed Silent Video Identity Conversion Intelligible Lip-to-Speech Synthesis with Speech Units

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.402047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.402047Z digest=sha256:3f984ae22fa9ac20e2f053b88374dcc8bb3739c967cfbec58d1b9b2065c952fb

Observation cdd914df-a71d-42fe-a688-a570af45917b · outbound

This paper cites S.; Nagrani, A.; and Zisserman, A.

MuteSwap: Visual-informed Silent Video Identity Conversion S.; Nagrani, A.; and Zisserman, A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.919670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.480178Z digest=sha256:61da45ee8b7614d82c9b4dd22e97a402f6ec50f85e6f255b418a5617040eaaff

Observation 9c492056-dde3-4364-b09b-f71204d78236 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

MuteSwap: Visual-informed Silent Video Identity Conversion Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.504954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.504954Z digest=sha256:0e25b9bc0586e594a255c5868ae19ddfc816f271c0b599eb86fba02d6ccb8e50

Observation cd2426c6-f426-4597-933d-b22d65292576 · outbound

This paper cites PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association.

MuteSwap: Visual-informed Silent Video Identity Conversion PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:19:53.919077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.555712Z digest=sha256:a2f17600c140a8ac9f5770d58af83d9f327339407d923706ab5d77178b3a6076

Observation 56aaea29-bba8-449d-a66d-fb3c40cdc900 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.720531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.633979Z digest=sha256:5772684dcb3e881fd6ca660b43815003b5636f54335bef277a225c7b4aa90917

Observation 3dd5f74d-27b0-4753-bda5-0a530fe82af2 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.479759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.670519Z digest=sha256:948fb788a9b5ae9eac0e2ce249c9203863d018669f7fbda52473f892bcf85982

Observation 72acfcf5-f43a-4759-97bd-069837665ac5 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.294575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.730095Z digest=sha256:bed302b1a479759b616dddd51601314a3ccc0801c9243ff34dcf64e5b87b3dbe

Observation ead1e5f4-0116-4f3a-9204-a73e4d863fef · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.

MuteSwap: Visual-informed Silent Video Identity Conversion HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.800306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.800306Z digest=sha256:260dc273a9c40b2f94a18df4c0db72a8795bcfc62ea7dbd37b4aa33cd5cd705e

Observation d44b3b0f-7301-4d36-8f68-2b9e7b233f15 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.104956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.840171Z digest=sha256:d68256fa7eb133fe4f14eac39eaf93702a2e04538d2676b7900ce88e2e07598c

Observation 56964d59-c821-4aad-80d6-d7417c6160e2 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:55.887753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.901215Z digest=sha256:ee5eb52546839d7e0f10aba173def9894349249923cb85f8c418273a001ff044

Observation 3f95c228-93b5-4a97-a1a6-c4029cf95e2f · outbound

This paper cites DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility.

MuteSwap: Visual-informed Silent Video Identity Conversion DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.799591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.982678Z digest=sha256:88e5c91962a2723cd50682acf2a51a4a1f56b7c484b69f0808f9482a5bc7396a

Observation b0ef3fd9-2231-4fd7-81ed-c7e16396d90c · outbound

This paper cites Decoupled Weight Decay Regularization.

MuteSwap: Visual-informed Silent Video Identity Conversion Decoupled Weight Decay Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.026803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.026803Z digest=sha256:921d48539e13ff0e0d4bb803b2a784b3bf883c014f902ad9dd6fa24c9d23a3f8

Observation 42454f16-fcba-4cd1-8889-8dc3634c7aaa · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:55.650955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.102249Z digest=sha256:34cfa90ef9676909156e70f08449492a40ff747e1c1491c762f76df25bdf0c7c

Observation 2349b402-a25c-4f30-8019-de0503723d04 · outbound

This paper cites SVTS: Scalable Video-to-Speech Synthesis.

MuteSwap: Visual-informed Silent Video Identity Conversion SVTS: Scalable Video-to-Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.154124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.154124Z digest=sha256:276a73ce0a740ed48b9c7254a27b42ecf456e646a389bfbf10d1d5f3fc2cb5ef

Observation 1ddb64d0-3d9b-4bab-bbd4-67d91f2768f6 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:55.361477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.238451Z digest=sha256:f26e135b49f4779998bb698ec6923852e0038e020893a53497cc78e6da3db7ef

Observation 8c51ab8f-2738-4bce-b9e7-13c903ae60b7 · outbound

This paper cites Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:19:53.678456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.300850Z digest=sha256:82a0825995b0f304a84fa45ce0a4aee6ad821ffc6f7bd54a933b6ca74ee92862

Observation aa72b93c-a3f2-4fc3-b831-72dac0197c3a · outbound

This paper cites R.; Mukhopadhyay, R.; Namboodiri, V.

MuteSwap: Visual-informed Silent Video Identity Conversion R.; Mukhopadhyay, R.; Namboodiri, V

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.151833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.373040Z digest=sha256:b85ebbe0e4d0d6f4a870e40de98b138ed3a3fb28aded3f44acab24f409e79543

Observation fc331158-e15a-4a77-a0e5-f0ce5bbd55f1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Transferable Visual Models From Natural Language Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.448733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.448733Z digest=sha256:5b11f0c29fbbc0c18731699ac6d2ce6a0e570f0f5d99479daff49b4cecb31c82

Observation 74b17056-4bc1-417d-8f27-8f6839385594 · outbound

This paper cites Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion.

MuteSwap: Visual-informed Silent Video Identity Conversion Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.571796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.494181Z digest=sha256:a85255f583cd0d7cad764d6b91f390a4c06da6c1f4b67c7b9b46fee65f48f1c4

Observation e1c57262-d27c-4f3f-85df-687f200d9f9c · outbound

This paper cites Fusion and Orthogonal Projection for Improved Face-Voice Association.

MuteSwap: Visual-informed Silent Video Identity Conversion Fusion and Orthogonal Projection for Improved Face-Voice Association

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.437781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.564281Z digest=sha256:7620d42553ad4ec298811e1f260106c91e5891d45396697b8437af515da18d2e

Observation b0cc4085-4273-41c9-9173-25cbc1c6259d · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.976495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.620787Z digest=sha256:f65ec1430a40ab3d2288bb83dfa71cdcea0fb802f82b3082e01bbf3ae1f48f31

Observation 3fcc39cd-9f1e-49b8-8764-0adcdd1345be · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.685389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.685389Z digest=sha256:87f40ced5ca5d0e1bdb94a85d26304f5ef2eddc70ea0e012d79889825e02b91c

Observation 8cb0660a-527a-42db-9ce8-43d9d600d0a5 · outbound

This paper cites Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.752908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.752908Z digest=sha256:0b4298fe340a356cb7dbdefd44639a2dab8d0df02ab182e3b1808d8d150f467a

Observation 883da9ac-52a1-46e5-a25f-59b1be6e6c70 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

MuteSwap: Visual-informed Silent Video Identity Conversion MUSAN: A Music, Speech, and Noise Corpus

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.819201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.819201Z digest=sha256:9475440e0765ce8a137426ee11eff888ad883ae58ec9b41faf6689671db1b7c9

Observation d725aa85-d2f9-4f6f-98c3-7f58e0c27f87 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.737554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.880819Z digest=sha256:fdf52c18d17a421af03b497d3f37dc1303e8728ca91041c2d2de6ac21f4d7dc0

Observation 867fd800-a343-4e75-b653-fd511c8127b0 · outbound

This paper cites T.; Chen, X.; Liu, X.; and Meng, H.

MuteSwap: Visual-informed Silent Video Identity Conversion T.; Chen, X.; Liu, X.; and Meng, H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.506179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.983030Z digest=sha256:1a40ff6e0288d20517dddb9ea18b11009cc7a54a3f7a31f07cc10df75a6ab939

Observation 0ff7a1f0-01db-4f74-b9cc-9a39661c1e70 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.402901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:53.056506Z digest=sha256:2f0a25293d93426e820f38d11e800d46bf15d1243f368463c7ee877c4ebc7826

Observation ef9ea61d-c62b-4b58-aeae-317c2391e166 · outbound

This paper cites LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading.

MuteSwap: Visual-informed Silent Video Identity Conversion LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.314800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:53.084783Z digest=sha256:f3118212dbe3f6fa558bd4e0fcc5a019b8f528573c807d7cef85c99b4cba24c8

Observation 9e976be1-c64c-4263-a750-d09811322450 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.271812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:53.157881Z digest=sha256:67cbbd80e9def1ab77d7fd83c51f62216739571a065be164ac687c2a9a74b6b4

Pith citing papers

No inbound Pith citation observations are available.