Pith. sign in

Paper Citation Record · LEDGER

MuteSwap: Visual-informed Silent Video Identity Conversion

As of 16 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2507.00498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00498 v3

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:53.157881Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact6
  • verified fuzzy3
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9291ee37-5d5e-4583-9f71-c7208b032c31 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

MuteSwap: Visual-informed Silent Video Identity Conversion , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.147531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.147531Z digest=sha256:688a01e264c491393c907eab3f08d8449fa64c446e1a2c870d8293551bd1ca21

Observation 5706bf15-168a-4b5c-b8d3-606e51028c32 · outbound

This paper cites write newline.

MuteSwap: Visual-informed Silent Video Identity Conversion write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.171962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.171962Z digest=sha256:0bdbaf37b543be8f00d659ba9b34489a2768a332931b253b3c6be0d0ffd3bda3

Observation 9d74eef3-79db-482a-95fc-45a46764d337 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

MuteSwap: Visual-informed Silent Video Identity Conversion LRS3-TED: a large-scale dataset for visual speech recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.211857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.211857Z digest=sha256:0b46fa6e781a58ee75b222a6e694de529fb795afda26b8ce64fa3fc7867fc711

Observation fe53fff1-4baf-44a5-878b-030fb9ff06e0 · outbound

This paper cites CLUB: A Contrastive Log-ratio Upper Bound of Mutual Information.

MuteSwap: Visual-informed Silent Video Identity Conversion CLUB: A Contrastive Log-ratio Upper Bound of Mutual Information

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.154005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.301107Z digest=sha256:f502209a89693c2dce404535465b6d552a6fb37724d1b146dcb20d76311906ea

Observation f1bd61be-3db6-4f8d-9f3e-756e67d0be9a · outbound

This paper cites DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embedding.

MuteSwap: Visual-informed Silent Video Identity Conversion DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embedding

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.028479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.343362Z digest=sha256:3c84764e2b79b10072e944eced589ea843f7451e2f7cd02c15f4bda5418d03a0

Observation d0dcf465-cbdc-4f90-8470-770e0df21ade · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units.

MuteSwap: Visual-informed Silent Video Identity Conversion Intelligible Lip-to-Speech Synthesis with Speech Units

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.402047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.402047Z digest=sha256:f9a026b3113b51aae79d659d6b9d4abadbb38b2ac2b9c3e13b4dcdce4038e056

Observation cdd914df-a71d-42fe-a688-a570af45917b · outbound

This paper cites S.; Nagrani, A.; and Zisserman, A.

MuteSwap: Visual-informed Silent Video Identity Conversion S.; Nagrani, A.; and Zisserman, A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.919670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.480178Z digest=sha256:98c570544ea02e7eb889a1375c447e250ee9dad9dd4c7400e7c027572c3bb5ea

Observation 9c492056-dde3-4364-b09b-f71204d78236 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

MuteSwap: Visual-informed Silent Video Identity Conversion Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.504954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.504954Z digest=sha256:3b051f449fe01968532a1b02944af99ed50139190cee457faac69ad8496149b9

Observation cd2426c6-f426-4597-933d-b22d65292576 · outbound

This paper cites PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association.

MuteSwap: Visual-informed Silent Video Identity Conversion PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:19:53.919077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.555712Z digest=sha256:885b5d1015e471de946678264e575b8983e2a2e0aa69b42afbf9cd4614790da4

Observation 56aaea29-bba8-449d-a66d-fb3c40cdc900 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.720531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.633979Z digest=sha256:1536f41930a0b82c0335a011a2cfb6f6ed7e9d9e9fd24ec48360c7556c41115b

Observation 3dd5f74d-27b0-4753-bda5-0a530fe82af2 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.479759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.670519Z digest=sha256:d39ed66e5fb28e9b93f5d34e3ece9ee9bd49e879f448604585caf7a3044425ef

Observation 72acfcf5-f43a-4759-97bd-069837665ac5 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.294575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.730095Z digest=sha256:d894df0f7c1c3e2abae699295730b98918e0e5ca00e2c30e59d6114a6e04e2d2

Observation ead1e5f4-0116-4f3a-9204-a73e4d863fef · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.

MuteSwap: Visual-informed Silent Video Identity Conversion HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.800306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.800306Z digest=sha256:18e7b819f8e46772d0d9d258d6bafb9bac5657e4698178372b8a8cb296ebe9a4

Observation d44b3b0f-7301-4d36-8f68-2b9e7b233f15 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:56.104956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.840171Z digest=sha256:13fdd39484140e25ebd7c9e9671873e4dbb6faa03ebf7126745d49eb5e2c74a6

Observation 56964d59-c821-4aad-80d6-d7417c6160e2 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:55.887753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.901215Z digest=sha256:c0f16ce5b931ff6d7120b82699929b6a21facb25de0db73fcfd23a61738fa0a2

Observation 3f95c228-93b5-4a97-a1a6-c4029cf95e2f · outbound

This paper cites DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility.

MuteSwap: Visual-informed Silent Video Identity Conversion DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.799591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:51.982678Z digest=sha256:8962c16d9b2fbf3b42fc58d6630eefd0e8a71ec6c3f4f744ea74e708e5b84c90

Observation b0ef3fd9-2231-4fd7-81ed-c7e16396d90c · outbound

This paper cites Decoupled Weight Decay Regularization.

MuteSwap: Visual-informed Silent Video Identity Conversion Decoupled Weight Decay Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.026803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.026803Z digest=sha256:651b64aaf9e6312b5c85fd759410fe20cf2e5dfc3148b799d65dd8baba59307d

Observation 42454f16-fcba-4cd1-8889-8dc3634c7aaa · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:55.650955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.102249Z digest=sha256:c55b87120a7c84924bfe7764e7bcb0e47e669fd75a9c55df1b105ee9b2044b49

Observation 2349b402-a25c-4f30-8019-de0503723d04 · outbound

This paper cites SVTS: Scalable Video-to-Speech Synthesis.

MuteSwap: Visual-informed Silent Video Identity Conversion SVTS: Scalable Video-to-Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.154124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.154124Z digest=sha256:d378ba991114386e5618d4280498d27401fcb6a10345668b305034579ee66915

Observation 1ddb64d0-3d9b-4bab-bbd4-67d91f2768f6 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:55.361477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.238451Z digest=sha256:4813c09f26f5b231026db31ede8d7ee0034db41f741afdacd0cd413f34f266fe

Observation 8c51ab8f-2738-4bce-b9e7-13c903ae60b7 · outbound

This paper cites Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:19:53.678456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.300850Z digest=sha256:6c02bcfa3a9303209d5395fb19d700afa57c703bc0d3f5c3607c8e8f475b9379

Observation aa72b93c-a3f2-4fc3-b831-72dac0197c3a · outbound

This paper cites R.; Mukhopadhyay, R.; Namboodiri, V.

MuteSwap: Visual-informed Silent Video Identity Conversion R.; Mukhopadhyay, R.; Namboodiri, V

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.151833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.373040Z digest=sha256:4de15e5e1a1db368d2e9f767613690649f3881f6a985bce730178df287c04d25

Observation fc331158-e15a-4a77-a0e5-f0ce5bbd55f1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Transferable Visual Models From Natural Language Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.448733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.448733Z digest=sha256:8271f3a7a00128e1787e61c44ccb40c7f59cb9277e4cbc80ad2e9bca8028df83

Observation 74b17056-4bc1-417d-8f27-8f6839385594 · outbound

This paper cites Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion.

MuteSwap: Visual-informed Silent Video Identity Conversion Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.571796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.494181Z digest=sha256:7168e79e09676d54fe6bb6803649f951da8e8a2b6162d353bc9b7f78a8b20db4

Observation e1c57262-d27c-4f3f-85df-687f200d9f9c · outbound

This paper cites Fusion and Orthogonal Projection for Improved Face-Voice Association.

MuteSwap: Visual-informed Silent Video Identity Conversion Fusion and Orthogonal Projection for Improved Face-Voice Association

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.437781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.564281Z digest=sha256:9117b0e834ff5f9f3b2b11c8b57d183be90543266917e19f5f4cda0044c11bdb

Observation b0cc4085-4273-41c9-9173-25cbc1c6259d · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.976495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.620787Z digest=sha256:adf2f3151f3188d2def9b1e68d86c7a22e683914db791fd46edd481d84ccd458

Observation 3fcc39cd-9f1e-49b8-8764-0adcdd1345be · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.685389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.685389Z digest=sha256:1de541e86c681a30abedff9ad0895f80e46f37755421f624f1da010005afd067

Observation 8cb0660a-527a-42db-9ce8-43d9d600d0a5 · outbound

This paper cites Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.752908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.752908Z digest=sha256:44e85ef0757e9fba86f22b6a33fd70846756d7a5216a21262a661a46c457310d

Observation 883da9ac-52a1-46e5-a25f-59b1be6e6c70 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

MuteSwap: Visual-informed Silent Video Identity Conversion MUSAN: A Music, Speech, and Noise Corpus

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.819201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.819201Z digest=sha256:ff2aa71d81331b606ee61c556c9022709e332674d4b4e4f7539b54c56ae8b61e

Observation d725aa85-d2f9-4f6f-98c3-7f58e0c27f87 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.737554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.880819Z digest=sha256:8da4927d52e0ce3e7e26c6b6f99919db00d9861c574bda0846b8652ab4cca57b

Observation 867fd800-a343-4e75-b653-fd511c8127b0 · outbound

This paper cites T.; Chen, X.; Liu, X.; and Meng, H.

MuteSwap: Visual-informed Silent Video Identity Conversion T.; Chen, X.; Liu, X.; and Meng, H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.506179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:52.983030Z digest=sha256:359ddd5afe71a850571d5152bbb823c8a93b434835c4f205f70baf891ef5d898

Observation 0ff7a1f0-01db-4f74-b9cc-9a39661c1e70 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.402901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:53.056506Z digest=sha256:7c9966d3a43eaf2c1e715566a85b4c7fd20655c26f90129b13880b80047c048a

Observation ef9ea61d-c62b-4b58-aeae-317c2391e166 · outbound

This paper cites LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading.

MuteSwap: Visual-informed Silent Video Identity Conversion LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:53.314800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:53.084783Z digest=sha256:71c27f692a38d87e46a92c638c812b8190ebf168d9d4d0540dd130fea349480f

Observation 9e976be1-c64c-4263-a750-d09811322450 · outbound

This paper cites an unresolved cited work.

MuteSwap: Visual-informed Silent Video Identity Conversion Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:19:54.271812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T21:19:53.157881Z digest=sha256:342043586bf4bbeebd7cfda0634d7e1d78d2cc995fb8133ac4f6f51bf2fbb69a

Pith citing papers

No inbound Pith citation observations are available.