Pith. sign in

Paper Citation Record · LEDGER

Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2201.02184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2201.02184 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:52:43.269182Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.728084Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 59f410ff-febe-4cbe-b011-09c8fe187403 · inbound

The Sound of Water: Inferring Physical Properties from Pouring Liquids cites this paper.

The Sound of Water: Inferring Physical Properties from Pouring Liquids Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T18:52:43.269182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:52:43.269182Z digest=sha256:2d0044d5b14df771275d48822a97e750419550340eecb972cc5a9b54c96d677f

Observation 9952afd4-d1d5-476f-8dfe-0356fbbcf237 · inbound

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning cites this paper.

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-12T14:04:00.708996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:04:00.708996Z digest=sha256:d5ce3b199650a31887efa44914427820ed7da186dbaf328cb58b5ee6cd5b7190

Observation 7f5c8314-bbe9-4a81-a4d7-f825b5eac8c6 · inbound

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies cites this paper.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T05:38:52.683006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:38:52.683006Z digest=sha256:ee066d5cf1183cfb40cf6ce7f313bff8d2cf5c83703104a439a5c609051f815b

Observation 035624cc-d8cf-4637-ad6c-ba3462df818c · inbound

MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI cites this paper.

MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:43.373283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:29:43.373283Z digest=sha256:2f72037272f8f339a2dbfa988d915cc615e398b1fec35f2d49e4fbdf45089288

Observation 66b601f3-6a97-4dd3-8729-ffc4361e204d · inbound

Negative to Positive Co-learning with Aggressive Modality Dropout cites this paper.

Negative to Positive Co-learning with Aggressive Modality Dropout Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:43:05.194883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:43:05.194883Z digest=sha256:9dce4a61dcffe2a1bbbd8bbaac4776cc684dbfedb904285ce4b24774aec9195b

Observation ac8670da-da16-4c4d-8a25-c7d4dcadcac0 · inbound

Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment cites this paper.

Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T00:35:48.921738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:35:48.921738Z digest=sha256:49d41d63d5d2983a9cbf8312e5f80ba1a33693bc9832567fec4a790eb903156a

Observation 2e50cd52-7bd4-4b48-8349-00991daf0a92 · inbound

Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions cites this paper.

Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T18:58:30.682618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:58:30.682618Z digest=sha256:b05d90602b25636c426349d829438b153e91ea50a4054e7c5cad931964ebb743

Observation e7a1966a-fbb3-4274-88fc-a3c1116e54fb · inbound

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition cites this paper.

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:44.345646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:44.345646Z digest=sha256:02aa078eabaa9416b5699764b68aae4924a1041dbcf6110eeb7a226341ea82fd

Observation fd50aea4-84af-4bc9-8022-d12eb1547d04 · inbound

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models cites this paper.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:50.391701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:50.391701Z digest=sha256:4f1a1daef75cb4f4268a495622cd7f43bc371a06512e105ea4ac6fc81cc79c23

Observation 3fcc39cd-9f1e-49b8-8764-0adcdd1345be · inbound

MuteSwap: Visual-informed Silent Video Identity Conversion cites this paper.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.685389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.685389Z digest=sha256:87f40ced5ca5d0e1bdb94a85d26304f5ef2eddc70ea0e012d79889825e02b91c

Observation 120056b9-0ef1-4f67-994f-127feb7cfcbe · inbound

Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring cites this paper.

Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:14:35.054556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:14:35.054556Z digest=sha256:c8b301f51a137350a5b500ebc4b986f664d100bc5ce6fc69f88e64249b2ce3de

Observation 808bfc5e-bf43-4fc8-9399-8e9f2e1dc2de · inbound

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation cites this paper.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.745331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:00.745331Z digest=sha256:38ca81ffce43af7d01aaa698e9a570df82e0f8d0c0215d81da300af8f7abd8ce

Observation e9117a31-9416-46fa-9714-2941e7c61f58 · inbound

HumanOmni-Speaker: Identifying Who said What and When cites this paper.

HumanOmni-Speaker: Identifying Who said What and When Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.524502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T00:57:12.239390Z digest=sha256:84c508b25af056267bdaacda385eed05ded0eb9d4f28584315cebfa6cbb35cd3

Observation 16998ea3-2045-49cb-80c7-b9342ea01be9 · inbound

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge cites this paper.

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:29.101624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-07T07:49:44.117839Z digest=sha256:a01d47322d8c1d030841e5984b68fd6b1377ac2d955b788e1f95f1e7f34448e9

Observation a8176aad-ec4b-4cb7-b5af-e25c1a36bbb5 · inbound

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge cites this paper.

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T19:21:19.583012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:21:19.583012Z digest=sha256:63037ef41d5cc0962cade85353874b023ca5b1a883bb3af9f85d25dfdf1d6fa0

Observation 709bdce4-0dfc-4a2d-b7a0-eb7505c97883 · inbound

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework cites this paper.

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:51:30.417122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T01:23:27.355803Z digest=sha256:998f27632d23b57f6fa3fe2a88ffb2966c5d7832a08c64ec9ec1d64ce00e496f

Observation caa96176-d73b-41ad-9fca-42357ca0e1d2 · inbound

Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization cites this paper.

Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:20:24.482019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-25T05:19:38.661190Z digest=sha256:af543ad9a57e4f603a82854ba8efe2cf658fbe9a69bfd6969f8037cd8476f2d7

Observation c1ac8eb4-36f1-4f64-a1c0-65dcd4a3abf9 · inbound

Your Multimodal Speech Model Says I Have a Face for Radio cites this paper.

Your Multimodal Speech Model Says I Have a Face for Radio Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:53:14.062161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T07:44:50.169335Z digest=sha256:6851a403e93b0a006a421797876bec49686dddd2b0e500b73c2eb699c9846ee4

Observation 994b35a5-0354-4e72-b202-f8ccd82040fe · inbound

Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography cites this paper.

Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:47:28.354383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T17:49:19.643889Z digest=sha256:d318ea5be6d10e48302a090b46f6d50e5dc4f9091d35bbf7f01356c87f4cdf90

Observation 641792ed-feaf-437a-8aec-8387dd5eddc8 · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.642562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:1c419746c617dd719cd227e1c7496fec21cfbe0eebac02a2b7118d9fb578942d

Observation 68ffcd41-0334-43b0-9014-0b48bc71b45c · inbound

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading cites this paper.

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:47:35.216339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T14:54:08.119264Z digest=sha256:cd2eebbe39e9d2d5ef9014aaf92df12d5b3f4d7dbb55a84d0b432bdac0b53baa

Observation e54e3b54-3554-46c1-b3f0-f9ec27863513 · inbound

Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning cites this paper.

Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.729646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T11:03:05.449385Z digest=sha256:5dca09742b6be80e334a5c34f33cda91a946547745c91054c6f951b63ed5f931

Observation 524a6f4f-0fc6-4ba4-a4ca-47e4a843d284 · inbound

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement cites this paper.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:49:01.368464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:d030935d093a17de0c9abb3d21c7a466e8753cc3c69b476ffd3f3bacf0dfdc31

Observation 47d8d709-c8a7-4fa2-9f76-942b31511c2b · inbound

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection cites this paper.

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T02:09:52.116351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:09:52.116351Z digest=sha256:8e7542d215f2f947397144f5cb0c0012e35704e8843af4b1915e4c18b746e240

Observation 00cdfd15-4ba4-4c9b-bfe7-2a6d6267fb59 · inbound

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE cites this paper.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T05:02:03.514200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:02:03.514200Z digest=sha256:782a7b9c1d003db14d5b8df7b2c5171cc48166a5a29e32681e2eb50450f9b733

Observation f48c3caa-2414-4616-9134-fb48b647a39f · inbound

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE cites this paper.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.387173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.387173Z digest=sha256:1e9c0ef9cf8fed6915c8edac825c582fd7fd9f880e9f59d6a00e68d743b497f7