Pith. sign in

Paper Citation Record · LEDGER

Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2201.02184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2201.02184 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:52:43.269182Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.728084Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 59f410ff-febe-4cbe-b011-09c8fe187403 · inbound

The Sound of Water: Inferring Physical Properties from Pouring Liquids cites this paper.

The Sound of Water: Inferring Physical Properties from Pouring Liquids Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T18:52:43.269182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:52:43.269182Z digest=sha256:cd67ed6d73f2acbe20e3ec3e5ae69893f85d47e5ccef0217175ad7bdbe993840

Observation 9952afd4-d1d5-476f-8dfe-0356fbbcf237 · inbound

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning cites this paper.

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-12T14:04:00.708996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:04:00.708996Z digest=sha256:9d724f74ba7a7892e2e7b65b929c559b71fe03c165304d8324af1d76c81ad456

Observation 7f5c8314-bbe9-4a81-a4d7-f825b5eac8c6 · inbound

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies cites this paper.

SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T05:38:52.683006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:38:52.683006Z digest=sha256:47b53666c02e947eb10a87028756b69c193748a07cf09029a03af1cca37529b4

Observation 035624cc-d8cf-4637-ad6c-ba3462df818c · inbound

MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI cites this paper.

MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:43.373283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:29:43.373283Z digest=sha256:d7306508599daf5bd186cce51e734eb352c4912fb1a7b9fa4af7e1e8a5f654a2

Observation 66b601f3-6a97-4dd3-8729-ffc4361e204d · inbound

Negative to Positive Co-learning with Aggressive Modality Dropout cites this paper.

Negative to Positive Co-learning with Aggressive Modality Dropout Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:43:05.194883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:43:05.194883Z digest=sha256:0c35d52d090248701ce48a281d8dd41fee64bf056b60cbee580712ab0a7386ea

Observation ac8670da-da16-4c4d-8a25-c7d4dcadcac0 · inbound

Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment cites this paper.

Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T00:35:48.921738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:35:48.921738Z digest=sha256:11cdfdc5a7dff1f4606ff4dfafeea2e6f92f751ee6386d2d7aa4b10406533d8f

Observation 2e50cd52-7bd4-4b48-8349-00991daf0a92 · inbound

Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions cites this paper.

Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T18:58:30.682618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:58:30.682618Z digest=sha256:8180f976be1b979a9668840e21a3da15e569f0cf140315fe46cdf978eeeeb333

Observation e7a1966a-fbb3-4274-88fc-a3c1116e54fb · inbound

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition cites this paper.

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:44.345646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:44.345646Z digest=sha256:3b25b83976a7778d257916fb4fa70b18dfcfc06b29df801c40df52b272762f38

Observation fd50aea4-84af-4bc9-8022-d12eb1547d04 · inbound

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models cites this paper.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:50.391701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:50.391701Z digest=sha256:86d62a74cb1e81491204206c01c62265c945129d0b1f42d2cf116a323f39915b

Observation 3fcc39cd-9f1e-49b8-8764-0adcdd1345be · inbound

MuteSwap: Visual-informed Silent Video Identity Conversion cites this paper.

MuteSwap: Visual-informed Silent Video Identity Conversion Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.685389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:52.685389Z digest=sha256:f903a86bb2d64b05376b9b5db6bcf746da22233504108f2c17456d28944460d4

Observation 120056b9-0ef1-4f67-994f-127feb7cfcbe · inbound

Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring cites this paper.

Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:14:35.054556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:14:35.054556Z digest=sha256:8bd09d3d75a929640fe26baebfd77a44f7f28d882a95be8523d45b2cb121ad12

Observation 808bfc5e-bf43-4fc8-9399-8e9f2e1dc2de · inbound

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation cites this paper.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.745331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:00.745331Z digest=sha256:6fbd68ff0b4ec202df655cb89113ef958b9b6e8b7445c43c8bd8421f60b7d59e

Observation e9117a31-9416-46fa-9714-2941e7c61f58 · inbound

HumanOmni-Speaker: Identifying Who said What and When cites this paper.

HumanOmni-Speaker: Identifying Who said What and When Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.524502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T00:57:12.239390Z digest=sha256:8f2ae85ea2a6b51201ff118a2b0783072f6323f67c8d9c017f3d0ef0c7f6df52

Observation 16998ea3-2045-49cb-80c7-b9342ea01be9 · inbound

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge cites this paper.

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:29.101624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-07T07:49:44.117839Z digest=sha256:faa99254ac3bab303f51adf591c09c542f6123df5a24cdf79beebbf8eb5db86f

Observation a8176aad-ec4b-4cb7-b5af-e25c1a36bbb5 · inbound

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge cites this paper.

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T19:21:19.583012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:21:19.583012Z digest=sha256:9bf968faaf002f65a6f154b01d6dd02372624c507cd7312298403ebc48c795fa

Observation 709bdce4-0dfc-4a2d-b7a0-eb7505c97883 · inbound

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework cites this paper.

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:51:30.417122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T01:23:27.355803Z digest=sha256:b28ac6452e6df9cfb034cdf29197df2f171355e9446a08afeaa9b626758aebb4

Observation caa96176-d73b-41ad-9fca-42357ca0e1d2 · inbound

Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization cites this paper.

Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:20:24.482019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-25T05:19:38.661190Z digest=sha256:748a555499e2cb6dbc7cb893d88a99c6e920313b7ce6072a3a8ad2f55c47fe9c

Observation c1ac8eb4-36f1-4f64-a1c0-65dcd4a3abf9 · inbound

Your Multimodal Speech Model Says I Have a Face for Radio cites this paper.

Your Multimodal Speech Model Says I Have a Face for Radio Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:53:14.062161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T07:44:50.169335Z digest=sha256:b4b2d26f782319797903916f228232d182615183a8c8441bf7ac8843a2841b9f

Observation 994b35a5-0354-4e72-b202-f8ccd82040fe · inbound

Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography cites this paper.

Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:47:28.354383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T17:49:19.643889Z digest=sha256:5a12433c6d1977060ddde9677db7476b1023732dfd3e40d80f9cdcd670f96863

Observation 641792ed-feaf-437a-8aec-8387dd5eddc8 · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.642562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:f577a144ac268b42930a6a29161555db1d1dd52442e679866fda95bc9a22588c

Observation 68ffcd41-0334-43b0-9014-0b48bc71b45c · inbound

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading cites this paper.

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:47:35.216339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T14:54:08.119264Z digest=sha256:a1a5e3d79d612ffbf7eae22b911576353ea50fc933efa004a9822126fc6a7af1

Observation e54e3b54-3554-46c1-b3f0-f9ec27863513 · inbound

Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning cites this paper.

Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.729646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T11:03:05.449385Z digest=sha256:b895cb92557487faeed1208c192040262e8c8d1147770aaee3c5381a1daca6a5

Observation 524a6f4f-0fc6-4ba4-a4ca-47e4a843d284 · inbound

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement cites this paper.

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:49:01.368464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T23:17:45.299833Z digest=sha256:47eb37480c1e94e23abb91561c5958db77e260c223b426ce3993ff5b7e8a33df

Observation 47d8d709-c8a7-4fa2-9f76-942b31511c2b · inbound

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection cites this paper.

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T02:09:52.116351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:09:52.116351Z digest=sha256:b8e7b0d400d8b1dd98d75bd8d6da6e119673c15f191861caef1e3844942b6acf

Observation 00cdfd15-4ba4-4c9b-bfe7-2a6d6267fb59 · inbound

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE cites this paper.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T05:02:03.514200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:02:03.514200Z digest=sha256:2f8dd3f3d478e505de8c1b78b4e26d1773361b01a5da2c0cc19a9777d2f4d2c6

Observation f48c3caa-2414-4616-9134-fb48b647a39f · inbound

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE cites this paper.

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:54.387173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:54.387173Z digest=sha256:c9175556db8a320ebe3bc89b8e858873ff1ebda9e7154e7251432457eed0fc95