Pith. sign in

Paper Citation Record · LEDGER

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

As of 9 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 1 inbound Pith citation observation for arXiv:2505.15233.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15233 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:27:00.996471Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T15:08:25.309094Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:13:25.040118Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact2
  • verified fuzzy53
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb4cd9f1-cff9-435c-89b3-9c71e68d3766 · outbound

This paper cites Mesonet: a compact facial video forgery detection network.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Mesonet: a compact facial video forgery detection network

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.335259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:54.074983Z digest=sha256:183c07ee3b2b146772e8afedf3da9dc0b3f142b47913566546d4c0ba36b92625

Observation c58e74d7-cdc7-43a4-8e66-492732a5fe8c · outbound

This paper cites A review of modern audio deepfake detection methods: challenges and future directions.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation A review of modern audio deepfake detection methods: challenges and future directions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.316636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:54.150530Z digest=sha256:e539b1bd6ca9783caebbeeadea5bf36ff6ad8e30b66dc9be8b6b9ca3890722da

Observation 1f187ee3-29bd-4605-9309-632285bf018e · outbound

This paper cites Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.292173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:54.322384Z digest=sha256:e93514f0dd772166ebc9e20a89a49d37b95f7313ae072352cdbafed0e9538603

Observation fe38a54f-d59a-4ce4-8404-a8aa568b0c69 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:54.433137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:54.433137Z digest=sha256:90e7bb37c6394db8cbcbe623947046f1461e1e7a55d1c0fa712de57c9aca5e1e

Observation d6ce7981-1a59-44ce-bea4-20bcf51e17cb · outbound

This paper cites an unresolved cited work.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:27:04.263567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:54.539628Z digest=sha256:5348199de850df8879067d9af4a1fc5616961b22ddbd1cf8fe123877f9e19960

Observation 933e19fd-3ba1-4275-9383-5013726f029d · outbound

This paper cites Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.246938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:54.649874Z digest=sha256:7e7873376fe050142025e3a0f4c12a2bdd0afe2c41349dce46de1957a039123e

Observation 32a37ffa-7e7c-4397-b206-ab43e5cbd069 · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation A Simple Framework for Contrastive Learning of Visual Representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:54.756360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:54.756360Z digest=sha256:9d3d2dd9e0f94c27954bac1629de6f2913d6e914be3a02b7061b9a427a2e42e8

Observation 50ac14fd-fba5-4306-a83f-f50b062b16be · outbound

This paper cites Exploring simple siamese representation learning.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Exploring simple siamese representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.230249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:54.856339Z digest=sha256:40837fc9eeeb38eaae79d3442d515a71d0e90b3b12215b5594d8ee49068d4c29

Observation a8f0832a-cbc0-436c-8860-320875445262 · outbound

This paper cites Sophia Koepke, Ying Shan, and Zeynep Akata.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Sophia Koepke, Ying Shan, and Zeynep Akata

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.213684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:54.977105Z digest=sha256:6444a37446146f2ac7b268077010737999ffc154832011a14eeb1ed3c250e24c

Observation 64487f5f-8e60-4cc6-bc87-df0abcfee9a4 · outbound

This paper cites V oice-face homogeneity tells deepfake.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation V oice-face homogeneity tells deepfake

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.197708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:55.048303Z digest=sha256:6f49c3ef6286cde9c0d64f91bd0c8313602e78359067f61a30a172fa7d6b62ac

Observation 5769a50b-9188-4a15-8163-a547698489f2 · outbound

This paper cites Can We Leave Deepfake Data Behind in Training Deepfake Detector?.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Can We Leave Deepfake Data Behind in Training Deepfake Detector?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:55.184286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:55.184286Z digest=sha256:c4c837b4abd61d3df968f2022372816625ce627e993d44188e97a5aaa9a38d0b

Observation 129e1880-4817-4737-ae70-a38b7652923e · outbound

This paper cites Xception: Deep learning with depthwise separable convolutions.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Xception: Deep learning with depthwise separable convolutions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.180876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:55.302595Z digest=sha256:aab306f15d18eed83ac59baeb7c6f4fed9f6f7ac3cff1d34228327c76f844f91

Observation 8284eb72-ad9e-472d-a7aa-63f5969c98d0 · outbound

This paper cites Not made for each other- audio-visual dissonance-based deepfake detection and localization.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Not made for each other- audio-visual dissonance-based deepfake detection and localization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.155779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:55.405611Z digest=sha256:0d20e8e3becf513632f3990ed2c3a1a2f7e9d16abeb01ed4dffc38c95f0d4531

Observation 56d9c05b-5efd-4049-9806-edda9d0c0262 · outbound

This paper cites Audio-visual person-of-interest deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Audio-visual person-of-interest deepfake detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.138476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:55.485300Z digest=sha256:ab40f4351faa7b64107f5314342dc5c1b3dd32b308f83a2d28235aa0b00ee129

Observation 5eb7d261-3663-4989-872a-9a2da6120e54 · outbound

This paper cites Dong, Jin Wang, Renhe Ji, Jiajun Liang, Haoqiang Fan, and Zheng Ge.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Dong, Jin Wang, Renhe Ji, Jiajun Liang, Haoqiang Fan, and Zheng Ge

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.119110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:55.597172Z digest=sha256:e8933a6137007b432beb3e2519c2eecb5ff2f267788ebe5530fa2852eee46c27

Observation 5b2b32de-9a52-488c-848a-68dffb596ebb · outbound

This paper cites www.github.com/MarekKowalski/FaceSwap Accessed 2021-04-24.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation www.github.com/MarekKowalski/FaceSwap Accessed 2021-04-24

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.099701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:55.697949Z digest=sha256:810b8dfff8fd9fb2c03af07eaec2b0e2eab899b9e57867b033622bd4555b3c8b

Observation 85c18813-1f05-4efb-ba7f-1188277e5ec9 · outbound

This paper cites Self-supervised video forensics by audio-visual anomaly detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Self-supervised video forensics by audio-visual anomaly detection

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.079789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:55.778876Z digest=sha256:7441ef51dc118b0ef8376a9d9d00edb9052b9baca2a7452dfd53128cc35a335d

Observation 2a551dd0-faf3-411d-88b2-5fbe681f7e98 · outbound

This paper cites Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:27:01.797723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:55.856641Z digest=sha256:60aa6d5a19d3814fc1bb88963a04f766455d1dc3fb2861a6e3a278d0253a28e0

Observation 910b8953-1a46-46c7-80d4-6a8afc3b5515 · outbound

This paper cites Imagebind one embedding space to bind them all.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Imagebind one embedding space to bind them all

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.058641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:55.927230Z digest=sha256:768570073a5058c25f90403003035cba86a77750c01c51007501f875f2be4f50

Observation 1d42e731-48c8-4c7c-adb2-ceddac4857b8 · outbound

This paper cites Leveraging real talking faces via self-supervision for robust forgery detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Leveraging real talking faces via self-supervision for robust forgery detection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.039846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:56.007650Z digest=sha256:2cdb167129103b78d4cc7e239c87e27636bb9d02ffcdd9e9122fd36c874530db

Observation 242ef530-74f6-4b75-955b-c98b0fde5954 · outbound

This paper cites Lips don’t lie: A generalisable and robust approach to face forgery detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Lips don’t lie: A generalisable and robust approach to face forgery detection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.016350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:56.106380Z digest=sha256:2db06e6d28f563d812486338b4fa4561b8d9e760a470e587ae2a7d688cec8a80

Observation 459ebd5f-63b3-4792-9178-d88286def08e · outbound

This paper cites Masked autoencoders are scalable vision learners.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Masked autoencoders are scalable vision learners

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.992759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:56.188215Z digest=sha256:036976032025fa5673c7b754c1de174af8bec28eca6b6f2400cb00924570b651

Observation c7900e5d-7aea-49b9-9312-8e74246eb0b9 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.972701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:56.329894Z digest=sha256:cbc994625fdac951bb7b9b8f98b589eb8618f9611847d2379155853f8f06f1ff

Observation 105c5295-9d6b-4502-bb53-817ee1220e08 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Lora: Low-rank adaptation of large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.400893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.400893Z digest=sha256:3c9a9c6eec4eb43f8fd8cf8c2a4ac52e955f41e9a3ddccc80dc8c6dac51680d1

Observation 02a2a4a0-ad58-4f6e-846c-7aec0d9ea89c · outbound

This paper cites Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio-visual deepfakes detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio-visual deepfakes detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.944903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:56.513786Z digest=sha256:8b34847cfbf95e7bb113956d32e1bcf342aecaa594605304285fcec136f97acf

Observation 4c94acf6-c13d-49e4-b28c-c07b57924b0c · outbound

This paper cites Information theory and statistical mechanics.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Information theory and statistical mechanics

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.596373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.596373Z digest=sha256:644de98e944a1d0a5eda11eb8ef1b775d803dc4d59d16261d614e5fd96bd5e87

Observation ceb33f51-45cd-4bcf-8c1c-5be15e06db80 · outbound

This paper cites an unresolved cited work.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:27:03.913566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:56.679484Z digest=sha256:2521c35a66571cb2dde8fd1e6bf92a4c28e5f2fd9313a9e6153833999fae3748

Observation 6a14a426-5563-4cd0-8978-90fd0c27ee70 · outbound

This paper cites FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.756989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.756989Z digest=sha256:06fce1e877c366cef09ff100c07035b43d2b5aa636dd984101849ebe5c05f11e

Observation aed45079-acd6-41b8-a7b5-9021e513db61 · outbound

This paper cites Fast face-swap using convolu- tional neural networks.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Fast face-swap using convolu- tional neural networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.893130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:56.853865Z digest=sha256:6a0a003f7d763795ce656e9092876e1119b6ff47456ae04a981cf87f4dfb28ab

Observation 9a0605a9-493f-4e9d-8144-64663846db63 · outbound

This paper cites Face x-ray for more general face forgery detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Face x-ray for more general face forgery detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.874270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:56.962771Z digest=sha256:e2f8d8aaae29bff5cce60612f0791b03d4682c4a54c3655197f8c469ad47d369

Observation 15b3d7be-6001-44bf-b67b-a03298ab009d · outbound

This paper cites A Survey on Speech Deepfake Detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation A Survey on Speech Deepfake Detection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.093958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.093958Z digest=sha256:2ee5e630e9882964d0504aef8f8b3f25b8852fed75081b5353c786add915207c

Observation a63c7869-feb6-493e-b625-b0cb81abfd36 · outbound

This paper cites Celeb-df: A large-scale challenging dataset for deepfake forensics.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Celeb-df: A large-scale challenging dataset for deepfake forensics

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.853254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:57.170497Z digest=sha256:f851e337f8d1a913a742f7552221394c25f4f554142d219ff37ebe2a8d4f8e2c

Observation 22291b3e-f02a-49b8-9700-1ad0cdd5a626 · outbound

This paper cites Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.836537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:57.249489Z digest=sha256:8122b272a2abb80d23525948a9850d1397e71193a7d734dd36d857e05415d351

Observation b27d098b-b827-4513-b161-e801418e734b · outbound

This paper cites Exploiting visual artifacts to expose deepfakes and face manipulations.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Exploiting visual artifacts to expose deepfakes and face manipulations

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.817396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:57.377989Z digest=sha256:8af43da6680702682251cee8c40c335e2903794bc0dcb78e30848d8daab9c44c

Observation ef0db15d-c892-49c4-9f39-138ae7df8e3c · outbound

This paper cites Emo- tions don’t lie: An audio-visual deepfake detection method using affective cues.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Emo- tions don’t lie: An audio-visual deepfake detection method using affective cues

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.798323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:57.468124Z digest=sha256:1d8460abb38c5e9881060e9bbe619bc9edc31d560cddc2e758682a58fa5fd443

Observation 9bc000c9-199a-4dc6-81f0-9058d7085044 · outbound

This paper cites Does audio deepfake detection generalize? arXiv preprint arXiv:2203.16263, 2022.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Does audio deepfake detection generalize? arXiv preprint arXiv:2203.16263, 2022

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.564874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.564874Z digest=sha256:9c21fecdd21c9a166d80b059180a3750146e3e081a8704afec8e5fddf28fb2bb

Observation 1761e696-aaa8-4fd5-9b5c-63c1bf77c4c1 · outbound

This paper cites Nguyen, Junichi Yamagishi, and Isao Echizen.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Nguyen, Junichi Yamagishi, and Isao Echizen

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.780765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:57.664207Z digest=sha256:cbcec31cfdc9473f71982ce29e1d0878acd50ce59ded3cced5e03949d9d877e1

Observation f62a7ef7-de00-4d25-b642-f086ef6c0a94 · outbound

This paper cites Towards universal fake image detectors that generalize across generative models.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Towards universal fake image detectors that generalize across generative models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.759973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:57.758259Z digest=sha256:65a717bff67dcfadbde66c38420c015d23fa32ff96c511dc817ff2e75f388293

Observation b33bf91a-4459-486a-a57f-7206d11da1d0 · outbound

This paper cites Avff: Audio-visual feature fusion for video deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Avff: Audio-visual feature fusion for video deepfake detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.738263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:57.831456Z digest=sha256:4ee4dc79a07f14a8c62172ee4b36108084cbf4c6f8d53e6e12ca4257ebd81aa6

Observation 2e33d572-82a6-49c0-a661-9a5ed21294cf · outbound

This paper cites Gpt-4: A large multimodal model.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Gpt-4: A large multimodal model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.709277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:57.907416Z digest=sha256:c0de9e10775788188be2c6b5dd42c319ac953448e2113169357aa38e13a86b5f

Observation a961ec6c-c6cd-4f8e-8810-3f959b1956e7 · outbound

This paper cites Deepfake generation and detection: A benchmark and survey.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Deepfake generation and detection: A benchmark and survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.992374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.992374Z digest=sha256:02f4affd7de692801cb173773d7676a5110643ee1d80470fd0e5fb12e76278fb

Observation 4939c70d-d9cf-4849-abff-6806947f0f82 · outbound

This paper cites Namboodiri, and C.V.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Namboodiri, and C.V

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.690146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.114333Z digest=sha256:4588b6c76e2dc62c7510c560383ef6de509d45bd45d2dd2947025854522a33df

Observation ce7db574-1fc5-4d78-8ee6-13722ead2071 · outbound

This paper cites Audio-visual deep neural network for robust person verification.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Audio-visual deep neural network for robust person verification

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.667273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.192243Z digest=sha256:bf28111ef898d478b585a5a726675b2b36833ec0221a3c8c8d25119e2cd113b9

Observation 998addff-3393-417d-b781-5e3d60ff255a · outbound

This paper cites Learning transferable visual models from natural language supervision.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Learning transferable visual models from natural language supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.644870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.255476Z digest=sha256:3075843b605fd9162b4329557bf315e41b3b4100c0ce87814b9874c991e84d62

Observation ed3b185b-2c25-44f5-8a5d-80711d7b9155 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Robust speech recognition via large-scale weak supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.327095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.327095Z digest=sha256:139bf1b799b9806e1ad38e3e8affa70955bd0b5e56bccf0b2424ec34b2ae9e4f

Observation 27b2ac9b-a357-416b-8236-577cc5b91d59 · outbound

This paper cites Faceforensics++: Learning to detect manipulated facial images.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Faceforensics++: Learning to detect manipulated facial images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.606717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.398378Z digest=sha256:bfe89cad8c2550416e7f399d3332a1cb3d8cef286a1aa20f60fd8897d3191756

Observation 4221d3cd-8174-470c-804c-b320f7db35c7 · outbound

This paper cites A comprehensive overview of deepfake: Generation, detection, datasets, and opportunities.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation A comprehensive overview of deepfake: Generation, detection, datasets, and opportunities

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.586152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.461992Z digest=sha256:1d972c90215f84d17a4814e47150373e4be6460194870d37d5223b82af9ba549

Observation 5bdf06cd-1658-4f2d-a4ef-2585d7efe552 · outbound

This paper cites Detecting deepfakes with self-blended images.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Detecting deepfakes with self-blended images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.518035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.518035Z digest=sha256:480c1a94b1c6d272157d7f63fdfb32430ce6175a768dc2814322cf82339a9b13

Observation 1a3c8a52-071e-4a1b-9382-492116dc065e · outbound

This paper cites Representative forgery mining for fake face detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Representative forgery mining for fake face detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.548967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.570297Z digest=sha256:ec2af1be4ae139f6a936fe98cdc9b26cdc12851fb09cd2788c7d132f44d49b26

Observation 97f193b6-0890-403e-8b58-e6db7fecf709 · outbound

This paper cites Exploring Depth Information for Detecting Manipulated Face Videos.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Exploring Depth Information for Detecting Manipulated Face Videos

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:27:01.335494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.647254Z digest=sha256:480fd042655315078c51b4a724e0af67a070d99ba1007d257e20d505aed1724b

Observation ae051bda-29c8-4dc1-a08a-1e6535fcd5a8 · outbound

This paper cites Tan, and Haizhou Li.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Tan, and Haizhou Li

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.522275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.706727Z digest=sha256:9ae5cfaaffa7c900c9b6fa1a5c44a41591b04bf6724250bf95bcf5d9c0dbd4ba

Observation 0ad396e1-9008-44d0-909e-bce6a7cfc442 · outbound

This paper cites Deep spatial gradient and temporal depth learning for face anti-spoofing.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Deep spatial gradient and temporal depth learning for face anti-spoofing

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.497856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.768804Z digest=sha256:c317bb71212cc223e4b726560fbeaed850118441d9ec8fd09559c2474e9a76cd

Observation 8654a873-3fb6-43db-b586-f4a3d6f549a9 · outbound

This paper cites Altfreezing for more general video face forgery detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Altfreezing for more general video face forgery detection

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.470186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.844117Z digest=sha256:53d736e3f3b6a825eccfd077c899843a4b5f4204b711e9b62655f0a3b484a3de

Observation febc692d-c351-4c3a-85a6-300c68ce4ca3 · outbound

This paper cites Deepfake Video Detection Using Convolutional Vision Transformer.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Deepfake Video Detection Using Convolutional Vision Transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.916792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.916792Z digest=sha256:24ff8a4fe7ada81d5ffe1ce47f326a307d61a4bdcf461f788174aa314f4958c2

Observation ddcbe21a-8038-433e-9d37-8d3b1f313f5f · outbound

This paper cites Binaural audio-visual localization.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Binaural audio-visual localization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.445286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:58.988970Z digest=sha256:e92c2e72cf4c5d4752ff25759340c1314834b76bac533f1676fd6f237ff1b3ee

Observation b4697dcc-b1a5-47a8-ad45-b2d8fffe439d · outbound

This paper cites Identity- driven multimedia forgery detection via reference assistance.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Identity- driven multimedia forgery detection via reference assistance

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.421904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:59.078297Z digest=sha256:af81f2f7eca0184d1999c2baba325ce3419c9e4b2bc329023f1b2efd67238c17

Observation 15ce0a97-bc90-4e7e-ac37-97abf628d8bb · outbound

This paper cites Tall: Thumbnail layout for deepfake video detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Tall: Thumbnail layout for deepfake video detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.390976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:59.161502Z digest=sha256:cec17e62a3f515dc15d02b74dfda13e6ae4c540fb38cdd421a6ceb6644292c00

Observation c6811e39-49e0-4c8c-af3b-f003898121cd · outbound

This paper cites Transcending forgery specificity with latent space augmentation for generalizable deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Transcending forgery specificity with latent space augmentation for generalizable deepfake detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.362476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:59.229390Z digest=sha256:603b0416d35d05c9a4e596ee30d2b947201cea1ef7fb7269e3b06deee885b19b

Observation 971ad1d4-6480-44f9-a694-e2f7f915d7e2 · outbound

This paper cites DF40: Toward Next-Generation Deepfake Detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation DF40: Toward Next-Generation Deepfake Detection

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.305225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.305225Z digest=sha256:23d0287b2033dfef49c5197f657258e0cfe3436470766f5da592d41c61f4c71c

Observation 1d877f55-ac55-496c-9d20-c55c0911f808 · outbound

This paper cites GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.364495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.364495Z digest=sha256:0a83edc8c7ec73b919c3c91de07e5494d8d9d7d363ab1df117a71bead77a2842

Observation d25b769b-48ca-4156-bb5f-d47bcb261003 · outbound

This paper cites Ucf: Uncovering common features for generalizable deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Ucf: Uncovering common features for generalizable deepfake detection

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.337860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:59.442188Z digest=sha256:5650165d7a9176a98d2b143afbf61a3ce52ae0456668b171456d421e90666cd8

Observation 9c056045-81dc-401a-a5b1-e90a2e2f039f · outbound

This paper cites Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.498474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.498474Z digest=sha256:98701d5cd40e62ecc67435e7fb708b8d7c265bafc25f1323a60b782da53b873a

Observation ad1cfbbd-6071-4c67-919d-5c751b1c3a31 · outbound

This paper cites Avoid-df: Audio-visual joint learning for detecting deepfake.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Avoid-df: Audio-visual joint learning for detecting deepfake

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.032256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:59.566845Z digest=sha256:0c454de4f1fa6a6343fe4ce2a06075a825827189db6243af000a775ce239f79c

Observation be58bad2-38dc-4944-941b-4c71e3c5c80e · outbound

This paper cites Exposing deep fakes using inconsistent head poses.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Exposing deep fakes using inconsistent head poses

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.912242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:59.676740Z digest=sha256:f5cbd40709b6d1739f44e6b5582abeb6b1b76cd5ad2be636af1a99e0efc58133

Observation 4939e96a-4e65-4f9e-b493-fec38e132990 · outbound

This paper cites Audio Deepfake Detection: A Survey.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Audio Deepfake Detection: A Survey

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.790961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.790961Z digest=sha256:07964c86e0141e3b91436fd772088dff41cd12a5ac15c67f352871213b1118ab

Observation f57e4aa7-b32c-4034-9cac-eafa94f7c76a · outbound

This paper cites Learning natural consistency representation for face forgery video detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Learning natural consistency representation for face forgery video detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.835435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:26:59.977050Z digest=sha256:936dec105b62f0a7b5552f90b5f0a35c8ff6fc63ae08761fdfccd37edd204f9f

Observation 0a1ded16-7871-485e-a9f1-31875ef21b89 · outbound

This paper cites Inclusion 2024 Global Multimedia Deepfake Detection Challenge: Towards Multi-dimensional Face Forgery Detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Inclusion 2024 Global Multimedia Deepfake Detection Challenge: Towards Multi-dimensional Face Forgery Detection

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.152871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.152871Z digest=sha256:dfc71b496a6a02c96f4cc6e7e8debfb57281cb71d0fa9f850bc63d452a89e203

Observation 7b594adb-5bb3-4472-b87f-5e8c0bae3f49 · outbound

This paper cites Multi-attentional deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Multi-attentional deepfake detection

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.717608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:27:00.352741Z digest=sha256:747a4a501700a657f848e611ecb5d20faa578250b0e2a7bbaeb453ac9efdc396

Observation e5f35ecf-5e29-4f4a-aa9d-05a896ab8679 · outbound

This paper cites Attention-based spatial-temporal multi-scale network for face anti-spoofing.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Attention-based spatial-temporal multi-scale network for face anti-spoofing

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.577494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:27:00.559647Z digest=sha256:25319a83531dfe07e191a2b89f786c8c3ab83f840a9568454a0e3172052469ac

Observation 98e6c56d-f524-4cca-8c56-5edb2bfa5436 · outbound

This paper cites Exploring temporal coherence for more general video face forgery detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Exploring temporal coherence for more general video face forgery detection

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.445239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:27:00.736070Z digest=sha256:14a32ac12e950e40c8004de8d373560b96a16e6e2c75aec43864ad46eef0c6a3

Observation ddf7a907-6417-4d79-bc02-88e3e060afbe · outbound

This paper cites Learning deep features for discriminative localization.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Learning deep features for discriminative localization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.349469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:27:00.875898Z digest=sha256:5af39f1ffa26603ae5aa63dadd5ed3eb5888b58d3e2b0d52849b05f885a198b6

Observation 80b7e914-261f-4337-9107-8595853eb3e8 · outbound

This paper cites Makelttalk: speaker-aware talking-head animation.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Makelttalk: speaker-aware talking-head animation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.294628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:27:00.979116Z digest=sha256:15444852b4fdabc75b88c7022fc8c0d118c979b672b6f76469b046bad5cfe2f6

Observation 80b73bbd-895d-4d40-88b8-38d16dff6ea2 · outbound

This paper cites Joint audio-visual deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Joint audio-visual deepfake detection

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.134020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:27:00.986411Z digest=sha256:a968e855ac2997e9571a9a353af8f5594316ca28b8bbc755c4daacc2fc0a86f8

Observation ec5cef83-b4fd-464f-a38d-b37ac7d8645e · outbound

This paper cites Lan- guagebind: Extending video-language pretraining to n-modality by language-based semantic alignment.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Lan- guagebind: Extending video-language pretraining to n-modality by language-based semantic alignment

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.020529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:27:00.996471Z digest=sha256:5a59ab4d85e06998b6dd88dde2d4cbab94625049da611d187dc752d0fdc1a52f

Pith citing papers

Observation 4929d710-214b-4f3a-b5a0-17abcb6b57a5 · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.042025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:bb31b2aae7d2f6e4ddb4fe3418f00c4413335db33dbe85a9bac250b943b3dbbb