Pith. sign in

Paper Citation Record · LEDGER

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

As of 14 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 11 inbound Pith citation observations for arXiv:2411.10193.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10193 v2

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:57:43.285908Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:44:30.005614Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T14:07:02.288478Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact1
  • verified fuzzy67
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba01d280-6c6d-4288-8824-d622934a0dc5 · outbound

This paper cites Mesonet: a compact facial video forgery detection network.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Mesonet: a compact facial video forgery detection network

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.926251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.926251Z digest=sha256:1f1a2df81198f5e94c0df9ab361b26e9eeba794145ad9c31a073fac738c8befa

Observation f1151922-57ea-4923-b5e7-fa222dbbfae5 · outbound

This paper cites Deep audio-visual speech recognition.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deep audio-visual speech recognition

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.931568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.931568Z digest=sha256:a7ce805c517cac394d10041226fc606f4c07028f7574dab6a30b0ee0c507894e

Observation 828b5171-0cd1-4e72-aec0-54453fecb232 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization LRS3-TED: a large-scale dataset for visual speech recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.935850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.935850Z digest=sha256:65084875662ede7703e6c4514702140c22e2153936fb6f86218a84f1d6495da3

Observation c509f8c8-f2bd-4b6f-85c7-63c96cfbf0fd · outbound

This paper cites Transitions in neural oscillations reflect pre- diction errors generated in audiovisual speech.Na- ture neuroscience, 14(6):797–801, 2011.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Transitions in neural oscillations reflect pre- diction errors generated in audiovisual speech.Na- ture neuroscience, 14(6):797–801, 2011

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.940463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.940463Z digest=sha256:07216f876c5d2df25edc721611b2f81cb223fdd1ea13fda3acd911d83eb763fc

Observation ec047c76-d001-42ee-b4db-eb30974fa3e7 · outbound

This paper cites Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.944640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.944640Z digest=sha256:d537b04f17c453b1d8465426119d34530cd5d80e8a6cd85715b8e1e0d3ed27f2

Observation ccd8201c-efa2-444d-bfad-649975a7270d · outbound

This paper cites Phoneme-to- viseme mappings: the good, the bad, and the ugly.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Phoneme-to- viseme mappings: the good, the bad, and the ugly

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.948696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.948696Z digest=sha256:486c7f447dc8f022bedc71ab3264edefb0a55d5eb1bb64c13bac301b25a63642

Observation 1477da99-97c7-4dfb-9fb4-a43a6fd92262 · outbound

This paper cites Lost in trans- lation: Lip-sync deepfake detection from audio- video mismatch.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Lost in trans- lation: Lip-sync deepfake detection from audio- video mismatch

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.952997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.952997Z digest=sha256:c442f091ea21a610eb89ee1fe679aa5ac18e6f33bd26e0d0317e438a6066cee7

Observation e1221b4f-f42b-4d64-b23b-02301d3d5f55 · outbound

This paper cites Av-deepfake1m: A large-scale llm-driven audio-visual deepfake dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Av-deepfake1m: A large-scale llm-driven audio-visual deepfake dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.957178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.957178Z digest=sha256:a909d3f0363bfc63639f3288ac454a7ae5f50762763f408fb0897b5d43f06d42

Observation e696f7b7-445d-47e5-9cb6-678e484b7745 · outbound

This paper cites Glitch in the matrix: A large scale benchmark for content driven audio–visual forgery detection and localization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Glitch in the matrix: A large scale benchmark for content driven audio–visual forgery detection and localization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.961279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.961279Z digest=sha256:7247efbf3beb5949e28c740bbdf507c59def3fe31b20ba52ecf26f3b32359504

Observation 821fc75e-e240-4207-962b-3a2555be233c · outbound

This paper cites Do you really mean that? con- tent driven audio-visual deepfake dataset and mul- timodal method for temporal forgery localiza- tion.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Do you really mean that? con- tent driven audio-visual deepfake dataset and mul- timodal method for temporal forgery localiza- tion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.965129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.965129Z digest=sha256:ae9b756d2b0f85ccd06d4cd347551895fb44476e5e79f637ee15a367faff73a6

Observation 4d4cbacb-e155-40f2-ad9e-501ea836ce7b · outbound

This paper cites End- to-end reconstruction-classification learning for face forgery detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization End- to-end reconstruction-classification learning for face forgery detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.968775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.968775Z digest=sha256:b30bda4beb9cdd7c6a5d17108dd2d70c471d9cd903408b66e8705d5f6434f9b8

Observation 9dd11b83-9da6-41c9-b8ea-f61fc5d1582d · outbound

This paper cites Chandrasekaran and A.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Chandrasekaran and A

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.972427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.972427Z digest=sha256:1bfd01bb800c4ac7215f88c2d5326e2dcc910393a3c837a7f9d2dc00121d4162

Observation ea3579cf-100f-4f19-b476-b0e23e09caba · outbound

This paper cites What you see depends on what you hear: Temporal averaging and crossmodal in- tegration.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization What you see depends on what you hear: Temporal averaging and crossmodal in- tegration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.976224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.976224Z digest=sha256:d5009d28f81fd81318a4442a2a8936e9c3642e1ef9f4bea25ea608f8e8f13c0e

Observation 7fb2ee97-c599-4f31-bf5d-81442ed91bad · outbound

This paper cites Voice-face ho- mogeneity tells deepfake.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Voice-face ho- mogeneity tells deepfake

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.980020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.980020Z digest=sha256:058dba09578cd7bd74b61ce661d8045f0ad4d0b7f9ddbecc7adfca763ea6b2aa

Observation c18a556a-c7e7-4820-a2b9-2aa1177a9adb · outbound

This paper cites Not made for each other-audio-visual dissonance-based deepfake de- tection and localization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Not made for each other-audio-visual dissonance-based deepfake de- tection and localization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.983840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.983840Z digest=sha256:7757e350fbc4c5c948e9eb688c1b4d5c84b285d140792bb215043239f3525dc6

Observation f699a059-1b0d-4958-929d-2cb7bd7f8bdc · outbound

This paper cites Vox- Celeb2: Deep speaker recognition.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Vox- Celeb2: Deep speaker recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.492884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:42.987741Z digest=sha256:7dbea8ef334406cff2908b086f6e8ec2421a387bc46c0bc64ebf0cf9f01bc355

Observation a1af069b-7321-490b-95e3-84634ff16a2c · outbound

This paper cites Lip reading in the wild.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Lip reading in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.482034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:42.991659Z digest=sha256:3886591cc29d6b030154c8ac7c80fc88e0f28ef38925e1cb4a844123e25e0564

Observation 6c21c2bd-d294-4363-8173-59661ba0919c · outbound

This paper cites Combin- ing efficientnet and vision transformers for video deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Combin- ing efficientnet and vision transformers for video deepfake detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.469704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:42.995406Z digest=sha256:5cc2d8809fe435908f37f572140da7758253315caf154a8d92844e4013145c29

Observation 69f40b8d-c460-40c3-ada1-7031d3d862b7 · outbound

This paper cites Audio-visual person-of-interest deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Audio-visual person-of-interest deepfake detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.457998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:42.998582Z digest=sha256:311292c5a5ccddd0318b43b6e9df44dddc48fa54b080b289bf0bd87c1bd894f2

Observation 0f0cdda5-73a8-4213-aa38-9ee2ff227120 · outbound

This paper cites Imagenet: A large-scale hi- erarchical image database.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Imagenet: A large-scale hi- erarchical image database

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.447053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.001811Z digest=sha256:eced64ee300103cda4bc9266019991f932a29c9fc255663c558661fab138ea70

Observation 00b02011-e2ab-4a8f-9e73-b72be070ad96 · outbound

This paper cites The DeepFake Detection Challenge (DFDC) Dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization The DeepFake Detection Challenge (DFDC) Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.005309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.005309Z digest=sha256:d0cb2bb2aa23456343f504bc0d4fd32144a6763bf11fcdbacf5790748bcb8eb0

Observation 22db688a-6f2d-411c-97be-e116aad56a93 · outbound

This paper cites Taming transformers for high-resolution im- age synthesis.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Taming transformers for high-resolution im- age synthesis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.434065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.008860Z digest=sha256:4812a67675b9a9e65205c182b87099501644a8a4056ceaabd5d9ab6df3b8b209

Observation baed580b-9dc3-48d1-a5c5-39454dc8f02a · outbound

This paper cites Self-supervised video forensics by audio-visual anomaly detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Self-supervised video forensics by audio-visual anomaly detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.419991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.012814Z digest=sha256:7544a6f1431ddd4ff7fb8058455bb04ea0ee6dca280931a9e919b985fb7108cc

Observation 9c06d2ac-5a8e-43b1-9045-1e8eb5568f08 · outbound

This paper cites Fast r-cnn.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Fast r-cnn

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.406731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.016582Z digest=sha256:468bee2878b852eb493c6a335ef1444be143eca6bad917329d75bc90b45cdbcc

Observation dab7ee5a-300a-458e-88a4-3dee2f931577 · outbound

This paper cites Rational decisions.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Rational decisions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.392238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.020842Z digest=sha256:3203232d071abeb6009de5623d25815924feb393fca1ac9f44ebcc1e6e160cde

Observation 1d1e77ac-7331-4cb8-96bf-a9322fe940b6 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Bootstrap your own latent-a new approach to self-supervised learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.379577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.025215Z digest=sha256:9b03fd01c4287571e4f16e724e621718586847c280395dbd9dd9468e6c5c5e60

Observation e0f1f88b-835e-48b7-92fa-4c93a9e0f5d9 · outbound

This paper cites Deepfake video detection using audio- visual consistency.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deepfake video detection using audio- visual consistency

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.368609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.029858Z digest=sha256:265001b7303645fb5ee7ac56a17b5b8f98e6e0a75e2c35a35aa1c154b8a09142

Observation 5ec760ee-13e0-4616-b656-f919a33436ee · outbound

This paper cites Delving into the local: Dynamic inconsistency learning for deep- fake video detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Delving into the local: Dynamic inconsistency learning for deep- fake video detection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.357235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.033876Z digest=sha256:6033c69fd1d2aa7e6aa45e93cd5a360e5ac7a7f57f975514c4ab7ab7a5a3b907

Observation 07767f8c-4def-4733-b97b-b8f0562a3030 · outbound

This paper cites Leveraging real talking faces via self-supervision for robust forgery detec- tion.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Leveraging real talking faces via self-supervision for robust forgery detec- tion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.347064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.037960Z digest=sha256:c60ab23b60d88607f6138ef352b3f14f41bff7115088071ac155e9b7a15df184

Observation 3caaede2-915e-46f2-a7fb-292e27946345 · outbound

This paper cites Lips don’t lie: A generalisable and robust approach to face forgery detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Lips don’t lie: A generalisable and robust approach to face forgery detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.333711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.041863Z digest=sha256:7b88f4c51b8c5e201212c8ba936a43ab74623505ab890fa03cb0d03e59e9fc01

Observation 23ead194-4691-43ea-8cea-a87b0e024c27 · outbound

This paper cites De- tection of fake images via the ensemble of deep representations from multi color spaces.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization De- tection of fake images via the ensemble of deep representations from multi color spaces

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.319906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.046264Z digest=sha256:5835779a1ed2aa65f508be1e8d100a44042d88b62de7593d45202520d87a9397

Observation 123db3e9-4e4d-48a2-a0aa-32c47a58ebb7 · outbound

This paper cites Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio- visual deepfakes detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio- visual deepfakes detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.307141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.051021Z digest=sha256:b782371ddf2f59bbdca94f1ac70486a2aed281878cd092fe13c56c738615038e

Observation 55b1aaac-a140-4697-b961-17b7222667d7 · outbound

This paper cites Deeperforensics-1.0: A large-scale dataset for real-world face forgery de- tection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deeperforensics-1.0: A large-scale dataset for real-world face forgery de- tection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.294608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.055087Z digest=sha256:26c3d2594f30ae0887eecf1c47b48d1057bc4891409f2fe37ac5ee58d3084b4f

Observation 72e2ad54-abcd-42d0-84d3-daf329e5ea0f · outbound

This paper cites Con- textual cross-modal attention for audio-visual deepfake detection and localization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Con- textual cross-modal attention for audio-visual deepfake detection and localization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.281106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.059518Z digest=sha256:022437ce9ab82703a1abea6c443fc6348d33f5b5e27e740ca4c3f572666346fa

Observation 34f77d1b-062f-4df0-a60f-8c76ce62a0c8 · outbound

This paper cites The Kinetics Human Action Video Dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization The Kinetics Human Action Video Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.063363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.063363Z digest=sha256:85f5b235ed049b221fbe4ce9dc509f3639b12edc82f8141681e212ca7550a601

Observation 33a3881c-034c-4b63-a1a7-85ad53ce3fbd · outbound

This paper cites Fakeavceleb: A novel audio-video multimodal deepfake dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Fakeavceleb: A novel audio-video multimodal deepfake dataset

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.269258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.067809Z digest=sha256:d3c6b4fc81e95e40c180d341d0d3772b530e8de5077a1369582925712b87792f

Observation 128aad35-bf40-4b2f-8a22-d537538d6ff6 · outbound

This paper cites Deep- fakes: a new threat to face recognition? assess- ment and detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deep- fakes: a new threat to face recognition? assess- ment and detection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.253639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.071622Z digest=sha256:8d47026aed9b260fe6e5530d9269dbd2a73c885de0ee5b945ca390ab016611ea

Observation 9446f9ec-fdb8-4fac-bc67-01faaf4750f7 · outbound

This paper cites Kodf: A large-scale korean deepfake detection dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Kodf: A large-scale korean deepfake detection dataset

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.240094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.075368Z digest=sha256:fdbad7de343326005efd5657eda494d7bba30b3ebbba990cb62c69abd76db764

Observation 9100865c-9700-49b6-bad7-8fe3530a42c8 · outbound

This paper cites Layer normalization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Layer normalization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.228890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.078998Z digest=sha256:bd2308f4399a6695c9e560255e3ec5469dd3439f3fac665057b8530c78cf7bdd

Observation 30998388-2085-4b43-bf51-7e0ca0e0bb0a · outbound

This paper cites Face x-ray for more general face forgery detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Face x-ray for more general face forgery detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.216358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.082609Z digest=sha256:ddf268ac416aad41e4d27dd4da1e5e45abcd5c0a49cf1aec878595f2f1a67f7d

Observation 2795419c-c1ad-4f57-9679-787517bac0a5 · outbound

This paper cites Spatio-temporal catcher: A self-supervised transformer for deepfake video detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Spatio-temporal catcher: A self-supervised transformer for deepfake video detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.203291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.086486Z digest=sha256:32dd16353a558e416da71a1714576113b7b21fbbd87b7a9888a9ce46e04a2a42

Observation bc8a5313-d2de-4a7f-98a3-7455c11e82ce · outbound

This paper cites Zero-shot fake video de- tection by audio-visual consistency.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Zero-shot fake video de- tection by audio-visual consistency

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.190739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.090412Z digest=sha256:a2af117e96746863ccf86f519291d404a1677e6301910d6d2e0794bdf33a6704

Observation 76094af0-f5b3-41ae-b72e-c166b7fd1d7b · outbound

This paper cites Celeb-df: A large-scale challenging dataset for deepfake forensics.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Celeb-df: A large-scale challenging dataset for deepfake forensics

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.176445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.093763Z digest=sha256:344875ccd2290e10e0da71b52d255a97d5b98c09d4ec8ddc6f3dd635d31e42e0

Observation d51e9110-c6bd-495a-b85b-027388a9ad0b · outbound

This paper cites Bmn: Boundary-matching network for temporal action proposal generation.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Bmn: Boundary-matching network for temporal action proposal generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.163658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.097499Z digest=sha256:b466122a06718b66ae59cc0c36529a4c7e78a2bdffcab31e7703a4c5f8f46865

Observation a17a86e6-b252-433a-9ef0-df14ea693e57 · outbound

This paper cites Focal loss for dense object detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Focal loss for dense object detection

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.151984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.101090Z digest=sha256:c60dade954b417154fc949290ae0e0b3d6efac33cc2752eb4c1cd8044772528a

Observation 7a519223-7193-4e09-97c5-753c7f2c48fe · outbound

This paper cites Visual speech recognition for multiple languages in the wild.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Visual speech recognition for multiple languages in the wild

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.135058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.104974Z digest=sha256:995b4f0a333f6131cc801c4d383ab72af35327948d632905c495a56a03eaccba

Observation b974cab0-f249-4722-9cf9-bcbcf2d2ac50 · outbound

This paper cites Deepfakes generation and detection: State- of-the-art, open challenges, countermeasures, and way forward.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deepfakes generation and detection: State- of-the-art, open challenges, countermeasures, and way forward

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.121575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.108472Z digest=sha256:7176dcb3e374a39a1d52b631cb84b432d8811f89bb4f569a42f57bfdd409cd9b

Observation a9bb3df7-8242-421b-afab-390bb0433827 · outbound

This paper cites Emotions don’t lie: An audio-visual deepfake de- tection method using affective cues.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Emotions don’t lie: An audio-visual deepfake de- tection method using affective cues

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.105024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.111885Z digest=sha256:780239db0fa0e486ce40920a8b7823819f1c4b69a28ef325e92d11f207140be8

Observation e1baa66b-138a-4536-b6a5-ea3083ca0c7a · outbound

This paper cites Df-platter: Multi-face heterogeneous deepfake dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Df-platter: Multi-face heterogeneous deepfake dataset

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.092989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.115608Z digest=sha256:9568f71060b958ac4cf145b57576d248a496e7cd99bdab7d96360494a8469259

Observation dfc75e8e-dbc3-4116-8f63-43df994f805f · outbound

This paper cites A neural basis for interindividual differences in the mcgurk effect, a multisensory speech illusion.Neu- roimage, 59(1):781–787, 2012.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization A neural basis for interindividual differences in the mcgurk effect, a multisensory speech illusion.Neu- roimage, 59(1):781–787, 2012

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.079001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.119372Z digest=sha256:4eca9fe9021073bf96dbe776cf81c19c63b544765306c1191739a082e1cb8aa6

Observation b1696474-fa3f-42fe-877b-1097f2b4397b · outbound

This paper cites Activity Graph Transformer for Temporal Action Localization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Activity Graph Transformer for Temporal Action Localization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.123104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.123104Z digest=sha256:09160d2f3ea5593a9a461257b22ee96e601ab704ba59cbcbc542e4a4d51b6e08

Observation cf3fd579-3604-42dc-9051-37c14cf3a53a · outbound

This paper cites Frade: Forgery-aware audio- distilled multimodal learning for deepfake detec- tion.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Frade: Forgery-aware audio- distilled multimodal learning for deepfake detec- tion

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.065982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.127280Z digest=sha256:09f934ff450f927de458760f4daf4ef30be01caa1564de97d718e88403118e78

Observation 15b9793b-a54f-47f6-857b-5463f88c8646 · outbound

This paper cites Seeing what you hear: Cross- modal illusions and perception.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Seeing what you hear: Cross- modal illusions and perception

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.053027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.131051Z digest=sha256:1946e94909ee48044febd5c15ec31836240afdf0c75b64ac6dfd88d33acc5221

Observation 6f2caed7-7260-4e84-9fd9-dbe25f435ce2 · outbound

This paper cites Avff: Audio-visual feature fusion for video deepfake de- tection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Avff: Audio-visual feature fusion for video deepfake de- tection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.040514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.134766Z digest=sha256:493095f4bf523ad456431c5578688f08a858ffbc207ecb8c89cefb037e2fc6b6

Observation f3ecac87-f8dd-4142-ba71-d8a010b2c416 · outbound

This paper cites Deep- fake generation and detection: A benchmark and survey.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deep- fake generation and detection: A benchmark and survey

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.138759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.138759Z digest=sha256:f39135077ee499074ba671270c40cd0b73a95145be7063161ea4a383fcaf99cc

Observation ba1211c5-7536-46e7-a7e9-9dea4da2337e · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Powerset multi-class cross entropy loss for neural speaker diarization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.142519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.142519Z digest=sha256:c8fa2c240c5ff00f0a39e32241b3e10bca94b86ed1823f1e03b46100b049ed86

Observation 80cb1353-d837-4424-95ee-54c02dfeea93 · outbound

This paper cites Thinking in frequency: Face forgery detection by mining frequency-aware clues.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Thinking in frequency: Face forgery detection by mining frequency-aware clues

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.027082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.146788Z digest=sha256:4488e7f5944c1a893932c5846764d8442e6398d41085228c692589d602b2425e

Observation f939d77a-9b46-493b-a82a-9045a855e781 · outbound

This paper cites Multimodaltrace: Deepfake detection us- ing audiovisual representation learning.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Multimodaltrace: Deepfake detection us- ing audiovisual representation learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.012935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.151257Z digest=sha256:c61a57f24a963388e8d543c3da066fedb52b185fae481f9e262a026e95d0097e

Observation 329e15fd-4d41-4a0e-86fe-881e18cf9f75 · outbound

This paper cites Detecting Deepfakes Without Seeing Any.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Detecting Deepfakes Without Seeing Any

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.155342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.155342Z digest=sha256:57685415bfd755574ef66bc9005869637ebe7bd8b1fbd9172a276644da2e56fd

Observation 91ddaaa7-7b49-44a8-8f89-047904d854d2 · outbound

This paper cites Faceforensics++: Learning to detect manipulated facial images.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Faceforensics++: Learning to detect manipulated facial images

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.998593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.159335Z digest=sha256:d27d28381f5d8e0bcd753fd76c7380d613ac5d69011609f00ffd0eb589591be5

Observation f6185d66-bcc0-4bbc-88dd-8a96f22f67ab · outbound

This paper cites Lip sync matters: A novel multimodal forgery detector.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Lip sync matters: A novel multimodal forgery detector

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.982817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.163041Z digest=sha256:50197165c547febf43107a2b94bf769db461d545eb7a5ebae94bd0b0c4f0b08e

Observation 9d1cf413-d3d9-4cb3-b389-8ca124c88ef2 · outbound

This paper cites Visual illusion induced by sound.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Visual illusion induced by sound

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.969553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.166706Z digest=sha256:bb087abe2635312c747cbb075269b6a4f64f4d45f04e682ba834d10f4968ab4f

Observation 303df92f-315a-472c-8734-3c2767c40494 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal clus- ter prediction.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Learning audio-visual speech representation by masked multimodal clus- ter prediction

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.953594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.170440Z digest=sha256:7997d3b86b26013603fa198f537574046c552a610bf47781805e16ef6f1c1653

Observation e037fbe6-c3dc-4e02-b3a3-0e3912b881d6 · outbound

This paper cites Tridet: Temporal ac- tion detection with relative boundary modeling.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Tridet: Temporal ac- tion detection with relative boundary modeling

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.939120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.174270Z digest=sha256:39ca1c0d83cabb7d37c9d444f86143b1cdd0186a259eb2dc29fb63868373c566

Observation e8caeb53-edcf-4e59-8dee-008f1461e8c7 · outbound

This paper cites De- tecting deepfakes with self-blended images.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization De- tecting deepfakes with self-blended images

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.923424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.178194Z digest=sha256:243ea99f9b58350164feee826af0b36f1fc64afd13a0c3ae35340c09b09cf658

Observation 75c2d7f8-85ad-4b4a-920c-e4dde18c964a · outbound

This paper cites Locate and ver- ify: A two-stream network for improved deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Locate and ver- ify: A two-stream network for improved deepfake detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.910567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.182180Z digest=sha256:7f01d8a79c1ac47ee53dced0cb44c2b92b0f1661f836e402269c16fcd171f845

Observation 874ad5b2-a553-4c21-85a1-bea3412b0123 · outbound

This paper cites The contribution of visual information to the perception of speech in noise with and without informative temporal fine structure.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization The contribution of visual information to the perception of speech in noise with and without informative temporal fine structure

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.898396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.185881Z digest=sha256:9052d52fa79de7d564789d1fa5036a9e31ab23daf7cf487bd5dac6039a003d24

Observation 5fe41125-7665-4625-939e-331c12026a91 · outbound

This paper cites Visual contribution to speech intelligibility in noise.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Visual contribution to speech intelligibility in noise

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.885985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.189720Z digest=sha256:b6ac0936764b5403c2b70e58fe83e90ae3de4f6dd962fcdd8d4d5b339f712a70

Observation 143b8ad7-dc2a-42f9-8d1e-2beda29d4862 · outbound

This paper cites Learning on gradients: Generalized artifacts representation for gan-generated images detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Learning on gradients: Generalized artifacts representation for gan-generated images detection

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.873276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.193519Z digest=sha256:87dd7de62431ee68654fdc9157470fddacc0ebcf5a5ceaf0d53646043b97f9d4

Observation 0ed2dd83-ab49-4660-aec6-7281ee692def · outbound

This paper cites Attention is all you need.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Attention is all you need

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.197257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.197257Z digest=sha256:3fe3a9536a5bf8f19c3d51c62c8ef20e05d2b2792d794a6e2912332af3793f85

Observation 86cfc1e0-69a2-4970-803c-3a463ef1ed7b · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Videomae v2: Scaling video masked autoencoders with dual masking

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.852680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.200837Z digest=sha256:93f22c096452ec7ab828fe92b6987a36df915ada4e8bbd752f9f72d47e33b7a7

Observation 81309a39-bd8b-44ff-b3e4-0ae69a8adf35 · outbound

This paper cites Building robust video-level deepfake detection via audio- visual local-global interactions.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Building robust video-level deepfake detection via audio- visual local-global interactions

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.835415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.204546Z digest=sha256:7d5c2d21f9b17968b5cf6e05043cd8cf8183a58555119178beed0665bb302319

Observation 6cb7def3-e3dd-465a-9962-4b2a978190b2 · outbound

This paper cites Audio-visual deep- fake detection using articulatory representation learning.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Audio-visual deep- fake detection using articulatory representation learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.822420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.208375Z digest=sha256:2db7d0de9a580a863c87432fe0286e0a0d99621419c7f5d6aa0bb74fd1d36975

Observation c3b3b261-43bb-46e4-8398-ff632abee857 · outbound

This paper cites What you see is what you hear: sounds alter the contents of visual per- ception.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization What you see is what you hear: sounds alter the contents of visual per- ception

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.810469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.212662Z digest=sha256:6ab3dcf786c9b4ca5818f5f2351dcdb3469811de1052350c955ccf96594575be

Observation acd175f2-4e6a-4f3a-bb1a-e9e2b52ac72a · outbound

This paper cites Avoid-df: Audio-visual joint learning for detecting deepfake.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Avoid-df: Audio-visual joint learning for detecting deepfake

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.799054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.216685Z digest=sha256:9788bc5b73b53b8a566c81d4d0b2c41c69ae25dd6bef5b200aefda97eb717270

Observation 6097e91f-0892-40c8-bf82-6e61e843dd01 · outbound

This paper cites Expos- ing deep fakes using inconsistent head poses.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Expos- ing deep fakes using inconsistent head poses

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.787515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.221186Z digest=sha256:6773f13638103d409397bf543f83a8f461a9a10f9f13806d7432c04632ba3bfd

Observation 91778712-caf1-48d4-9c23-be4d264ebf1b · outbound

This paper cites Pvass-mdd: predictive visual-audio alignment self-supervision for multimodal deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Pvass-mdd: predictive visual-audio alignment self-supervision for multimodal deepfake detection

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.773917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.224964Z digest=sha256:2df316f8bf7f7ec746085eda0bd3bf5ea99dde63811f224c392c1375cdff13b1

Observation 4c63d14e-9419-4b77-bf3b-2e33525c25e8 · outbound

This paper cites Ac- tionformer: Localizing moments of actions with transformers.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Ac- tionformer: Localizing moments of actions with transformers

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.760264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.228298Z digest=sha256:86d2874da9633f9a12f511709c7aedb535bed8468f754b06eb27115aadf5fea2

Observation dca4b936-a2e0-4a10-b38c-09c9f991b6c0 · outbound

This paper cites Video- llama: An instruction-tuned audio-visual lan- guage model for video understanding.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Video- llama: An instruction-tuned audio-visual lan- guage model for video understanding

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.746605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.231704Z digest=sha256:c79e2fb776ebbc8fb9b1bd5e622db6c2d2932adf74795557df2e8daf6dffc3a1

Observation 140244de-1acb-40f0-af06-5e4ebbacff0b · outbound

This paper cites Um- maformer: A universal multimodal-adaptive transformer framework for temporal forgery local- ization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Um- maformer: A universal multimodal-adaptive transformer framework for temporal forgery local- ization

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.735455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.235779Z digest=sha256:9a3149537f2a96d0e07b8ef7200a3cb9801b275ec275f40c5bedfa924d096e29

Observation 1b9a6618-3e17-480e-940e-b685275af55b · outbound

This paper cites Joint audio-visual attention with contrastive learning for more general deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Joint audio-visual attention with contrastive learning for more general deepfake detection

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.723760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.239365Z digest=sha256:b8b990349dba01ebdcaaa6b713f19d198df4b62a8fc1af201bae28a7a9fa9742

Observation 556ae367-c25b-4bba-a206-dd7c2c67cf69 · outbound

This paper cites Exploring temporal co- herence for more general video face forgery de- tection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Exploring temporal co- herence for more general video face forgery de- tection

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.710125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.243054Z digest=sha256:7fe9c8a0f842baf7854cba2c37ceb43288dfdbd9477a3bfdca5fa002a070551f

Observation 186c60e4-8edd-49de-a06f-0e8b077fa5e3 · outbound

This paper cites Distance-iou loss: Faster and better learning for bounding box regression.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Distance-iou loss: Faster and better learning for bounding box regression

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.697029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.247172Z digest=sha256:d268cd7fd1ba9486ce1cdf3b431bb8a8fd4a8697eb99eac308d36ebcaffb0340

Observation 2ccd7527-90c6-4436-8597-75225c08094f · outbound

This paper cites Joint audio- visual deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Joint audio- visual deepfake detection

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.683747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.250958Z digest=sha256:a235d20def0e9954fbfbdf1ad1d570891db71b29afd2f1ce948ee371c4a20004

Observation 970e70eb-b674-44f6-ab45-932773b56d88 · outbound

This paper cites Wilddeepfake: A chal- lenging real-world dataset for deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Wilddeepfake: A chal- lenging real-world dataset for deepfake detection

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.672699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.254182Z digest=sha256:2ebc570f37d778eeb9ecbfc8e49aa46c0cebb20f0e0d047508191853197e62c7

Observation 1d2e6818-3e17-4b91-945f-d50179b9694f · outbound

This paper cites Cross- modality and within-modality regularization for audio-visual deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Cross- modality and within-modality regularization for audio-visual deepfake detection

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.661608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.257554Z digest=sha256:a5d6e4fc851691adb2f17a425a2ee2b329e8f706f522cbc49512b9af5502fa77

Observation 40f4ca2d-d003-428c-8292-3bb3bfa32533 · outbound

This paper cites an unresolved cited work.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:57:43.649278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.261344Z digest=sha256:a27e3425fa40b43ebce8b51bd5870380010c9b68e87a57631387244d9a7b1d7d

Observation a51072a7-9b24-4a70-b4d8-53248fabc08a · outbound

This paper cites an unresolved cited work.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:57:43.633490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.265660Z digest=sha256:c4c1659d3858c88c974d95a6871bb828e42d9370ba625f436f5e7ad1a6a90ceb

Observation 310d7c0c-8b90-4791-a1d4-b8dfcf45e59c · outbound

This paper cites Since we apply the manual screening process on synthesized videos, the final video count is more than 20,000.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Since we apply the manual screening process on synthesized videos, the final video count is more than 20,000

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.619113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.269730Z digest=sha256:65e4596eb311441247e6a6acd323fd9820f99da10d164ded600119f4ef38ee79

Observation c71b1ba0-f73b-4698-b244-9d98710a070e · outbound

This paper cites an unresolved cited work.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:57:43.604141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.274100Z digest=sha256:6bf2ffb44e953f17dcf11b197cf9b008d1937120625aa9d1f76ea745826b3af1

Observation 222fe795-9d65-45a4-85ec-9a3f8a14d386 · outbound

This paper cites In addition, DiMoDif outperforms A VFF un- der all perturbation scenarios and at all intensity levels.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization In addition, DiMoDif outperforms A VFF un- der all perturbation scenarios and at all intensity levels

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.590814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.277915Z digest=sha256:826457b97d50e750ad4fa92ffc8dbeb791284fd13dcf2c5eefc180595ebb1bdd

Observation 9a41d238-2c84-4614-b317-998cf3c071cf · outbound

This paper cites Table 12 presents the corresponding results.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Table 12 presents the corresponding results

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.578772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.281841Z digest=sha256:70a801654a324761ca24491cf5eaaba6ad5d17cfed2c3510b0e3beedfe4f8674

Observation bcb532c3-57a7-4c9a-a5e0-490816961776 · outbound

This paper cites an unresolved cited work.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Unresolved cited work

Reference 93

Resolution
verified exact
raw_fallback, observed 2026-08-12T19:57:43.390782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:57:43.285908Z digest=sha256:5e9d477fc2a586005f9bc6bcb4fbede34c29a7430b4fb2939b1be4d11023eb9f

Pith citing papers

Observation 2f3969ed-538d-4505-85f6-cf0859757533 · inbound

Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning cites this paper.

Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:30.005614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:30.005614Z digest=sha256:d80f21479433f3cb5dda44dfec17553d320aec689651ddadb78b06abc2ce7723

Observation e2840c89-2355-444f-bdac-136f94fd0a5d · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 163

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:23.115682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:23.115682Z digest=sha256:409f11e8cdbf024553d8bcb710bbaabd92bee4fdcbf81567b83b3a64c485a43e

Observation 96394362-d0dc-478a-bf7d-d7b9b080d66c · inbound

DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection cites this paper.

DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:37.830717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:37.830717Z digest=sha256:721a8ea50adcef702ddc0780f574f11000539dd934d71be5e4ca3fb0148287a0

Observation a89fe335-7651-4d5e-a693-ba01739559a8 · inbound

Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization cites this paper.

Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:36.277635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:36.277635Z digest=sha256:0e8b9dd9cb947cad8aec5cfc141306a3e00a52df062fef933a501ca5ad1c4d32

Observation cb3528dc-8b71-41c0-b153-3295687cd0cb · inbound

Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems cites this paper.

Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 180

Resolution
unresolved
no resolver link, observed 2026-08-06T14:34:11.191723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:34:11.191723Z digest=sha256:33a98744a5ff1e13a8db50fd18cd9b3fb8a655414d3b4797429fdbde32237b4f

Observation 4c29502d-5ae7-445d-8642-809ed8fb6505 · inbound

Generalizing Video DeepFake Detection by Self-generated Audio-Visual Pseudo-Fakes cites this paper.

Generalizing Video DeepFake Detection by Self-generated Audio-Visual Pseudo-Fakes DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:25:59.992133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:39:31.212977Z digest=sha256:7ee1d0424ae548d0bb86ceb47718c2ebf3206ebe1869753270996815ad32d2b1

Observation b034ae34-662b-41a5-ac29-7e8303e9d928 · inbound

Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization cites this paper.

Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:20:24.495845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T05:19:38.661190Z digest=sha256:a265ce810eab644663178b123d4f46e75d860434bc0699f66916212fcd8cd39a

Observation 739272c8-217b-465e-8123-fb9e5cf0c772 · inbound

MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization cites this paper.

MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:07:02.289977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T14:03:46.585001Z digest=sha256:4a5f45e7238ccbe97903176c766a561312debab9cf1d276f28e177c5d0b212a5

Observation 37442a88-3105-480f-9aa7-a64e3f0ae0ed · inbound

EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration cites this paper.

EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T18:55:21.867598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:55:21.867598Z digest=sha256:ea727d5aee28797aee640e3e2ef2866f51918f63d65ec11ac1cde9233d6342ff

Observation ed579bdb-c456-4701-acfd-a3a3eded52d6 · inbound

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization cites this paper.

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T18:34:29.549966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:34:29.549966Z digest=sha256:1d2093fc56ab3bca968d7c5d60d032c603d57ca6daa93750af7763587c10a4de

Observation bd5aa6ed-9377-4bee-8c26-4f1972756867 · inbound

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization cites this paper.

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:52.084084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:40:52.084084Z digest=sha256:15b8adc75895231cd8f7f85c193a25a4cfa48e18a04df64c29d482c87d77747d