Pith. sign in

REVIEW 5 cited by

Interpretable Detection of Out-of-Context Misinformation with Neural-Symbolic-Enhanced Large Multimodal Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.07633 v2 pith:LHVU5GVO submitted 2023-04-15 cs.CL cs.LG

classification cs.CLcs.LG
keywords misinformationdetectionmodelinterpretablecross-modalfakehelpfulimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent years have witnessed the sustained evolution of misinformation that aims at manipulating public opinions. Unlike traditional rumors or fake news editors who mainly rely on generated and/or counterfeited images, text and videos, current misinformation creators now more tend to use out-of-context multimedia contents (e.g. mismatched images and captions) to deceive the public and fake news detection systems. This new type of misinformation increases the difficulty of not only detection but also clarification, because every individual modality is close enough to true information. To address this challenge, in this paper we explore how to achieve interpretable cross-modal de-contextualization detection that simultaneously identifies the mismatched pairs and the cross-modal contradictions, which is helpful for fact-check websites to document clarifications. The proposed model first symbolically disassembles the text-modality information to a set of fact queries based on the Abstract Meaning Representation of the caption and then forwards the query-image pairs into a pre-trained large vision-language model select the ``evidences" that are helpful for us to detect misinformation. Extensive experiments indicate that the proposed methodology can provide us with much more interpretable predictions while maintaining the accuracy same as the state-of-the-art model on this task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection

    cs.LG 2025-02 conditional novelty 6.0 of 10

    E2LVLM improves out-of-context misinformation detection by having a vision-language model rerank and rewrite retrieved evidence, then fine-tune on LVLM-generated explanations.

  2. LLM-based Contrastive Self-Supervised AMR Learning with Masked Graph Autoencoders for Fake News Detection

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A self-supervised fake news detector fusing Abstract Meaning Representation graphs with propagation network features, using an LLM to pick contrastive negatives, reports SOTA results on PolitiFact and GossipCop.

  3. COVE: COntext and VEracity prediction for out-of-context images

    cs.CL 2025-02 conditional novelty 5.0 of 10

    COVE predicts an image's true context before judging caption veracity, improving real-world out-of-context detection and providing a reusable context artifact for human verifiers.

  4. E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs

    cs.MM 2025-06 conditional novelty 4.0 of 10

    A training-free pipeline using image and text retrieval plus two-stage Gemini and GPT-4o mini reasoning reaches 90.0% accuracy on NewsCLIPpings out-of-context detection, but code, prompts, and error bars are missing.

  5. Zero-Shot Warning Generation for Misinformative Multimodal Content

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A multimodal fact-checking pipeline combines web-evidence consistency scores with zero-shot MiniGPT-4 prompting to detect out-of-context image-caption pairs and generate contextualized warnings.

Pith tools