Pith. sign in

REVIEW 5 cited by

AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11697 v2 pith:XZA4BMIX submitted 2024-05-19 cs.CY

classification cs.CY
keywords misinformationmedia-basedonlineimagemethodsai-basedammebaclaims
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The prevalence and harms of online misinformation is a perennial concern for internet platforms, institutions and society at large. Over time, information shared online has become more media-heavy and misinformation has readily adapted to these new modalities. The rise of generative AI-based tools, which provide widely-accessible methods for synthesizing realistic audio, images, video and human-like text, have amplified these concerns. Despite intense public interest and significant press coverage, quantitative information on the prevalence and modality of media-based misinformation remains scarce. Here, we present the results of a two-year study using human raters to annotate online media-based misinformation, mostly focusing on images, based on claims assessed in a large sample of publicly-accessible fact checks with the ClaimReview markup. We present an image typology, designed to capture aspects of the image and manipulation relevant to the image's role in the misinformation claim. We visualize the distribution of these types over time. We show the rise of generative AI-based content in misinformation claims, and that its commonality is a relatively recent phenomenon, occurring significantly after heavy press coverage. We also show "simple" methods dominated historically, particularly context manipulations, and continued to hold a majority as of the end of data collection in November 2023. The dataset, Annotated Misinformation, Media-Based (AMMeBa), is publicly-available, and we hope that these data will serve as both a means of evaluating mitigation methods in a realistic setting and as a first-of-its-kind census of the types and modalities of online misinformation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset

    cs.CL 2025-10 conditional novelty 7.0 of 10

    M4FC is a 4,982-image, 6,980-claim, ten-language dataset covering six multimodal fact-checking tasks with baseline results.

  2. Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking

    cs.HC 2025-02 conditional novelty 7.0 of 10

    Fact-checkers want automated fact-checking explanations that trace the reasoning path, cite checkable evidence, and clearly flag uncertainty and information gaps, not just confidence scores.

  3. Detecting Text Manipulation in Images using Vision Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    In zero-shot benchmarks, GPT-4o outperforms open-source VLMs and specialized manipulation detectors on text tampering detection in scene images and fantasy ID documents.

  4. Dataset of News Articles with Provenance Metadata for Media Relevance Assessment

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark dataset and two tasks let researchers test whether AI systems can judge if a news image's recorded location and date match the article, with current chatbots scoring 64-81% on location but 42-58% on date.

  5. Large Language Models and Provenance Metadata for Determining the Relevance of Images and Videos in News Stories

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A prototype combines LLM reasoning with C2PA provenance metadata to classify news images and videos as relevant or not, without any benchmark evaluation.

Pith tools