Pith. sign in

REVIEW 7 cited by

r/Fakeddit: A New Multimodal Benchmark Dataset for Fine-grained Fake News Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03854 v2 pith:XXLFXHJV submitted 2019-11-10 cs.CL cs.CYcs.IR

classification cs.CLcs.CYcs.IR
keywords fakenewsclassificationdatasetfakedditfine-grainedmultimodalcategories
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Fake news has altered society in negative ways in politics and culture. It has adversely affected both online social network systems as well as offline communities and conversations. Using automatic machine learning classification models is an efficient way to combat the widespread dissemination of fake news. However, a lack of effective, comprehensive datasets has been a problem for fake news research and detection model development. Prior fake news datasets do not provide multimodal text and image data, metadata, comment data, and fine-grained fake news categorization at the scale and breadth of our dataset. We present Fakeddit, a novel multimodal dataset consisting of over 1 million samples from multiple categories of fake news. After being processed through several stages of review, the samples are labeled according to 2-way, 3-way, and 6-way classification categories through distant supervision. We construct hybrid text+image models and perform extensive experiments for multiple variations of classification, demonstrating the importance of the novel aspect of multimodality and fine-grained classification unique to Fakeddit.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. YTCommentVerse: A Multi-Category Multi-Lingual YouTube Comment Corpus

    cs.SI 2025-09 conditional novelty 6.0 of 10

    YTCommentVerse is a release of 32 million YouTube comments from 178,000 videos across 15 categories and 50 languages, with upvotes and anonymized identifiers.

  2. XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs

    cs.CL 2025-08 conditional novelty 6.0 of 10

    XFacta is a new real-world, post-January-2024 multimodal misinformation dataset from X, and evaluations show that MLLM detectors need external evidence, especially image-to-text evidence, with multi-step reasoning per...

  3. Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline

    cs.AI 2025-09 conditional novelty 5.0 of 10

    A unified detector with category-aware mixture-of-experts and attribution chain-of-thought reaches 86.7% accuracy on a new combined human-crafted + AI-generated misinformation benchmark.

  4. An Audit and Analysis of LLM-Assisted Health Misinformation Jailbreaks Against LLMs

    cs.CL 2025-08 conditional novelty 5.0 of 10

    LLM-generated jailbreak prompts elicited health misinformation from GPT-3.5, Llama 3.1-8B, and Gemini 2.0 Flash at high rates, and both LLM judges and simple classifiers detected the resulting texts with high accuracy.

  5. WISE: Web Information Satire and Fakeness Evaluation

    cs.CL 2025-12 conditional novelty 4.0 of 10

    Ten transformers are compared on satire-vs-fake news headlines; MiniLM reaches 87.58% accuracy, beating larger baselines, while RoBERTa has the best ROC-AUC (95.42%).

  6. A Comprehensive Dataset for Human vs. AI Generated Text Detection

    cs.CL 2025-10 reject novelty 4.0 of 10

    A dataset of ~58k NYT articles plus AI rewrites from six LLMs, evaluated with a rewrite-distance baseline reaching 58.35% detection and 8.92% attribution accuracy.

  7. A Survey on False Information Detection: From A Perspective of Propagation on Social Networks

    cs.SI 2025-06 conditional novelty 3.0 of 10

    A survey that organizes propagation-based false information detection into homogeneous and heterogeneous categories, summarizing datasets, methods, and future directions.

Pith tools