Pith. sign in

REVIEW 2 cited by

MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09761 v2 pith:VH6PTWUN submitted 2023-11-16 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords annotationfallacymafaldabenchmarkclassificationfallaciesmanualaligns
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce MAFALDA, a benchmark for fallacy classification that merges and unites previous fallacy datasets. It comes with a taxonomy that aligns, refines, and unifies existing classifications of fallacies. We further provide a manual annotation of a part of the dataset together with manual explanations for each annotation. We propose a new annotation scheme tailored for subjective NLP tasks, and a new evaluation method designed to handle subjectivity. We then evaluate several language models under a zero-shot learning setting and human performances on MAFALDA to assess their capability to detect and classify fallacies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AMELIA: A Family of Multi-task End-to-end Language Models for Argumentation

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A single LoRA fine-tuned Llama-3.1-8B-Instruct model trained jointly on eight argument-mining tasks across 19 datasets matches or beats task-specific models, and merged models offer a cheaper compromise.

  2. SLURG: Investigating the Feasibility of Generating Synthetic Online Fallacious Discourse

    cs.CL 2025-04 conditional novelty 4.0 of 10

    DeepHermes-3-Mistral-24B can generate synthetic forum comments with plausible syntactic and vocabulary patterns, and few-shot prompting moves the generated text closer to real Reddit and 4chan data in vocabulary diversity.

Pith tools