REVIEW 2 cited by
MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce MAFALDA, a benchmark for fallacy classification that merges and unites previous fallacy datasets. It comes with a taxonomy that aligns, refines, and unifies existing classifications of fallacies. We further provide a manual annotation of a part of the dataset together with manual explanations for each annotation. We propose a new annotation scheme tailored for subjective NLP tasks, and a new evaluation method designed to handle subjectivity. We then evaluate several language models under a zero-shot learning setting and human performances on MAFALDA to assess their capability to detect and classify fallacies.
Forward citations
Cited by 2 Pith papers
-
AMELIA: A Family of Multi-task End-to-end Language Models for Argumentation
A single LoRA fine-tuned Llama-3.1-8B-Instruct model trained jointly on eight argument-mining tasks across 19 datasets matches or beats task-specific models, and merged models offer a cheaper compromise.
-
SLURG: Investigating the Feasibility of Generating Synthetic Online Fallacious Discourse
DeepHermes-3-Mistral-24B can generate synthetic forum comments with plausible syntactic and vocabulary patterns, and few-shot prompting moves the generated text closer to real Reddit and 4chan data in vocabulary diversity.
Discussion (0). Continue with ORCID to comment.