REVIEW 5 cited by
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they are, the datasets used for evaluation are insufficiently challenging-lacking variations in sampling strategy, adversarial attacks, and open-source generative models. In this work we present RAID: the largest and most challenging benchmark dataset for machine-generated text detection. RAID includes over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies. Using RAID, we evaluate the out-of-domain and adversarial robustness of 8 open- and 4 closed-source detectors and find that current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models. We release our data along with a leaderboard to encourage future research.
Forward citations
Cited by 5 Pith papers
-
MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark
Adding human-alignment augmentation (roleplaying, BPO, self-refine, RLDF) to machine-generated text both fools existing detectors and improves the generalization of detectors fine-tuned on it.
-
Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors
Fine-tuning LLMs with DPO to push generated news and abstracts toward human style substantially reduces the F1 scores of state-of-the-art machine-generated text detectors.
-
Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters
LoRA-adapted 0.5B-7B language models all reach the same automatic rewriting score (0.69), indicating model size does not change measured quality for this single-user style-rewriting task.
-
Hijacking Text Heritage: Hiding the Human Signature through Homoglyphic Substitution
Replacing characters in at least 37.5% of words with visual homoglyphs degrades authorship verification scores enough to obfuscate style, with diminishing returns past 50%.
-
A Comprehensive Dataset for Human vs. AI Generated Text Detection
A dataset of ~58k NYT articles plus AI rewrites from six LLMs, evaluated with a rewrite-distance baseline reaching 58.35% detection and 8.92% attribution accuracy.
Discussion (0). Continue with ORCID to comment.