REVIEW 13 cited by
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Detecting text generated by modern large language models is thought to be hard, as both LLMs and humans can exhibit a wide range of complex behaviors. However, we find that a score based on contrasting two closely related language models is highly accurate at separating human-generated and machine-generated text. Based on this mechanism, we propose a novel LLM detector that only requires simple calculations using a pair of pre-trained LLMs. The method, called Binoculars, achieves state-of-the-art accuracy without any training data. It is capable of spotting machine text from a range of modern LLMs without any model-specific modifications. We comprehensively evaluate Binoculars on a number of text sources and in varied situations. Over a wide range of document types, Binoculars detects over 90% of generated samples from ChatGPT (and other LLMs) at a false positive rate of 0.01%, despite not being trained on any ChatGPT data.
Forward citations
Cited by 13 Pith papers
-
Attacks on Machine-Text Detectors Retain Stylistic Fingerprints
A style-aware paraphrasing attack evades all nine tested AI-text detectors at the single-document level, but multi-document analysis makes the attack detectable again.
-
UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors
Structural register and narrative-form shifts make AI-written text evade adversarially retrained detectors, winning the ELOQUENT 2026 Voight-Kampff competition.
-
AEyeDE: An Attention-Based Attribution Framework for AI-Generated Text Detection
Attention attribution maps from a white-box proxy Transformer, classified by a lightweight CNN, provide a competitive and interpretable signal for AI-generated text detection.
-
MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark
Adding human-alignment augmentation (roleplaying, BPO, self-refine, RLDF) to machine-generated text both fools existing detectors and improves the generalization of detectors fine-tuned on it.
-
DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection
DEER, a disentangled mixture-of-experts detector with RL-based instance routing, reports F1 gains of about 1.4 in-domain and 5.3 points out-of-domain over prior MGT detectors.
-
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
A new modern Chinese poetry detection benchmark shows most current AI-text detectors are unreliable, particularly when LLMs imitate a human style.
-
WETBench: A Benchmark for Detecting Task-Specific Machine-Generated Text on Wikipedia
WETBench shows that existing machine-generated text detectors, particularly zero-shot methods, underperform on task-specific Wikipedia editing scenarios, with supervised detectors averaging 78% accuracy and zero-shot ...
-
PhantomHunter: Detecting Unseen Privately-Tuned LLM-Generated Text via Family-Aware Learning
PhantomHunter detects text from privately fine-tuned LLMs by learning shared token-probability traits within LLaMA, Gemma and Mistral families, reporting F1 above 96% on held-out derivatives.
-
Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors
Fine-tuning LLMs with DPO to push generated news and abstracts toward human style substantially reduces the F1 scores of state-of-the-art machine-generated text detectors.
-
Human-LLM Coevolution: Evidence from Academic Writing
After ChatGPT-style words were publicly flagged in early 2024, their frequency in arXiv abstracts dropped, while other common LLM-favored words kept rising, suggesting authors are adapting their writing to avoid detection.
-
GPT Editors, Not Authors: The Stylistic Footprint of LLMs in Academic Preprints
Across 2,408 arXiv preprints, LLM-typical word usage does not cluster in any section, indicating that AI assistance, when used, is uniform rather than limited to specific parts of a paper.
-
DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
DP-MGTD claims that applying differential-privacy entity sanitization amplifies human-vs-machine text separability, reaching F1 > 0.99 on MGTBench-2.0 while satisfying an epsilon-DP guarantee.
-
A Comprehensive Dataset for Human vs. AI Generated Text Detection
A dataset of ~58k NYT articles plus AI rewrites from six LLMs, evaluated with a rewrite-distance baseline reaching 58.35% detection and 8.92% attribution accuracy.
Discussion (0). Continue with ORCID to comment.