Pith. sign in

REVIEW 3 cited by

On the Possibilities of AI-Generated Text Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.04736 v3 pith:X3VHHN36 submitted 2023-04-10 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords textdetectionacrossai-generateddetectorsempiricalincludingsample
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Our work addresses the critical issue of distinguishing text generated by Large Language Models (LLMs) from human-produced text, a task essential for numerous applications. Despite ongoing debate about the feasibility of such differentiation, we present evidence supporting its consistent achievability, except when human and machine text distributions are indistinguishable across their entire support. Drawing from information theory, we argue that as machine-generated text approximates human-like quality, the sample size needed for detection increases. We establish precise sample complexity bounds for detecting AI-generated text, laying groundwork for future research aimed at developing advanced, multi-sample detectors. Our empirical evaluations across multiple datasets (Xsum, Squad, IMDb, and Kaggle FakeNews) confirm the viability of enhanced detection methods. We test various state-of-the-art text generators, including GPT-2, GPT-3.5-Turbo, Llama, Llama-2-13B-Chat-HF, and Llama-2-70B-Chat-HF, against detectors, including oBERTa-Large/Base-Detector, GPTZero. Our findings align with OpenAI's empirical data related to sequence length, marking the first theoretical substantiation for these observations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Attacks on Machine-Text Detectors Retain Stylistic Fingerprints

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A style-aware paraphrasing attack evades all nine tested AI-text detectors at the single-document level, but multi-document analysis makes the attack detectable again.

  2. Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback

    cs.CL 2025-08 conditional novelty 6.0 of 10

    People prefer text containing the words that an instruction-tuned model uses far more than its base version, linking human feedback training to LLM word overuse.

  3. Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training

    cs.AI 2025-05 conditional novelty 5.0 of 10

    Reinforcing the two experts most correlated with thinking tokens improves reasoning accuracy and efficiency in MoE large reasoning models, with gains of up to 10 points on AIME benchmarks.

Pith tools