Pith. sign in

REVIEW 6 cited by

All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.00061 v2 pith:5AUE5C3C submitted 2021-06-30 cs.CL

classification cs.CL
keywords textevaluatorshumandomainsevaluationsthreeacrossevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human evaluations are typically considered the gold standard in natural language generation, but as models' fluency improves, how well can evaluators detect and judge machine-generated text? We run a study assessing non-experts' ability to distinguish between human- and machine-authored text (GPT2 and GPT3) in three domains (stories, news articles, and recipes). We find that, without training, evaluators distinguished between GPT3- and human-authored text at random chance level. We explore three approaches for quickly training evaluators to better identify GPT3-authored text (detailed instructions, annotated examples, and paired examples) and find that while evaluators' accuracy improved up to 55%, it did not significantly improve across the three domains. Given the inconsistent results across text domains and the often contradictory reasons evaluators gave for their judgments, we examine the role untrained human evaluations play in NLG evaluation and provide recommendations to NLG researchers for improving human evaluations of text generated from state-of-the-art models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Efficient and Effective Alignment of Large Language Models

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A thesis presenting Lion, WebR, LTE, BMC, and FollowBench, five empirical methods that together address LLM alignment data, training, and evaluation.

  2. Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"

    cs.CY 2025-09 conditional novelty 6.0 of 10

    Across seven AI chatbots, Reddit users report mostly reliability failures, with each chatbot showing a distinct pattern of safety, privacy, and security complaints.

  3. Is Your LLM-Based Multi-Agent a Reliable Real-World Planner? Exploring Fraud Detection in Travel Planning

    cs.MA 2025-05 conditional novelty 6.0 of 10

    Multi-agent LLM travel planners are frequently deceived by injected fake listings, coordinated fake reviews, and multi-round scam conversations, and a simple anti-fraud reviewer helps only some models.

  4. The unintended consequences of large language models as a labor-augmenting technology in science

    physics.soc-ph 2026-07 conditional novelty 5.0 of 10

    LLM speed-ups in discovery or production raise the opportunity cost of researcher time, making scientists more selective in some phases, less thorough in most, and only deeper when follow-up work itself is accelerated.

  5. Psychology-Driven Enhancement of Humour Translation

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A decomposition-and-recomposition prompt method for humor translation reports large gains on LLM-based metrics, but the evaluation lacks human validation and statistical checks.

  6. LLM Encoder vs. Decoder: Robust Detection of Chinese AI-Generated Text with LoRA

    cs.CL 2025-08 conditional novelty 3.0 of 10

    On the NLPCC 2025 Chinese AI-text detection benchmark, LoRA-adapted Qwen2.5-7B reaches 95.94% test accuracy, beating BERT-large (79.3%), RoBERTa-large (76.3%), and FastText (83.5%).

Pith tools