REVIEW 6 cited by
All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Human evaluations are typically considered the gold standard in natural language generation, but as models' fluency improves, how well can evaluators detect and judge machine-generated text? We run a study assessing non-experts' ability to distinguish between human- and machine-authored text (GPT2 and GPT3) in three domains (stories, news articles, and recipes). We find that, without training, evaluators distinguished between GPT3- and human-authored text at random chance level. We explore three approaches for quickly training evaluators to better identify GPT3-authored text (detailed instructions, annotated examples, and paired examples) and find that while evaluators' accuracy improved up to 55%, it did not significantly improve across the three domains. Given the inconsistent results across text domains and the often contradictory reasons evaluators gave for their judgments, we examine the role untrained human evaluations play in NLG evaluation and provide recommendations to NLG researchers for improving human evaluations of text generated from state-of-the-art models.
Forward citations
Cited by 6 Pith papers
-
Towards Efficient and Effective Alignment of Large Language Models
A thesis presenting Lion, WebR, LTE, BMC, and FollowBench, five empirical methods that together address LLM alignment data, training, and evaluation.
-
Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"
Across seven AI chatbots, Reddit users report mostly reliability failures, with each chatbot showing a distinct pattern of safety, privacy, and security complaints.
-
Is Your LLM-Based Multi-Agent a Reliable Real-World Planner? Exploring Fraud Detection in Travel Planning
Multi-agent LLM travel planners are frequently deceived by injected fake listings, coordinated fake reviews, and multi-round scam conversations, and a simple anti-fraud reviewer helps only some models.
-
The unintended consequences of large language models as a labor-augmenting technology in science
LLM speed-ups in discovery or production raise the opportunity cost of researcher time, making scientists more selective in some phases, less thorough in most, and only deeper when follow-up work itself is accelerated.
-
Psychology-Driven Enhancement of Humour Translation
A decomposition-and-recomposition prompt method for humor translation reports large gains on LLM-based metrics, but the evaluation lacks human validation and statistical checks.
-
LLM Encoder vs. Decoder: Robust Detection of Chinese AI-Generated Text with LoRA
On the NLPCC 2025 Chinese AI-text detection benchmark, LoRA-adapted Qwen2.5-7B reaches 95.94% test accuracy, beating BERT-large (79.3%), RoBERTa-large (76.3%), and FastText (83.5%).
Discussion (0). Sign in to comment.