Pith. sign in

hub

How close is chatgpt to human experts? comparison corpus, evaluation, and detection

26 Pith papers cite this work, alongside 293 external citations. Polarity classification is still indexing.

26 Pith papers citing it
293 external citations · Pith
abstract

The introduction of ChatGPT has garnered widespread attention in both academic and industrial communities. ChatGPT is able to respond effectively to a wide range of human questions, providing fluent and comprehensive answers that significantly surpass previous public chatbots in terms of security and usefulness. On one hand, people are curious about how ChatGPT is able to achieve such strength and how far it is from human experts. On the other hand, people are starting to worry about the potential negative impacts that large language models (LLMs) like ChatGPT could have on society, such as fake news, plagiarism, and social security issues. In this work, we collected tens of thousands of comparison responses from both human experts and ChatGPT, with questions ranging from open-domain, financial, medical, legal, and psychological areas. We call the collected dataset the Human ChatGPT Comparison Corpus (HC3). Based on the HC3 dataset, we study the characteristics of ChatGPT's responses, the differences and gaps from human experts, and future directions for LLMs. We conducted comprehensive human evaluations and linguistic analyses of ChatGPT-generated content compared with that of humans, where many interesting results are revealed. After that, we conduct extensive experiments on how to effectively detect whether a certain text is generated by ChatGPT or humans. We build three different detection systems, explore several key factors that influence their effectiveness, and evaluate them in different scenarios. The dataset, code, and models are all publicly available at https://github.com/Hello-SimpleAI/chatgpt-comparison-detection.

hub tools

citation-role summary

dataset 2 background 1

citation-polarity summary

representative citing papers

PA-User: Simulating Trust and Verification under AI-Generated Content

cs.IR · 2026-06-22 · unverdicted · novelty 6.0

PA-User simulates user trust and verification in AI-generated content scenarios using effort budgets, Beta trust beliefs, and decision rules, showing lower trust-calibration error and regret than ablations on the HC3 corpus.

LLM Self-Recognition: Steering and Retrieving Activation Signatures

cs.AI · 2026-06-04 · unverdicted · novelty 6.0

Steering LLM residual streams with random sparse vectors creates detectable self-recognition fingerprints that enable over 98% accurate attribution of generated text to specific models without degrading output quality.

MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text

cs.CL · 2026-05-07 · unverdicted · novelty 6.0

MELD is a multi-task AI-text detector using auxiliary heads, uncertainty-weighted losses, EMA distillation, and pairwise ranking that reaches 99.9% TPR at 1% FPR on a new held-out benchmark while remaining competitive on the RAID leaderboard.

A Unified Detection Framework for AI-Related Content and Artifacts

stat.ML · 2026-07-08 · conditional · novelty 4.0

The paper develops joint casewise and cellwise MCD estimators for multi-class robust covariance estimation and applies them within a Mahalanobis-distance-based detection framework across four AI-content detection tasks, with convergence and breakdown-point guarantees.

citing papers explorer

Showing 26 of 26 citing papers.