REVIEW 7 cited by
ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Many NLP applications require manual data annotations for a variety of tasks, notably to train classifiers or evaluate the performance of unsupervised models. Depending on the size and degree of complexity, the tasks may be conducted by crowd-workers on platforms such as MTurk as well as trained annotators, such as research assistants. Using a sample of 2,382 tweets, we demonstrate that ChatGPT outperforms crowd-workers for several annotation tasks, including relevance, stance, topics, and frames detection. Specifically, the zero-shot accuracy of ChatGPT exceeds that of crowd-workers for four out of five tasks, while ChatGPT's intercoder agreement exceeds that of both crowd-workers and trained annotators for all tasks. Moreover, the per-annotation cost of ChatGPT is less than $0.003 -- about twenty times cheaper than MTurk. These results show the potential of large language models to drastically increase the efficiency of text classification.
Forward citations
Cited by 7 Pith papers
-
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
With a two-question interface, full-pool LLM annotation (GPT-5.2 or Qwen3.5) beats human-supervised classifiers for anti-immigrant hostility detection at ~1/10 cost; AL does not reliably beat random sampling.
-
Towards Efficient and Effective Alignment of Large Language Models
A thesis presenting Lion, WebR, LTE, BMC, and FollowBench, five empirical methods that together address LLM alignment data, training, and evaluation.
-
SERUM: State Extraction and Refinement for User Modeling
A multi-pass VLM annotation pipeline extracts action and intent state machines from egocentric video, then fits Markov models that improve on frequency baselines by a few points after normalization (activity top-1 48....
-
Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks
Annotators presented with LLM label suggestions adopt them heavily, shifting gold labels and inflating reported LLM F1 scores by up to .35, even when humans review each suggestion.
-
Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation
Prompting LLMs for candidate labels and distilling them into a small model improves annotation accuracy and noise tolerance over single-label annotation.
-
Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution
Prompt injection can make LLM agents leak personal data they observed while executing tasks, with measured attack success rates around 15-20 percent and password leakage much rarer.
-
Using AI to replicate human experimental results: a motion study
Across four motion-verb psycholinguistic tasks, ChatGPT o1 responses correlated strongly with human judgments (Spearman rho around .69 to .96), but methodological gaps limit the strength of the replication claim.
Discussion (0). Sign in to comment.