Pith. sign in

REVIEW 7 cited by

ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.15056 v2 pith:45NSUTS7 submitted 2023-03-27 cs.CL cs.CY

classification cs.CLcs.CY
keywords taskschatgptcrowd-workersannotatorsexceedsmodelsmturkoutperforms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Many NLP applications require manual data annotations for a variety of tasks, notably to train classifiers or evaluate the performance of unsupervised models. Depending on the size and degree of complexity, the tasks may be conducted by crowd-workers on platforms such as MTurk as well as trained annotators, such as research assistants. Using a sample of 2,382 tweets, we demonstrate that ChatGPT outperforms crowd-workers for several annotation tasks, including relevance, stance, topics, and frames detection. Specifically, the zero-shot accuracy of ChatGPT exceeds that of crowd-workers for four out of five tasks, while ChatGPT's intercoder agreement exceeds that of both crowd-workers and trained annotators for all tasks. Moreover, the per-annotation cost of ChatGPT is less than $0.003 -- about twenty times cheaper than MTurk. These results show the potential of large language models to drastically increase the efficiency of text classification.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 71 citations worldwide. Full citation record

  1. Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection

    cs.CL 2026-04 conditional novelty 7.0 of 10

    With a two-question interface, full-pool LLM annotation (GPT-5.2 or Qwen3.5) beats human-supervised classifiers for anti-immigrant hostility detection at ~1/10 cost; AL does not reliably beat random sampling.

  2. Towards Efficient and Effective Alignment of Large Language Models

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A thesis presenting Lion, WebR, LTE, BMC, and FollowBench, five empirical methods that together address LLM alignment data, training, and evaluation.

  3. SERUM: State Extraction and Refinement for User Modeling

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A multi-pass VLM annotation pipeline extracts action and intent state machines from egocentric video, then fits Markov models that improve on frequency baselines by a few points after normalization (activity top-1 48....

  4. Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks

    cs.CY 2025-07 conditional novelty 6.0 of 10

    Annotators presented with LLM label suggestions adopt them heavily, shifting gold labels and inflating reported LLM F1 scores by up to .35, even when humans review each suggestion.

  5. Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Prompting LLMs for candidate labels and distilling them into a small model improves annotation accuracy and noise tolerance over single-label annotation.

  6. Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution

    cs.CR 2025-06 conditional novelty 5.0 of 10

    Prompt injection can make LLM agents leak personal data they observed while executing tasks, with measured attack success rates around 15-20 percent and password leakage much rarer.

  7. Using AI to replicate human experimental results: a motion study

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Across four motion-verb psycholinguistic tasks, ChatGPT o1 responses correlated strongly with human judgments (Spearman rho around .69 to .96), but methodological gaps limit the strength of the replication claim.

Pith tools