REVIEW 7 cited by
Automated Annotation with Generative AI Requires Validation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generative large language models (LLMs) can be a powerful tool for augmenting text annotation procedures, but their performance varies across annotation tasks due to prompt quality, text data idiosyncrasies, and conceptual difficulty. Because these challenges will persist even as LLM technology improves, we argue that any automated annotation process using an LLM must validate the LLM's performance against labels generated by humans. To this end, we outline a workflow to harness the annotation potential of LLMs in a principled, efficient way. Using GPT-4, we validate this approach by replicating 27 annotation tasks across 11 datasets from recent social science articles in high-impact journals. We find that LLM performance for text annotation is promising but highly contingent on both the dataset and the type of annotation task, which reinforces the necessity to validate on a task-by-task basis. We make available easy-to-use software designed to implement our workflow and streamline the deployment of LLMs for automated annotation.
Forward citations
Cited by 7 Pith papers
-
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
With a two-question interface, full-pool LLM annotation (GPT-5.2 or Qwen3.5) beats human-supervised classifiers for anti-immigrant hostility detection at ~1/10 cost; AL does not reliably beat random sampling.
-
Auditing Differential Visibility of Political Content on TikTok
Account-level analysis finds no evidence of moderate-to-large reach suppression on TikTok for three political topics; an apparent pooled gap is a statistical artifact, while oppositional content earns more engagement ...
-
Grounded Event Extraction from SEC 8-K Filings with a Fine-Grained Taxonomy
Schema-constrained, quote-grounded LLM extraction plus a second-pass quality score yields 601k auditable 8-K event tags whose precision and market reactions both improve with the score.
-
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
A mixed-initiative system where LLMs flag low-confidence annotations, cluster them, and propose codebook rules that human experts review and iterate on.
-
Structuring Radiology Reports: Challenging LLMs with Lightweight Models
Fully finetuned T5 and BERT2BERT models match or beat prompt-adapted LLMs up to 70B parameters on radiology report structuring, at less than 1% of the inference cost.
-
A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol
The teaching-feedback classification protocol remains durable across three representation generations and transfers to English sentiment, so model choice is a deployment decision.
-
Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
LLMs alone annotate search clarifications unreliably; adding confidence-based selective human review cuts effort 24-45% in simulation, but the evaluation is partly built from the ground truth it predicts.
Discussion (0). Continue with ORCID to comment.