Pith. sign in

REVIEW 4 cited by

Want To Reduce Labeling Cost? GPT-3 Can Help

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.13487 v1 pith:437CQS27 submitted 2021-08-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords datagpt-3labelslabelingmanytasksmodelperformance
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Data annotation is a time-consuming and labor-intensive process for many NLP tasks. Although there exist various methods to produce pseudo data labels, they are often task-specific and require a decent amount of labeled data to start with. Recently, the immense language model GPT-3 with 175 billion parameters has achieved tremendous improvement across many few-shot learning tasks. In this paper, we explore ways to leverage GPT-3 as a low-cost data labeler to train other models. We find that, to make the downstream model achieve the same performance on a variety of NLU and NLG tasks, it costs 50% to 96% less to use labels from GPT-3 than using labels from humans. Furthermore, we propose a novel framework of combining pseudo labels from GPT-3 with human labels, which leads to even better performance with limited labeling budget. These results present a cost-effective data labeling methodology that is generalizable to many practical applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning to Select Visual In-Context Demonstrations

    cs.LG 2026-03 reject novelty 5.0 of 10

    A Dueling-DQN agent selects visual in-context demonstrations and outperforms kNN retrieval on objective regression benchmarks but not on subjective preference tasks, per the paper's main table.

  2. The Impostor is Among Us: Can Large Language Models Capture the Complexity of Human Personas?

    cs.HC 2025-01 conditional novelty 5.0 of 10

    Participants distinguished human-written from GPT-4o-generated personas, rating AI personas higher on informativeness, positivity, consistency, and clarity but also higher on stereotypicality.

  3. DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization

    cs.CL 2024-12 reject novelty 4.0 of 10

    Pre-training BART on OCR text and augmenting input with LLM-generated question-answer pairs yields small metric gains on administrative document summarization, measured against LLM-written references.

  4. A Comprehensive Survey of Synthetic Tabular Data Generation

    cs.LG 2025-04 conditional novelty 3.0 of 10

    A structured survey that categorizes synthetic tabular data generation into traditional, diffusion, and LLM-based methods, with a comparative benchmark and a taxonomy of post-processing and evaluation.

Pith tools