Pith. sign in

REVIEW 2 cited by

Towards Zero-Label Language Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.09193 v1 pith:VJDODCYW submitted 2021-09-19 cs.CL cs.LG

classification cs.CLcs.LG
keywords datamodelslanguagelearningtrainingzero-labelapproachbetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper explores zero-label learning in Natural Language Processing (NLP), whereby no human-annotated data is used anywhere during training and models are trained purely on synthetic data. At the core of our framework is a novel approach for better leveraging the powerful pretrained language models. Specifically, inspired by the recent success of few-shot inference on GPT-3, we present a training data creation procedure named Unsupervised Data Generation (UDG), which leverages few-shot prompts to synthesize high-quality training data without real human annotations. Our method enables zero-label learning as we train task-specific models solely on the synthetic data, yet we achieve better or comparable results from strong baseline models trained on human-labeled data. Furthermore, when mixed with labeled data, our approach serves as a highly effective data augmentation procedure, achieving new state-of-the-art results on the SuperGLUE benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Gradient Descent Simulate Prompting?

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A MAML-style meta-training objective makes a single gradient step on new text recover part of the performance that prompting achieves, on reversal-curse and passage-QA tasks.

  2. DeepThink: Aligning Language Models with Domain-Specific User Intents

    cs.CL 2025-02 conditional novelty 6.0 of 10

    DeepThink improves domain-specific QA by synthesizing conversation-based training data and refining answers with retrieval-augmented feedback, beating a GPT-4-turbo+RAG assistant by 7.92% on advertising-domain real us...

Pith tools