Pith. sign in

REVIEW 2 cited by

Cold-start Active Learning through Self-supervised Language Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.09535 v2 pith:7RZ4SHTB submitted 2020-10-19 cs.CL cs.LG

classification cs.CLcs.LG
keywords activelanguagelearningmodelclassificationlossmodelingcold-start
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Active learning strives to reduce annotation costs by choosing the most critical examples to label. Typically, the active learning strategy is contingent on the classification model. For instance, uncertainty sampling depends on poorly calibrated model confidence scores. In the cold-start setting, active learning is impractical because of model instability and data scarcity. Fortunately, modern NLP provides an additional source of information: pre-trained language models. The pre-training loss can find examples that surprise the model and should be labeled for efficient fine-tuning. Therefore, we treat the language modeling loss as a proxy for classification uncertainty. With BERT, we develop a simple strategy based on the masked language modeling loss that minimizes labeling costs for text classification. Compared to other baselines, our approach reaches higher accuracy within less sampling iterations and computation time.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ALPET: Active Few-shot Learning for Citation Worthiness Detection in Low-Resource Wikipedia Languages

    cs.CL 2025-02 conditional novelty 6.0 of 10

    ALPET, an active-learning plus PET pipeline, detects citation-worthy sentences in Catalan, Basque and Albanian while needing roughly 58-72% fewer labeled examples than its CCW baseline.

  2. Reducing Labeling Effort in Architecture Technical Debt Detection through Active Learning and Explainable AI

    cs.SE 2026-03 conditional novelty 5.0 of 10

    Combining keyword pre-filtering with Breaking-Ties active learning labels 51% of a Jira dataset to detect architecture technical debt at F1 0.72; domain experts prefer LIME over SHAP for explaining predictions.

Pith tools