Pith. sign in

REVIEW 10 cited by

It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.07118 v2 pith:HGCCVCZP submitted 2020-09-15 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagemodelsfew-shotgpt-3performancerequiredsmallachieve
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When scaled to hundreds of billions of parameters, pretrained language models such as GPT-3 (Brown et al., 2020) achieve remarkable few-shot performance. However, enormous amounts of compute are required for training and applying such big models, resulting in a large carbon footprint and making it difficult for researchers and practitioners to use them. We show that performance similar to GPT-3 can be obtained with language models that are much "greener" in that their parameter count is several orders of magnitude smaller. This is achieved by converting textual inputs into cloze questions that contain a task description, combined with gradient-based optimization; exploiting unlabeled data gives further improvements. We identify key factors required for successful natural language understanding with small language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Machine Understanding of Scientific Language

    cs.CL 2025-06 conditional novelty 7.0 of 10

    The thesis defines and evaluates tasks and datasets for automatic fact checking, cite-worthiness, exaggeration detection, and information change measurement in science communication, culminating in SPICED, a cross-med...

  2. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  3. ArgCMV: An Argument Summarization Benchmark for the LLM-era

    cs.CL 2025-08 conditional novelty 6.0 of 10

    ArgCMV is a new LLM-curated benchmark of about 12,000 arguments from r/ChangeMyView, and current key point extraction methods transfer poorly to it.

  4. Tiny Reward Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    TinyRM shows that 400M-parameter bidirectional masked language models, tuned with FLAN-style prompting, DoRA, and layer freezing, outperform a 70B reward model on RewardBench reasoning and come close on safety.

  5. S$^2$GPT-PINNs: Sparse and Small models for PDEs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    S2GPT-PINN sparsifies GPT-PINN's collocation points through empirical interpolation and residual selection, matching GPT-PINN accuracy on four parametric PDEs with much smaller sparse grids.

  6. BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning

    cs.LG 2026-08 reject novelty 4.0 of 10

    A 90%-pruned few-shot Bengali model is reported to rival larger baselines on some tasks, but the reported F1 scores contradict the paper's own precision and recall values.

  7. Interfaze: The Future of AI is built on Task-Specific Small Models

    cs.AI 2026-02 reject novelty 4.0 of 10

    Interfaze-Beta uses small specialist models and tools to build a compact context that a general-purpose LLM answers from, reporting competitive benchmark scores without reproducible evidence.

  8. Domain-Adaptive Small Language Models for Structured Tax Code Prediction

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Fine-tuning a small encoder-decoder T5 model with hierarchical tax-code tokens and constrained beam search improves HSN/SAC code prediction over flat classifiers and other SLM architectures in the authors' internal benchmark.

  9. LLM-Guided Agentic Object Detection for Open-World Understanding

    cs.CV 2025-07 conditional novelty 4.0 of 10

    An LLM generates scene-specific object names that are fed to YOLO-World, enabling label-free open-world detection evaluated with new CAAP and SNAP metrics.

  10. Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs

    cs.LG 2025-02

Pith tools