REVIEW 10 cited by
It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
When scaled to hundreds of billions of parameters, pretrained language models such as GPT-3 (Brown et al., 2020) achieve remarkable few-shot performance. However, enormous amounts of compute are required for training and applying such big models, resulting in a large carbon footprint and making it difficult for researchers and practitioners to use them. We show that performance similar to GPT-3 can be obtained with language models that are much "greener" in that their parameter count is several orders of magnitude smaller. This is achieved by converting textual inputs into cloze questions that contain a task description, combined with gradient-based optimization; exploiting unlabeled data gives further improvements. We identify key factors required for successful natural language understanding with small language models.
Forward citations
Cited by 10 Pith papers
-
Machine Understanding of Scientific Language
The thesis defines and evaluates tasks and datasets for automatic fact checking, cite-worthiness, exaggeration detection, and information change measurement in science communication, culminating in SPICED, a cross-med...
-
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.
-
ArgCMV: An Argument Summarization Benchmark for the LLM-era
ArgCMV is a new LLM-curated benchmark of about 12,000 arguments from r/ChangeMyView, and current key point extraction methods transfer poorly to it.
-
Tiny Reward Models
TinyRM shows that 400M-parameter bidirectional masked language models, tuned with FLAN-style prompting, DoRA, and layer freezing, outperform a 70B reward model on RewardBench reasoning and come close on safety.
-
S$^2$GPT-PINNs: Sparse and Small models for PDEs
S2GPT-PINN sparsifies GPT-PINN's collocation points through empirical interpolation and residual selection, matching GPT-PINN accuracy on four parametric PDEs with much smaller sparse grids.
-
BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning
A 90%-pruned few-shot Bengali model is reported to rival larger baselines on some tasks, but the reported F1 scores contradict the paper's own precision and recall values.
-
Interfaze: The Future of AI is built on Task-Specific Small Models
Interfaze-Beta uses small specialist models and tools to build a compact context that a general-purpose LLM answers from, reporting competitive benchmark scores without reproducible evidence.
-
Domain-Adaptive Small Language Models for Structured Tax Code Prediction
Fine-tuning a small encoder-decoder T5 model with hierarchical tax-code tokens and constrained beam search improves HSN/SAC code prediction over flat classifiers and other SLM architectures in the authors' internal benchmark.
-
LLM-Guided Agentic Object Detection for Open-World Understanding
An LLM generates scene-specific object names that are fed to YOLO-World, enabling label-free open-world detection evaluated with new CAAP and SNAP metrics.
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
Discussion (0). Continue with ORCID to comment.