REVIEW 18 cited by
Inverse Scaling: When Bigger Isn't Better
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Work on scaling laws has found that large language models (LMs) show predictable improvements to overall loss with increased scale (model size, training data, and compute). Here, we present evidence for the claim that LMs may show inverse scaling, or worse task performance with increased scale, e.g., due to flaws in the training objective and data. We present empirical evidence of inverse scaling on 11 datasets collected by running a public contest, the Inverse Scaling Prize, with a substantial prize pool. Through analysis of the datasets, along with other examples found in the literature, we identify four potential causes of inverse scaling: (i) preference to repeat memorized sequences over following in-context instructions, (ii) imitation of undesirable patterns in the training data, (iii) tasks containing an easy distractor task which LMs could focus on, rather than the harder real task, and (iv) correct but misleading few-shot demonstrations of the task. We release the winning datasets at https://inversescaling.com/data to allow for further investigation of inverse scaling. Our tasks have helped drive the discovery of U-shaped and inverted-U scaling trends, where an initial trend reverses, suggesting that scaling trends are less reliable at predicting the behavior of larger-scale models than previously understood. Overall, our results suggest that there are tasks for which increased model scale alone may not lead to progress, and that more careful thought needs to go into the data and objectives for training language models.
Forward citations
Cited by 18 Pith papers
-
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
MoReBench evaluates LLM moral reasoning processes with 1,000 expert-rubric scenarios and finds current models are partial toward Utilitarian and Deontological frameworks.
-
On the Fitness Landscape in the $NK$ Model
For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.
-
Towards Efficient and Effective Alignment of Large Language Models
A thesis presenting Lion, WebR, LTE, BMC, and FollowBench, five empirical methods that together address LLM alignment data, training, and evaluation.
-
Derivational Morphology Reveals Analogical Generalization in Large Language Models
GPT-J's adjective nominalization behavior is better explained by an exemplar-based analogical model than by a rule-based model, with word frequency effects even for regular forms.
-
Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation
Stacking 1 to 20 verifiable instructions into one prompt drops follow rates from about 96% to as low as 20%, and a one-shot rewrite recovers up to +11 points, mainly for weaker models.
-
Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator
A generator that produces oracle-verified, deliberately messy code in five languages with orthogonal difficulty labels and a contamination-resistant re-minting protocol.
-
The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini Models
Continuous sycophancy scoring reveals that 27% of Gemini responses contain substantial social compliance missed by binary filters, with sycophancy correlating with hallucination and simple guardrails outperforming com...
-
Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
HumorBench scores LLM explanations of cartoon jokes against expert-written objective elements and finds reasoning skills transfer from STEM benchmarks, while extra thinking tokens help only some models.
-
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
PCPO selects preference pairs by combining correct-answer status with token-level probability consistency, then trains with a weighted DPO+NLL loss, yielding small and partly inconsistent gains over outcome-only metho...
-
A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation
Introduces a multi-role red teaming framework using attacker and jury models that increases attack success rates by up to 7.9% on LLM faithfulness in question-answering tasks.
-
Quantifier Scope Interpretation in Language Learners and LLMs
Across seven LLMs, surface-scope readings dominate "every...a" sentences while "a...every" sentences favor inverse scope, and only BERT-family models reproduce the English-Chinese contrast seen in humans.
-
Departures from Standard Disk Predictions in Intensive Ground-Based Monitoring of Three AGN
Based only on the abstract, the paper reports that Mrk 509 inter-band continuum lags scale as wavelength to the 2.17 power, steeper than the thin-disk prediction, but the appended full text is a different article.
-
Toward the Axiomatization of Intelligence: Structure, Time, and Existence
An entity is intelligent, by this axiomatic definition, exactly when it has input, processing, and output structures realized as cardinality-changing subsets inside a time-evolving universe.
-
VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models
A new graph-image benchmark generator shows that six large vision-language models are sensitive to layout, labeling, and visual defects across seven graph tasks.
-
VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation
VLN-Game maps a room with CLIP features and uses an equilibrium-seeking vision-language model to find objects described by names or natural language, achieving state-of-the-art object-goal navigation on HM3D.
-
Causality for Natural Language Processing
A dissertation assembling the author's prior publications on causality for NLP, centered on two LLM causal reasoning benchmarks and their implications.
-
A Survey on Data Security in Large Language Models
A survey of data security risks in LLMs that organizes threats, defenses, and evaluation datasets, with notable factual errors in its tables.
-
Foundations of GenIR
A survey chapter proposing that generative AI reshapes information access through two paradigms, information generation and information synthesis.
Discussion (0). Continue with ORCID to comment.