Pith. sign in

REVIEW 18 cited by

Inverse Scaling: When Bigger Isn't Better

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.09479 v2 pith:C7MNKQ5A submitted 2023-06-15 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords scalinginversedatatasktrainingdatasetsincreasedmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Work on scaling laws has found that large language models (LMs) show predictable improvements to overall loss with increased scale (model size, training data, and compute). Here, we present evidence for the claim that LMs may show inverse scaling, or worse task performance with increased scale, e.g., due to flaws in the training objective and data. We present empirical evidence of inverse scaling on 11 datasets collected by running a public contest, the Inverse Scaling Prize, with a substantial prize pool. Through analysis of the datasets, along with other examples found in the literature, we identify four potential causes of inverse scaling: (i) preference to repeat memorized sequences over following in-context instructions, (ii) imitation of undesirable patterns in the training data, (iii) tasks containing an easy distractor task which LMs could focus on, rather than the harder real task, and (iv) correct but misleading few-shot demonstrations of the task. We release the winning datasets at https://inversescaling.com/data to allow for further investigation of inverse scaling. Our tasks have helped drive the discovery of U-shaped and inverted-U scaling trends, where an initial trend reverses, suggesting that scaling trends are less reliable at predicting the behavior of larger-scale models than previously understood. Overall, our results suggest that there are tasks for which increased model scale alone may not lead to progress, and that more careful thought needs to go into the data and objectives for training language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

    cs.CL 2025-10 conditional novelty 7.0 of 10

    MoReBench evaluates LLM moral reasoning processes with 1,000 expert-rubric scenarios and finds current models are partial toward Utilitarian and Deontological frameworks.

  2. On the Fitness Landscape in the $NK$ Model

    math.PR 2025-08 unverdicted novelty 7.0 of 10

    For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.

  3. Towards Efficient and Effective Alignment of Large Language Models

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A thesis presenting Lion, WebR, LTE, BMC, and FollowBench, five empirical methods that together address LLM alignment data, training, and evaluation.

  4. Derivational Morphology Reveals Analogical Generalization in Large Language Models

    cs.CL 2024-11 conditional novelty 7.0 of 10

    GPT-J's adjective nominalization behavior is better explained by an exemplar-based analogical model than by a rule-based model, with word frequency effects even for regular forms.

  5. Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation

    cs.SE 2026-07 conditional novelty 6.0 of 10

    Stacking 1 to 20 verifiable instructions into one prompt drops follow rates from about 96% to as low as 20%, and a one-shot rewrite recovers up to +11 points, mainly for weaker models.

  6. Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A generator that produces oracle-verified, deliberately messy code in five languages with orthogonal difficulty labels and a contamination-resistant re-minting protocol.

  7. The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini Models

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Continuous sycophancy scoring reveals that 27% of Gemini responses contain substantial social compliance missed by binary filters, with sycophancy correlating with hallucination and simple guardrails outperforming com...

  8. Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

    cs.CL 2025-07 conditional novelty 6.0 of 10

    HumorBench scores LLM explanations of cartoon jokes against expert-written objective elements and finds reasoning skills transfer from STEM benchmarks, while extra thinking tokens help only some models.

  9. Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    PCPO selects preference pairs by combining correct-answer status with token-level probability consistency, then trains with a weighted DPO+NLL loss, yielding small and partly inconsistent gains over outcome-only metho...

  10. A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Introduces a multi-role red teaming framework using attacker and jury models that increases attack success rates by up to 7.9% on LLM faithfulness in question-answering tasks.

  11. Quantifier Scope Interpretation in Language Learners and LLMs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Across seven LLMs, surface-scope readings dominate "every...a" sentences while "a...every" sentences favor inverse scope, and only BERT-family models reproduce the English-Chinese contrast seen in humans.

  12. Departures from Standard Disk Predictions in Intensive Ground-Based Monitoring of Three AGN

    astro-ph.GA 2025-08 unverdicted novelty 5.0 of 10

    Based only on the abstract, the paper reports that Mrk 509 inter-band continuum lags scale as wavelength to the 2.17 power, steeper than the thin-disk prediction, but the appended full text is a different article.

  13. Toward the Axiomatization of Intelligence: Structure, Time, and Existence

    cs.AI 2025-04 reject novelty 5.0 of 10

    An entity is intelligent, by this axiomatic definition, exactly when it has input, processing, and output structures realized as cardinality-changing subsets inside a time-evolving universe.

  14. VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A new graph-image benchmark generator shows that six large vision-language models are sensitive to layout, labeling, and visual defects across seven graph tasks.

  15. VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation

    cs.RO 2024-11 conditional novelty 5.0 of 10

    VLN-Game maps a room with CLIP features and uses an equilibrium-seeking vision-language model to find objects described by names or natural language, achieving state-of-the-art object-goal navigation on HM3D.

  16. Causality for Natural Language Processing

    cs.CL 2025-04 conditional novelty 3.0 of 10

    A dissertation assembling the author's prior publications on causality for NLP, centered on two LLM causal reasoning benchmarks and their implications.

  17. A Survey on Data Security in Large Language Models

    cs.CR 2025-08 conditional novelty 2.0 of 10

    A survey of data security risks in LLMs that organizes threats, defenses, and evaluation datasets, with notable factual errors in its tables.

  18. Foundations of GenIR

    cs.IR 2025-01 unverdicted novelty 1.0 of 10

    A survey chapter proposing that generative AI reshapes information access through two paradigms, information generation and information synthesis.

Pith tools