Pith. sign in

REVIEW 10 cited by

Overtrained Language Models Are Harder to Fine-Tune

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.19206 v2 pith:VDYYBT3S submitted 2025-03-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsperformancepre-trainedpre-trainingassumptioncatastrophicdownstreamfine-tune
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models are pre-trained on ever-growing token budgets under the assumption that better pre-training performance translates to improved downstream models. In this work, we challenge this assumption and show that extended pre-training can make models harder to fine-tune, leading to degraded final performance. We term this phenomenon catastrophic overtraining. For example, the instruction-tuned OLMo-1B model pre-trained on 3T tokens leads to over 2% worse performance on multiple standard LLM benchmarks than its 2.3T token counterpart. Through controlled experiments and theoretical analysis, we show that catastrophic overtraining arises from a systematic increase in the broad sensitivity of pre-trained parameters to modifications, including but not limited to fine-tuning. Our findings call for a critical reassessment of pre-training design that considers the downstream adaptability of the model.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging Compute- and Data-Optimal Pretraining

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Pretraining loss obeys a single law in which repeated or paraphrased tokens count as η(N, data-per-parameter, expansion-ratio) fresh tokens, with total effective data saturating as derived tokens grow.

  2. Understanding Reasoning from Pretraining to Post-Training

    cs.LG 2026-07 conditional novelty 7.0 of 10

    A joint scaling law: post-RL chess and math performance is predictable from pretraining loss, RL improvement rate grows with pretraining tokens, and RL both amplifies and discovers moves depending on difficulty.

  3. Disentangling Geometry, Performance, and Training in Language Models

    cs.CL 2026-02 conditional novelty 7.0 of 10

    Effective rank of the unembedding matrix mainly reflects hyperparameters like batch size and weight decay and is not a reliable predictor of language-model performance.

  4. The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions

    cs.LG 2025-06 conditional novelty 7.0 of 10

    A single-weight perturbation at the very start of training makes otherwise identical neural networks diverge to different loss basins, and this sensitivity drops sharply within the first fraction of training.

  5. Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.

  6. A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A skill-graph random-walk model gives closed-form accuracy-versus-compute formulas for four reasoning strategies and connects them to training scaling.

  7. Manifold-Constrained Hyper-Connections for Parameter-Efficient Finetuning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Applying mHC as a PEFT method shows that learned residual mixing is unnecessary — even harmful — in finetuning, and mHC+LoRA combinations give small task-dependent gains.

  8. Reinitializing weights vs units for maintaining plasticity in neural networks

    cs.NE 2025-07 conditional novelty 5.0 of 10

    Selective weight reinitialization, which resets the least useful weights, maintains plasticity in small and layer-normalized networks where unit-level reinitialization methods fail.

  9. Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Training-free amplification of selected last-layer activations, combined with 'wait' token insertion, elicits long chain-of-thought reasoning in base LLMs and improves accuracy on math and science benchmarks.

  10. The wall confronting large language models

    cs.AI 2025-07 conditional novelty 4.0 of 10

    LLM scaling exponents near 0.1 imply that reducing loss tenfold would need 10^10 more compute, making scientific-grade reliability unreachable by brute-force scaling.

Pith tools