REVIEW 3 major objections 5 minor 55 references
Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
T0 review · 3 major / 5 minor · reviewed 2026-07-31 · grok-4.5
Pith's one-line read What a model is pretrained on last shapes how much of its SFT refusal survives the next alignment update, even when post-SFT scores match.
desk verdict Clean controlled result: SFT-matched 1B forks diverge under identical DPO/GRPO when the final pretraining window is safety-last; the reporting pitch outruns the dose evidence at saturated scale. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Refusal erosion: the drop in harmful-request refusal rate from the shared SFT checkpoint to the post-training endpoint under a fixed update. Protection is web-branch erosion minus branch erosion, isolating how much of the SFT-installed refusal the shared stage removes rather than how high refusal started.
What would settle it
Repeat the matched-fork design at larger scale or later in training with the same relative final-window dose: if safety-last and web-last branches that still match after SFT no longer separate on refusal erosion under identical DPO or GRPO, the claimed imprint does not hold where it would matter most.
Extended reading notes
Core claim
Checkpoints matched after identical SFT on instruction following, refusal, and capability still diverge under the same post-training update. A final pretraining window of safety transformation text yields substantially lower refusal erosion under UltraFeedback DPO and under GRPO with a verifiable reward than a generic web window, even though the safety branch does not start higher after SFT. The protection is content-selective, requires safety data to come last, and is not a generic benefit of extra late tokens.
Load-bearing premise
That how much refusal is lost on a few harmful-request suites under one fixed helpfulness DPO or arithmetic RL recipe, on small 1B forks, is a fair stand-in for general alignment plasticity and for how production checkpoints should be judged.
Editorial extensions
If this is right
- Two checkpoints with matching post-SFT scores are not interchangeable for the next alignment stage.
- Release cards and model reports should include what data the model saw last in pretraining, not only behavior scores.
- Safety-oriented text placed in the final pretraining window can change how much SFT refusal survives later preference or RL updates that never reward refusal.
- Order matters: the same safety content earlier in late pretraining does not give the same protection as placing it last.
- The effect is dose-relative: as the final window becomes a vanishing fraction of prior tokens, the divergence fades.
Reading between the lines
- Upstream data teams and downstream aligners may need shared contracts about the final pretraining mixture, not only about SFT and preference datasets.
- If last-window imprints generalize beyond refusal, other SFT-installed behaviors (style, tool use, calibration) could also erode differently across behavior-matched checkpoints.
- Reporting last-window provenance would let evaluators test whether apparent alignment gains or losses are really post-training effects or inherited plasticity differences.
- Curriculum design that deliberately ends pretraining on task-proximal critique text, rather than only mixing it earlier, becomes a testable lever for alignment stability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether two checkpoints that are behaviorally matched after identical SFT can nonetheless respond differently to the same post-training update, as a function of the final window of pretraining. Six branches fork from a shared OLMo-2-1B checkpoint at ~49B tokens and differ only in a 500M-token continued-pretraining window (web, DCLM, normative discourse, safety transformation text, math, synthetic education). After identical Tulu-style SFT, the branches are matched within ~1 point on refusal, capability, and IFEval; under identical UltraFeedback DPO they diverge, with the safety-text branch losing substantially less refusal (overall protection ~8.2 pp vs. web, Table 2). The effect is selective to safety content (§4.2), requires that content to come last (order control, §4.3), reproduces under GRPO with verifiable rewards on two tasks (§4.4, App. H), survives a lower DPO learning rate, a preference-data swap to Chatbot Arena, a no-CPT baseline, and a Pythia-1B replication (§4.5, App. G). The authors define refusal erosion E_b(c) and protection P_b(c) stage-aware (Eqs. 4–6), decontaminate evaluation prompts against the safety corpus (App. B), cross-validate the lexical refusal detector against WildGuard (App. A), and report costs honestly: elevated OR-Bench over-refusal and a ~0.9 pp capability cost (Fig. 6, Table 7). Figure 5 shows the protection attenuating from 9.1 → 3.0 → ~0.5 pp as the fixed 500M window shrinks from 1.02% to 0.0125% of prior training. The pap
Significance. If the result holds, it identifies a genuine blind spot in standard checkpoint evaluation: post-SFT behavioral matching does not certify equal readiness for post-training, and the final pretraining window is a load-bearing variable. This is a clean, well-controlled demonstration of path dependence in post-training, with several features that raise confidence beyond the norm for empirical work at this scale: three-seed means with seed SDs and paired protection SDs, Welch tests and ANOVA (Table 9), an explicit order control that rules out a pure exposure account, four negative-content corpora, decontamination with an explicit overlap audit (BeaverTails 196/1000 overlap disclosed and removed), detector cross-validation against a non-lexical classifier on identical completions, a second model family, and a second post-training algorithm class (GRPO with verifiable reward, two tasks). The authors also report the effect's costs and boundary conditions rather than hiding them, which makes the paper useful even to readers who doubt the practical recommendation. The contribution is a reproducible experimental design and a falsifiable phenomenon, not a fitted narrative.
major comments (3)
- [§4.6 / Fig. 5 / App. D (Table 10)] The attenuation result is interpreted as a relative-dose boundary, but the design does not include the discriminating cell. The 4T saturated fork is probed only at 0.0125% relative dose (fixed 500M window). If a window scaled to ~1% of prior training at the 4T fork (~40B tokens) recovered large protection, the relative-dose reading is confirmed and the §5 reporting recommendation stands for production-scale checkpoints; if it did not, the phenomenon requires early-training plasticity (the 49B fork is ~2% of the way through OLMo-2's schedule) and is confined to a regime no shipped checkpoint occupies. The fixed-49B shrinking-window control (Fig. 5, right) varies dose at fixed plasticity and therefore cannot separate these hypotheses. As written, the paper's practical upshot ('what a model was trained on last should be reported') rests on an untested cell. Either run the scaled-window expe
- [App. D, Table 10] The 4T comparison is partly uninformative for a second reason the paper only partially addresses: at the saturated fork the branches sit at ceiling on AdvBench (99.7/99.9% post-DPO), so erosion E_b is compressed toward zero by construction on that benchmark, and the XSTest gap actually reverses sign (-4.0 pp). The authors note the ceiling issue for AdvBench, but the '~0.5 pp overall protection' number in §4.6 averages across benchmarks with very different headroom, mixing a true null with a measurement artifact. A cleaner read of the 4T cell would report per-benchmark protection with SFT starting points and headroom stated, or restrict the dose-response curve (Fig. 5, left) to benchmarks off ceiling. This matters because Fig. 5 is the evidence base for the dose-boundary claim that feeds the recommendation.
- [§3 / §4.6 / Fig. 6] The headline quantity is protection of refusal, but Fig. 6 and the WildGuard harm-axis result (§4.6: 'at most a weak difference... on the harm axis') show that what is retained is a broad refusal prior, including over-refusal of benign prompts, rather than calibrated harm refusal. The framing throughout (e.g., 'the protection requires that safety content arrive last') invites a safety-benefit reading that the paper's own measurements do not support; the retained behavior is refusal plasticity, not safety. This does not undermine the path-dependence claim, which is the real contribution, but the abstract and §5 should state plainly that the protected quantity is refusal rate (with an over-refusal cost), so that practitioners do not read the result as a free alignment intervention.
minor comments (5)
- [§4.2 / Fig. 2] The C_synth partial effect (strong on XSTest, weak on AdvBench) is noted but not discussed. Since C_synth is the one non-safety branch with a clear positive signal, a sentence on why synthetic educational text might partially protect refusal (or a corpus-diagnostic comparison via App. I) would sharpen the selectivity claim.
- [App. F, Table 12] The mechanism diagnostics rest on a single seed and the refusal-direction projection probe 'did not cleanly separate' the branches. This is fine as a negative result, but the section would benefit from stating explicitly that no mechanism is claimed, and that the update-norm equivalence only rules out the trivial frozen-model explanation.
- [§3, Refusal measurement] The lexical detector uses ten patterns matched in the first 300 characters; please state whether the 64-new-token generation budget (Table 4) ever truncates before the prefix region for non-refusing completions, and whether AdvBench/XSTest/BeaverTails use identical generation settings in all figures (App. A uses 256 tokens, which is noted, but the main-text reader has to cross-reference).
- [Table 1] The safety corpus is cycled 2.3 passes to reach 500M tokens while C_web/C_dclm are single-pass samples; the repetition control (App. B, Table 8) addresses this, but a forward pointer from Table 1 to that control would help readers who notice the asymmetry early.
- [References] Several 2026-dated citations (Baek et al., Feng et al., Akter et al., Li et al. 2026) are arXiv preprints central to the motivation; please verify identifiers (e.g., arXiv:2603.16177, 2605.12705) resolve, since these anchor the claim that percent-scale windows are known to matter.
Circularity Check
No significant circularity: empirical fork-and-compare design with erosion defined as a measured difference, not a quantity forced by construction.
full rationale
The paper’s load-bearing claim is experimental, not a first-principles derivation. Six branches fork from one checkpoint, differ only in a fixed 500M-token final window, then receive identical SFT and post-training; refusal erosion is defined as Eb(c)=Rb(θS_c)−Rb(θPT_c) on external benchmarks and protection as the difference of those erosions versus Cweb. That definition isolates the change caused by the shared update; it does not fit a parameter to produce the headline gap, nor does any equation reduce the outcome to the input by construction. Order swap, non-safety corpora, dose attenuation, Pythia replication, GRPO, and WildGuard cross-checks are independent controls, not self-referential closures. Citations are to external corpora, algorithms, and benchmarks; there is no self-citation uniqueness theorem, smuggled ansatz, or renaming of a known law presented as a new derivation. Metric choices (lexical refusal patterns) are validated against an independent classifier and do not force the branch ordering. Score 0 is therefore appropriate.
Assumptions & free parameters
free parameters (4)
- final_window_token_budget_T =
500M tokens (main); 50M/5M and later forks in dose study
- DPO_learning_rate_and_beta =
lr=5e-7, β=0.1 (main)
- lexical_refusal_pattern_set =
10 patterns, 300-char prefix window
- SFT_and_DPO_data_subset_sizes =
100k SFT / 60k DPO pairs
assumptions (6)
- domain assumption Post-SFT behavioral match on refusal, IFEval, and a small capability suite is a sufficient operational notion of 'interchangeable checkpoints' for the paper's contrast.
- domain assumption Refusal of harmful requests installed by a fixed SFT safety component is a valid probe of alignment plasticity under subsequent non-safety-targeted updates.
- domain assumption Lexical string-matching refusal labels are comparable across branches when the generation path is identical (with WildGuard as supporting validation).
- domain assumption Continued pretraining for fixed token budget T on corpus Dc produces a well-defined branch state θc comparable across content types when optimizer settings match.
- domain assumption Standard DPO pairwise objective and GRPO-with-KL-to-SFT are representative instances of 'post-training' for the generalization claim.
- standard math Statistical comparisons with three seeds and Welch/ANOVA tests adequately support branch separation claims at the reported effect sizes.
invented entities (2)
-
refusal erosion Eb(c) and protection Pb(c)
independent evidence
-
final-window pretraining imprint (path-dependent plasticity for alignment)
Cite this review
Pith. "Pith review of Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT." pith.science (2026). https://pith.science/paper/IE33TVTG
@misc{pith2026260725063,
author = {Pith},
title = {Pith review of: Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT},
year = {2026},
howpublished = {\url{https://pith.science/paper/IE33TVTG}},
note = {Machine review of arXiv:2607.25063}
}
read the original abstract
Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same across relevant benchmarks are treated as interchangeable, equally ready for the next alignment stage, typically preference optimization. We ask whether this judgment misses a pretraining imprint: a difference that no post-SFT benchmark reveals, yet that decides how each checkpoint responds to further training. To find out, we run a controlled experiment on the final window of pretraining, the last data trained on before instruction tuning. Six branches fork from one partially pretrained checkpoint and differ only in this window: 500 million tokens, 0.1% to 1% of the tokens that precede it. Each branch trains its window on a single data source: generic web text, filtered web text, normative discourse, safety text, mathematical text, or synthetic educational text. SFT and post-training are then identical. After SFT the branches behave near-identically, within about one point on instruction following, refusal, and capability, yet the same post-training carries them to very different endpoints, under both a direct preference optimization update and a reinforcement learning update with a verifiable reward. We measure this deviation through refusal of harmful requests: when post-training begins the safety text branch refuses no more than the web text branch, yet by the end it has lost far less of its refusal. The other four branches gain little or no protection, so the effect is selective to what the window contained. The protection requires the safety text to arrive last rather than earlier in pretraining, and it reproduces on a second model family. What a model is pretrained on last shapes how it reacts to alignment. Therefore, a checkpoint should not be evaluated by its post-SFT behavior alone, and what it was trained on last should be reported with it.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2001.08361 , year=
Scaling laws for neural language models , author=. arXiv preprint arXiv:2001.08361 , year=
arXiv 2001
-
[2]
Advances in neural information processing systems , volume=
An empirical analysis of compute-optimal large language model training , author=. Advances in neural information processing systems , volume=
-
[3]
arXiv preprint arXiv:2101.00027 , year=
The pile: An 800gb dataset of diverse text for language modeling , author=. arXiv preprint arXiv:2101.00027 , year=
-
[4]
International conference on machine learning , pages=
Pythia: A suite for analyzing large language models across training and scaling , author=. International conference on machine learning , pages=. 2023 , organization=
2023
-
[5]
Groeneveld, Dirk and Beltagy, Iz and Walsh, Evan and Bhagia, Akshita and Kinney, Rodney and Tafjord, Oyvind and Jha, Ananya and Ivison, Hamish and Magnusson, Ian and Wang, Yizhong and others , booktitle=
-
[6]
International Conference on Learning Representations , volume=
Language models scale reliably with over-training and on downstream tasks , author=. International Conference on Learning Representations , volume=
-
[7]
Li, Jeffrey and Fang, Alex and Smyrnis, Georgios and Ivgi, Maor and Jordan, Matt and Gadre, Samir and Bansal, Hritik and Guha, Etash and Keh, Sedrick and Arora, Kushal and others , journal=
-
[8]
arXiv preprint arXiv:2603.16177 , year=
The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data , author=. arXiv preprint arXiv:2603.16177 , year=
Show all 55 references
-
[9]
arXiv preprint arXiv:2605.12705 , year=
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning , author=. arXiv preprint arXiv:2605.12705 , year=
-
[10]
Advances in neural information processing systems , volume=
Training language models to follow instructions with human feedback , author=. Advances in neural information processing systems , volume=
-
[11]
Constitutional
Bai, Yuntao and Kadavath, Saurav and Kundu, Sandipan and Askell, Amanda and Kernion, Jackson and Jones, Andy and Chen, Anna and Goldie, Anna and Mirhoseini, Azalia and McKinnon, Cameron and others , journal=. Constitutional
-
[12]
Advances in neural information processing systems , volume=
Direct preference optimization: Your language model is secretly a reward model , author=. Advances in neural information processing systems , volume=
-
[13]
Advances in Neural Information Processing Systems , volume=
How far can camels go? exploring the state of instruction tuning on open resources , author=. Advances in Neural Information Processing Systems , volume=
-
[14]
Proceedings of the 41st International Conference on Machine Learning , articleno =
Cui, Ganqu and Yuan, Lifan and Ding, Ning and Yao, Guanming and He, Bingxiang and Zhu, Wei and Ni, Yuan and Xie, Guotong and Xie, Ruobing and Lin, Yankai and Liu, Zhiyuan and Sun, Maosong , title =. Proceedings of the 41st International Conference on Machine Learning , article...
2024
-
[15]
arXiv preprint arXiv:2502.14768 , year=
Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning , author=. arXiv preprint arXiv:2502.14768 , year=
-
[16]
2023 , eprint=
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena , author=. 2023 , eprint=
2023
-
[17]
arXiv preprint arXiv:2307.15043 , year=
Universal and transferable adversarial attacks on aligned language models , author=. arXiv preprint arXiv:2307.15043 , year=
-
[18]
XST est: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
R. XST est: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. doi:10...
2024 doi
-
[19]
Ji, Jiaming and Liu, Mickel and Dai, Josef and Pan, Xuehai and Zhang, Chi and Bian, Ce and Chen, Boyuan and Sun, Ruiyang and Wang, Yizhou and Yang, Yaodong , journal=
-
[20]
2025 , url=
Justin Cui and Wei-Lin Chiang and Ion Stoica and Cho-Jui Hsieh , booktitle=. 2025 , url=
2025
-
[21]
Advances in Neural Information Processing Systems , volume=
The fineweb datasets: Decanting the web for the finest text data at scale , author=. Advances in Neural Information Processing Systems , volume=
-
[22]
International Conference on Learning Representations , volume=
Openwebmath: An open dataset of high-quality mathematical web text , author=. International Conference on Learning Representations , volume=
-
[23]
Ben Allal, Loubna and Lozhkov, Anton and Penedo, Guilherme and Wolf, Thomas and von Werra, Leandro , title =
-
[24]
arXiv preprint arXiv:2204.05862 , year=
Training a helpful and harmless assistant with reinforcement learning from human feedback , author=. arXiv preprint arXiv:2204.05862 , year=
-
[25]
International Conference on Learning Representations , year=
Measuring Massive Multitask Language Understanding , author=. International Conference on Learning Representations , year=
-
[26]
Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin , booktitle=
-
[27]
arXiv preprint arXiv:1803.05457 , year=
Think you have solved question answering? try arc, the ai2 reasoning challenge , author=. arXiv preprint arXiv:1803.05457 , year=
-
[28]
2021 , publisher=
Sakaguchi, Keisuke and Bras, Ronan Le and Bhagavatula, Chandra and Choi, Yejin , journal=. 2021 , publisher=
2021
-
[29]
arXiv preprint arXiv:2407.21783 , year=
The llama 3 herd of models , author=. arXiv preprint arXiv:2407.21783 , year=
-
[30]
arXiv preprint arXiv:2501.00656 , year=
2 OLMo 2 Furious , author=. arXiv preprint arXiv:2501.00656 , year=
-
[31]
Advances in neural information processing systems , volume=
Openassistant conversations-democratizing large language model alignment , author=. Advances in neural information processing systems , volume=
-
[32]
Advances in Neural Information Processing Systems , volume=
The PRISM alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models , author=. Advances in Neural Information Processing Systems , volume=
-
[33]
Wang, Zhilin and Dong, Yi and Delalleau, Olivier and Zeng, Jiaqi and Shen, Gerald and Egert, Daniel and Zhang, Jimmy J and Sreedhar, Makesh N and Kuchaiev, Oleksii , journal=
-
[34]
Kim, Seungone and Shin, Jamin and Cho, Yejin and Jang, Joel and Longpre, Shayne and Lee, Hwaran and Yun, Sangdoo and Shin, Seongjin and Kim, Sungdong and Thorne, James and Seo, Minjoon , booktitle=
-
[35]
arXiv preprint arXiv:2206.02841 , year=
Researching Alignment Research: Unsupervised Analysis , author=. arXiv preprint arXiv:2206.02841 , year=
-
[36]
Ji, Jiaming and Hong, Donghai and Zhang, Borong and Chen, Boyuan and Dai, Josef and Zheng, Boren and Qiu, Tianyi Alex and Zhou, Jiayi and Wang, Kaile and Li, Boxun and others , booktitle=
-
[37]
Findings of the association for computational linguistics: EMNLP 2020 , pages=
Realtoxicityprompts: Evaluating neural toxic degeneration in language models , author=. Findings of the association for computational linguistics: EMNLP 2020 , pages=
2020
-
[38]
Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
2023
-
[39]
Aligning
Hendrycks, Dan and Burns, Collin and Basart, Steven and Critch, Andrew and Li, Jerry and Song, Dawn and Steinhardt, Jacob , journal=. Aligning
-
[40]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
The multilingual alignment prism: Aligning global and local preferences to reduce harm , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
2024
-
[41]
arXiv preprint arXiv:2605.02087 , year=
Model spec midtraining: Improving how alignment training generalizes , author=. arXiv preprint arXiv:2605.02087 , year=
-
[42]
arXiv preprint arXiv:2311.07911 , year=
Instruction-Following Evaluation for Large Language Models , author=. arXiv preprint arXiv:2311.07911 , year=
-
[43]
2025 , howpublished =
2025
-
[44]
Han, Seungju and Rao, Kavel and Ettinger, Allyson and Jiang, Liwei and Lin, Bill Yuchen and Lambert, Nathan and Choi, Yejin and Dziri, Nouha , journal=
-
[45]
International Conference on Learning Representations , volume=
Fine-tuning aligned language models compromises safety, even when users do not intend to! , author=. International Conference on Learning Representations , volume=
-
[46]
Advances in Neural Information Processing Systems , volume=
Refusal in language models is mediated by a single direction , author=. Advances in Neural Information Processing Systems , volume=
-
[47]
Findings of the Association for Computational Linguistics: ACL 2023 , pages=
Discovering Language Model Behaviors with Model-Written Evaluations , author=. Findings of the Association for Computational Linguistics: ACL 2023 , pages=. 2023 , publisher=
2023
-
[48]
International Conference on Learning Representations , volume=
Towards understanding sycophancy in language models , author=. International Conference on Learning Representations , volume=
-
[49]
Zenodo , year=
A framework for few-shot language model evaluation , author=. Zenodo , year=
-
[50]
Advances in Neural Information Processing Systems (NeurIPS) , year=
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=
-
[51]
International conference on machine learning , pages=
Linear mode connectivity and the lottery ticket hypothesis , author=. International conference on machine learning , pages=. 2020 , organization=
2020
-
[52]
Advances in neural information processing systems , volume=
What is being transferred in transfer learning? , author=. Advances in neural information processing systems , volume=
-
[53]
International Conference on Machine Learning , pages=
Understanding plasticity in neural networks , author=. International Conference on Machine Learning , pages=. 2023 , organization=
2023
-
[54]
Nature , volume=
Loss of plasticity in deep continual learning , author=. Nature , volume=. 2024 , publisher=
2024
-
[55]
The Fourteenth International Conference on Learning Representations , year=
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data , author=. The Fourteenth International Conference on Learning Representations , year=
Reviewed July 31, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.