Pith. sign in

REVIEW 3 major objections 5 minor 55 references

Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT

T0 review · 3 major / 5 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read What a model is pretrained on last shapes how much of its SFT refusal survives the next alignment update, even when post-SFT scores match.

desk verdict Clean controlled result: SFT-matched 1B forks diverge under identical DPO/GRPO when the final pretraining window is safety-last; the reporting pitch outruns the dose evidence at saturated scale. read the letter →

arxiv 2607.25063 v1 pith:IE33TVTG submitted 2026-07-27 cs.AI

classification cs.AI
keywords final-windowpretrainingpost-trainingpathdependencerefusalerosionsupervisedfine-tuningdirectpreferenceoptimizationGRPOalignmentplasticitycheckpointevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Developers treat two checkpoints as interchangeable once they score about the same after supervised fine-tuning on instruction following, refusal, and capability. This paper argues that judgment misses a pretraining imprint set by the final data window before instruction tuning. In a controlled fork from one partially pretrained 1B checkpoint, six branches receive the same 500M-token final window budget on different single sources, then identical SFT and post-training. After SFT they sit within about one point on the usual release criteria, yet the same preference update and the same reinforcement-learning update drive them to different endpoints. The safety-transformation branch does not start with higher refusal; it simply loses far less of it. The effect is selective to that content, requires the safety text to arrive last rather than earlier, weakens as the window shrinks to a tiny fraction of prior training, and appears on a second model family. If true, post-SFT behavior alone is not enough to certify readiness for alignment, and the last pretraining data should be reported with the checkpoint.

What carries the argument

Refusal erosion: the drop in harmful-request refusal rate from the shared SFT checkpoint to the post-training endpoint under a fixed update. Protection is web-branch erosion minus branch erosion, isolating how much of the SFT-installed refusal the shared stage removes rather than how high refusal started.

What would settle it

Repeat the matched-fork design at larger scale or later in training with the same relative final-window dose: if safety-last and web-last branches that still match after SFT no longer separate on refusal erosion under identical DPO or GRPO, the claimed imprint does not hold where it would matter most.

Watch

Extended reading notes

Core claim

Checkpoints matched after identical SFT on instruction following, refusal, and capability still diverge under the same post-training update. A final pretraining window of safety transformation text yields substantially lower refusal erosion under UltraFeedback DPO and under GRPO with a verifiable reward than a generic web window, even though the safety branch does not start higher after SFT. The protection is content-selective, requires safety data to come last, and is not a generic benefit of extra late tokens.

Load-bearing premise

That how much refusal is lost on a few harmful-request suites under one fixed helpfulness DPO or arithmetic RL recipe, on small 1B forks, is a fair stand-in for general alignment plasticity and for how production checkpoints should be judged.

Editorial extensions

If this is right

  • Two checkpoints with matching post-SFT scores are not interchangeable for the next alignment stage.
  • Release cards and model reports should include what data the model saw last in pretraining, not only behavior scores.
  • Safety-oriented text placed in the final pretraining window can change how much SFT refusal survives later preference or RL updates that never reward refusal.
  • Order matters: the same safety content earlier in late pretraining does not give the same protection as placing it last.
  • The effect is dose-relative: as the final window becomes a vanishing fraction of prior tokens, the divergence fades.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Upstream data teams and downstream aligners may need shared contracts about the final pretraining mixture, not only about SFT and preference datasets.
  • If last-window imprints generalize beyond refusal, other SFT-installed behaviors (style, tool use, calibration) could also erode differently across behavior-matched checkpoints.
  • Reporting last-window provenance would let evaluators test whether apparent alignment gains or losses are really post-training effects or inherited plasticity differences.
  • Curriculum design that deliberately ends pretraining on task-proximal critique text, rather than only mixing it earlier, becomes a testable lever for alignment stability.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper asks whether two checkpoints that are behaviorally matched after identical SFT can nonetheless respond differently to the same post-training update, as a function of the final window of pretraining. Six branches fork from a shared OLMo-2-1B checkpoint at ~49B tokens and differ only in a 500M-token continued-pretraining window (web, DCLM, normative discourse, safety transformation text, math, synthetic education). After identical Tulu-style SFT, the branches are matched within ~1 point on refusal, capability, and IFEval; under identical UltraFeedback DPO they diverge, with the safety-text branch losing substantially less refusal (overall protection ~8.2 pp vs. web, Table 2). The effect is selective to safety content (§4.2), requires that content to come last (order control, §4.3), reproduces under GRPO with verifiable rewards on two tasks (§4.4, App. H), survives a lower DPO learning rate, a preference-data swap to Chatbot Arena, a no-CPT baseline, and a Pythia-1B replication (§4.5, App. G). The authors define refusal erosion E_b(c) and protection P_b(c) stage-aware (Eqs. 4–6), decontaminate evaluation prompts against the safety corpus (App. B), cross-validate the lexical refusal detector against WildGuard (App. A), and report costs honestly: elevated OR-Bench over-refusal and a ~0.9 pp capability cost (Fig. 6, Table 7). Figure 5 shows the protection attenuating from 9.1 → 3.0 → ~0.5 pp as the fixed 500M window shrinks from 1.02% to 0.0125% of prior training. The pap

Significance. If the result holds, it identifies a genuine blind spot in standard checkpoint evaluation: post-SFT behavioral matching does not certify equal readiness for post-training, and the final pretraining window is a load-bearing variable. This is a clean, well-controlled demonstration of path dependence in post-training, with several features that raise confidence beyond the norm for empirical work at this scale: three-seed means with seed SDs and paired protection SDs, Welch tests and ANOVA (Table 9), an explicit order control that rules out a pure exposure account, four negative-content corpora, decontamination with an explicit overlap audit (BeaverTails 196/1000 overlap disclosed and removed), detector cross-validation against a non-lexical classifier on identical completions, a second model family, and a second post-training algorithm class (GRPO with verifiable reward, two tasks). The authors also report the effect's costs and boundary conditions rather than hiding them, which makes the paper useful even to readers who doubt the practical recommendation. The contribution is a reproducible experimental design and a falsifiable phenomenon, not a fitted narrative.

major comments (3)
  1. [§4.6 / Fig. 5 / App. D (Table 10)] The attenuation result is interpreted as a relative-dose boundary, but the design does not include the discriminating cell. The 4T saturated fork is probed only at 0.0125% relative dose (fixed 500M window). If a window scaled to ~1% of prior training at the 4T fork (~40B tokens) recovered large protection, the relative-dose reading is confirmed and the §5 reporting recommendation stands for production-scale checkpoints; if it did not, the phenomenon requires early-training plasticity (the 49B fork is ~2% of the way through OLMo-2's schedule) and is confined to a regime no shipped checkpoint occupies. The fixed-49B shrinking-window control (Fig. 5, right) varies dose at fixed plasticity and therefore cannot separate these hypotheses. As written, the paper's practical upshot ('what a model was trained on last should be reported') rests on an untested cell. Either run the scaled-window expe
  2. [App. D, Table 10] The 4T comparison is partly uninformative for a second reason the paper only partially addresses: at the saturated fork the branches sit at ceiling on AdvBench (99.7/99.9% post-DPO), so erosion E_b is compressed toward zero by construction on that benchmark, and the XSTest gap actually reverses sign (-4.0 pp). The authors note the ceiling issue for AdvBench, but the '~0.5 pp overall protection' number in §4.6 averages across benchmarks with very different headroom, mixing a true null with a measurement artifact. A cleaner read of the 4T cell would report per-benchmark protection with SFT starting points and headroom stated, or restrict the dose-response curve (Fig. 5, left) to benchmarks off ceiling. This matters because Fig. 5 is the evidence base for the dose-boundary claim that feeds the recommendation.
  3. [§3 / §4.6 / Fig. 6] The headline quantity is protection of refusal, but Fig. 6 and the WildGuard harm-axis result (§4.6: 'at most a weak difference... on the harm axis') show that what is retained is a broad refusal prior, including over-refusal of benign prompts, rather than calibrated harm refusal. The framing throughout (e.g., 'the protection requires that safety content arrive last') invites a safety-benefit reading that the paper's own measurements do not support; the retained behavior is refusal plasticity, not safety. This does not undermine the path-dependence claim, which is the real contribution, but the abstract and §5 should state plainly that the protected quantity is refusal rate (with an over-refusal cost), so that practitioners do not read the result as a free alignment intervention.
minor comments (5)
  1. [§4.2 / Fig. 2] The C_synth partial effect (strong on XSTest, weak on AdvBench) is noted but not discussed. Since C_synth is the one non-safety branch with a clear positive signal, a sentence on why synthetic educational text might partially protect refusal (or a corpus-diagnostic comparison via App. I) would sharpen the selectivity claim.
  2. [App. F, Table 12] The mechanism diagnostics rest on a single seed and the refusal-direction projection probe 'did not cleanly separate' the branches. This is fine as a negative result, but the section would benefit from stating explicitly that no mechanism is claimed, and that the update-norm equivalence only rules out the trivial frozen-model explanation.
  3. [§3, Refusal measurement] The lexical detector uses ten patterns matched in the first 300 characters; please state whether the 64-new-token generation budget (Table 4) ever truncates before the prefix region for non-refusing completions, and whether AdvBench/XSTest/BeaverTails use identical generation settings in all figures (App. A uses 256 tokens, which is noted, but the main-text reader has to cross-reference).
  4. [Table 1] The safety corpus is cycled 2.3 passes to reach 500M tokens while C_web/C_dclm are single-pass samples; the repetition control (App. B, Table 8) addresses this, but a forward pointer from Table 1 to that control would help readers who notice the asymmetry early.
  5. [References] Several 2026-dated citations (Baek et al., Feng et al., Akter et al., Li et al. 2026) are arXiv preprints central to the motivation; please verify identifiers (e.g., arXiv:2603.16177, 2605.12705) resolve, since these anchor the claim that percent-scale windows are known to matter.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical fork-and-compare design with erosion defined as a measured difference, not a quantity forced by construction.

full rationale

The paper’s load-bearing claim is experimental, not a first-principles derivation. Six branches fork from one checkpoint, differ only in a fixed 500M-token final window, then receive identical SFT and post-training; refusal erosion is defined as Eb(c)=Rb(θS_c)−Rb(θPT_c) on external benchmarks and protection as the difference of those erosions versus Cweb. That definition isolates the change caused by the shared update; it does not fit a parameter to produce the headline gap, nor does any equation reduce the outcome to the input by construction. Order swap, non-safety corpora, dose attenuation, Pythia replication, GRPO, and WildGuard cross-checks are independent controls, not self-referential closures. Citations are to external corpora, algorithms, and benchmarks; there is no self-citation uniqueness theorem, smuggled ansatz, or renaming of a known law presented as a new derivation. Metric choices (lexical refusal patterns) are validated against an independent classifier and do not force the branch ordering. Score 0 is therefore appropriate.

Assumptions & free parameters 4 free parameters · 6 assumptions · 2 invented entities

The paper is an empirical systems experiment. Load-bearing commitments are domain assumptions about measurement and training setup rather than free physical constants or invented particles. The main quantities (erosion, protection) are operational definitions, not fitted laws. Free choices that affect magnitude include window size, fork point, DPO LR/β, and the lexical refusal rule—none of which are tuned to invent the sign of the effect, but all bound external validity.

free parameters (4)
  • final_window_token_budget_T = 500M tokens (main); 50M/5M and later forks in dose study
    Fixed at 500M tokens (~1% of 49B-fork history); dose sweeps later vary T or fork point. Chosen as a probe scale, not fit to maximize protection.
  • DPO_learning_rate_and_beta = lr=5e-7, β=0.1 (main)
    Main lr 5e-7, β=0.1; lower-lr check at 2e-7. Standard recipe knobs that change absolute erosion but are shown to preserve the gap direction.
  • lexical_refusal_pattern_set = 10 patterns, 300-char prefix window
    Ten fixed AdvBench-style prefix patterns in first 300 characters define Rb(θ). Absolute rates depend on this detector; branch gaps are cross-checked with WildGuard.
  • SFT_and_DPO_data_subset_sizes = 100k SFT / 60k DPO pairs
    100k Tulu-style SFT examples and 60k UltraFeedback pairs held fixed across branches; sizes are design choices of the shared post-training recipe.
assumptions (6)
  • domain assumption Post-SFT behavioral match on refusal, IFEval, and a small capability suite is a sufficient operational notion of 'interchangeable checkpoints' for the paper's contrast.
    Stated in intro and Fig. 1(a); the claim that benchmarks miss an imprint depends on this evaluation surface being what developers actually use.
  • domain assumption Refusal of harmful requests installed by a fixed SFT safety component is a valid probe of alignment plasticity under subsequent non-safety-targeted updates.
    Methodology §3 and evaluation design; erosion isolates change caused by shared PT.
  • domain assumption Lexical string-matching refusal labels are comparable across branches when the generation path is identical (with WildGuard as supporting validation).
    §3 Refusal measurement and Appendix A; absolute rates may be biased but differences are treated as meaningful.
  • domain assumption Continued pretraining for fixed token budget T on corpus Dc produces a well-defined branch state θc comparable across content types when optimizer settings match.
    Eq. (1)–(2); standard CPT assumption underlying the fork design.
  • domain assumption Standard DPO pairwise objective and GRPO-with-KL-to-SFT are representative instances of 'post-training' for the generalization claim.
    §3–4.4; two algorithms support 'not DPO-specific' but not all alignment methods.
  • standard math Statistical comparisons with three seeds and Welch/ANOVA tests adequately support branch separation claims at the reported effect sizes.
    Appendix Table 9 and seed SDs throughout results tables.
invented entities (2)
  • refusal erosion Eb(c) and protection Pb(c) independent evidence
    purpose: Stage-aware metrics that separate starting refusal from loss under shared post-training, avoiding endpoint confounds.
    Defined in §3 Eqs. (4)–(6). Operational definitions, not ontological new objects; listed because the headline quantitative claim is stated in these terms.
  • final-window pretraining imprint (path-dependent plasticity for alignment)
    purpose: Name the hypothesized latent difference invisible post-SFT but causal for post-training response.
    Interpretive construct in intro/conclusion; evidenced only via behavioral divergence under controlled forks, not via an independent neural or theoretical identifier (mechanism probes in Appendix F are largely null).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT." pith.science (2026). https://pith.science/paper/IE33TVTG

@misc{pith2026260725063,
  author       = {Pith},
  title        = {Pith review of: Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IE33TVTG}},
  note         = {Machine review of arXiv:2607.25063}
}
read the original abstract

Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same across relevant benchmarks are treated as interchangeable, equally ready for the next alignment stage, typically preference optimization. We ask whether this judgment misses a pretraining imprint: a difference that no post-SFT benchmark reveals, yet that decides how each checkpoint responds to further training. To find out, we run a controlled experiment on the final window of pretraining, the last data trained on before instruction tuning. Six branches fork from one partially pretrained checkpoint and differ only in this window: 500 million tokens, 0.1% to 1% of the tokens that precede it. Each branch trains its window on a single data source: generic web text, filtered web text, normative discourse, safety text, mathematical text, or synthetic educational text. SFT and post-training are then identical. After SFT the branches behave near-identically, within about one point on instruction following, refusal, and capability, yet the same post-training carries them to very different endpoints, under both a direct preference optimization update and a reinforcement learning update with a verifiable reward. We measure this deviation through refusal of harmful requests: when post-training begins the safety text branch refuses no more than the web text branch, yet by the end it has lost far less of its refusal. The other four branches gain little or no protection, so the effect is selective to what the window contained. The protection requires the safety text to arrive last rather than earlier in pretraining, and it reproduces on a second model family. What a model is pretrained on last shapes how it reacts to alignment. Therefore, a checkpoint should not be evaluated by its post-SFT behavior alone, and what it was trained on last should be reported with it.

Figures

Figures reproduced from arXiv: 2607.25063 by the authors.

Figure 1
Figure 1. Similar models learn differently. Two branches differ only in their final pretraining window, [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Protection against DPO erosion relative to [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Order control erosion over three seeds. Both [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Robustness checks for the Cweb–Csafety erosion contrast (bars show DPO erosion, lower is better). Left: lower learning rate DPO, 2×10−7 versus the main 5×10−7 , on the OLMo 49B fork branches. Right: the same Cweb/Csafety final window contrast on Pythia-1B-deduped (Bide…
Figure 6
Figure 6. Figure 6: OR-Bench refusal tradeoff (Cui et al. 2025). Points [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 8
Figure 8. Figure 8: Repetition control for the safety corpus. Lower [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: All branches, matched after SFT and divergent under DPO. [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]
Figure 10
Figure 10. Figure 10: Pythia replication of the matched then divergent pattern for the [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: The same math reinforcement learning (GRPO [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 12
Figure 12. Figure 12: Corpus diagnostics. Higher values mean more [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]
Figure 13
Figure 13. Figure 13: Exploratory value probes, both null under a change [PITH_FULL_IMAGE:figures/full_fig_p016_13.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

55 extracted references · 12 linked inside Pith

  1. [1]

    arXiv preprint arXiv:2001.08361 , year=

    Scaling laws for neural language models , author=. arXiv preprint arXiv:2001.08361 , year=

  2. [2]

    Advances in neural information processing systems , volume=

    An empirical analysis of compute-optimal large language model training , author=. Advances in neural information processing systems , volume=

  3. [3]

    arXiv preprint arXiv:2101.00027 , year=

    The pile: An 800gb dataset of diverse text for language modeling , author=. arXiv preprint arXiv:2101.00027 , year=

  4. [4]

    International conference on machine learning , pages=

    Pythia: A suite for analyzing large language models across training and scaling , author=. International conference on machine learning , pages=. 2023 , organization=

  5. [5]

    Groeneveld, Dirk and Beltagy, Iz and Walsh, Evan and Bhagia, Akshita and Kinney, Rodney and Tafjord, Oyvind and Jha, Ananya and Ivison, Hamish and Magnusson, Ian and Wang, Yizhong and others , booktitle=

  6. [6]

    International Conference on Learning Representations , volume=

    Language models scale reliably with over-training and on downstream tasks , author=. International Conference on Learning Representations , volume=

  7. [7]

    Li, Jeffrey and Fang, Alex and Smyrnis, Georgios and Ivgi, Maor and Jordan, Matt and Gadre, Samir and Bansal, Hritik and Guha, Etash and Keh, Sedrick and Arora, Kushal and others , journal=

  8. [8]

    arXiv preprint arXiv:2603.16177 , year=

    The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data , author=. arXiv preprint arXiv:2603.16177 , year=

Show all 55 references
  1. [9]

    arXiv preprint arXiv:2605.12705 , year=

    Early Data Exposure Improves Robustness to Subsequent Fine-Tuning , author=. arXiv preprint arXiv:2605.12705 , year=

  2. [10]

    Advances in neural information processing systems , volume=

    Training language models to follow instructions with human feedback , author=. Advances in neural information processing systems , volume=

  3. [11]

    Constitutional

    Bai, Yuntao and Kadavath, Saurav and Kundu, Sandipan and Askell, Amanda and Kernion, Jackson and Jones, Andy and Chen, Anna and Goldie, Anna and Mirhoseini, Azalia and McKinnon, Cameron and others , journal=. Constitutional

  4. [12]

    Advances in neural information processing systems , volume=

    Direct preference optimization: Your language model is secretly a reward model , author=. Advances in neural information processing systems , volume=

  5. [13]

    Advances in Neural Information Processing Systems , volume=

    How far can camels go? exploring the state of instruction tuning on open resources , author=. Advances in Neural Information Processing Systems , volume=

  6. [14]

    Proceedings of the 41st International Conference on Machine Learning , articleno =

    Cui, Ganqu and Yuan, Lifan and Ding, Ning and Yao, Guanming and He, Bingxiang and Zhu, Wei and Ni, Yuan and Xie, Guotong and Xie, Ruobing and Lin, Yankai and Liu, Zhiyuan and Sun, Maosong , title =. Proceedings of the 41st International Conference on Machine Learning , article...

  7. [15]

    arXiv preprint arXiv:2502.14768 , year=

    Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning , author=. arXiv preprint arXiv:2502.14768 , year=

  8. [16]

    2023 , eprint=

    Judging LLM-as-a-judge with MT-Bench and Chatbot Arena , author=. 2023 , eprint=

  9. [17]

    arXiv preprint arXiv:2307.15043 , year=

    Universal and transferable adversarial attacks on aligned language models , author=. arXiv preprint arXiv:2307.15043 , year=

  10. [18]

    XST est: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

    R. XST est: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. doi:10...

  11. [19]

    Ji, Jiaming and Liu, Mickel and Dai, Josef and Pan, Xuehai and Zhang, Chi and Bian, Ce and Chen, Boyuan and Sun, Ruiyang and Wang, Yizhou and Yang, Yaodong , journal=

  12. [20]

    2025 , url=

    Justin Cui and Wei-Lin Chiang and Ion Stoica and Cho-Jui Hsieh , booktitle=. 2025 , url=

  13. [21]

    Advances in Neural Information Processing Systems , volume=

    The fineweb datasets: Decanting the web for the finest text data at scale , author=. Advances in Neural Information Processing Systems , volume=

  14. [22]

    International Conference on Learning Representations , volume=

    Openwebmath: An open dataset of high-quality mathematical web text , author=. International Conference on Learning Representations , volume=

  15. [23]

    Ben Allal, Loubna and Lozhkov, Anton and Penedo, Guilherme and Wolf, Thomas and von Werra, Leandro , title =

  16. [24]

    arXiv preprint arXiv:2204.05862 , year=

    Training a helpful and harmless assistant with reinforcement learning from human feedback , author=. arXiv preprint arXiv:2204.05862 , year=

  17. [25]

    International Conference on Learning Representations , year=

    Measuring Massive Multitask Language Understanding , author=. International Conference on Learning Representations , year=

  18. [26]

    Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin , booktitle=

  19. [27]

    arXiv preprint arXiv:1803.05457 , year=

    Think you have solved question answering? try arc, the ai2 reasoning challenge , author=. arXiv preprint arXiv:1803.05457 , year=

  20. [28]

    2021 , publisher=

    Sakaguchi, Keisuke and Bras, Ronan Le and Bhagavatula, Chandra and Choi, Yejin , journal=. 2021 , publisher=

  21. [29]

    arXiv preprint arXiv:2407.21783 , year=

    The llama 3 herd of models , author=. arXiv preprint arXiv:2407.21783 , year=

  22. [30]

    arXiv preprint arXiv:2501.00656 , year=

    2 OLMo 2 Furious , author=. arXiv preprint arXiv:2501.00656 , year=

  23. [31]

    Advances in neural information processing systems , volume=

    Openassistant conversations-democratizing large language model alignment , author=. Advances in neural information processing systems , volume=

  24. [32]

    Advances in Neural Information Processing Systems , volume=

    The PRISM alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models , author=. Advances in Neural Information Processing Systems , volume=

  25. [33]

    Wang, Zhilin and Dong, Yi and Delalleau, Olivier and Zeng, Jiaqi and Shen, Gerald and Egert, Daniel and Zhang, Jimmy J and Sreedhar, Makesh N and Kuchaiev, Oleksii , journal=

  26. [34]

    Kim, Seungone and Shin, Jamin and Cho, Yejin and Jang, Joel and Longpre, Shayne and Lee, Hwaran and Yun, Sangdoo and Shin, Seongjin and Kim, Sungdong and Thorne, James and Seo, Minjoon , booktitle=

  27. [35]

    arXiv preprint arXiv:2206.02841 , year=

    Researching Alignment Research: Unsupervised Analysis , author=. arXiv preprint arXiv:2206.02841 , year=

  28. [36]

    Ji, Jiaming and Hong, Donghai and Zhang, Borong and Chen, Boyuan and Dai, Josef and Zheng, Boren and Qiu, Tianyi Alex and Zhou, Jiayi and Wang, Kaile and Li, Boxun and others , booktitle=

  29. [37]

    Findings of the association for computational linguistics: EMNLP 2020 , pages=

    Realtoxicityprompts: Evaluating neural toxic degeneration in language models , author=. Findings of the association for computational linguistics: EMNLP 2020 , pages=

  30. [38]

    Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

    Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

  31. [39]

    Aligning

    Hendrycks, Dan and Burns, Collin and Basart, Steven and Critch, Andrew and Li, Jerry and Song, Dawn and Steinhardt, Jacob , journal=. Aligning

  32. [40]

    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=

    The multilingual alignment prism: Aligning global and local preferences to reduce harm , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=

  33. [41]

    arXiv preprint arXiv:2605.02087 , year=

    Model spec midtraining: Improving how alignment training generalizes , author=. arXiv preprint arXiv:2605.02087 , year=

  34. [42]

    arXiv preprint arXiv:2311.07911 , year=

    Instruction-Following Evaluation for Large Language Models , author=. arXiv preprint arXiv:2311.07911 , year=

  35. [43]

    2025 , howpublished =

  36. [44]

    Han, Seungju and Rao, Kavel and Ettinger, Allyson and Jiang, Liwei and Lin, Bill Yuchen and Lambert, Nathan and Choi, Yejin and Dziri, Nouha , journal=

  37. [45]

    International Conference on Learning Representations , volume=

    Fine-tuning aligned language models compromises safety, even when users do not intend to! , author=. International Conference on Learning Representations , volume=

  38. [46]

    Advances in Neural Information Processing Systems , volume=

    Refusal in language models is mediated by a single direction , author=. Advances in Neural Information Processing Systems , volume=

  39. [47]

    Findings of the Association for Computational Linguistics: ACL 2023 , pages=

    Discovering Language Model Behaviors with Model-Written Evaluations , author=. Findings of the Association for Computational Linguistics: ACL 2023 , pages=. 2023 , publisher=

  40. [48]

    International Conference on Learning Representations , volume=

    Towards understanding sycophancy in language models , author=. International Conference on Learning Representations , volume=

  41. [49]

    Zenodo , year=

    A framework for few-shot language model evaluation , author=. Zenodo , year=

  42. [50]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  43. [51]

    International conference on machine learning , pages=

    Linear mode connectivity and the lottery ticket hypothesis , author=. International conference on machine learning , pages=. 2020 , organization=

  44. [52]

    Advances in neural information processing systems , volume=

    What is being transferred in transfer learning? , author=. Advances in neural information processing systems , volume=

  45. [53]

    International Conference on Machine Learning , pages=

    Understanding plasticity in neural networks , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  46. [54]

    Nature , volume=

    Loss of plasticity in deep continual learning , author=. Nature , volume=. 2024 , publisher=

  47. [55]

    The Fourteenth International Conference on Learning Representations , year=

    Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data , author=. The Fourteenth International Conference on Learning Representations , year=

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.