{"id":"888f94b9-fa97-44ba-8ccb-d35603429511","arxiv_id":"2607.25063","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"A 500M-token final pretraining window of safety text leaves matched post-SFT models that lose far less refusal under identical DPO or GRPO than web-text counterparts.","lead":"Two language-model checkpoints that look the same after instruction tuning can still react very differently to the next alignment step, depending only on the last slice of pretraining data. The finding challenges the common practice of treating post-SFT scores as enough to decide whether a checkpoint is ready for preference optimization.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The internal 1B result is well-supported, but the paper's practical upshot (\"report what a model was trained on last\") rests on an untested cell: the 4T saturated fork was only ever probed at 0.0125% relative dose, so we cannot tell whether relative dose or early-training plasticity governs the effe","rationale":"The reader's weakest_assumption identifies the same load-bearing concern I arrive at independently: the strongest claim's internal component (matched SFT, divergent post-training, selective and order-dependent) is thoroughly controlled, but the external/practical component is undermined by the paper's own attenuation data. I am sharpening rather than replacing the reader's point: the attenuation section (§4.6, Fig. 5) tests relative dose cleanly only at the 49B fork, and tests the 4T fork only at a trivially small dose, leaving the governing variable — relative dose vs. early-stage plasticity — unresolved. This matters because the two hypotheses have opposite implications for the paper's headline recommendation: if relative dose governs, late-stage annealing windows at ~1% scale (which real pipelines do use) would still imprint, and reporting last-window data is well-motivated; if early plasticity governs, the recommendation addresses a regime that never occurs in production. This does not warrant a verdict change: the reader's CONDITIONAL already prices in exactly this external-validity gap, the paper discloses the attenuation and the over-refusal/capability costs with unusual candor, and the detector validation, decontamination audit, and Pythia replication are genuine supporting evidence. The recommended test is cheap relative to the study's total compute and would convert the paper's honest limitation into a settled boundary condition.","tokens_in":23876,"tokens_out":3441,"duration_ms":151065,"concrete_test":"Run the missing cell: fork the saturated OLMo-2-1B 4T checkpoint and train a safety-last window at ~1% relative dose (~40B tokens, subsampled/cycled from the same Csafety corpus), plus a matched Cweb window, then apply the identical SFT/DPO pipeline and measure decontaminated erosion on BeaverTails and XSTest-unsafe (AdvBench is ceiling-saturated at 4T). If protection recovers toward the ~9 pp seen at the 49B fork, relative dose governs and the reporting recommendation stands for any stage with a ~1% late window; if protection stays near zero, the effect is an early-training-plasticity phenomenon and §5's practice claim should be scoped accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two tiers: (a) the experimental result — matched-after-SFT branches diverge under identical DPO/GRPO, selective to safety-last content — and (b) the practice recommendation in §5 that last-window data should be reported because it shapes alignment. Tier (a) is convincingly supported: three seeds, order controls (§4.3), negative corpora (§4.2), detector cross-validation against WildGuard (App. A), a second model family (App. G), and honest cost reporting (OR-Bench over-refusal, Fig. 6). I find no internal inconsistency there.\n\nThe soft spot is the bridge to tier (b), and it is narrower than generic \"1B doesn't scale\" skepticism. Figure 5 shows protection falling 9.1 → 3.0 → ~0.5 pp as the fork moves 49B → 504B → 4T with a fixed 500M window. The authors interpret this as a relative-dose boundary, and the fixed-49B shrinking-window control (right panel) is consistent with that reading. But the design never runs the discriminating experiment: a window scaled to ~1% of prior training at the saturated fork (~40B tokens at 4T). If relative dose is the governing variable, that cell should recover large protection; if the phenomenon instead requires early-training plasticity (the 49B fork is only ~2% of the way through OLMo-2's schedule), it should not. Under the latter outcome, the effect is real but confined to a regime no shipped checkpoint occupies, and the reporting recommendation loses its motivation. A secondary wrinkle: at the 4T fork AdvBench sits at 99.7% post-DPO (Table 10), so erosion there is ceiling-limited and the \"~0.5 pp\" attenuation figure is partly a measurement artifact — though BeaverTails/XSTest still show the gap vanishing, so this does not rescue the effect.","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper asks whether two checkpoints that are behaviorally matched after identical SFT can nonetheless respond differently to the same post-training update, as a function of the final window of pretraining. Six branches fork from a shared OLMo-2-1B checkpoint at ~49B tokens and differ only in a 500M-token continued-pretraining window (web, DCLM, normative discourse, safety transformation text, math, synthetic education). After identical Tulu-style SFT, the branches are matched within ~1 point on refusal, capability, and IFEval; under identical UltraFeedback DPO they diverge, with the safety-text branch losing substantially less refusal (overall protection ~8.2 pp vs. web, Table 2). The effect is selective to safety content (§4.2), requires that content to come last (order control, §4.3), reproduces under GRPO with verifiable rewards on two tasks (§4.4, App. H), survives a lower DPO learning rate, a preference-data swap to Chatbot Arena, a no-CPT baseline, and a Pythia-1B replication (§4.5, App. G). The authors define refusal erosion E_b(c) and protection P_b(c) stage-aware (Eqs. 4–6), decontaminate evaluation prompts against the safety corpus (App. B), cross-validate the lexical refusal detector against WildGuard (App. A), and report costs honestly: elevated OR-Bench over-refusal and a ~0.9 pp capability cost (Fig. 6, Table 7). Figure 5 shows the protection attenuating from 9.1 → 3.0 → ~0.5 pp as the fixed 500M window shrinks from 1.02% to 0.0125% of prior training. The pap","tokens_in":24282,"tokens_out":2372,"duration_ms":96961,"significance":"If the result holds, it identifies a genuine blind spot in standard checkpoint evaluation: post-SFT behavioral matching does not certify equal readiness for post-training, and the final pretraining window is a load-bearing variable. This is a clean, well-controlled demonstration of path dependence in post-training, with several features that raise confidence beyond the norm for empirical work at this scale: three-seed means with seed SDs and paired protection SDs, Welch tests and ANOVA (Table 9), an explicit order control that rules out a pure exposure account, four negative-content corpora, decontamination with an explicit overlap audit (BeaverTails 196/1000 overlap disclosed and removed), detector cross-validation against a non-lexical classifier on identical completions, a second model family, and a second post-training algorithm class (GRPO with verifiable reward, two tasks). The authors also report the effect's costs and boundary conditions rather than hiding them, which makes the paper useful even to readers who doubt the practical recommendation. The contribution is a reproducible experimental design and a falsifiable phenomenon, not a fitted narrative.","major_comments":[{"comment":"The attenuation result is interpreted as a relative-dose boundary, but the design does not include the discriminating cell. The 4T saturated fork is probed only at 0.0125% relative dose (fixed 500M window). If a window scaled to ~1% of prior training at the 4T fork (~40B tokens) recovered large protection, the relative-dose reading is confirmed and the §5 reporting recommendation stands for production-scale checkpoints; if it did not, the phenomenon requires early-training plasticity (the 49B fork is ~2% of the way through OLMo-2's schedule) and is confined to a regime no shipped checkpoint occupies. The fixed-49B shrinking-window control (Fig. 5, right) varies dose at fixed plasticity and therefore cannot separate these hypotheses. As written, the paper's practical upshot ('what a model was trained on last should be reported') rests on an untested cell. Either run the scaled-window expe","section":"§4.6 / Fig. 5 / App. D (Table 10)"},{"comment":"The 4T comparison is partly uninformative for a second reason the paper only partially addresses: at the saturated fork the branches sit at ceiling on AdvBench (99.7/99.9% post-DPO), so erosion E_b is compressed toward zero by construction on that benchmark, and the XSTest gap actually reverses sign (-4.0 pp). The authors note the ceiling issue for AdvBench, but the '~0.5 pp overall protection' number in §4.6 averages across benchmarks with very different headroom, mixing a true null with a measurement artifact. A cleaner read of the 4T cell would report per-benchmark protection with SFT starting points and headroom stated, or restrict the dose-response curve (Fig. 5, left) to benchmarks off ceiling. This matters because Fig. 5 is the evidence base for the dose-boundary claim that feeds the recommendation.","section":"App. D, Table 10"},{"comment":"The headline quantity is protection of refusal, but Fig. 6 and the WildGuard harm-axis result (§4.6: 'at most a weak difference... on the harm axis') show that what is retained is a broad refusal prior, including over-refusal of benign prompts, rather than calibrated harm refusal. The framing throughout (e.g., 'the protection requires that safety content arrive last') invites a safety-benefit reading that the paper's own measurements do not support; the retained behavior is refusal plasticity, not safety. This does not undermine the path-dependence claim, which is the real contribution, but the abstract and §5 should state plainly that the protected quantity is refusal rate (with an over-refusal cost), so that practitioners do not read the result as a free alignment intervention.","section":"§3 / §4.6 / Fig. 6"}],"minor_comments":[{"comment":"The C_synth partial effect (strong on XSTest, weak on AdvBench) is noted but not discussed. Since C_synth is the one non-safety branch with a clear positive signal, a sentence on why synthetic educational text might partially protect refusal (or a corpus-diagnostic comparison via App. I) would sharpen the selectivity claim.","section":"§4.2 / Fig. 2"},{"comment":"The mechanism diagnostics rest on a single seed and the refusal-direction projection probe 'did not cleanly separate' the branches. This is fine as a negative result, but the section would benefit from stating explicitly that no mechanism is claimed, and that the update-norm equivalence only rules out the trivial frozen-model explanation.","section":"App. F, Table 12"},{"comment":"The lexical detector uses ten patterns matched in the first 300 characters; please state whether the 64-new-token generation budget (Table 4) ever truncates before the prefix region for non-refusing completions, and whether AdvBench/XSTest/BeaverTails use identical generation settings in all figures (App. A uses 256 tokens, which is noted, but the main-text reader has to cross-reference).","section":"§3, Refusal measurement"},{"comment":"The safety corpus is cycled 2.3 passes to reach 500M tokens while C_web/C_dclm are single-pass samples; the repetition control (App. B, Table 8) addresses this, but a forward pointer from Table 1 to that control would help readers who notice the asymmetry early.","section":"Table 1"},{"comment":"Several 2026-dated citations (Baek et al., Feng et al., Akter et al., Li et al. 2026) are arXiv preprints central to the motivation; please verify identifiers (e.g., arXiv:2603.16177, 2605.12705) resolve, since these anchor the claim that percent-scale windows are known to matter.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The experimental core is unusually careful for this venue and I want to be clear that my major-revision vote is not about the internal 1B result, which I find convincing. It is about the bridge to the practice recommendation: the dose-vs-plasticity confound in Fig. 5 is the one experiment whose outcome could flip the paper's practical message, and the authors are best placed to run it (they already have the 4T fork infrastructure). If they instead choose to scope the claim and soften §5, I would be comfortable with a lighter revision. The work fits the journal's scope on training-dynamics and alignment methodology."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The load-bearing finding is real and carefully isolated. Six branches fork from one OLMo-2-1B checkpoint, differ only in a 500M-token final window, then get identical Tulu SFT and UltraFeedback DPO (plus GRPO on two arithmetic tasks). After SFT they sit within about a point on refusal, IFEval, and capability; under the same post-training the safety-last branch loses far less refusal (~8 pp overall protection vs web) even though it does not start higher. Order control, non-safety negatives (web/DCLM/math/normative/synth), Base-without-CPT, one-pass repetition, WildGuard cross-check, Chatbot Arena swap, lower LR, and a Pythia replication all point the same way. Erosion/protection is the right metric: it separates starting level from how much the shared update removes.\n\nWhat is new relative to Baek/Feng/Li midtraining and plasticity work is exactly that matched-after-SFT / divergent-under-identical-PT design, plus the necessity that safety content arrive last. Citations look honest; no circular fitting; three seeds with SDs and Welch/ANOVA in the appendix.\n\nSoft spots in proportion. Scale is 1B and the fork is early (~49B). Figure 5 shows protection collapsing as the same absolute window becomes a tinier fraction of prior tokens, and the 4T cell was only ever run at ~0.0125% relative dose—never the discriminating ~1%-of-4T window—so we cannot yet separate relative dose from early-training plasticity. That weakens the §5 practice claim (“report what it was trained on last”) more than the experimental claim. Mechanism is open; refusal is a convenient probe, not a full map of alignment plasticity; they already report the over-refusal and small capability costs. None of that undoes the internal result.\n\nFor anyone who builds or evaluates open post-training pipelines, this is worth an hour. I would send it to referees: the experiment is sharp enough to deserve pushback on generalization, not a desk reject. Engage; cite the controlled finding if you work on data order or refusal stability; treat the reporting recommendation as a hypothesis still needing the scaled-window saturated-fork cell.","headline":"Clean controlled result: SFT-matched 1B forks diverge under identical DPO/GRPO when the final pretraining window is safety-last; the reporting pitch outruns the dose evidence at saturated scale.","tokens_in":25231,"tokens_out":580,"would_cite":true,"duration_ms":23608,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"What a model is pretrained on last shapes how much of its SFT refusal survives the next alignment update, even when post-SFT scores match.","keywords":["final-window pretraining","post-training path dependence","refusal erosion","supervised fine-tuning","direct preference optimization","GRPO","alignment plasticity","checkpoint evaluation"],"falsifier":"Repeat the matched-fork design at larger scale or later in training with the same relative final-window dose: if safety-last and web-last branches that still match after SFT no longer separate on refusal erosion under identical DPO or GRPO, the claimed imprint does not hold where it would matter most.","tokens_in":24861,"feed_emoji":"🧩","tokens_out":979,"duration_ms":15182,"temperature":0.7,"pith_summary":"Developers treat two checkpoints as interchangeable once they score about the same after supervised fine-tuning on instruction following, refusal, and capability. This paper argues that judgment misses a pretraining imprint set by the final data window before instruction tuning. In a controlled fork from one partially pretrained 1B checkpoint, six branches receive the same 500M-token final window budget on different single sources, then identical SFT and post-training. After SFT they sit within about one point on the usual release criteria, yet the same preference update and the same reinforcement-learning update drive them to different endpoints. The safety-transformation branch does not start with higher refusal; it simply loses far less of it. The effect is selective to that content, requires the safety text to arrive last rather than earlier, weakens as the window shrinks to a tiny fraction of prior training, and appears on a second model family. If true, post-SFT behavior alone is not enough to certify readiness for alignment, and the last pretraining data should be reported with the checkpoint.","feed_headline":"Last pretraining data decides how SFT refusal survives alignment","feed_subtitle":"Matched after SFT, models still diverge under the same DPO and RL updates","key_machinery":"Refusal erosion: the drop in harmful-request refusal rate from the shared SFT checkpoint to the post-training endpoint under a fixed update. Protection is web-branch erosion minus branch erosion, isolating how much of the SFT-installed refusal the shared stage removes rather than how high refusal started.","core_discovery":"Checkpoints matched after identical SFT on instruction following, refusal, and capability still diverge under the same post-training update. A final pretraining window of safety transformation text yields substantially lower refusal erosion under UltraFeedback DPO and under GRPO with a verifiable reward than a generic web window, even though the safety branch does not start higher after SFT. The protection is content-selective, requires safety data to come last, and is not a generic benefit of extra late tokens.","pith_inferences":["Upstream data teams and downstream aligners may need shared contracts about the final pretraining mixture, not only about SFT and preference datasets.","If last-window imprints generalize beyond refusal, other SFT-installed behaviors (style, tool use, calibration) could also erode differently across behavior-matched checkpoints.","Reporting last-window provenance would let evaluators test whether apparent alignment gains or losses are really post-training effects or inherited plasticity differences.","Curriculum design that deliberately ends pretraining on task-proximal critique text, rather than only mixing it earlier, becomes a testable lever for alignment stability."],"forward_implications":["Two checkpoints with matching post-SFT scores are not interchangeable for the next alignment stage.","Release cards and model reports should include what data the model saw last in pretraining, not only behavior scores.","Safety-oriented text placed in the final pretraining window can change how much SFT refusal survives later preference or RL updates that never reward refusal.","Order matters: the same safety content earlier in late pretraining does not give the same protection as placing it last.","The effect is dose-relative: as the final window becomes a vanishing fraction of prior tokens, the divergence fades."],"fun_headline_variants":["Final pretrain window steers post-SFT refusal under same DPO and RL","Safety-last pretraining cuts refusal loss after matched SFT checkpoints","Same SFT scores, different alignment paths from 0.1–1% final tokens","What you pretrain last shapes how refusal holds through DPO and GRPO","Post-SFT twins diverge: last-window data decides alignment endpoints"],"cache_read_input_tokens":16384,"weakest_assumption_plain":"That how much refusal is lost on a few harmful-request suites under one fixed helpfulness DPO or arithmetic RL recipe, on small 1B forks, is a fair stand-in for general alignment plasticity and for how production checkpoints should be judged.","fun_headline_variants_meta":{"raw":{"variants":["Final pretrain window steers post-SFT refusal under same DPO and RL","Safety-last pretraining cuts refusal loss after matched SFT checkpoints","Same SFT scores, different alignment paths from 0.1–1% final tokens","What you pretrain last shapes how refusal holds through DPO and GRPO","Post-SFT twins diverge: last-window data decides alignment endpoints"]},"model":"grok-4.5","effort":"low","cost_usd":0.002005,"raw_usage":{"total_tokens":1018,"prompt_tokens":917,"num_sources_used":0,"completion_tokens":86,"cost_in_usd_ticks":20048000,"prompt_tokens_details":{"text_tokens":917,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":15,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":917,"tokens_out":86,"duration_ms":2830,"temperature":1.0,"reasoning_tokens":15,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T02:12:23.729013+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Repeat the matched-fork design at larger scale or later in training with the same relative final-window dose: if safety-last and web-last branches that still match after SFT no longer separate on refusal erosion under identical DPO or GRPO, the claimed imprint does not hold where it would matter most.","supporting_citations":[],"review_version":1}