{"id":"d556d16d-ed45-441d-96ca-bbfd00729578","arxiv_id":"2608.08542","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Static refusal tests overstate the safety of skill-merged LLMs: models with identical static safety differ sharply under adaptive attack, and a task-vector overlap with a safety subspace flags only same-recipe abliteration-style erosion.","lead":"This paper benchmarks merged LLMs and finds that a model's refusal rate on fixed harmful prompts does not predict how easily it is jailbroken by adaptive attacks. The authors propose a geometric pre-merge screen and a projection-based repair that only covers one family of safety-eroding merges.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the static-adaptive decoupling is well supported and the reader's subspace concern is bounded to the repair tool, not the central claim.","rationale":"After reading the paper in good faith, I find the central empirical claim robust. The decoupling is established on the main three bases with matched static and adaptive evaluations on the same 50-behavior subset, paired McNemar tests and cluster bootstrap CIs, three attack families, and a judge-invariance check. The six-base extension is presented with an explicit scope: the separator is the template attack, and GCG under-transfer is disclosed in Table S2, so the paper does not overclaim an attack-invariant fragile/robust taxonomy. The reader's weakest assumption about S is correct and well-documented in Table S6, but it only affects Contributions C and the geometric signal; the benchmark decoupling is independent of S. The most plausible challenge, that the fragile/robust ordering might be attack-relative, is already acknowledged internally and does not undermine the negative claim that low static ASR is not evidence of adaptive safety. A held-out attack-family replication would further confirm the ordering, but its absence is not a flaw given the paper's scoping. I recommend no change to the reader's ACCEPT.","tokens_in":21695,"tokens_out":14418,"duration_ms":162298,"concrete_test":"Hold out one attack family not used to establish the ordering, e.g., run a stronger PAIR with a GPT-4-class attacker or HarmBench's full attack suite on the Qwen-7B and Llama-8B math merges (lambda=0.6, same seeded 50 behaviors, same two-judge AND rule) and check that Qwen's static-adaptive gap remains positive with paired significance while Llama's stays near zero; this would settle whether the decoupling generalizes beyond the specific GCG/template/PAIR configurations reported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that static refusal scores do not predict adaptive jailbreak robustness. This is supported by paired per-behavior statistics on the same 50-behavior subset (Table 1), by replication across three attack families on the main bases, and by the judge-invariance check; the six-base extension is explicitly scoped to the template probe and the paper discloses that GCG under-transfers non-uniformly (Table S2), so the fragile/robust labeling is attack-relative rather than a model-level fact. The reader's identified weakest assumption, the safety-subspace estimator S (Section 6, Table S6), is genuinely recipe-dependent, but it bounds only the geometric screen and SubSafe-Merge (Contribution C); Contribution B's decoupling does not use S and is not affected. The disclosed correction and the principled deviation from the pre-registered decision rule are mechanical data-dependent revisions, not hidden flaws. I therefore find no load-bearing flaw in the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that static refusal tests misrepresent the safety of skill-merged LLMs and introduces SkillSafe-Bench, a factorial protocol (base model x skill x merging method x merge coefficient) that jointly reports static refusal (400 HarmBench behaviors, two-judge AND rule), adaptive jailbreak success (GCG, best-of-6 semantic templates, PAIR, and a multi-turn Crescendo arm), and capability retention. The central empirical claim (Contribution B) is that static safety does not predict adaptive robustness: across six open-weight bases spanning five families, statically safe merges on Qwen2.5-7B/14B and Gemma-2-9B are attacked at 60-76% under the template probe while Llama-3.1-8B and Phi-4-mini remain robust, and on the three main bases the Qwen-fragile/Llama-robust ordering replicates under GCG at 100 and 500 steps, templates, and PAIR on a matched 50-behavior subset with paired McNemar tests and cluster-bootstrap CIs. Contribution C is a data-free geometric signal (the overlap of a task vector with a per-layer safety subspace S estimated from a public abliterated checkpoint), which separates a same-recipe uncensored donor (overlap 0.99) from genuine skills (overlap 0.001), plus SubSafe-Merge, a projection-based repair. The paper pre-registers its decision rule, discloses a principled deviation and a data correction, and explicitly bounds the attack-relativity and recipe-dependence of its findings.","tokens_in":21869,"tokens_out":32564,"duration_ms":306617,"significance":"If the central claim holds, it reframes merging-safety evaluation: a low static refusal score carries little information about vulnerability to attack, and the statically cleanest bases (e.g., Gemma-2-9B, static ASR 0.02, template ASR 0.60-0.64) can be among the most exposed. The methodological assets are real: paired per-behavior statistics on identical behavior subsets, three attack families plus a multi-turn arm on the main bases, a human-labeled judge audit (AND rule kappa = 0.66) with the fragile/robust ordering shown to be judge-invariant, a 500-step GCG control against under-optimization, an out-of-sample code-LoRA prediction validating the low-overlap/no-erosion direction of the geometric signal, a disclosed pre-registration deviation and corrected earlier subset, and a code/data supplement in which every numeric table is generated from backing result files. The provisional part is the geometric screen and SubSafe-Merge: the authors' own Table S6 shows the safety-subspace estimator is recipe-dependent (overlap ranges 0.99 down to 0.001 across six public checkpoints, five of them dangerous), and the repair's high-overlap direction is demonstrated in-sample.","major_comments":[{"comment":"The six-base fragile/robust taxonomy is a single-probe result whose qualifier appears only later in the text: under the template probe, the Gemma math-merge template ASR is 0.60 while its GCG ASR is 0.04, and Phi-4 shows the reverse pattern (template 0.14, GCG 0.32), so the labels 'fragile' and 'robust' would invert under a different choice of attack. The body is explicit that GCG ordering of the extension bases is not claimed and that multi-attack invariance holds only for the three main bases, but the abstract's sentence 'Across six open-weight bases ... static safety does not predict robustness to attack' does not carry this scope. I recommend adding one sentence to the abstract or conclusion stating that the six-base fragile/robust labeling is relative to the template probe, and that the multi-attack invariance is established for the three main bases only.","section":"Abstract and Section 5.2 (Table S2)"},{"comment":"The manuscript discloses the same-recipe scope of the geometric screen and SubSafe-Merge, but two claims read more strongly than the evidence supports. First, Section 6 opens by calling the subspace overlap 'the feature that cleanly separates skills,' whereas Table S6 shows five of six public refusal-removal checkpoints are dangerous (unmerged static ASR 0.42-0.68) yet have overlap at or near the genuine-skill baseline (0.001-0.003), so the feature separates one lineage rather than refusal-removal in general. Second, the out-of-sample validation (the independent code LoRA) confirms only the low-overlap/no-erosion direction; the high-overlap/repair direction is demonstrated in-sample with the same-recipe donor, which is partly by construction since S is estimated from that donor's abliteration lineage. I recommend that the contribution summary and Section 7 state explicitly that the screen is a same-recipe detector, that it misses SFT/DPO-decensored donors (a documented false negative, Orion, overlap 0.001, static 0.63), and that the repair direction is not out-of-sample validated.","section":"Sections 6-7, Tables 2 and S6"}],"minor_comments":[{"comment":"The discordant-pair counts reported in Section S5 for the Qwen +math lambda=0.6 cell (b=8, c=0) do not add up to the marginals in Table 1: 5/50 static-unsafe plus 8 new flips gives 13/50 GCG-unsafe, while Table 1 reports 14/50 (Wilson CI [0.17, 0.42]). Please reconcile the count; the conclusion is unaffected since b=9 with c=0 still gives McNemar p < 0.01.","section":"Section S5 vs Table 1"},{"comment":"The claim that 'the four methods agree to within 0.035 with identical ordering' should state the aggregation (range across methods per skill and coefficient, or across the whole grid) and the exact quantity compared, since Table S1 shows the four methods varying by up to 0.033 at some cells (e.g., Qwen math lambda=0.2, DARE-TIES 0.147 versus TIES 0.180).","section":"Section 5"},{"comment":"Because the six-base replication and the abstract's headline numbers rest on the best-of-6 template ensemble, which the manuscript itself classifies as non-adaptive, the term 'adaptive' should be defined at its first appearance (abstract or introduction) so that readers do not attribute adaptiveness to the template probe.","section":"Section 4 and footnote 1"},{"comment":"The Crescendo arm compares base (0.20), plain merge (0.14), and SubSafe (0.16) at n=50 with a deliberately weakened attacker, and the differences are within run noise; the claim that 'neither merge adds multi-turn risk' should state explicitly that no significance test is applied to these three arms.","section":"Section 7 and Table S8 (Crescendo)"},{"comment":"The geometry measures used to document Limitation 5 vary by only a few percent across bases (participation ratio 1.013-1.020, spectral entropy 1.089-1.143); consider reporting additional precision or uncertainty so the reader can distinguish a genuine null from a resolution-limited comparison.","section":"Table S3"}],"recommendation":"minor_revision","confidential_remarks":"I found the manuscript unusually well-disclosed: the pre-registered decision rule with a flagged deviation, the corrected earlier subset, the per-cell multiplicity caveat, the judge audit, and the recipe-dependence survey (Table S6) are all in the paper rather than hidden. The remaining work is local, mainly qualifying the abstract and the Contribution-C framing. I see no citation-pattern or novelty-disclosure concern; the 2026 references are consistent with the manuscript's arXiv date. Fit with cs.LG is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know this paper delivers on its headline. Static refusal tests do not predict adaptive jailbreak robustness for skill-merged LLMs, and the evidence is stronger than usual for this area. The authors build a controlled factorial benchmark, run it on three bases fully and replicate the core decoupling on three more, and back the claim with paired McNemar tests, cluster bootstrap CIs, three attack families, and two judges. The Qwen-fragile / Llama-robust contrast at matched static safety is convincing, and the template attack being stronger than GCG is a nice twist—it makes the decoupling a lower bound rather than an artifact of one attack.\n\nWhat is genuinely new: the decoupling itself, the base-conditional static effect, and the data-free subspace-overlap screen. The screen has real out-of-sample validation—a code LoRA trained without ever touching the safety subspace lands at 0.001 overlap and behaves as predicted. Good science habits throughout: they disclose a correction of an earlier inflated gap, state their decision rule before the results, and flag their deviation from it. The limitations section is honest, not decorative.\n\nWhere I would push back: the safety subspace S is estimated from one public abliterated model per base, and the paper itself shows it is recipe-dependent. An SFT/DPO-decensored donor is genuinely uncensored but near-orthogonal to S, and SubSafe-Merge cannot repair it. That bounds the repair tool, not the central claim—the decoupling never uses S—but it does mean SubSafe-Merge should not be read as a general safety fix. Also, the six-base replication rests on a single non-adaptive template probe; cross-family GCG magnitudes are uninformative, so the fragile/robust labeling is attack-relative. That is an honest scope, but narrower than the abstract wants you to believe. The n=50 adaptive subsets are small, though the paired tests handle that decently.\n\nWho this is for: anyone working on model merging, alignment evaluation, or red-teaming open-weight models. The central finding should change how merge safety is assessed. It deserves a serious referee; the empirical core is solid and the claims are measured. I would send it out.","headline":"A careful empirical paper that actually delivers on its central claim—static refusal does not predict adaptive robustness for merged models—while the subspace repair tool is honestly bounded to same-recipe abliteration.","tokens_in":22407,"tokens_out":1871,"would_cite":true,"duration_ms":18695,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Static safety does not predict adaptive jailbreak robustness in skill-merged LLMs.","keywords":["model merging","safety alignment","adaptive jailbreak robustness","static refusal evaluation","task vectors","safety subspace","SkillSafe-Bench","SubSafe-Merge"],"falsifier":"Take the paper's six bases and push each adaptive attack well beyond the reported budgets: a GCG run of 1,000+ steps with a wider suffix search, a GPT-4-class attacker in the PAIR loop, and the same best-of-6 template set, then rank the bases by adaptive ASR. If Llama-3.1-8B and Phi-4-mini, the paper's attack-resistant bases, break at rates comparable to Qwen and Gemma under any of these stronger probes, the fragile/robust ordering is an artifact of attack strength; if the ordering survives at every budget, the conclusion that static refusal screening cannot tell which merges are safe to deploy is confirmed.","tokens_in":2834,"feed_emoji":"🛡️","tokens_out":8205,"duration_ms":233726,"temperature":0.7,"pith_summary":"The paper argues that the standard measure of safety for skill-merged language models — static refusal tests, which present fixed harmful prompts and check for compliance — systematically overstates how safe a merge is. Because refusal is shallow, concentrated in the first few generated tokens, a merged model can keep its static refusal clean while an adaptive attacker that searches over inputs still breaks it. Using a new controlled benchmark, SkillSafe-Bench, the authors show across six open-weight bases (five model families) that static safety does not predict robustness to attack: safe-looking merges on Qwen (7B and 14B) and Gemma are jailbroken 60–76% of the time by a fixed template attack, while Llama and Phi-4 merges at indistinguishable static safety stay attack-resistant. The paper also identifies a data-free geometric signal — the overlap of a task vector with the safety subspace $S$ estimated from an abliterated model — that separates eroding refusal-removal skills from genuine skills before any merge, and sketches SubSafe-Merge, which projects that overlap away to restore safety at held capability. A reader should care because merging is the default way open-weight models gain new skills, and the models that most need adaptive checking are exactly the ones that look safe under current static screening.","feed_headline":"60-76% of static-safe merged LLMs fall to a template jailbreak","feed_subtitle":"Qwen and Gemma merges pass static checks yet jailbreak 60-76% under a template attack; Llama and Phi-4 hold.","key_machinery":"The load-bearing object is the safety subspace $S$. For each base, the authors form the alignment vector $\\tau_{\\mathrm{safe}} = \\theta_{\\mathrm{base}} - \\theta_{\\mathrm{abliterated}}$, the weight difference between the aligned base and a public abliterated version of it (refusal direction removed), and define $S^{(\\ell)}$, per layer, as the span of the top-$k$ left singular vectors of $\\tau_{\\mathrm{safe}}^{(\\ell)}$; projection onto $S$ and its complement are $P_S$ and $P_{S^\\perp}$. Two things are done with it. First, a data-free pre-merge screen: the subspace overlap of a candidate task vector $\\tau_i$, the energy fraction of $\\tau_i$ inside $S$, acts as a binary detector of the merge type that erodes safety, with same-recipe abliteration-style refusal-removal vectors at $\\approx 0.99$ and genuine skills at $\\approx 0.001$. Second, a repair: SubSafe-Merge computes $\\theta_{\\mathrm{merge}} = \\theta_{\\mathrm{base}} + \\sum_i \\lambda_i P_{S^\\perp} f(\\tau_i)$, projecting each method-processed task vector off $S$ so the eroding component is removed before the merge while the orthogonal capability component is retained. The construction is validated by a control (removing the top refusal direction from Qwen raises static ASR from 0.20 to 0.64 while GSM8K and MMLU hold), and its scope is bounded by the estimator: $S$ comes from one public abliteration per base, and the paper's own survey of six public uncensored checkpoints shows overlap is recipe-dependent, with genuinely dangerous SFT-decensored donors sitting at the 0.001 skill baseline.","core_discovery":"The paper's central claim is that static refusal screening is not a proxy for the safety of skill-merged models under real attack. On three bases run over a full grid of merging method × skill × coefficient, two strongly-aligned bases that look identically safe under fixed prompts — Qwen2.5-7B (static ASR 0.24 on the evaluation subset) and Llama-3.1-8B (0.20) — separate sharply under GCG: Qwen's benign math merges rise from 0.10–0.24 static to 0.28–0.38 adaptive (pooled gap +17pp, 95% CI [0.09, 0.25], paired McNemar $p<0.01$), while Llama's merges stay at or below 0.12. The fragile/robust ordering replicates under a semantic template attack (Qwen 0.66–0.76 vs Llama 0.22–0.24), under PAIR, under a 500-step GCG budget, and with a single judge dropped, and it extends to six bases: Qwen2.5-14B and Gemma-2-9B are fragile (Gemma, the safest base statically at 0.02, is jailbroken 60% of the time under the template attack), while Phi-4-mini is robust. The static effect of merging is itself base-conditional: on strongly-aligned bases a benign math skill lowers static ASR while an uncensored skill raises it, but on weakly-aligned Mistral any merge raises it. Complementing the benchmark, the paper shows that the energy fraction of a task vector inside the safety subspace $S$ (estimated from an abliterated model, no harmful data) separates same-recipe refusal-removal vectors (overlap $\\approx 0.99$) from genuine skills including an independently-trained code specialist ($\\approx 0.001$), validated out-of-sample; and that projecting that overlap off (SubSafe-Merge) restores both static and adaptive ASR to base on strongly-aligned bases at unchanged GSM8K, while honestly failing on out-of-$S$ erosion and on bases that are fragile before merging.","pith_inferences":["If shallow alignment is the mechanism, the decoupling should generalize beyond merging to other weight-space edits that leave the first-token refusal layer intact (LoRA fine-tuning, pruning, quantization); a testable prediction is that such edits reproduce the Qwen-like static–adaptive gap and spare Llama-like bases.","The overlap screen is best read as a lineage detector, not a general safety test: the paper's six-checkpoint survey places genuinely uncensored SFT/DPO donors at the 0.001 skill baseline and a dangerous independent abliteration at 0.32, so a practitioner adopting the screen should first verify that their donor's refusal-removal recipe matches the recipe used to build $S$.","Because static repair scales with overlap (none at 0.001, partial at 0.32, full at 0.99), a natural extension is to calibrate an acceptance threshold on overlap plus unmerged donor static ASR to turn the binary detector into a graded pre-merge risk score; the paper's leave-one-base-out regression at three bases found no significant predictor, so that calibration awaits more bases and skills.","The multi-turn Crescendo arms show merging does not push adaptive risk beyond the base's own ceiling, which suggests the binding constraint for a merged model is the base's intrinsic robustness; a cheaper mitigation than post-merge repair would be to select robust bases (Llama, Phi-4) at the start."],"forward_implications":["A low static refusal score on a merged model carries no information about its adaptive robustness: safe-looking Qwen and Gemma merges are jailbroken 60–76% of the time by a single fixed template attack, so a practitioner who screens only statically will deploy the very models that most need hardening.","Adaptive evaluation is necessary, not optional, for merged models: the shortfall of static screening is largest for strongly-aligned, static-safe merges, precisely the ones a practitioner would most trust.","The safety cost of merging is base-conditional rather than algorithm-conditional: all four merging methods agree to within 0.035 static ASR with identical ordering, while a weakly-aligned base (Mistral) is eroded by even a benign math merge and fragile bases (Gemma) inherit fragility from the unmerged base itself.","A data-free geometric screen can be applied before any merge or attack: the overlap of the candidate task vector with the safety subspace $S$ separates same-recipe refusal-removal from genuine skills, and projecting that overlap off restores safety to base at held capability on strongly-aligned bases.","The repair is honestly bounded: SubSafe-Merge removes only in-$S$ erosion, cannot fix out-of-$S$ benign erosion (an Alpaca fine-tune still raises static ASR from 0.20 to 0.28), and cannot make a model more adaptively robust than its base (Mistral's adaptive ASR remains $\\approx 0.66$)."],"supporting_citations":[{"why":"Defines task vectors and task arithmetic, the merging formalism the benchmark operates on.","marker":"Ilharco et al. 2023"},{"why":"Provides GCG, the gradient-search attack whose 100- and 500-step runs yield the adaptive ASR numbers.","marker":"Zou et al. 2023"},{"why":"Supplies the 400 HarmBench behaviors and the classifier judge used in the two-judge AND rule.","marker":"Mazeika et al. 2024"},{"why":"The shallow-alignment result (refusal lives in the first few generated tokens) that explains why static and adaptive safety decouple.","marker":"Qi et al. 2025"},{"why":"Evidence that refusal is mediated by an approximately one-dimensional direction, supporting the safety-subspace construction.","marker":"Arditi et al. 2024"},{"why":"The closest prior work showing merge-induced safety loss with a static signal, which this paper extends to adaptive attacks.","marker":"Hammoud et al. 2024"},{"why":"Shows that fine-tuning can efficiently undo safety training, a premise for the fragility of alignment.","marker":"Lermen, Rogers-Smith, and Ladish 2023"},{"why":"Provides PAIR, the attacker-LLM family that confirms the fragile/robust ordering directionally.","marker":"Chao et al. 2023"},{"why":"Provides Crescendo, the multi-turn attack used to check that merging adds no adaptive risk beyond the base's ceiling.","marker":"Russinovich, Salem, and Eldan 2025"},{"why":"SafeMERGE, the selective-layer baseline whose head-to-head with SubSafe-Merge bounds the subspace-projection claim.","marker":"Djuhera et al. 2026"}],"fun_headline_variants":["Safe-looking merges jailbroken 60-76% under template attack","Static safety check misses 60-76% jailbreaks on Qwen, Gemma","Qwen, Gemma merges look safe but fail 60-76% template jailbreak","Safe static score, broken under attack: merged LLM fragility (Qwen, Gemma)","Skill-merge static safety is not adaptive robustness (Qwen/Gemma prove it)"],"cache_read_input_tokens":24576,"weakest_assumption_plain":"The safety subspace $S$, estimated from a single public abliterated model per base, faithfully represents the directions through which merging erodes safety; the paper's own survey shows the estimate is recipe-dependent, so a donor that removes refusal by SFT/DPO decensoring rather than abliteration is invisible to the geometric screen (overlap 0.001 despite unmerged static ASR 0.63) and is not repaired by SubSafe-Merge.","fun_headline_variants_meta":{"raw":{"variants":["Safe-looking merges jailbroken 60-76% under template attack","Static safety check misses 60-76% jailbreaks on Qwen, Gemma","Qwen, Gemma merges look safe but fail 60-76% template jailbreak","Safe static score, broken under attack: merged LLM fragility (Qwen, Gemma)","Skill-merge static safety is not adaptive robustness (Qwen/Gemma prove it)"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000958,"raw_usage":{"total_tokens":4264,"prompt_tokens":1309,"completion_tokens":2955,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":925,"completion_tokens_details":{"reasoning_tokens":2844}},"tokens_in":925,"tokens_out":2955,"duration_ms":21235,"temperature":1.0,"reasoning_tokens":2844,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:32:28.348310+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the paper's six bases and push each adaptive attack well beyond the reported budgets: a GCG run of 1,000+ steps with a wider suffix search, a GPT-4-class attacker in the PAIR loop, and the same best-of-6 template set, then rank the bases by adaptive ASR. If Llama-3.1-8B and Phi-4-mini, the paper's attack-resistant bases, break at rates comparable to Qwen and Gemma under any of these stronger probes, the fragile/robust ordering is an artifact of attack strength; if the ordering survives at every budget, the conclusion that static refusal screening cannot tell which merges are safe to deploy is confirmed.","supporting_citations":[],"review_version":1}