Pith. sign in

REVIEW 2 major objections 3 minor 55 references

Around a trained multitask LoRA checkpoint, a single task perturbation is first-order predictable through scale 1e-2, but two sequential task updates stop commuting at a pair-specific scale set by a curvature commutator with onset ηκ≈0.1.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 19:50 UTC pith:YKU5H7TK

load-bearing objection A careful, honestly-scoped empirical atlas showing that one-step local adaptation is first-order predictable to 1e-2 while two-step order effects can break inside that window; the Lie-bracket law is solid in fixed LoRA coordinates, but gauge-dependence and the failed cross-fit limit its reach. the 2 major comments →

arxiv 2607.16821 v1 pith:YKU5H7TK submitted 2026-07-18 cs.LG cs.AI

First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers

classification cs.LG cs.AI
keywords local linearityLie bracketupdate order sensitivitytask arithmeticLoRAHessian-vector productsprobe losssequential fine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Around one multitask LoRA checkpoint in nine transformers (82M–7B), the paper measures eight local properties that task arithmetic, sequential fine-tuning, activation steering, and random search all assume. It finds a shared one-direction window: on every model median, a perturbation's effect on a fixed 32-example probe loss is essentially its projection onto the gradient, through the tested scale 1e-2. Pairwise structure is much more fragile: on more than a third of (model, task-pair) cells, order sensitivity between two task-gradient updates sets in strictly inside that window, and the leading term is the Lie bracket H_B g_A − H_A g_B. With matched minibatches, the order defect follows c(η)=ηκ+O(η²) with median ratio 1.002 and onset η†≈0.10/κ, so two Hessian-vector products can warn about a pair. The paper also reports several registered properties that fail—mixture-gradient rank, tangent-subspace stability, full-scale activation additivity, a global weight-to-steering bridge—and a cross-sample forecast that fails to transfer.

Core claim

The central discovery is that the local geometry around a trained checkpoint separates cleanly by symmetry into three objects with different scales: self-curvature governs single-direction antisymmetry, a symmetric mixed derivative governs simultaneous activation additivity, and the antisymmetric Lie bracket H_B g_A − H_A g_B governs the order of two sequential memoryless gradient updates. In the fixed LoRA factor coordinates used by the optimizer, the normalized endpoint defect is c(η)=ηκ+O(η²) with κ = ‖H_B g_A − H_A g_B‖/‖g_A+g_B‖; when both the two-step paths and the Hessian-vector products use the same frozen minibatches, finite-difference and HVP routes agree at median ratio 1.002, and

What carries the argument

The Lie bracket of the two task-gradient vector fields, H_B g_A − H_A g_B, normalized by ‖g_A+g_B‖. It is the leading order-dependent difference between the two update orders: whichever task goes second takes its step from ground the first task has already moved, so the expansion produces the bracket at order η². Two Hessian-vector products at the operating point compute κ, and the onset prediction η†≈0.10/κ is parameter-free. The measurements all live in the fixed Euclidean coordinates of the LoRA factors, which the paper states explicitly is gauge-dependent.

Load-bearing premise

Every measured radius, window, and onset scale is quoted in the fixed Euclidean coordinates of the LoRA factors, and those coordinates are not unique: the same network can be represented with different factor scalings, so the headline numbers might be an artifact of that coordinate choice.

What would settle it

Re-run the fine-grid commutator protocol after applying a function-preserving gauge transformation to the LoRA factors (e.g., multiply A by c and B by 1/c) and recompute the onset for the same task pairs; if the products η†κ scatter outside ~0.10 and the onset moves materially, the law is coordinate-specific. A second falsifier is to compute the bracket in a function-space or Fisher metric and check whether the median ratio 1.002 survives.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • For a single intended direction, probe-loss changes are first-order predictable through the tested grid on every model median, so random proposals can be ranked by their gradient projection before running any of them.
  • For two sequential memoryless gradient steps, update-order sensitivity is set by κ: two Hessian-vector products give the warning onset η†≈0.10/κ, and pair-to-pair differences in κ reach up to ~15× within one model.
  • No universal radius exists: on over a third of measured model–task-pair combinations, order sensitivity starts inside the single-update window, so composition claims must be checked per pair and per scale.
  • Activation additivity at full task-vector scale fails on several models, including both held-out 7B models; additivity holds on all models only for fractions α≲0.3 under this probe.
  • The bracket is a same-probe diagnostic, not a transferable predictor: cross-sample prediction of endpoints or task effects from the bracket was not validated.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: because all scales are quoted in the fixed LoRA factor coordinates, a function-preserving rescaling of the factors could move the measured windows and onsets; the 1e-2 and 0.1/κ numbers should be expected to change under a gauge-invariant metric.
  • Beyond the paper: the matched-route success and cross-sample failure together suggest that minibatch stochasticity, not just curvature, is what breaks transfer; a stochastic-bracket estimator with covariance correction is a natural next test.
  • Beyond the paper: the symmetry split (self-curvature vs symmetric mixed derivative vs antisymmetric bracket) predicts that each adaptation tool needs its own diagnostic; a single 'linear regime' radius is unlikely to exist for any model.
  • Beyond the paper: the η†≈0.10/κ law is cheap to check on new architectures and task pairs; if it reproduces, update-order sensitivity can be screened without running the two-step optimization.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper measures eight local properties of the loss landscape around a fixed multitask LoRA operating point on nine transformers (82M–7B), using a prospectively registered property list, thresholds, development/held-out split, and a uniform harness. The main positive finding is single-update first-order predictability: along a task's gradient, probe-loss changes are dominated by the linear term through the tested scale 10^-2 on every model median, and first-order scores rank random perturbations. The main negative finding is pairwise fragility: order sensitivity of two sequential task-gradient steps sets in at pair-specific scales, sometimes well inside the single-update window. For two memoryless gradient steps, the leading order-dependent endpoint difference is the Lie bracket H_B g_A - H_A g_B, and with matched minibatches the normalized defect satisfies c(eta)=eta*kappa+O(eta^2) at median ratio 1.002; the onset eta_dag defined by c(eta)=0.10 gives eta_dag*kappa close to 0.10. The paper also reports several registered failures: the mixture-gradient covariance low-rank test is non-falsifiable at m=32 and shows no plateau up to m=256, activation additivity fails at full task-vector scale on several models including both held-out 7B models, the global mean-vector weight-to-steering correspondence fails on all model medians, and the cross-fitted functional forecast from the bracket does not validate.

Significance. The paper is methodologically strong in several respects: the pre-registration, explicit reporting of failed controls, per-seed tables, threshold-sensitivity analysis, and the clean shared-minibatch HVP/finite-difference agreement at median ratio 1.002 are all commendable. The atlas across nine models, with a single frozen harness, is a useful empirical resource, and the honest reporting of the failed cross-fitted forecast is scientifically valuable. If the results hold, the practical message—check the diagnostic that matches your operation at your scale and on your pair—is sensible. However, the central quantitative claim about the onset scale eta_dag ~ 0.10/kappa and its three-order-of-magnitude span across models and pairs is coordinate-specific and partly definitional, so its interpretation as a statement about task geometry requires additional support or careful reframing.

major comments (2)
  1. [§6.2, Table 3/Fig. 9 and §3 'Coordinate convention'] The central quantitative claim—that kappa is a pair-specific warning signal whose scale and ordering span three orders of magnitude across models—is measured entirely in fixed Euclidean coordinates of the LoRA factors. The paper acknowledges the gauge freedom (B,A) -> (BQ, Q^{-1}A) and states in Limitation 1 that no gauge-rescaling, balanced-gauge, Fisher, or output-KL control is reported. This is not merely a philosophical caveat: gradient, Hessian, and the HVP commutator are all coordinate-dependent, so both kappa and the measured onset eta_dag can change under a function-preserving reparametrization. As written, the abstract's claim that eta_dag ~ 0.10/kappa 'spans three orders of magnitude across models and task pairs' and the Section 8 advice to use kappa as a pair-specific order-sensitivity diagnostic are not yet supported as statements about task geometry; they are statements abou
  2. [§6.1, Eq. (10) and §6.2] The 'law' eta_dag*kappa ~ 0.10 is a consequence of the definition of eta_dag as the first point with c(eta) >= 0.10 together with the leading-order expansion c(eta) = eta*kappa + O(eta^2); it is not an independent empirical discovery. The paper acknowledges this ('Because eta_dag is read from the same defect curve... the substantive content is route-to-route consistency'), but the abstract and Section 6.2 headline the onset span as a main result. The actual empirical content is the pointwise agreement between finite-difference defects and HVP commutators at median ratio 1.002 (Fig. 8b), which is strong. Please restructure the presentation so that the route-to-route consistency is the claim, and the product eta_dag*kappa ~ 0.10 is presented as a consistency check of the threshold definition rather than as a separate falsifiable law.
minor comments (3)
  1. [Abstract and §5.4] The one-direction window is right-censored at 10^-2 on every model median, so the phrase 'through the tested scale 10^-2' is accurate but should be read as a lower bound. The common-ruler analysis in §5.4 is a descriptive reanalysis and is clearly labeled as such; consider stating this in the abstract as well.
  2. [Table 2 and §6] DistilGPT-2 and OPT-1.3B are excluded from the fine-grid commutator measurements because of failed HVP reconstruction checks, but Table 2 lists P8 onsets for them on the coarse grid. This should be stated more prominently, as the 'three orders of magnitude' span is over the seven passing models, not all nine.
  3. [§6.5] The lambda_max association is explicitly exploratory and based on seven model medians, with a failed out-of-sample forecast. The text already qualifies this appropriately; the one seed-level reproducibility mismatch (Pythia-160M seed 1) should also be mentioned in the main text rather than only in the scoring-script flag.

Circularity Check

1 steps flagged

Central Lie-bracket expansion is genuine and self-contained; the onset law eta^dag~0.10/kappa is a disclosed arithmetic corollary of the definition of eta^dag, not an independent prediction.

specific steps
  1. self definitional [Section 3 (P8: 'we report the onset η† = min{η : c(η) ≥ 0.10}'); Section 6.2; Table 3]
    "we report the onset η† = min{η : c(η) ≥ 0.10} ... Because η† is read from the same defect curve, η†κ≈ 0.10 is arithmetic once the leading expansion holds pointwise; the substantive content is route-to-route consistency."

    η† is defined as the crossing of the measured defect curve c(η) at 0.10, and κ is the measured slope of that same curve (c(η)=ηκ, verified with matched finite-difference/HVP estimators). At the crossing, 0.10=c(η†)≈η†·κ, so the pre-registered 'prediction' η†·κ≈0.10 (Table 3: every product in [0.096,0.109]) is an arithmetic corollary of the definition plus the verified slope law, not an independent test. The paper discloses this explicitly. Independent content survives in the route-to-route slope agreement (κ from double-backward HVPs vs finite-difference two-step endpoints, median ratio 1.002, 94% of mid-range points within 10%), and the three-orders-of-magnitude span is a real measured property of κ; but the onset relation itself adds no measurement beyond rescaling κ by the fixed thresho

full rationale

The core derivation is not circular: Section 6.1 re-derives the Lie-bracket leading term by elementary Taylor expansion, and the median ratio 1.002 between the HVP-computed κ and the finite-difference defect is a genuine two-route consistency check (no fitted parameters; the routes compute different objects). The single partially definitional element is the onset law η†≈0.10/κ, which reduces to the definition of η† plus the verified pointwise law; the paper itself states this ('arithmetic once the leading expansion holds pointwise'), so it is disclosed, partial, and a corollary rather than the central claim. No load-bearing self-citation: the companion best-of-N formula [31] is re-derived and corrected here with standard normal order statistics, and the bracket attribution to [34,39] does not carry the derivation since Section 6.1 re-derives it. The paper's own reported failures and scoping statements reduce circularity concerns: the cross-fitted forecast fails (R2_zero=-560.4), the λmax forecast is undecidable, P2's registered bar is exposed as algebraically non-falsifiable, P6 is disclosed as in-sample, and Limitation 1 concedes the fixed-coordinate, gauge-dependent scope. Gauge dependence is a validity/robustness concern about what the numbers mean, not a reduction of the derivation to its inputs, and per the reviewing rules it does not count as circularity. Score 3 reflects the one disclosed, definitional corollary while recognizing that the substantive two-route check stands.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

No new physical or architectural entities are introduced; 'primed' is a descriptive label for a checkpoint that still has residual gradients. The central quantitative claims rest on hand-set thresholds, the fixed-coordinate gauge convention, a 32-example probe, smoothness assumptions, and memoryless gradient steps. The thresholds are registered, not fitted, but several verdicts are definitional once those thresholds are fixed.

free parameters (4)
  • commutator threshold c>=0.10 = 0.10
    Hand-set threshold that defines eta_dag = min{eta: c(eta)>=0.10}; it makes eta_dag*kappa ~ 0.10 arithmetic at leading order if the expansion holds.
  • P1 antisymmetry threshold A<0.10 = 0.10
    Hand-set bar for the one-direction linearity window; changing it would move the P1 radius and the '1e-2 window' claim.
  • P4 additivity threshold epsilon_add<0.15 = 0.15
    Hand-set bar for activation additivity pass/fail; sensitivity analysis shows the held-out verdict remains a drop at full scale under nearby bars.
  • P5 correspondence cosine bar = 0.30
    Hand-set bar for weight-to-CAA correspondence; no model median reaches it.
axioms (5)
  • ad hoc to paper Euclidean geometry in the fixed LoRA factor coordinates is the right metric for the measured radii and curvatures.
    Section 3 'Coordinate convention' states all norms, gradients, Hessians, and kappa are in fixed factor coordinates. Function-preserving gauge transforms can change every quantitative claim, so this is a load-bearing, paper-specific convention.
  • domain assumption The loss and activation maps are smooth enough for second-order Taylor expansions.
    The Lie-bracket derivation and P4's mixed-derivative prediction assume C^3 / twice differentiable behavior, which is not guaranteed for ReLU networks.
  • domain assumption The fixed 32-example multitask probe is representative of the task loss, gradient, and Hessian.
    All of L, g, H, P1, P6, P7, and the commutator measurements use this probe; P6 scores and evaluates on the same probe, so it is an in-sample consistency check.
  • domain assumption Memoryless gradient steps without optimizer state are the relevant update model for P8.
    The bracket derivation assumes no momentum or AdamW state; the paper notes that fixed-clock optimizer state can introduce a leading O(eta) order term.
  • domain assumption One 300-step multitask LoRA operating point represents adaptation behavior across model families and scales.
    The atlas measures nine models but at a single operating point, one task mixture, and three task pairs; generalization beyond this operating point is extrapolation.

pith-pipeline@v1.3.0-alltime-deepseek · 31751 in / 12098 out tokens · 127130 ms · 2026-08-01T19:50:16.422926+00:00 · methodology

0 comments
read the original abstract

Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space. We measure 8 such properties with the same harness around a multitask LoRA operating point, on 9 transformers (82M-7B), with a prospectively registered property list, thresholds, and test split. We find a shared one-direction validity window up to the tested scale $10^{-2}$, but no universal radius for pairwise composition or update ordering. Along individual directions, changes of the probe loss remain first-order predictable throughout the grid: a perturbation's effect on the loss is essentially its projection onto the gradient, which is also what makes local random search work. Pairwise structure, however, proves to be far more fragile: on over a third of the measured (model, task pair) combinations, two-update order sensitivity sets in strictly inside that window; task-gradient subspaces rotate within tens of steps; additivity under our fixed activation probe fails at full task-vector scale on several models, including both held-out 7B models; and no model median passes the registered global mean-vector weight-to-steering correspondence bar. For two sequential task-gradient steps, the leading order-dependent term is the Lie bracket $H_B\textbf{g}_A-H_A\textbf{g}_B$; its normalized prediction $c(\eta)=\eta\kappa+O(\eta^2)$ tracks the measured defect at median ratio 1.002, while the onset scale $\eta^\dagger\approx0.10/\kappa$ spans three orders of magnitude across models and task pairs.

Figures

Figures reproduced from arXiv: 2607.16821 by Irina Piontkovskaia, Sergey Nikolenko.

Figure 1
Figure 1. Figure 1: Single-update predictability versus pairwise order dependence. All parameter-space [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Extended loss sweeps at θ ⋆ (one seed, three models): per-task probe loss along ±gˆt (colored), two fixed random unit directions (gray), and the largest contiguous tested segment with At(s) < 0.10 (dotted verticals). Protocol and registered predictions in Section 4. reported every failure, and touched the three held-out models only once (see Appendix A for detailed provenance). Before any statistical resul… view at source ↗
Figure 3
Figure 3. Figure 3: P2 sample-count control: r90 against m for the seven models (thin lines, seed 0), their median (bold), and the matched isotropic null (dashed). 5.1 What fails The mixture-gradient covariance shows no low-rank plateau. The substantive P2 question is whether the mixture-gradient covariance concentrates in a small subspace whose dimension stops growing once enough samples are seen. It does not: in a seed-0, s… view at source ↗
Figure 4
Figure 4. Figure 4: P3 on the BoolQ trajectory (seed 0): overlap with the step-8 subspace against the same-checkpoint resampling floor. 0.03 0.1 0.3 1.0 task-vector fraction 10 2 10 1 co mposition error add( ) additivity bar = 0.15 norm-matched random null (median) distilgpt2 pythia-160m pythia-410m neo-1.3B opt-1.3b TinyLlama-1.1B pythia-1.4B (held-out) Qwen2.5-7B (held-out) OLMo-7B (held-out) [PITH_FULL_IMAGE:figures/full_… view at source ↗
Figure 5
Figure 5. Figure 5: P4 composition error εadd(α), per-model median over pairs and seeds (log–log); held-out models dashed, norm-matched random null in gray. The registered global mean-vector correspondence does not generalize. For P5, model￾median cosines between one SST-2 gradient step’s mean activation shadow and one labelled-contrast CAA mean vector range from −0.22 to +0.22: no model median reaches the 0.30 bar, including… view at source ↗
Figure 6
Figure 6. Figure 6: P7 per-seed |gˆ ⊤Hgˆ/r ⊤Hr|; open circles mark seeds failing the finite-difference HVP reconstruction check, amber bars are per-model medians of |ratio|. question: on the seven models passing the reconstruction check, the task-direction ratio to the median random quotient is 1.4 · 104–5.9 · 105 , and the more stable comparator ∥Hv∥ gives 71×–590× against its median. The sign is a different matter: the sign… view at source ↗
Figure 7
Figure 7. Figure 7: All measured boundaries on the common displacement ruler, models ordered by median [PITH_FULL_IMAGE:figures/full_fig_p015_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Fine-grid defect curves for all 42 (pair, seed) cells on the seven models passing the HVP reconstruction check: (a) raw c(η); (b) the same curves against ηκ, collapsing onto the identity (dashed) [PITH_FULL_IMAGE:figures/full_fig_p017_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Measured onset η † against the two-HVP prediction 0.1/κ for the 41 measured fine-grid cells, colored by task pair (circles seed 0, triangles seed 1); the band marks ×1.5 [PITH_FULL_IMAGE:figures/full_fig_p018_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Model-median κ against the probe-loss λmax on the seven models passing the HVP reconstruction check (log–log). 6.5 An exploratory association, and a forecast that failed Across the seven model medians that pass the HVP reconstruction check, κ is rank-associated with the probe-loss λmax at Spearman +1.0 with log–log slope 0.85 ( [PITH_FULL_IMAGE:figures/full_fig_p021_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Base-point per-task along-gradient loss reduction over the tested [PITH_FULL_IMAGE:figures/full_fig_p023_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

55 extracted references · 38 linked inside Pith

  1. [1]

    Intrinsic dimensionality explains the effectiveness of language model fine-tuning

    Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. InAnnual Meeting of the Association for Computational Linguistics (ACL), 2021

  2. [2]

    Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices.Annals of Probability, 33:1643–1697, 2005

    Jinho Baik, Gerard Ben Arous, and Sandrine Peche. Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices.Annals of Probability, 33:1643–1697, 2005

  3. [3]

    The singular values and vectors of low rank perturbations of large rectangular random matrices.Advances in Mathematics (arXiv:0910.2120), 2011

    Florent Benaych-Georges and Raj Rao Nadakuditi. The singular values and vectors of low rank perturbations of large rectangular random matrices.Advances in Mathematics (arXiv:0910.2120), 2011

  4. [4]

    Pythia: A suite for analyzing large language models across training and scaling

    Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. Pythia: A suite for analyzing large language models across training and scaling. InInternational Conference on Machine Learning (...

  5. [5]

    GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow

    Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow. Zenodo, 2021. URL https: //doi.org/10.5281/zenodo.5297715

  6. [6]

    BoolQ: Exploring the surprising difficulty of natural yes/no questions

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In NAACL-HLT, 2019

  7. [7]

    Think you have solved question answering? try ARC, the AI2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try ARC, the AI2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018

  8. [8]

    da Silva, Mohammed Adnan, Felix Dangel, and Sageev Oore

    Marvin F. da Silva, Mohammed Adnan, Felix Dangel, and Sageev Oore. Generalizing the geometry of model merging through Fréchet averages.arXiv preprint arXiv:2604.27155, 2026. URLhttps://arxiv.org/abs/2604.27155

  9. [9]

    Sharp minima can generalize for deep nets.arXiv:1703.04933 (ICML), 2017

    Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. Sharp minima can generalize for deep nets.arXiv:1703.04933 (ICML), 2017

  10. [10]

    Hamprecht

    Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred A. Hamprecht. Essentially no barriers in neural network energy landscape. InInternational Conference on Machine Learning (ICML), 2018. URLhttps://arxiv.org/abs/1803.00885

  11. [11]

    Sharpness-aware mini- mization for efficiently improving generalization.arXiv:2010.01412 (ICLR), 2021

    Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware mini- mization for efficiently improving generalization.arXiv:2010.01412 (ICLR), 2021

  12. [12]

    Deep ensembles: A loss landscape perspective.arXiv preprint arXiv:1912.02757, 2019

    Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan. Deep ensembles: A loss landscape perspective.arXiv preprint arXiv:1912.02757, 2019

  13. [13]

    Neural thickets: Diverse task experts are dense around pretrained weights.arXiv preprint arXiv:2603.12228, 2026

    Yulu Gan and Phillip Isola. Neural thickets: Diverse task experts are dense around pretrained weights.arXiv preprint arXiv:2603.12228, 2026. URLhttps://arxiv.org/abs/2603.12228

  14. [14]

    Loss surfaces, mode connectivity, and fast ensembling of DNNs

    Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry Vetrov, and Andrew Gordon Wilson. Loss surfaces, mode connectivity, and fast ensembling of DNNs. InAdvances in Neural Information Processing Systems (NeurIPS), 2018. URLhttps://arxiv.org/abs/1802.10026

  15. [15]

    An investigation into neural net optimization via hessian eigenvalue density

    Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao. An investigation into neural net optimization via hessian eigenvalue density. InICML, 2019

  16. [16]

    OLMo: Accelerating the science of language models

    Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. OLMo: Accelerating the science of language models. InAnnual Meeting of the Association for Computational Linguistics (ACL), 2024. URLhttps://arxiv.org/abs/2402.00838

  17. [17]

    Roberts, and Ethan Dyer

    Guy Gur-Ari, Daniel A. Roberts, and Ethan Dyer. Gradient descent happens in a tiny subspace. arXiv:1812.04754, 2018

  18. [18]

    LoRA: Low-rank adaptation of large language models

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. InInternational Conference on Learning Representations (ICLR), 2022. URLhttps://arxiv.org/abs/2106. 09685. 27

  19. [19]

    DistilGPT2 model card

    Hugging Face. DistilGPT2 model card. Hugging Face model repository, 2019. URLhttps: //huggingface.co/distilbert/distilgpt2. Accessed 2026-07-10

  20. [20]

    Editing models with task arithmetic

    Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic. InInternational Conference on Learning Representations (ICLR), 2023. URLhttps://arxiv.org/abs/2212.04089

  21. [21]

    On large-batch training for deep learning: Generalization gap and sharp minima

    Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On large-batch training for deep learning: Generalization gap and sharp minima. InInternational Conference on Learning Representations (ICLR), 2017. URLhttps: //arxiv.org/abs/1609.04836

  22. [22]

    Measuring the intrinsic dimension of objective landscapes

    Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski. Measuring the intrinsic dimension of objective landscapes. InInternational Conference on Learning Representations (ICLR), 2018. URLhttps://arxiv.org/abs/1804.08838

  23. [23]

    Visualizing the loss landscape of neural nets

    Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. Visualizing the loss landscape of neural nets. InAdvances in Neural Information Processing Systems (NeurIPS), 2018

  24. [24]

    Fine-tuning language models with just forward passes

    Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora. Fine-tuning language models with just forward passes. InAdvances in Neural Information Processing Systems (NeurIPS), 2023

  25. [25]

    Locating and editing factual associations in gpt.NeurIPS (arXiv:2202.05262), 2022

    Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in gpt.NeurIPS (arXiv:2202.05262), 2022

  26. [26]

    Phase transitions for feature learning in neural networks

    Andrea Montanari and Zihao Wang. Phase transitions for feature learning in neural networks. arXiv:2602.01434, 2026. URLhttps://arxiv.org/abs/2602.01434

  27. [27]

    Dome: Im- proving signal-to-noise in stochastic gradient descent via sharp-direction subspace filtering

    Julien Nicolas, Mohamed Maouche, Sonia Ben Mokhtar, and Mark Coates. Dome: Im- proving signal-to-noise in stochastic gradient descent via sharp-direction subspace filtering. arXiv:2507.03545, 2025. URLhttps://arxiv.org/abs/2507.03545

  28. [28]

    Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians.arXiv:1901.08244 (ICML), 2019

    Vardan Papyan. Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians.arXiv:1901.08244 (ICML), 2019

  29. [29]

    The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size.arXiv:1811.07062, 2019

    Vardan Papyan. The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size.arXiv:1811.07062, 2019

  30. [30]

    Asymptotics of sample eigenstructure for a large dimensional spiked covariance model.Statistica Sinica, 17:1617–1642, 2007

    Debashis Paul. Asymptotics of sample eigenstructure for a large dimensional spiked covariance model.Statistica Sinica, 17:1617–1642, 2007

  31. [31]

    Recoverable but not stationary: Local linear structures in weights and activations.arXiv preprint arXiv:2606.10929, 2026

    Irina Piontkovskaia and Sergey Nikolenko. Recoverable but not stationary: Local linear structures in weights and activations.arXiv preprint arXiv:2606.10929, 2026. URL https: //arxiv.org/abs/2606.10929

  32. [32]

    Steering llama 2 via contrastive activation addition

    Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. Steering llama 2 via contrastive activation addition. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15504–15522,

  33. [33]

    Discretization drift in two-player games

    Mihaela Rosca, Yan Wu, Benoit Dherin, and David G T Barrett. Discretization drift in two-player games. InICML, 2021. arXiv:2105.13922

  34. [35]

    Ugur Guney, Yann Dauphin, and Leon Bottou

    Levent Sagun, Utku Evci, V. Ugur Guney, Yann Dauphin, and Leon Bottou. Empirical analysis of the hessian of over-parametrized neural networks.arXiv:1706.04454, 2017

  35. [36]

    Evolution strategies as a scalable alternative to reinforcement learning.arXiv preprint arXiv:1703.03864, 2017

    Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. Evolution strategies as a scalable alternative to reinforcement learning.arXiv preprint arXiv:1703.03864, 2017

  36. [37]

    On the origin of implicit regularization in stochastic gradient descent

    Samuel L Smith, Benoit Dherin, David G T Barrett, and Soham De. On the origin of implicit regularization in stochastic gradient descent. InICLR, 2021. arXiv:2101.12176

  37. [38]

    Does sgd really happen in tiny subspaces? In International Conference on Learning Representations (ICLR),2025

    Minhak Song, Kwangjun Ahn, and Chulhee Yun. Does sgd really happen in tiny subspaces? In International Conference on Learning Representations (ICLR),2025. URLhttps://openreview. net/forum?id=v6iLQBoIJw

  38. [39]

    The geometry of sequential learning: Lie-bracket prediction of transfer order

    John Sweeney. The geometry of sequential learning: Lie-bracket prediction of transfer order. Proceedings of the 43rd International Conference on Machine Learning, 2026. URLhttps: //arxiv.org/abs/2606.24993. arXiv:2606.24993

  39. [40]

    Optimizer memory makes shuffle order a first-order source of fine-tuning noise

    John Sweeney. Optimizer memory makes shuffle order a first-order source of fine-tuning noise. arXiv preprint arXiv:2606.29554, 2026. URLhttps://arxiv.org/abs/2606.29554

  40. [41]

    Parameter efficient multi-task model fusion with partial linearization.arXiv preprint arXiv:2310.04742, 2023

    Anke Tang, Li Shen, Yong Luo, Yibing Zhan, Han Hu, Bo Du, Yixin Chen, and Dacheng Tao. Parameter efficient multi-task model fusion with partial linearization.arXiv preprint arXiv:2310.04742, 2023. URLhttps://arxiv.org/abs/2310.04742

  41. [42]

    Predicting mergeability of parameter-efficient fine-tuning updates.arXiv preprint arXiv:2606.19549, 2026

    Lin Tang, Wei Zhang, Jing Li, Hongyu Chen, Ming Zhao, and Yuxuan Wang. Predicting mergeability of parameter-efficient fine-tuning updates.arXiv preprint arXiv:2606.19549, 2026. URLhttps://arxiv.org/abs/2606.19549

  42. [43]

    2 OLMo 2 furious

    Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, Nathan Lambert, et al. 2 OLMo 2 furious. arXiv preprint arXiv:2501.00656, 2024. URLhttps://arxiv.org/abs/2501.00656

  43. [44]

    Li, Arnab Sen Sharma, Aaron Mueller, Byron C

    Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau. Function vectors in large language models. InInternational Conference on Learning Representations (ICLR), 2024. URLhttps://arxiv.org/abs/2310.15213

  44. [45]

    Steering language models with activation engineering.arXiv preprint arXiv:2308.10248, 2023

    Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering.arXiv preprint arXiv:2308.10248, 2023. URLhttps://arxiv.org/abs/2308.10248

  45. [46]

    Introduction to the non-asymptotic analysis of random matrices

    Roman Vershynin. Introduction to the non-asymptotic analysis of random matrices. arXiv:1003.2990 (in: Compressed Sensing, CUP 2012), 2010

  46. [47]

    Model soups: Averaging weights of multiple fine-tuned models improves 29 accuracy without increasing inference time

    Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo- Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt. Model soups: Averaging weights of multiple fine-tuned models improves 29 accuracy without increasing inference time. InInternational Conference on Machine Learning...

  47. [48]

    Manning, and Christopher Potts

    Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D. Manning, and Christopher Potts. ReFT: Representation finetuning for language models. In Advances in Neural Information Processing Systems (NeurIPS), 2024. URLhttps://arxiv. org/abs/2404.03592

  48. [49]

    TIES-merging: Resolving interference when merging models

    Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal. TIES-merging: Resolving interference when merging models. InAdvances in Neural Information Processing Systems (NeurIPS), 2023. URLhttps://arxiv.org/abs/2306.01708

  49. [50]

    Qwen2.5 technical report.arXiv preprint arXiv:2412.15115, 2024

    An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2.5 technical report.arXiv preprint arXiv:2412.15115, 2024. URLhttps://arxiv.org/abs/2412.15115

  50. [51]

    HellaSwag: Can a machine really finish your sentence? InAnnual Meeting of the Association for Computational Linguistics (ACL), 2019

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? InAnnual Meeting of the Association for Computational Linguistics (ACL), 2019

  51. [52]

    TinyLlama: An open-source small language model.arXiv preprint arXiv:2401.02385, 2024

    Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, and Wei Lu. TinyLlama: An open-source small language model.arXiv preprint arXiv:2401.02385, 2024. URLhttps://arxiv.org/abs/ 2401.02385

  52. [53]

    OPT: Open pre-trained transformer language models.arXiv preprint arXiv:2205.01068, 2022

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. OPT: Open pre-trained transformer language models.arXiv preprint arXiv:2205.01068...

  53. [54]

    registered

    Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. Representation engineering: A top- down approach to ...

  54. [2024]

    URLhttps://aclanthology.org/2024.acl-long.828/. 28

  55. [2025]

    URLhttps://arxiv.org/abs/2501.15556