REVIEW 3 major objections 7 minor 22 references
Eigenvalues miss finite-depth attention dynamics; sinks damp transients and a routing minority holds reserve that eigenvalues cannot see.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 04:10 UTC pith:IUGUYSCM
load-bearing objection Resolvent tools on the attention propagator give a real, pre-registered win over eigenvalues for depth-transient persistence and routing identity, with the 7–11-order failure carefully scoped to the attention skeleton. the 3 major comments →
Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On trained transformers the eigenvalue picture of attention depth products is not a loose bound but the wrong account: surviving deviations exceed the eigenvalue heuristic by 10^7 to 10^11, while mass-matched nulls stay calibrated. Learned non-normality is signed—routing heads hold excess transient reserve, sinks damp it below chance—and resolvent features are required, under pre-registered cross-validated rules, for depth-transient persistence and routing identity.
What carries the argument
Exact Perron deflation: with P the mask-forced projector 1 e_0^T, A^k = P + A_dev^k and the depth deviation product factors as Π_L − P = product of the deflated layer-mean operators. All transient metrics and the eigenvalue-failure comparison are computed on this deflated skeleton against row-permutation nulls.
Load-bearing premise
The claim rests on treating head-mean attention-only depth products (without value maps, MLPs, or residual paths except a separate rollout check) as a faithful enough skeleton of how tokens mix.
What would settle it
Re-run the pre-registered feature contest and the depth-product comparison after replacing the attention-only head-mean stack with the full residual-stream Jacobian or with products that include OV and MLP maps; if eigenvalue-side features then match or beat resolvent features on half-depth persistence and the 10^7–10^11 gap disappears, the central claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that causal attention matrices are non-normal by construction, so finite-depth token mixing is governed by resolvent quantities (pseudospectra, Kreiss constants) rather than eigenvalues. After establishing that the mask pins the raw Kreiss constant at √n and that Perron deflation yields an exact product factorization of depth deviation dynamics, it reports a signed census across GPT-2, Pythia-410m, and Llama-3-8B: a routing minority carries excess transient reserve correlated with previous-token function, while the sink majority is suppressed below mass-matched nulls and acts as a transient damper. On head-mean depth products the eigenvalue heuristic underpredicts surviving deviations by 10^7–10^11 on trained stacks (an error absent in matched nulls). Supporting results include state-conditional reserve in induction heads, post-formation consolidation on Pythia checkpoints, an architecture-conditional clamping intervention linking massive activations to sink damping on Llama-3-8B, and a pre-registered cross-validated contest in which resolvent features are required for depth-transient persistence and routing identity but no single-operator summary predicts per-head causal criticality.
Significance. If the claims hold within the stated scope, the work supplies a concrete, falsifiable reason to stop treating spectral gaps and eigenvalue decay as default accounts of attention over depth, and it gives an operator-level function for attention sinks (transient dampers) that is measurable, state-conditional, and architecture-conditional. Strengths that raise the bar for the field include: exact structural propositions (mask-trivial Kreiss pinning; Perron deflation identity), systematic mass-matched nulls with FDR control, multi-architecture replication, bit-compatible checkpoint censuses, explicit scoring of the near-tautological Kreiss observable as negative, dual-preprocessing pre-registered contest rules with full negative reporting, and a causal intervention rather than pure correlation. The companion division of labor with the static QK operator is also a useful organizational contribution. The practical readings on outlier-aware compression and sink retention in long-context inference are proportionate to the evidence.
major comments (3)
- [§4, Table 2, Fig. 2] §4 and Table 2: the headline 10^7–10^11 eigenvalue-failure magnitude is measured on head-mean attention-only products Π_L = ∏ Ā^(l)_dev (and a residual-path rollout check). Proposition 2 makes the factorization exact for that skeleton, and the learned-vs-null contrast is clean, but transfer of the reported orders of magnitude to full forward-pass dynamics (with OV maps, MLPs, and residual paths) is load-bearing for the claim that “eigenvalue reasoning about depth fails on trained stacks.” §10 already lists the abstraction as a limitation; the manuscript should either (i) add at least one controlled check that re-inserts OV/MLP structure on a subset of layers/heads and reports how D_L/E_L changes, or (ii) restate the abstract and §4 claim more narrowly as a result about the attention skeleton, with the full-network implication demoted to discussion. Without one of those, the magnitude cl
- [§7, Fig. 5] §7 intervention: the causal chain on Llama-3-8B (three massive dimensions → BOS sink collapse 0.675 o0.049 → K_dev 1.41 o7.44) is the only causal evidence for the damper mechanism, yet it rests on eight sequences and whole-stream channel clamps that the text itself notes are stronger than pointwise edits. Random-dimension controls are appropriate, but the sample is too small to support the architecture-conditional mechanism claim at the same weight as the multi-model census. Either enlarge the intervention set and report sequence-level variability, or qualify the causal claim as a pilot demonstration whose architecture contrast (RMSNorm/RoPE vs LayerNorm) remains correlational outside Llama.
- [§8, Table 3] §8 contest, Table 3: support is declared only when R²(P)>R²(E) (or the union increment) holds under both raw-standardized and rank-normal preprocessings. That dual rule is strict and well-motivated, but preprocessing sensitivity is material: previous-token identity on Pythia fails under raw features and is recorded negative, while GPT-2 margins differ sharply between preprocessings (ΔR² +0.128 vs +0.566). For the central “resolvent features are required” verdict on persistence (three of three models), report the actual R² pairs and confidence intervals for all three models under both preprocessings in the main text or a table, not only the Llama head-to-head numbers. Without those numbers the pre-registered win is hard to audit at the level the protocol promises.
minor comments (7)
- [Abstract, §1] Abstract and §1: “err by seven to eleven orders of magnitude” should be explicitly scoped to head-mean attention products so that abstract readers do not infer a full-network result before §4 and §10.
- [Table 1] Table 1: the Llama induction correlation (+0.02, n.s.) is correctly flagged; consider adding a one-line note that induction excess is state-conditional (§5) so the dormant-census entry is not misread as a null result on induction structure.
- [Fig. 1] Fig. 1 right panel: the dual y-axis (median K_dev and BOS mass) is dense; a small annotation of the layer where real K_dev peels from the null would help.
- [§6] §6: the 160m “formal miss” on pattern-precedes-function (79% vs 90% bar) is handled honestly; move the absolute-onset secondary readout fully into a short appendix table so the main text stays on the pre-registered half-final criterion.
- [Appendix D] Appendix D / Fig. 8: clarify that annotated K_dev values are single-matrix illustrations, not the census medians, already in the caption but worth a sentence in the main text when the figure is first referenced.
- [References] References: Fernando & Guitchounts (2026) is marked “PREPRINT, VERIFY”; either verify the citation or drop the flag before camera-ready.
- [§2] Notation: A_dev, K_dev, and ρ_ε excess are introduced cleanly in §2; a short symbol table would still help readers jumping to §3–§4.
Circularity Check
No significant circularity: mask-pinned Kreiss and exact deflation are standard operator facts; the sole near-tautology is explicitly flagged and scored negative; companion citations supply parallel annotations and weight-space trajectories, not definitional premises for the A-side claims.
specific steps
-
self citation load bearing
[§3 Table 1 caption; §6; §8 feature families; §10]
"Head-class scores (previous-token, induction) are behavioral and come from the companion paper’s probe taxonomy. ... the companion paper’s population Dhead trajectory crosses its own half-depth at the same checkpoint ... Family E contains ... and the companion paper’s weight-space eigenvalue features. Family P contains ... and the weight-space non-normality scalars."
Behavioral targets and some weight-space covariates are imported from the same-author companion rather than measured independently inside this study. The step is only weakly circular: the labels are used as prediction targets (not as definitional inputs that force the resolvent excesses), the companion analyzes a different operator (M), and the A-side signed census, depth-product failure, and contest wins for persistence/routing stand without them. Scored as minor non-load-bearing self-citation.
full rationale
The derivation chain is self-contained. Propositions 1–2 follow immediately from the causal row-stochastic structure (lower-triangular, mask-forced Perron pair P=1e0⊤, ∥P∥2=√n) and are verified numerically; they are not fitted to behavioral data. All subsequent claims are excesses over mass-matched row-permutation (and Dirichlet) nulls, measured transients, depth-product comparisons (DL vs EL), checkpoint trajectories, a clamping intervention, and a pre-registered cross-validated contest between feature families. The one place where a resolvent quantity is related to an observable by theorem rather than by transformer-specific fact (supk∥Akdev∥2≈Kdev) is flagged in §8, excluded from positive credit, and recorded as formally negative. Companion-paper citations supply head-class labels, weight-space Dhead trajectories, and a constrained-training pilot; these are parallel analyses of the static scoring operator M and external annotations, not uniqueness theorems or load-bearing premises that force the A-side resolvent results. No fitted parameter is renamed a prediction, no ansatz is smuggled via self-citation, and no known empirical pattern is merely re-coordinatized. The paper therefore scores at most a minor self-citation that is not load-bearing for the central subdomain claim.
Axiom & Free-Parameter Ledger
free parameters (5)
- sequence length n =
128
- sequences per head/state =
8 to 32
- BH-FDR q =
0.001
- pseudospectral ε grid =
1e-1, 1e-2, 1e-3
- ridge contest regularization and preprocessing
axioms (5)
- standard math For non-normal operators, finite-horizon growth is controlled by resolvent quantities (ε-pseudospectrum, Kreiss constant) with K ≤ sup‖A^k‖ ≤ e n K.
- domain assumption Causal post-softmax attention is lower-triangular and row-stochastic, forcing the Perron pair (1, e0) and P = 1 e0^T.
- domain assumption Row permutation within causal support preserves mass profile while destroying learned column/diagonal structure, so excess over this null measures learning.
- ad hoc to paper Head-mean attention products (without OV/MLP) are a sufficient token-mixing skeleton for depth-transient claims.
- domain assumption Behavioral head-class scores (previous-token, induction) from the companion probe taxonomy correctly label function for correlation and contest targets.
invented entities (2)
-
transient reserve (Kreiss excess of A_dev)
independent evidence
-
sink dampers
independent evidence
read the original abstract
The attention matrix of a causal transformer is row-stochastic, iterated over depth, and non-normal by construction. For non-normal operators, eigenvalues control only asymptotic behavior; finite-depth behavior is controlled by resolvent quantities such as pseudospectra and Kreiss constants. We test, under pre-registered criteria, whether this resolvent view predicts anything about trained transformers that eigenvalues miss. Two structural facts organize the analysis: the mask pins the Kreiss constant of every causal stochastic matrix at $\sqrt{n}$, and deflating the mask-forced Perron projector factorizes the depth deviation dynamics exactly into a product of deflated operators. Across GPT-2, Pythia-410m, and Llama-3-8B, learned non-normality proves to be signed. A routing minority carries excess transient reserve that tracks previous-token function and doubles when induction heads engage, while the sink majority is suppressed below matched shuffle nulls, so that attention sinks act as transient dampers. On depth products, eigenvalue predictions of surviving deviations err by seven to eleven orders of magnitude, an error absent in matched nulls. Checkpoint censuses date this organization to a consolidation phase after circuit formation, and a clamping intervention on Llama-3-8B establishes a causal chain from three massive activation dimensions through sink attention to transient damping; LayerNorm models implement the same functions elsewhere. A cross-validated contest concludes that resolvent features are required for depth-transient persistence and routing-head identity, and that no single-operator summary of any kind predicts per-head causal criticality.
Figures
Reference graph
Works this paper leans on
-
[1]
Quantifying attention flow in transformers
Samira Abnar and Willem Zuidema. Quantifying attention flow in transformers. InACL, 2020. arXiv:2005.00928
Pith/arXiv arXiv 2020
-
[2]
Pseudospectral bounds for transient amplification in coupled gradient descent, 2026
Ahanaf Hasan Ariq. Pseudospectral bounds for transient amplification in coupled gradient descent, 2026. arXiv:2606.04031; HiLD workshop, ICML 2026
Pith/arXiv arXiv 2026
-
[3]
Attention is not all you need: pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: pure attention loses rank doubly exponentially with depth. InICML, 2021. arXiv:2103.03404
Pith/arXiv arXiv 2021
-
[4]
Dynamics of the transformer residual stream: Coupling spectral geometry to network topology, 2026
Jesseba Fernando and Grigori Guitchounts. Dynamics of the transformer residual stream: Coupling spectral geometry to network topology, 2026. arXiv:2605.14258 — PREPRINT, VERIFY
Pith/arXiv arXiv 2026
-
[5]
The Pile: an 800GB dataset of diverse text for language modeling, 2020
Leo Gao et al. The Pile: an 800GB dataset of diverse text for language modeling, 2020. arXiv:2101.00027
Pith/arXiv arXiv 2020
-
[6]
The emergence of clusters in self-attention dynamics
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. The emergence of clusters in self-attention dynamics. InNeurIPS, 2023. arXiv:2305.05465
Pith/arXiv arXiv 2023
-
[7]
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. A mathematical perspective on transformers.Bulletin of the American Mathematical Society, 62:427–479, 2025. arXiv:2312.10794. 13
Pith/arXiv arXiv 2025
-
[8]
Non-normal spectral signatures of instability in neural network training dynam- ics, 2026
Souvik Ghosh. Non-normal spectral signatures of instability in neural network training dynam- ics, 2026. arXiv:2605.23476
Pith/arXiv arXiv 2026
-
[9]
Mark S. Goldman. Memory without feedback in a neural network.Neuron, 61(4):621–634, 2009
2009
-
[10]
When attention sink emerges in language models: an empirical view
Xiangming Gu, Tianyu Pang, Chao Du, Qian Liu, Fengzhuo Zhang, Cunxiao Du, Ye Wang, and Min Lin. When attention sink emerges in language models: an empirical view. InICLR,
-
[11]
Jordan, and Song Mei
Tianyu Guo, Druv Pai, Yu Bai, Jiantao Jiao, Michael I. Jordan, and Song Mei. Active- dormant attention heads: mechanistically demystifying extreme-token phenomena in LLMs,
-
[12]
Non-normalrecurrentneuralnetwork(nnRNN):learning long time dependencies while improving expressivity with transient dynamics
Giancarlo Kerg, Kyle Goyette, Maximilian Puelma Touzel, Gauthier Gidel, Eugene Vorontsov, YoshuaBengio, andGuillaumeLajoie. Non-normalrecurrentneuralnetwork(nnRNN):learning long time dependencies while improving expressivity with transient dynamics. InNeurIPS,
-
[13]
Hengyu Li. When the complex spectrum of attention does work: architecture-conditional non-Hermitian structure in the QK operator, 2026. Companion paper, arXiv:2607.06621
Pith/arXiv arXiv 2026
-
[14]
Computing the Kreiss constant of a matrix, 2020
Tim Mitchell. Computing the Kreiss constant of a matrix, 2020. arXiv:1907.06537
Pith/arXiv arXiv 2020
-
[15]
Murphy and Kenneth D
Brendan K. Murphy and Kenneth D. Miller. Balanced amplification: a new mechanism of selective amplification of neural activity patterns.Neuron, 61(4):635–648, 2009
2009
-
[16]
Mind the gap: a spectral analysis of rank collapse and signal propagation in attention layers
Thiziri Nait Saada, Alireza Naderi, and Jared Tanner. Mind the gap: a spectral analysis of rank collapse and signal propagation in attention layers. InICML, 2025. arXiv:2410.07799
Pith/arXiv arXiv 2025
-
[17]
Attention sinks and compression valleys in LLMs are two sides of the same coin, 2025
Enrique Queipo-de Llano, Álvaro Arroyo, Federico Barbero, Xiaowen Dong, Michael Bronstein, Yann LeCun, and Ravid Shwartz-Ziv. Attention sinks and compression valleys in LLMs are two sides of the same coin, 2025. arXiv:2510.06477
arXiv 2025
-
[18]
Mingjie Sun, Xinlei Chen, J. Zico Kolter, and Zhuang Liu. Massive activations in large language models. InCOLM, 2024. arXiv:2402.17762
Pith/arXiv arXiv 2024
-
[19]
Trefethen
Lloyd N. Trefethen. Computation of pseudospectra.Acta Numerica, 8:247–295, 1999
1999
-
[20]
Trefethen and Mark Embree.Spectra and Pseudospectra: The Behavior of Nonnormal Matrices and Operators
Lloyd N. Trefethen and Mark Embree.Spectra and Pseudospectra: The Behavior of Nonnormal Matrices and Operators. Princeton University Press, 2005
2005
-
[21]
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. InICLR, 2024. arXiv:2309.17453. A Validation of theσ min core The grid core reproduces the Grcar pseudospectrum (Fig. 6) and the Jordan-blockε1/n law (ρε = 0.253againstε 1/n = 0.251atn=10,ε= 10 −6). It agrees with dense SVD to10 −...
Pith/arXiv arXiv 2024
-
[22]
and writes it nearly uniformly across downstream positions (|u| ≈n−1/2). C The 160M real-corpus implementation-invariance pilot Six Pythia-160M-architecture models (12 layers×12 heads,d=768, rotary fraction0.25) were trained from scratch on 4B tokens of real corpus under the companion paper’s constrained-training protocol [13]: two free seeds; two seeds w...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.