REVIEW 4 minor 2 cited by
The geometric effort of reaching a useful preconditioner is a path-space value whose responses are generated by one Green–Jacobi inverse.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 15:49 UTC pith:PXLVNNNO
load-bearing objection Solid path-space theory for restricted preconditioner metrics: global Green–Jacobi and bordered hard-target response under Hadamard/convexity hypotheses, with clean closed-form checks.
Restricted Dynamic Geometric Complexity: Path-Space Reduction and M\"obius--Jacobi Response
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Restricted dynamic geometric complexity is a genuine path-space value: its global elimination law is Bellman composition of affine-invariant lengths, its smooth local response is the Schur complement of a coercive Jacobi operator, and its finite intervention effects are exact integrals of one Green or bordered-Green kernel. On a Hadamard manifold with geodesically convex potentials the Green inverse is global on a whole intervention cube; on a regular active spectral stratum the bordered inverse differentiates the moving hard condition target and explains why hard interactions need not share the unconstrained sign.
What carries the argument
The Green–Jacobi inverse of the path second variation (and its bordered Jacobi–KKT counterpart for hard targets). One uniformly coercive operator solves bulk and terminal forcing, yields the value Hessian as a negative Gram, and by recursion generates every finite-order Möbius response; the bordered version jointly differentiates path, moving projection endpoint, and multiplier.
Load-bearing premise
All bulk and terminal potentials must stay geodesically convex on a whole neighborhood of the intervention cube, and the state space must be Hadamard (nonpositive curvature), so that a single coercive Green inverse exists cube-wide; hard-target claims further need a fixed smooth active spectral stratum with simple extreme eigenvalues.
What would settle it
In the two-dimensional determinant-one diagonal model, compute the interaction matrix of the forced kinetic action both from the closed-form second derivatives and by finite differences of the value; if the matrix is not strictly negative definite with the stated determinant 1/2160, or if the sequential three-dimensional protocol fails to exceed ambient AIRM distance by the factor 2/√3, the explicit claims collapse.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper defines restricted dynamic geometric complexity (RDGC) as the least affine-invariant length of an admissible metric path whose endpoint meets a Hessian-relative generalized-eigenvalue condition target. Path elimination yields an exact min-plus semigroup and Bellman principle; fixed-horizon kinetic energy equals squared complexity over twice the horizon. On a Hadamard state space with geodesically convex bulk and terminal potentials, a single uniformly coercive Green–Jacobi inverse produces the value Hessian, force-to-curvature bounds, exact Möbius pair/conditional effects, and arbitrary finite-order responses (Theorem 4.11, Corollary 4.12). For hard condition targets on a regular active spectral stratum, a bordered Jacobi–KKT theorem differentiates the moving projection, multiplier, and constrained Hessian, allowing interaction signs that need not match the unconstrained Gram law (Theorem 4.13). The theory specializes to AIRM geometry, with closed-form diagonal models (Theorems 5.1–5.4) and a sequential protocol whose RDGC strictly exceeds ambient projection distance (Proposition 5.3). Appendices supply detailed proofs; Section 6 reports deterministic numerical verification of discrete identities.
Significance. If the results hold under the stated hypotheses, the paper supplies a coherent path-space foundation for measuring geometric effort of structured preconditioners, together with explicit differential response laws (Green and bordered) that convert path elimination into Möbius interaction formulas. Strengths include the clean separation of algebraic (min-plus), variational (Tonelli/HJB), and differential (Jacobi–Schur) layers; the global Hadamard derivation of cube-wide uniqueness and coercivity; closed-form diagonal realizations that instantiate both unconstrained and hard-target responses with both interaction signs; and the sequential gap showing that the declared path class can change the value. The appendices give full proofs, and Section 6 provides reproducible deterministic checks of discrete algebraic identities. The work is self-contained at the path-space level and correctly scopes its classical differentiability to smooth active strata.
minor comments (4)
- The companion papers [29,30] are cited as arXiv preprints with 2026 dates; ensure final bibliographic entries are complete and that the self-contained claim for the path-space layer is preserved if those works remain unpublished.
- Notation for the relative spectrum and κ_gen is introduced carefully, but occasional shorthand κ(G^{-1}H) could still be flagged once more in the introduction to prevent misreading as the Euclidean condition number of a nonnormal product.
- Table 1 and the verification script are useful; a one-sentence note that the coordinate eigenvalue is mesh-scaled (already present) could be repeated in the table caption for readers who skip the surrounding text.
- A few long sentences in the abstract and contributions list could be split for readability without changing content.
Circularity Check
No significant circularity: path-elimination, length–energy, and global Green/bordered responses are derived under stated hypotheses; companion [30] is infrastructure only and is re-proved in path-Hilbert form.
specific steps
-
self citation load bearing
[Introduction Contributions / §4.4–4.7 (Thm 4.9, Cor 4.15); companion [30]]
"The finite-dimensional companion [30] proves generic Schur response and affine-intervention Gram formulas for hidden controller states. Those formulas are the finite-dimensional template here. The present paper’s new response results begin with their path-Hilbert realization… The local formulas in Theorem 4.9 and Corollary 4.15 are their Hilbert-path lift and are used here as infrastructure."
Same-author arXiv [30] is cited as the finite-dimensional Schur/Gram template for the local path-response formulas. This is not load-bearing for the paper’s central claims: Thm 4.9 is proved independently via the Hilbert IFT and coercive second variation, Cor 4.15 is derived from it, and the global cube-wide Green and bordered hard-target theorems are new under Hadamard/KKT hypotheses with full proofs. Mild self-citation only; does not force the main results by construction.
full rationale
The paper’s derivation chain is self-contained. RDGC is defined as an infimum of affine-invariant path length to a spectral target; the min-plus semigroup (Thm 4.2) follows from path concatenation and action additivity alone; the length–energy identity (Thm 4.4) is Cauchy–Schwarz plus reparameterization, not a smuggled prediction. The global Möbius–Jacobi theorem (Thm 4.11) and bordered hard-target theorem (Thm 4.13) are proved from geodesic convexity (MJ2), Hadamard nonpositive curvature, and KKT regularity on a fixed active stratum, with full appendix proofs (implicit-function, Lax–Milgram, envelope, Boolean integration). Closed-form diagonal models (Thms 5.1–5.2, 5.4; Prop 5.3) instantiate those formulas by direct calculation without fitted parameters. Companion [30] supplies the finite-dimensional Schur/Gram template; the present paper re-proves the path-Hilbert lift (Thm 4.9, Cor 4.15) and explicitly separates infrastructure from the new Hadamard/cube-wide and bordered results. No self-definitional loop, no fitted-input-as-prediction, and no uniqueness imported solely from an unverified self-citation. Score 1 only for the minor same-author template citation, which is not load-bearing for the central claims.
Axiom & Free-Parameter Ledger
axioms (6)
- domain assumption State space is a finite-dimensional Hadamard manifold (nonpositive curvature, complete, simply connected), so geodesics and projections behave well and curvature contributes nonnegatively to the index form.
- domain assumption Bulk and terminal potentials are geodesically convex (Hess W_u ⪰ 0, Hess Φ_u ⪰ 0) with C^{r+2} regularity and uniform base-point bounds on a neighborhood of the intervention cube (MJ2).
- standard math Admissible paths form a composable system with additive action (Assumption 4.1), enabling exact min-plus path elimination without smoothness.
- domain assumption Fixed-horizon path class is closed under absolutely continuous reparameterization so length and kinetic energy share minimizers (Theorem 4.4).
- domain assumption For hard targets, the active condition-number stratum has simple extreme eigenvalues, LICQ, positive multiplier, and bordered invertibility on the face of interest.
- standard math Structured families used in corollaries (diagonal, fixed block-diagonal, det-one diagonal) are complete simply connected totally geodesic AIRM submanifolds.
invented entities (2)
-
Restricted dynamic geometric complexity (RDGC) D_{K,F}(S_0;H)
no independent evidence
-
Intervention-cube Green–Jacobi / bordered Jacobi–KKT response operators for metric paths
no independent evidence
read the original abstract
Structured preconditioners restrict optimization to a small family of positive metrics, but endpoint condition-number reachability does not measure the geometric effort required to reach a useful metric. We formulate this effort as a path-space value problem. Restricted dynamic geometric complexity is the least affine-invariant length of an admissible metric path whose endpoint reaches a Hessian-relative generalized-eigenvalue condition target. Path elimination gives an exact min-plus semigroup and Bellman principle, while fixed-horizon kinetic energy is exactly squared complexity divided by twice the horizon. The main response result is global on a Hadamard state space: geodesic convexity produces a smooth intervention-cube path branch and a uniformly coercive Jacobi form, while one Green inverse generates the value Hessian, two-sided force-to-curvature bounds, exact M\"obius effects, and arbitrary prescribed finite-order responses. For the hard condition target, a bordered Jacobi--KKT theorem differentiates the moving projection endpoint and multiplier on every regular active spectral stratum; its indefinite inverse also explains why hard-target interactions need not share the unconstrained sign. The theory specializes to affine-invariant positive-definite geometry. A determinant-one two-dimensional diagonal model has an exact target interval, a closed-form forced path, and a strictly negative-definite interaction matrix. A moving diagonal Hessian gives a closed-form hard-target projection, multiplier, and pair effects of either sign, while a coordinate-sequential three-dimensional protocol yields an exact path metric strictly larger than the ambient projection distance. Thus the global Green and bordered hard-target responses are explicit laws of restricted metric-path elimination built on Bellman composition.
Forward citations
Cited by 2 Pith papers
-
Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions
Under fixed innovation coupling, finite-horizon optimizers admit minimal pathwise realizations and incidence-identifiable Möbius effects, with a five-term readout transfer from hidden relaxation and a closed reduced-v...
-
Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions
A modular calculus decomposes optimizer updates into geometric preconditioning plus structured nongeometric mechanisms, with a direction-expressivity theorem showing full SPD geometry captures exactly strict descent d...
Reference graph
Works this paper leans on
-
[1]
A. A. Agrachev and Y. L. Sachkov,Control Theory from the Geometric Viewpoint, vol. 87 of Encyclopaedia of Mathematical Sciences, Springer, Berlin, 2004, https://doi.org/10.1007/ 978-3-662-06404-7
2004
-
[2]
Amari,Natural gradient works efficiently in learning, Neural Computation, 10 (1998), pp
S.-i. Amari,Natural gradient works efficiently in learning, Neural Computation, 10 (1998), pp. 251–276, https://doi.org/10.1162/089976698300017746
-
[3]
B. Amos and J. Z. Kolter,OptNet: Differentiable optimization as a layer in neural networks, in Proceedings of the 34th International Conference on Machine Learning, vol. 70 of Proceedings of Machine Learning Research, PMLR, 2017, pp. 136–145, https://arxiv.org/abs/1703.00443
Pith/arXiv arXiv 2017
-
[4]
R. Anil, V. Gupta, T. Koren, K. Regan, and Y. Singer,Scalable second order optimization for deep learning, arXiv preprint arXiv:2002.09018, (2020), https://arxiv.org/abs/2002. 09018
Pith/arXiv arXiv 2002
-
[5]
M. Baˇcak,Convex Analysis and Optimization in Hadamard Spaces, De Gruyter, 2014, https: //doi.org/10.1515/9783110361629
-
[6]
A. Beck and M. Teboulle,Mirror descent and nonlinear projected subgradient methods for convex optimization, Operations Research Letters, 31 (2003), pp. 167–175, https: //doi.org/10.1016/S0167-6377(02)00231-6. [7]R. Bhatia,Positive Definite Matrices, Princeton University Press, 2007
-
[7]
M. Blondel, Q. Berthet, M. Cuturi, R. Frostig, S. Hoyer, F. Llinares-L ´opez, F. Pe- dregosa, and J.-P. Vert,Efficient and modular implicit differentiation, in Advances in Neural Information Processing Systems, vol. 35, 2022, https://arxiv.org/abs/2105.15183
Pith/arXiv arXiv 2022
-
[8]
J. F. Bonnans and A. Shapiro,Perturbation Analysis of Optimization Problems, Springer New York, 2000, https://doi.org/10.1007/978-1-4612-1394-9
-
[9]
Dacorogna,Direct Methods in the Calculus of Variations, vol
B. Dacorogna,Direct Methods in the Calculus of Variations, vol. 78 of Applied Mathe- matical Sciences, Springer New York, New York, 2 ed., 2008, https://doi.org/10.1007/ 978-0-387-55249-1
2008
-
[10]
T. Dasgupta, N. S. Pillai, and D. B. Rubin,Causal inference from2 K factorial designs 20Z. LI by using potential outcomes, Journal of the Royal Statistical Society: Series B (Statistical Methodology), 77 (2015), pp. 727–753, https://doi.org/10.1111/rssb.12085
-
[11]
K. Dhamdhere, M. Sundararajan, and A. Agarwal,The Shapley–Taylor interaction index, in Proceedings of the 37th International Conference on Machine Learning, vol. 119 of Proceedings of Machine Learning Research, PMLR, 2020, pp. 2709–2718, https://arxiv.org/ abs/1902.05622
Pith/arXiv arXiv 2020
-
[12]
S. Diamond and S. Boyd,Stochastic matrix-free equilibration, Journal of Optimization Theory and Applications, 172 (2016), pp. 436–454, https://doi.org/10.1007/s10957-016-0990-2
-
[13]
M. L. Do˘gan, A. Erg¨ur, and E. Tsigaridas,Optimal preconditioning is a geodesically convex optimization problem, 2025, https://arxiv.org/abs/2512.06618, https://arxiv.org/abs/2512. 06618
arXiv 2025
-
[14]
A. L. Dontchev and R. T. Rockafellar,Implicit Functions and Solution Mappings, Springer Series in Operations Research and Financial Engineering, Springer New York, 2014, https: //doi.org/10.1007/978-1-4939-1037-3
-
[15]
Duchi, E
J. Duchi, E. Hazan, and Y. Singer,Adaptive subgradient methods for online learning and stochastic optimization, Journal of Machine Learning Research, 12 (2011), pp. 2121–2159, https://jmlr.org/papers/v12/duchi11a.html
2011
-
[16]
W. H. Fleming and H. M. Soner,Controlled Markov Processes and Viscosity Solutions, vol. 25 of Stochastic Modelling and Applied Probability, Springer, New York, 2 ed., 2006, https://doi.org/10.1007/0-387-31071-1
-
[17]
L. Franceschi, P. Frasconi, S. Salzo, R. Grazzi, and M. Pontil,Bilevel programming for hyperparameter optimization and meta-learning, in Proceedings of the 35th International Conference on Machine Learning, vol. 80 of Proceedings of Machine Learning Research, PMLR, 2018, pp. 1568–1577, https://arxiv.org/abs/1806.04910
Pith/arXiv arXiv 2018
-
[18]
G. H. Golub and V. Pereyra,The differentiation of pseudo-inverses and nonlinear least squares problems whose variables separate, SIAM Journal on Numerical Analysis, 10 (1973), pp. 413–432, https://doi.org/10.1137/0710036
doi:10.1137/0710036 1973
-
[19]
R. Grosse and J. Martens,A Kronecker-factored approximate Fisher matrix for convolution layers, in Proceedings of the 33rd International Conference on Machine Learning, vol. 48 of Proceedings of Machine Learning Research, PMLR, 2016, pp. 573–582, https://proceedings. mlr.press/v48/grosse16.html, https://arxiv.org/abs/1602.01407
Pith/arXiv arXiv 2016
-
[20]
Gupta, T
V. Gupta, T. Koren, and Y. Singer,Shampoo: Preconditioned stochastic tensor opti- mization, in Proceedings of the 35th International Conference on Machine Learning, vol. 80 of Proceedings of Machine Learning Research, PMLR, 2018, pp. 1842–1850, https://proceedings.mlr.press/v80/gupta18a.html
2018
-
[21]
N. J. Higham,Functions of Matrices: Theory and Computation, Society for Industrial and Applied Mathematics, 2008, https://doi.org/10.1137/1.9780898717778
-
[22]
J. D. Janizek, P. Sturmfels, and S.-I. Lee,Explaining explanations: Axiomatic feature interactions for deep networks, Journal of Machine Learning Research, 22 (2021), pp. 1–54, https://arxiv.org/abs/2002.04138
Pith/arXiv arXiv 2021
-
[23]
D. P. Kingma and J. Ba,Adam: A method for stochastic optimization, in International Confer- ence on Learning Representations, 2015, https://arxiv.org/abs/1412.6980. arXiv:1412.6980
Pith/arXiv arXiv 2015
-
[24]
Krichene, A
W. Krichene, A. M. Bayen, and P. L. Bartlett,Accelerated mirror de- scent in continuous and discrete time, in Advances in Neural Information Processing Systems, 2015, https://proceedings.neurips.cc/paper/2015/hash/ f60bb6bb4c96d4df93c51bd69dcc15a0-Abstract.html
2015
-
[25]
A. S. Lewis,Derivatives of spectral functions, Mathematics of Operations Research, 21 (1996), pp. 576–588, https://doi.org/10.1287/moor.21.3.576
-
[26]
A. S. Lewis,The mathematics of eigenvalue optimization, Mathematical Programming, 97 (2003), pp. 155–176, https://doi.org/10.1007/s10107-003-0441-3
-
[27]
A. S. Lewis and H. S. Sendov,Twice differentiable spectral functions, SIAM Journal on Matrix Analysis and Applications, 23 (2001), pp. 368–386, https://doi.org/10.1137/ S089547980036838X
2001
-
[28]
Li,Information-Induced Training Geometry: Exact Reduction, Canonical Completion, and Structured Expressivity, 2026
Z. Li,Information-Induced Training Geometry: Exact Reduction, Canonical Completion, and Structured Expressivity, 2026. arXiv preprint
2026
-
[29]
Z. Li,Optimization Geometrodynamics: Variational Reduction and Interaction Curvature, 2026, https://arxiv.org/abs/2607.06723, https://arxiv.org/abs/2607.06723. arXiv:2607.06723
Pith/arXiv arXiv 2026
-
[30]
Lim,Best approximation in riemannian geodesic submanifolds of positive definite matri- ces, Canadian Journal of Mathematics, 56 (2004), pp
Y. Lim,Best approximation in riemannian geodesic submanifolds of positive definite matri- ces, Canadian Journal of Mathematics, 56 (2004), pp. 776–793, https://doi.org/10.4153/ CJM-2004-035-5
2004
-
[31]
Lu and T
Z. Lu and T. K. Pong,Minimizing condition number via convex programming, SIAM Journal on Matrix Analysis and Applications, 32 (2011), pp. 1193–1211, https://doi.org/10.1137/ RESTRICTED DYNAMIC GEOMETRIC COMPLEXITY21 100795097
2011
-
[32]
P. Mar´echal and J. J. Ye,Optimizing condition numbers, SIAM Journal on Optimization, 20 (2009), pp. 935–947, https://doi.org/10.1137/080740544
-
[33]
Martens and R
J. Martens and R. Grosse,Optimizing neural networks with Kronecker-factored approximate curvature, in Proceedings of the 32nd International Conference on Machine Learning, vol. 37 of Proceedings of Machine Learning Research, PMLR, 2015, pp. 2408–2417, https: //proceedings.mlr.press/v37/martens15.html
2015
-
[34]
Pennec, P
X. Pennec, P. Fillard, and N. Ayache,A riemannian framework for tensor computing, International Journal of Computer Vision, 66 (2006), pp. 41–66, https://doi.org/10.1007/ s11263-005-3222-z
2006
-
[35]
F. Pukelsheim,Optimal Design of Experiments, Society for Industrial and Applied Mathematics, Philadelphia, 2006, https://doi.org/10.1137/1.9780898719109
-
[36]
Z. Qu, W. Gao, O. Hinder, Y. Ye, and Z. Zhou,Optimal diagonal preconditioning, Operations Research, 73 (2025), pp. 1479–1495, https://doi.org/10.1287/opre.2022.0592
-
[37]
Shazeer and M
N. Shazeer and M. Stern,Adafactor: Adaptive learning rates with sublinear memory cost, in Proceedings of the 35th International Conference on Machine Learning, vol. 80 of Proceedings of Machine Learning Research, PMLR, 2018, pp. 4596–4604, https://proceedings.mlr.press/ v80/shazeer18a.html
2018
-
[38]
Sra and R
S. Sra and R. Hosseini,Conic geometric optimization on the manifold of positive definite matrices, SIAM Journal on Optimization, 25 (2015), pp. 713–739, https://doi.org/10.1137/ 140978168
2015
-
[39]
W. Su, S. Boyd, and E. J. Cand `es,A differential equation for modeling Nesterov’s accelerated gradient method: Theory and insights, Journal of Machine Learning Research, 17 (2016), pp. 1–43, https://arxiv.org/abs/1503.01243
Pith/arXiv arXiv 2016
-
[40]
Tanaka and K
M. Tanaka and K. Nakata,Positive definite matrix approximation with condition num- ber constraint, Optimization Letters, 8 (2014), pp. 939–947, https://doi.org/10.1007/ s11590-013-0632-7. Published online in 2013
2014
-
[41]
van der Sluis,Condition numbers and equilibration of matrices, Numerische Mathematik, 14 (1969), pp
A. van der Sluis,Condition numbers and equilibration of matrices, Numerische Mathematik, 14 (1969), pp. 14–23, https://doi.org/10.1007/BF02165096
-
[42]
A. Wibisono, A. C. Wilson, and M. I. Jordan,A variational perspective on accelerated methods in optimization, Proceedings of the National Academy of Sciences, 113 (2016), pp. E7351–E7358, https://arxiv.org/abs/1603.04245. Appendix A. Proofs for Path-Space Reduction. Proof of Theorem4.2. Fix an admissible path γ∈P s,t(x, z) and put y = γ(r). Restriction an...
Pith/arXiv arXiv 2016
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.