Pith. sign in

REVIEW 2 major objections 5 minor 174 references

Optimizer interventions have a complete causal calculus: minimal pathwise state, Möbius pure effects, and a five-term law that transfers hidden relaxation to any smooth readout.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 15:49 UTC pith:GBXMSZ43

load-bearing objection Solid experimental outer layer for optimizer interventions: clean pathwise/Möbius gauge, a correct five-term readout transfer, and an unusually tight reduced-value factorial closure—general nonconvex readout still untested. the 2 major comments →

arxiv 2607.07206 v2 pith:GBXMSZ43 submitted 2026-07-08 cs.LG math.OC

Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions

classification cs.LG math.OC
keywords optimizer calculuscausal interventionsMöbius effectshidden relaxationreadout transferfactorial designresponse classesidentifiability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Optimizer experiments see only responses to configurations, not the hidden mechanisms that produced them. This paper builds a causal interaction calculus that separates three layers: a pathwise minimal state under a fixed innovation coupling, a Möbius decomposition of pure interventional effects over lower-finite posets, and a linear incidence operator that is the complete observational gauge for any finite effect support and design. That gauge yields exact identifiability, sharp quotient stability, held-out predictions, and noiseless configuration complexity equal to the support size. Taking smooth hidden-state relaxation as a structural input, the paper proves that any smooth update or trace readout inherits an explicit five-term interaction through the first and second hidden responses; Boolean effects remain exact integrals of that continuous curvature and are therefore accessible to factorial designs, even though general readouts need not inherit the negative-semidefinite sign of the reduced optimal value. Gaussian quotient minimax risk, exact confidence sets, residualized sector tests, misspecification decomposition, certified decisions, and optimal replication complete the statistical layer. A controlled strongly convex logistic experiment closes the reduced-value chain to high numerical precision; neural audits show that declared geometric response classes remain informative in nonconvex training.

Core claim

Under a fixed innovation coupling every finite-horizon innovation-driven optimizer admits a behaviorally minimal pathwise realization; for any finite effect support and design the incidence operator is the complete observational gauge with exact identifiability and noiseless complexity equal to support size; and any smooth readout FR(u)=R(u,σ⋆(u)) inherits an explicit five-term mixed derivative through first and second hidden responses whose Boolean pair effect is the exact integral of that curvature, with no universal sign beyond the reduced-value case.

What carries the argument

The incidence (design) operator XD,K together with the five-term observable-readout transfer: ∂²uiuj FR = Ruiuj + Ruiσ vj + Rσ uj vi + Rσσ[vi,vj] + Rσ wij. The operator is the complete observational gauge; the transfer carries hidden Schur curvature to experimentally chosen updates and traces so that Boolean effects remain integrable and identifiable.

Load-bearing premise

Hidden relaxation must have a unique interior minimizer with positive-definite vertical Hessian, and the experiment must fix the innovation coupling so that pathwise predictive equivalence is a well-defined congruence.

What would settle it

A factorial experiment on a smooth optimizer whose Boolean pair effects disagree with independently integrated five-term curvature of a declared readout, or a finite support and design whose incidence matrix has full column rank yet fails to recover held-out configuration responses within noise.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper develops a causal calculus for optimizer experiments that separates pathwise realization under fixed innovation coupling, Möbius pure-effect coordinates on lower-finite intervention posets, and incidence-based observational identification. Theorem 1 gives a behaviorally minimal pathwise/unifilar state, unique pure effects, the complete design gauge, exact identifiability when rank equals support size, sharp quotient stability, held-out falsification, and noiseless configuration complexity |K|. Taking the companion P1 Schur law as input, Theorem 3 transfers first and second hidden responses to an arbitrary smooth readout FR(u)=R(u,σ⋆(u)), yielding an explicit five-term mixed derivative whose Boolean pair effect is the integral of that curvature and need not inherit the reduced-value negative-semidefinite sign. Theorem 4 supplies Gaussian quotient minimax risk, residualized sector tests, misspecification decomposition, certified decisions, and exact continuous A-optimal replication. A controlled 65-dimensional strongly convex logistic experiment closes the reduced-value chain (Boolean vs independent curvature integrals within 4.21e-11; nine held-out continuous intensities within 8.88e-13; Gaussian coverage/power as predicted; real-minibatch rejection of order-two support). Neural trace audits test geometric response-class residuals in nonconvex training.

Significance. If the package holds, the paper supplies a usable experimental outer layer for optimizer mechanism studies: exact incidence gauges and configuration complexity, a transfer law that predicts which observable effects factorial designs can identify, and finite-sample inference/design formulas specialized to intervention masks. Strengths that raise the contribution above a pure restatement of classical incidence algebra and Schur sensitivity include: full proofs in Appendices A–D; an independent Boolean-versus-quadrature closure with machine-auditable artifacts and sub-1e-10 agreement; exact Gaussian risk/coverage Monte Carlo confirmation; and explicit separation of response-class residual tests from causal attribution. The five-term readout law is the main novel interface between hidden relaxation geometry and the traces practitioners actually log.

major comments (2)
  1. The abstract and Contributions package Theorem 1 with Theorem 3 as the experimental interface for 'arbitrary smooth update or trace readouts.' Appendix B.4 derives the five-term law correctly under Assumption 1, but §§8.1–8.5 close only the reduced-value special case F(u)=min_w E(w,u) on a globally strongly convex logistic model. §8.6 and the Discussion explicitly state that the neural audits lack factorial masks and the hidden first/second responses required by Eq. (13). The load-bearing claim that general update/trace readouts inherit identifiable Boolean effects via the same calculus therefore rests on an untested transfer step. Either (i) add a factorial campaign that measures an actual update readout together with the quantities in Eqs. (12)–(14), or (ii) restate the abstract/contributions so that the validated claim is the reduced-value chain plus the analytical transfer theorem, w
  2. Assumption 1 (unique interior minimizer, positive-definite vertical Hessian) and the fixed innovation coupling of §3.1/Theorem 1(i) are load-bearing for both the five-term transfer and the canonical pathwise minimal state. The Discussion notes nonsmooth/multi-valued/hysteretic cases require generalized derivatives, but the manuscript does not quantify how often these fail for standard adaptive optimizers (Adam-style second-moment state, Muon-style operators, clipping). A short scope paragraph or counter-example class would make the boundary of Theorems 1 and 3 precise rather than leaving it as a regularity caveat.
minor comments (5)
  1. Notation for the Hilbert-valued response F(a) versus scalar readouts Fφ and FR is introduced carefully in §3.1 but is easy to lose later; a one-line glossary or consistent subscript convention would help.
  2. Table 1 and Figure 1 report absolute errors of order 1e-11–1e-15; stating the floating-point / solver tolerance used for the Newton solves would make the agreement easier to interpret.
  3. The neural audit tables (Tables 2–6) mix accuracy and TGER residuals; a clearer caption that these are response-class diagnostics, not optimizer rankings, would reduce misreading (the Discussion already says this, but the tables themselves do not).
  4. Cross-references to the companion P1/P2/P4 papers are frequent; a short dependency diagram or table of which structural lemmas are imported versus proved here would improve self-containment for readers who have not read the series.
  5. In §8.3 the Gaussian channel uses σa=0.004(1+0.2|a|); a one-sentence justification for this mild heteroscedasticity model would be useful.

Circularity Check

1 steps flagged

Minor concurrent same-author self-citation for the P1 Schur structural input; Theorem 1, the five-term readout transfer, incidence gauge, and factorial numerics are independently derived or checked, not forced by fit or definition.

specific steps
  1. self citation load bearing [Abstract; §1; §5 Theorem 2 (restated from P1); series position §2]
    "Building on this structural law, we prove an observable-readout transfer theorem... The companion P1 paper already answers the structural reduced-value question... it proves that the reduced intervention Hessian is D²_uu Ē = −G*H⁻¹G, and that Boolean reduced-value contrasts integrate this curvature (Li, 2026b). ... Taking the P1 Schur–Möbius theorem as a structural input, we prove an observable readout interaction law."

    The abstract and framing present inverse-stiffness interaction curvature as an established structural law imported from concurrent same-author P1 (and path-space P2), then build the experimental calculus on top. That is a self-citation chain for the reduced-value sign/Gram story. It is only minor/non-load-bearing here because Appendix B fully reproduces the P1 reduction proof, Theorem 3 itself needs only Assumption 1 + chain rule (not the P1 Gram sign), and the factorial numerics independently recompute both sides of the identity rather than fitting free parameters from P1.

full rationale

The load-bearing novel layers do not reduce to their inputs by construction. Theorem 1 (pathwise Nerode-style quotient, Möbius expansion, incidence kernel, exact noiseless complexity |K|) is proved from first principles in Appendix A under the fixed-innovation protocol; Möbius inversion and rank-nullity are classical, not renamed empirical patterns. Theorem 3 is a local chain-rule / IFT calculation (Appendix B.4) under Assumption 1: FR(u)=R(u,σ⋆(u)) yields the explicit five-term mixed derivative whose Boolean pair effect is the FTC integral; this is not defined in terms of the experimental targets. Theorem 2 (reduced-value Schur Gram −G*H−1G and Boolean integral) is attributed to concurrent same-author P1, which is the only self-citation of note, but Appendix B reproduces the full proof so the present manuscript is mathematically self-contained rather than resting on an unverified external uniqueness claim. The digits factorial closure computes Boolean corner effects and independent Gauss–Legendre curvature integrals on separate paths (“Neither path uses the output of the other”) and reports numerical agreement; that is consistency under a declared strongly convex additive model, not a parameter fitted on a subset and re-labeled as prediction. Held-out continuous intensities, Gaussian coverage/power Monte Carlo, and minibatch bootstrap rejection of order-two support are likewise external checks of the incidence/inference layer. Neural audits are explicitly scoped as response-class sensitivity, not as circular confirmation of Theorem 3. No self-definitional loop, no fitted-input-as-prediction of the central claims, and no uniqueness theorem imported solely to forbid alternatives. Score 2 reflects only the concurrent P1 framing of the structural Schur input, which is not load-bearing for the paper’s own proofs or experiments.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 3 invented entities

Universality rests on protocol assumptions (fixed innovation coupling, lower-finite poset, standard Borel histories) and standard incidence algebra. Smooth interaction rests on unique positive-definite hidden relaxation (Assumption 1) imported structurally from P1. Inference rests on finite-dimensional linear Gaussian experiments with known V≻0. Experiments add hand-chosen regularization, intervention scales, and design budgets that affect numerical validation but not the abstract theorems. Invented entities are mostly definitional coordinate systems (effect support K, response classes MF,B) rather than new physical mediators.

free parameters (5)
  • ridge λ = 0.15
    Logistic energy uses λ=0.15 to enforce strong convexity Eww ⪰ 0.15 I; chosen for the controlled experiment, not derived.
  • intervention loss scale = 0.75
    Additive interventions enter as 0.75 ∑ ui Δi(w); scale is a design choice for the digits cube.
  • Gaussian channel noise model σa=0.004(1+0.2|a|) = 0.004 base, 0.2 slope
    Heteroscedastic noise law used for power calibration and 100k campaigns; pilot-driven experimental parameter.
  • replication budgets (3780 main / 720 held-out) = 3780 / 720
    Scaled from pilot (420/80) by first integer reaching predeclared 0.95 power; allocation affects empirical confirmation, not theorems.
  • bounded-diagonal audit box [1e-6, 1] = [1e-6, 1]
    Response-class family bounds for neural TGER audits; hand-set geometry constraints.
axioms (6)
  • domain assumption Fixed innovation coupling: exogenous randomness is coupled across configurations so pathwise predictive equivalence is well-defined.
    Theorem 1(i) and §3.1; without it only distributional realization is available, which the paper explicitly does not claim.
  • standard math Lower-finite intervention poset with Möbius inversion of expected responses.
    Classical incidence algebra; used in eq. (2) and Theorem 1(ii).
  • domain assumption Assumption 1: Cr+1 energy, local uniform inf-compactness, unique interior minimizer, H=D²σσE ≻ 0.
    Required for Theorems 2–3 and the five-term transfer; stated before Theorem 2.
  • domain assumption Finite-dimensional response with known positive-definite observation covariance for exact Gaussian quotient minimax, χ² sets, and A-optimal replication.
    Theorem 4 setup; paper flags estimated covariance and dependent traces as outside exact claims.
  • domain assumption P1 Schur reduced-value law D²uu Ē = −G* H^{-1} G as structural input.
    Imported from companion Li 2026b; restated not re-derived as primary novelty.
  • ad hoc to paper Affine intervention entry before reduction for the negative Gram law and analytical examples.
    Eq. (8); sufficient for sign law and digits model, not necessary for general readout transfer.
invented entities (3)
  • Behaviorally minimal pathwise/unifilar optimizer state under fixed innovation coupling independent evidence
    purpose: Canonical causal realization for finite-horizon configuration responses.
    Quotient of extended histories by pathwise predictive equivalence; standard minimal-realization idea specialized to the protocol.
  • Observable-readout five-term interaction transfer independent evidence
    purpose: Map hidden first/second responses to arbitrary smooth update or trace scalars.
    Theorem 3; falsifiable via factorial integrals vs independent hidden-model predictions when H,G measured.
  • Geometric response class MF,B with TGER residual no independent evidence
    purpose: Passive membership test for declared positive maps along a trace under variation budget B.
    Definitional audit set; informative in neural tables but not a new physical mechanism.

pith-pipeline@v1.1.0-grok45 · 33086 in / 4121 out tokens · 43797 ms · 2026-07-14T15:49:35.625512+00:00 · methodology

0 comments
read the original abstract

Optimizer experiments observe responses to algorithmic configurations without uniquely revealing hidden mechanisms. We develop a causal optimizer interaction calculus that separates pathwise realization, Mobius decomposition, and experimental identification. Under a fixed innovation coupling, every finite-horizon innovation-driven optimizer admits a behaviorally minimal pathwise realization. For any finite effect support and intervention design, an incidence operator gives the complete observational gauge, exact identifiability, sharp quotient stability, held-out predictions, and exact noiseless configuration complexity. Smooth hidden relaxation generates interactions through inverse hidden-state stiffness. Building on this structural law, we prove an observable-readout transfer theorem: arbitrary smooth update or trace readouts inherit an explicit five-term interaction through first and second hidden responses. Unlike the reduced optimal value, a general readout has no universal interaction sign. Its Boolean effects remain exact integrals of continuous interaction curvature and can therefore be identified by factorial interventions. We also derive Gaussian quotient minimax risk, exact confidence sets and tests, misspecification decomposition, certified downstream decisions, and optimal replication. A controlled real-data experiment on a 65-dimensional strongly convex logistic model validates the complete reduced-value chain. Boolean effects and independently integrated curvature agree within 4.21e-11, while nine held-out continuous intensities agree within 8.88e-13. Gaussian campaigns attain the predicted coverage and power, and 4,500 real-minibatch observations reject an order-two interaction model. Neural trace audits provide complementary evidence that the declared response classes remain informative in nonconvex training.

Figures

Figures reproduced from arXiv: 2607.07206 by Zavier Li.

Figure 1
Figure 1. Figure 1: Boolean pair effects and independently integrated Schur curvature. The largest absolute discrepancy [PITH_FULL_IMAGE:figures/full_fig_p013_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Convergence of finite-horizon pair effects to the quasistatic effects. The maximum error decreases [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Response-class sensitivity. Layer scalar is nested inside bounded diagonal geometry, and the [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

174 extracted references · 5 canonical work pages · 4 internal anchors

  1. [1]

    A Differential Equation for Modeling

    Su, Weijie and Boyd, Stephen and Cand. A Differential Equation for Modeling. Journal of Machine Learning Research , year =

  2. [2]

    Proceedings of the National Academy of Sciences , year =

    A Variational Perspective on Accelerated Methods in Optimization , author =. Proceedings of the National Academy of Sciences , year =

  3. [3]

    Advances in Neural Information Processing Systems , year =

    Accelerated Mirror Descent in Continuous and Discrete Time , author =. Advances in Neural Information Processing Systems , year =

  4. [4]

    Neural Computation , year =

    Natural Gradient Works Efficiently in Learning , author =. Neural Computation , year =

  5. [5]

    IEEE Transactions on Information Theory , year =

    The Information Geometry of Mirror Descent , author =. IEEE Transactions on Information Theory , year =. doi:10.1109/TIT.2015.2391243 , url =

  6. [6]

    Proceedings of the 24th International Conference on Machine Learning , year =

    Information-Theoretic Metric Learning , author =. Proceedings of the 24th International Conference on Machine Learning , year =. doi:10.1145/1273496.1273523 , url =

  7. [7]

    Optimizing Neural Networks with

    Martens, James and Grosse, Roger , booktitle =. Optimizing Neural Networks with. 2015 , series =

  8. [8]

    Grosse, Roger and Martens, James , booktitle =. A. 2016 , series =. 1602.01407 , archivePrefix =

  9. [9]

    Fast Approximate Natural Gradient Descent in a

    George, Thomas and Laurent, C. Fast Approximate Natural Gradient Descent in a. Advances in Neural Information Processing Systems , year =

  10. [10]

    Journal of Machine Learning Research , volume =

    New Insights and Perspectives on the Natural Gradient Method , author =. Journal of Machine Learning Research , volume =. 2020 , url =

  11. [11]

    Advances in Neural Information Processing Systems , volume =

    Limitations of the Empirical Fisher Approximation for Natural Gradient Descent , author =. Advances in Neural Information Processing Systems , volume =. 2019 , url =

  12. [12]

    and Pitsianis, Nikos , booktitle =

    Van Loan, Charles F. and Pitsianis, Nikos , booktitle =. Approximation with. 1993 , doi =

  13. [13]

    Dutilleul, Pierre , journal =. The. 1999 , doi =

  14. [14]

    On Estimation of Covariance Matrices with

    Werner, Karl and Jansson, Magnus and Stoica, Petre , journal =. On Estimation of Covariance Matrices with. 2008 , doi =

  15. [15]

    IEEE Transactions on Signal Processing , volume =

    Geodesic Convexity and Covariance Estimation , author =. IEEE Transactions on Signal Processing , volume =. 2012 , doi =

  16. [16]

    On the Convexity in

    Wiesel, Ami , booktitle =. On the Convexity in. 2012 , doi =

  17. [17]

    Information Geometry and Asymptotics for

    McCormack, Andrew and Hoff, Peter , year =. Information Geometry and Asymptotics for. 2308.02260 , archivePrefix =

  18. [18]

    2025 , eprint =

    Geodesic Variational Bayes for Multiway Covariances , author =. 2025 , eprint =

  19. [19]

    Bouchard, Florent and Breloy, Arnaud and Mian, Ammar and Ginolhac, Guillaume , booktitle =. On-line. 2021 , doi =

  20. [20]

    Journal of Machine Learning Research , year =

    Adaptive Subgradient Methods for Online Learning and Stochastic Optimization , author =. Journal of Machine Learning Research , year =

  21. [21]

    International Conference on Learning Representations , year =

    Adam: A Method for Stochastic Optimization , author =. International Conference on Learning Representations , year =

  22. [22]

    International Conference on Learning Representations , year =

    Decoupled Weight Decay Regularization , author =. International Conference on Learning Representations , year =

  23. [23]

    Proceedings of the 35th International Conference on Machine Learning , year =

    Shampoo: Preconditioned Stochastic Tensor Optimization , author =. Proceedings of the 35th International Conference on Machine Learning , year =

  24. [24]

    and Janson, Lucas , year =

    Morwani, Depen and Shapira, Itai and Vyas, Nikhil and Malach, Eran and Kakade, Sham M. and Janson, Lucas , year =. A New Perspective on. 2406.17748 , archivePrefix =

  25. [25]

    arXiv preprint arXiv:2002.09018 , year =

    Scalable Second Order Optimization for Deep Learning , author =. arXiv preprint arXiv:2002.09018 , year =

  26. [26]

    2023 , eprint =

    A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale , author =. 2023 , eprint =

  27. [28]

    Machine Learning , year =

    Logarithmic Regret Algorithms for Online Convex Optimization , author =. Machine Learning , year =

  28. [29]

    Proceedings of the 20th International Conference on Machine Learning , year =

    Online Convex Programming and Generalized Infinitesimal Gradient Ascent , author =. Proceedings of the 20th International Conference on Machine Learning , year =

  29. [30]

    Proceedings of the 35th International Conference on Machine Learning , year =

    Adafactor: Adaptive Learning Rates with Sublinear Memory Cost , author =. Proceedings of the 35th International Conference on Machine Learning , year =

  30. [31]

    Advances in Neural Information Processing Systems , year =

    Memory Efficient Adaptive Optimization , author =. Advances in Neural Information Processing Systems , year =

  31. [32]

    2023 , eprint =

    Symbolic Discovery of Optimization Algorithms , author =. 2023 , eprint =

  32. [33]

    2023 , eprint =

    Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training , author =. 2023 , eprint =

  33. [34]

    2024 , howpublished =

    Muon: An Optimizer for Hidden Layers in Neural Networks , author =. 2024 , howpublished =

  34. [35]

    2025 , eprint =

    Practical Efficiency of Muon for Pretraining , author =. 2025 , eprint =

  35. [36]

    2025 , eprint =

    An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants , author =. 2025 , eprint =

  36. [37]

    Proceedings of the 36th International Conference on Machine Learning , year =

    Efficient Full-Matrix Adaptive Regularization , author =. Proceedings of the 36th International Conference on Machine Learning , year =

  37. [38]

    arXiv preprint arXiv:1912.02928 , year =

    Bregman Dynamics, Contact Transformations and Convex Optimization , author =. arXiv preprint arXiv:1912.02928 , year =

  38. [39]

    arXiv preprint arXiv:2505.12553 , year =

    Hamiltonian Descent Algorithms for Optimization: Accelerated Rates via Randomized Integration Time , author =. arXiv preprint arXiv:2505.12553 , year =

  39. [40]

    arXiv preprint arXiv:2606.17260 , year =

    Accelerated Convex Optimization via Hamiltonian Dynamics with Deterministic Integration Time , author =. arXiv preprint arXiv:2606.17260 , year =

  40. [41]

    arXiv preprint arXiv:2412.02291 , year =

    Conformal Symplectic Optimization for Stable Reinforcement Learning , author =. arXiv preprint arXiv:2412.02291 , year =

  41. [42]

    Journal of Machine Learning Research , year =

    Learning Discretized Neural Networks under Ricci Flow , author =. Journal of Machine Learning Research , year =

  42. [43]

    arXiv preprint arXiv:2509.22362 , year =

    Neural Feature Geometry Evolves as Discrete Ricci Flow , author =. arXiv preprint arXiv:2509.22362 , year =

  43. [44]

    International Conference on Machine Learning , year =

    Revisiting Over-smoothing and Over-squashing Using Ollivier-Ricci Curvature , author =. International Conference on Machine Learning , year =

  44. [45]

    Physical Review Research , year =

    Improving Gradient Methods via Coordinate Transformations: Applications to Quantum Machine Learning , author =. Physical Review Research , year =

  45. [46]

    Journal of Machine Learning Research , year =

    The Z-Gromov-Wasserstein Distance , author =. Journal of Machine Learning Research , year =

  46. [47]

    Optimization Algorithms on Matrix Manifolds , author =

  47. [48]

    International Conference on Learning Representations , year =

    Riemannian Adaptive Optimization Methods , author =. International Conference on Learning Representations , year =

  48. [49]

    2023 , doi =

    An Introduction to Optimization on Smooth Manifolds , author =. 2023 , doi =

  49. [50]

    An Introduction to Morse Theory , author =

  50. [51]

    2004 , url =

    Convex Optimization , author =. 2004 , url =

  51. [52]

    SIAM Review , volume =

    Semidefinite Programming , author =. SIAM Review , volume =. 1996 , doi =

  52. [53]

    Convex Analysis , author =

  53. [54]

    Positive Definite Matrices , author =

  54. [55]

    Arsigny, Vincent and Fillard, Pierre and Pennec, Xavier and Ayache, Nicholas , journal =. Log-. 2006 , doi =

  55. [56]

    2008 , doi =

    Functions of Matrices: Theory and Computation , author =. 2008 , doi =

  56. [57]

    1997 , doi =

    Matrix Analysis , author =. 1997 , doi =

  57. [58]

    International Journal of Computer Vision , volume =

    A Riemannian Framework for Tensor Computing , author =. International Journal of Computer Vision , volume =. 2006 , doi =

  58. [59]

    Biometrics , volume =

    Covariance Selection , author =. Biometrics , volume =. 1972 , doi =

  59. [60]

    Linear Algebra and its Applications , volume =

    Positive Definite Completions of Partial Hermitian Matrices , author =. Linear Algebra and its Applications , volume =. 1984 , doi =

  60. [61]

    Journal of Differential Geometry , volume =

    Three-Manifolds with Positive Ricci Curvature , author =. Journal of Differential Geometry , volume =. 1982 , doi =

  61. [62]

    Journal of Differential Geometry , volume =

    Deforming Metrics in the Direction of Their Ricci Tensors , author =. Journal of Differential Geometry , volume =. 1983 , doi =

  62. [63]

    USSR Computational Mathematics and Mathematical Physics , volume =

    Some Methods of Speeding Up the Convergence of Iteration Methods , author =. USSR Computational Mathematics and Mathematical Physics , volume =. 1964 , doi =

  63. [64]

    Doklady Akademii Nauk SSSR , volume =

    A Method for Solving the Convex Programming Problem with Convergence Rate \(O(1/k^2)\) , author =. Doklady Akademii Nauk SSSR , volume =

  64. [65]

    Problem Complexity and Method Efficiency in Optimization , author =

  65. [66]

    Operations Research Letters , volume =

    Mirror Descent and Nonlinear Projected Subgradient Methods for Convex Optimization , author =. Operations Research Letters , volume =. 2003 , doi =

  66. [68]

    Foundations and Trends in Optimization , volume =

    Proximal Algorithms , author =. Foundations and Trends in Optimization , volume =. 2014 , doi =

  67. [69]

    2006 , doi =

    Numerical Optimization , author =. 2006 , doi =

  68. [70]

    SIAM Review , volume =

    Optimization Methods for Large-Scale Machine Learning , author =. SIAM Review , volume =. 2018 , doi =

  69. [71]

    SIAM Journal on Optimization , volume =

    Analysis and Design of Optimization Algorithms via Integral Quadratic Constraints , author =. SIAM Journal on Optimization , volume =. 2016 , doi =

  70. [72]

    Mathematical Programming , volume =

    Performance of First-Order Methods for Smooth Convex Minimization: A Novel Approach , author =. Mathematical Programming , volume =. 2014 , doi =

  71. [73]

    Mathematical Programming , volume =

    Smooth Strongly Convex Interpolation and Exact Worst-Case Performance of First-Order Methods , author =. Mathematical Programming , volume =. 2017 , doi =

  72. [74]

    Advances in Neural Information Processing Systems , volume =

    Learning to Learn by Gradient Descent by Gradient Descent , author =. Advances in Neural Information Processing Systems , volume =. 2016 , url =

  73. [75]

    Journal of Machine Learning Research , volume =

    Learning to Optimize: A Primer and A Benchmark , author =. Journal of Machine Learning Research , volume =. 2022 , url =

  74. [77]

    IEEE Signal Processing Magazine , volume =

    Algorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing , author =. IEEE Signal Processing Magazine , volume =. 2021 , doi =

  75. [78]

    2024 , eprint =

    The Road Less Scheduled , author =. 2024 , eprint =

  76. [79]

    Evolutionary Computation , volume =

    Automated Algorithm Selection: Survey and Perspectives , author =. Evolutionary Computation , volume =. 2019 , doi =

  77. [80]

    Operations Research , volume =

    Inverse Optimization , author =. Operations Research , volume =. 2001 , doi =

  78. [81]

    Foundations of Computational Mathematics , volume =

    The Convex Geometry of Linear Inverse Problems , author =. Foundations of Computational Mathematics , volume =. 2012 , doi =

  79. [82]

    TEST , volume =

    Exact Testing with Random Permutations , author =. TEST , volume =. 2018 , doi =

  80. [83]

    Statistical Applications in Genetics and Molecular Biology , volume =

    Permutation p -values Should Never Be Zero: Calculating Exact p -values When Permutations Are Randomly Drawn , author =. Statistical Applications in Genetics and Molecular Biology , volume =. 2010 , doi =

Showing first 80 references.