REVIEW 3 major objections 6 minor 46 references
Langevin dynamics along the zero set of real-analytic potentials
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read For nonnegative real-analytic potentials, the large-β Langevin diffusion has a candidate limit that descends through progressively more singular strata and is biased toward them.
desk verdict Solid per-stratum Dirichlet-form limits for Langevin near real-analytic zero sets; the advertised global 'bias' is explicitly left open, but the theorems are real and worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine of the paper is the Dirichlet form of the diffusion, Eβ(f,g) = ∫ ∇f·∇g $e^{{-βV}}$ dx, viewed as a Laplace integral. Resolution of singularities monomializes V, so the Laplace asymptotics can be extracted by residues of localized zeta functions; the local learning coefficient λ and multiplicity m of each point of the zero set determine the leading decay $β^{{-λ}}$(log β)^{-(m-1)} and index the strata Z_{λ,m}. The limiting measures μ_{λ,m} are pushforwards of explicit densities in resolved coordinates, whose blowup rates encode the pull toward deeper strata. A Whitney stratification with a local integrability condition on each stratum makes the limiting forms closable, so standard Dirichlet-form theory yields the associated Markov processes.
What would settle it
Take a potential with two singularity strata, such as V(x,y)=$x^{{2k_1}}$ $y^{{2k_2}}$, and compute the limit measure μ_{λ,m} in resolved coordinates on a small ball around a point that maps to the deeper stratum. The paper's absorption picture predicts the density is nonintegrable there, so the limiting process must hit the deeper stratum and stop; if the computation yields an integrable density, the descent mechanism would fail. The paper itself notes it is not proved that μ_{λ,m} assigns infinite mass to neighborhoods of frontier points, making this the concrete check.
Extended reading notes
Core claim
The central discovery is Theorem 1.1: for each singularity type (λ,m), with cβ = β^λ (log β)^(1-m), and for $C^{1}$ test functions whose gradient is tangent to the stratum Z_{λ,m} and whose support avoids deeper strata, the rescaled Dirichlet forms satisfy cβ Eβ(f,g) → ∫_{Z_{λ,m}} ∇f·∇g dμ_{λ,m}, where μ_{λ,m} is a Radon measure constructed through resolution of singularities. The closure of this limiting form is a strongly local regular Dirichlet form, which corresponds to a Markov process with continuous trajectories in the one-point compactification of the stratum. The dynamics on each stratum can leave only by reaching the frontier, and points on the frontier lie in deeper strata; the explicit density of μ_{λ,m} in resolved coordinates shows nonintegrable blowup toward those deeper strata, which is interpreted as absorption. The upshot is a candidate limiting evolution on the whole zero set that descends through progressively more singular strata and is strongly biased toward them.
Load-bearing premise
The advertised single limiting process that descends through progressively more singular strata depends on gluing the separately constructed stratum processes into one process on the whole zero set, which the paper states is left open.
Editorial extensions
If this is right
- Every singularity stratum carries a well-defined Markov process with continuous paths in its one-point compactification, so the zero set acquires a hierarchical family of candidate limit dynamics rather than a single SDE.
- Because trajectories can leave a stratum only through its frontier, and frontier points lie in deeper strata, the limiting evolution is biased toward deeper (more singular) strata; in the two-axis example this reproduces Bessel-type absorption at the origin.
- For the SGLD/SGD interpretation, the result supplies a mechanism for the observed preference for singular, well-generalizing solutions: the effective drift toward more singular strata is encoded in the limiting measure's density and diverges as the process approaches deeper strata.
- The paper explicitly does not prove weak convergence of X^β or concatenation of the per-stratum processes into a single process on all of Z; these remain open, and the advertised descent should be read as a candidate limit unless one of them is established.
- The theory extends immediately to a tilted version with weight decay, where the limiting measures are multiplied by e^{-γ|x|^2/2} and the limiting dynamics acquires a tangential drift -γ τ(x) dt.
Reading between the lines
- If a future tightness and Mosco-convergence argument closes the gap, the limiting process would be the concatenation of the per-stratum processes, and the singularity-type hierarchy would then govern the long-run distribution of SGLD on the empirical-loss landscape—directly connecting the fixed-sample dynamics to Bayesian free-energy asymptotics of singular learning theory.
- The explicit density of μ_{λ,m} yields a quantitative prediction that can be tested before any convergence theorem: near a stratum front, the effective one-dimensional drift should diverge like a power of the distance, with an exponent determined by the ratio of local learning coefficients, so simulated SGLD trajectories in toy potentials could be used to estimate that exponent.
- The restriction to symmetric, isotropic-noise diffusions suggests the singularity bias may be special to SGLD-type algorithms; a natural test is whether anisotropic noise (e.g. noise covariance ∇²V) preserves the same hierarchy or shifts the preferred stratum, which would distinguish this mechanism from more generic flat-minima biases.
- One could ask whether the candidate limiting occupation measure on the deepest reachable stratum matches the Bayesian posterior's concentration region for the same empirical loss; if it does, the 'generalization puzzle' would be explained by the same singularities that govern Bayesian free energy, now appearing dynamically.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the Langevin SDE dX_t = -β∇V(X_t)dt + √2 dB_t for a nonnegative real-analytic potential V in the large-β limit, starting on the zero set Z=V^{-1}(0). Using Hironaka resolution of singularities, the author decomposes Z into finitely many strata Z_{λ,m} indexed by the local learning coefficient λ and multiplicity m, ordered so that smaller λ (and larger m for equal λ) corresponds to more singular behavior. The main technical result, Theorem 3.1, establishes rescaled Laplace asymptotics: for test functions supported away from deeper strata, c_β ∫ f e^{-βV} dx converges to ∫_{Z_{λ,m}} f dμ_{λ,m}, where μ_{λ,m} is a Radon measure pushforward of an explicit density in resolved coordinates. Theorem 4.1 then shows that the corresponding gradient form E_{λ,m}(f,g)=∫_{Z_{λ,m}} ∇f·∇g dμ_{λ,m} is closable on a suitable domain and its closure is a strongly local regular Dirichlet form, hence corresponds to a Markov process with continuous trajectories on the one-point compactification of Z_{λ,m}. The paper is explicit that weak convergence of X^β to these processes is not proved, and that concatenating the per-stratum processes into a single process on the whole zero set is left open. The global narrative—that SGLD/SGD dynamics descend through progressively more singular strata—is presented as a heuristic consequence of the per-stratum forms.
Significance. If the per-stratum results are correct, the paper makes a substantial contribution: it gives a rigorous, resolution-based description of the limiting Dirichlet forms associated with a large family of degenerate Langevin diffusions, and it connects this to the singularity theory of statistical learning. The proof chain is unusually detailed and self-contained, and the paper is careful to label the global process as a candidate rather than a proved limit. The main theorems are genuinely new and go beyond the smooth-manifold results of Li et al. (2022). The paper also avoids any fitted parameters: the normalization c_β and the limiting measures are determined by the intrinsic invariants λ and m. The principal weakness is that the advertised global mechanism—the strong bias toward more singular strata—is not itself a theorem, and the manuscript acknowledges this. The significance of the paper is therefore conditional on the per-stratum convergence being regarded as the main contribution, with the global evolution treated as a well-motivated conjecture.
major comments (3)
- [§1.1 (Limitations bullet), Remark 3.4] The abstract and Section 1.1 describe the result as a hierarchy of Dirichlet forms corresponding to a stochastic evolution that is 'strongly biased toward higher-dimensional, or more singular, strata.' As the paper itself notes in Section 1.1, rigorously concatenating the per-stratum processes into a single process on Z is left open, and Remark 3.4 states that it is not even proved that μ_{λ,m} assigns infinite mass to neighborhoods of frontier points. This gap is load-bearing for the global narrative: per-stratum Dirichlet forms do not determine whether the process continues at a frontier point, and the two-axis example V=x_1^{2k_1}x_2^{2k_2} shows that the form on the starting axis can have an absorbing point at the origin, so continuation on the other axis is an additional gluing prescription rather than a consequence of Theorem 1.1. The authors should reframe the global 'descent through strata' statement explicitly as a conjecture, and adjust the abstract and title so that the proved per-stratum convergence is the primary claim.
- [§4, Eq. (4.1), Remark 3.4] The heuristic that nonintegrable density blowup at frontier points 'corresponds to an exploding drift strong enough to force absorption' is not established by the proved results. Pointwise blowup of the density in (3.2), or even infinite total mass of μ_{λ,m} near a point, does not by itself imply that the associated Dirichlet form process is absorbed there; this depends on capacity and recurrence properties of the form. Since the absorption mechanism is the basis for the advertised bias toward deeper strata, it should be stated as a heuristic or proved directly. The current text in the introduction and Remark 3.4 is appropriately hedged, but the abstract's unconditional wording does not match this.
- [Theorem 1.1 and Theorem 4.1] Theorem 1.1 states convergence of c_β E_β(f,g) to E_{λ,m}(f,g) for f,g in a domain that depends on a Whitney stratification of Z_{λ,m}. The theorem statement does not mention that the tangency condition depends on the specific stratification constructed in Lemma 4.2, and that different stratifications could give different domains. This is not a technical error, but it is a presentation issue that should be clarified in the statement of Theorem 1.1, since the domain D_{λ,m} is defined before the stratification is introduced.
minor comments (6)
- [Abstract] The phrase 'strongly biased toward higher-dimensional, or "more singular", strata' is misleading: in the example V=x_1^{2k}x_2^{2k}, the deepest stratum is the origin, which is 0-dimensional, while the less singular strata are the punctured axes. The ordering by singularity type is not an ordering by stratum dimension, so the abstract should say only 'more singular'.
- [Page 3, paragraph after Eq. (1.2)] There is a typo: 'using on argument based on Itô’s formula' should read 'using an argument based on Itô’s formula'.
- [Section 3, around Eq. (3.2)] The notation y_{-I} in the density formula (3.2) is used but not defined; it should be stated that this denotes the coordinates with indices outside I.
- [Theorem 1.1] The statement 'whose gradient is tangent to Z_{λ,m}' depends on a Whitney stratification, but the theorem does not specify that the stratification is the one constructed in Theorem 4.1. This should be made explicit to avoid ambiguity.
- [Section 3, proof of Theorem 3.1] The notation c' for the shifted contour is close to the notation c_U for Laurent coefficients; renaming the contour parameter would improve readability.
- [References] The paper cites Drusvyatskiy and Larsson (2015) for an auxiliary approximation result in Lemma B.2(iii). This self-citation is appropriate and should be kept, but a brief sentence noting the result's role would help readers unfamiliar with it.
Circularity Check
No circularity: the limiting measures and scalings are derived from Laplace asymptotics, not fitted or defined by the claimed limit.
full rationale
The paper's derivation chain is self-contained in the relevant sense. Theorem 3.1 obtains the measure mu_{lambda,m} and the scaling beta^lambda (log beta)^(1-m) from a genuine asymptotic analysis of the Laplace integral over f e^{-beta V} dx, using Hironaka resolution and Mellin-Barnes residues; the limiting measure is given by an explicit pushforward density in equation (3.2). The Dirichlet form limit E_beta -> E_{lambda,m} is then a consequence rather than an input: E_{lambda,m} is defined through this independently derived mu_{lambda,m}, and no parameter is fitted to the target dynamics. The only occurrence of a self-citation is Drusvyatskiy and Larsson (2015) in Lemma B.2(iii), used to prove existence of cutoff functions with stratified gradients; that is an external published approximation theorem whose assumptions do not include the target result, so it constitutes independent support and does not raise the circularity score. Limitations are openly acknowledged rather than masked: Section 1.1 states that weak convergence of X^beta is not proved and that rigorously concatenating the per-stratum forms into a single process on Z is left open, and Remark 3.4 notes that it is not shown that mu_{lambda,m} assigns infinite mass to neighborhoods of frontier points. These are gaps or heuristic extrapolations, not circular reductions. The headline 'bias toward more singular strata' is a heuristic reading of the density blowup, not a quantity fitted to match the answer. The rigorous content of the paper, the per-stratum form convergence, does not reduce by construction to its inputs.
Assumptions & free parameters
assumptions (8)
- standard math Hironaka resolution of singularities for real-analytic functions with monomialization (Theorem 2.1)
- domain assumption Real analyticity and nonnegativity of V with nonempty proper zero set
- standard math Finite number of singularity types on compact subanalytic D
- standard math Subanalytic Whitney stratification theory (Appendix B, items (vii) through (ix))
- standard math Coarea formula for submanifolds (Lemma A.1)
- standard math Drusvyatskiy and Larsson approximation theorem
- standard math Correspondence between strongly local regular Dirichlet forms and continuous Markov processes
- domain assumption Starting point x0 lies on the zero set
Cite this review
Pith. "Pith review of Langevin dynamics along the zero set of real-analytic potentials." pith.science (2026). https://pith.science/paper/SVH3Q6P4
@misc{pith2026260809840,
author = {Pith},
title = {Pith review of: Langevin dynamics along the zero set of real-analytic potentials},
year = {2026},
howpublished = {\url{https://pith.science/paper/SVH3Q6P4}},
note = {Machine review of arXiv:2608.09840}
}
abstract
We consider the Langevin diffusion $dX_t = - \beta \nabla V(X_t) dt + \sqrt{2} dB_t$ for a general nonnegative real-analytic potential $V$ and a large parameter $\beta$. In the large-$\beta$ limit the process is confined to the zero set of $V$, assuming that it starts there. We derive a candidate limiting evolution on the zero set. To do so, the zero set is partitioned into strata according to a measure of local codimension known as the local learning coefficient and its multiplicity. It is then shown that the Dirichlet form associated with $X$ converges in a certain sense to a hierarchy of Dirichlet forms corresponding to a stochastic evolution on the strata. This evolution is strongly biased toward higher-dimensional, or "more singular", strata. This result is motivated by a question from Watanabe's singular learning theory regarding the learning dynamics of overparameterized statistical models and the generalization puzzle in deep learning. The result suggests a mechanism for the observation that stochastic gradient methods tend to be biased toward singular solutions that generalize well.
Reference graph
Works this paper leans on
-
[1]
ockner. Dirichlet forms and generalized S chr\
Sergio Albeverio, Johannes Brasche, and Michael R\"ockner. Dirichlet forms and generalized S chr\"odinger operators. In Schr\"odinger operators ( S nderborg, 1988) , volume 345 of Lecture Notes in Phys., pages 1--42. Springer, Berlin, 1989. ISBN 3-540-51783-9. doi:10.1007/3-540-51783-9\_15. URL https://doi.org/10.1007/3-540-51783-9_15
-
[2]
V. I. Arnold, S. M. Gusein-Zade, and A. N. Varchenko. Singularities of differentiable maps. V olume 2 . Modern Birkh\"auser Classics. Birkh\"auser/Springer, New York, 2012. ISBN 978-0-8176-8342-9. Monodromy and asymptotics of integrals, Translated from the Russian by Hugh Porteous and revised by the authors and James Montaldi, Reprint of the 1988 translation
work page 2012
-
[3]
K. B. Athreya and Chii-Ruey Hwang. Gibbs measures asymptotics. Sankhya A, 72 0 (1): 0 191--207, 2010. ISSN 0976-836X,0976-8378. doi:10.1007/s13171-010-0006-5. URL https://doi.org/10.1007/s13171-010-0006-5
-
[4]
M. F. Atiyah. Resolution of singularities and division of distributions. Comm. Pure Appl. Math., 23: 0 145--150, 1970. ISSN 0010-3640,1097-0312. doi:10.1002/cpa.3160230202. URL https://doi.org/10.1002/cpa.3160230202
-
[5]
Edward Bierstone and Pierre D. Milman. Semianalytic and subanalytic sets. Inst. Hautes \'Etudes Sci. Publ. Math., 0 (67): 0 5--42, 1988. ISSN 0073-8301,1618-1913. URL http://www.numdam.org/item?id=PMIHES_1988__67__5_0
work page 1988
-
[6]
Edward Bierstone and Pierre D. Milman. Canonical desingularization in characteristic zero by blowing up the maximum strata of a local invariant. Inventiones Mathematicae, 128 0 (2): 0 207--302, 1997. doi:10.1007/s002220050141
-
[7]
Implicit regularization for deep neural networks driven by an O rnstein-- U hlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant. Implicit regularization for deep neural networks driven by an O rnstein-- U hlenbeck like process. In Jacob Abernethy and Shivani Agarwal, editors, Proceedings of Thirty Third Conference on Learning Theory (COLT), volume 125 of Proceedings of Machine Learning Research, pages 483--513. PMLR, 2020. U...
work page 2020
-
[8]
Metastability in reversible diffusion processes
Anton Bovier, Michael Eckhoff, V\'eronique Gayrard, and Markus Klein. Metastability in reversible diffusion processes. I . S harp asymptotics for capacities and exit times. J. Eur. Math. Soc. (JEMS), 6 0 (4): 0 399--424, 2004. ISSN 1435-9855,1435-9863. doi:10.4171/JEMS/14. URL https://doi.org/10.4171/JEMS/14
doi:10.4171/jems/14 2004
Show all 46 references
-
[9]
Metastability in reversible diffusion processes
Anton Bovier, V\'eronique Gayrard, and Markus Klein. Metastability in reversible diffusion processes. II . P recise asymptotics for small eigenvalues. J. Eur. Math. Soc. (JEMS), 7 0 (1): 0 69--99, 2005. ISSN 1435-9855,1435-9863. doi:10.4171/JEMS/22. URL https://doi.org/10.4171/JEMS/22
2005 doi
-
[10]
Convergence rates of G ibbs measures with degenerate minimum
Pierre Bras. Convergence rates of G ibbs measures with degenerate minimum. Bernoulli, 28 0 (4): 0 2431--2458, 2022. ISSN 1350-7265,1573-9759. doi:10.3150/21-bej1424. URL https://doi.org/10.3150/21-bej1424
2022 doi
-
[11]
Dynamical versus B ayesian phase transitions in a toy model of superposition, 2023
Zhongtian Chen, Edmund Lau, Jake Mendel, Susan Wei, and Daniel Murfet. Dynamical versus B ayesian phase transitions in a toy model of superposition, 2023. URL https://arxiv.org/abs/2310.06301
2023 arXiv
-
[12]
Alex Damian, Tengyu Ma, and Jason D. Lee. Label noise SGD provably prefers flat global minimizers. In Marc'Aurelio Ranzato, Alina Beygelzimer, Yann Dauphin, Percy S. Liang, and Jenn Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 2...
2021
-
[13]
Drusvyatskiy and M
D. Drusvyatskiy and M. Larsson. Approximating functions on stratified sets. Trans. Amer. Math. Soc., 367 0 (1): 0 725--749, 2015. ISSN 0002-9947,1088-6850. doi:10.1090/S0002-9947-2014-06412-X. URL https://doi.org/10.1090/S0002-9947-2014-06412-X
2015 doi
-
[14]
Geometric measure theory, volume Band 153 of Die Grundlehren der mathematischen Wissenschaften
Herbert Federer. Geometric measure theory, volume Band 153 of Die Grundlehren der mathematischen Wissenschaften. Springer-Verlag New York, Inc., New York, 1969
1969
-
[15]
Gerald B. Folland. Real analysis. Pure and Applied Mathematics (New York). John Wiley & Sons, Inc., New York, second edition, 1999. ISBN 0-471-31716-0. Modern techniques and their applications, A Wiley-Interscience Publication
1999
-
[16]
M. I. Freidlin and A. D. Wentzell. Random perturbations of dynamical systems, volume 260 of Grundlehren der mathematischen Wissenschaften [Fundamental Principles of Mathematical Sciences]. Springer-Verlag, New York, 1984. ISBN 0-387-90858-7. doi:10.1007/978-1-4684-0176-9. URL ...
1984 doi
-
[17]
Freidlin and Alexander D
Mark I. Freidlin and Alexander D. Wentzell. Diffusion processes on graphs and the averaging principle. Ann. Probab., 21 0 (4): 0 2215--2245, 1993. ISSN 0091-1798,2168-894X. URL http://links.jstor.org/sici?sici=0091-1798(199310)21:4<2215:DPOGAT>2.0.CO;2-G&origin=MSN
1993
-
[18]
Dirichlet forms and symmetric M arkov processes , volume 19 of De Gruyter Studies in Mathematics
Masatoshi Fukushima, Yoichi Oshima, and Masayoshi Takeda. Dirichlet forms and symmetric M arkov processes , volume 19 of De Gruyter Studies in Mathematics. Walter de Gruyter & Co., Berlin, extended edition, 2011. ISBN 978-3-11-021808-4
2011
-
[19]
Estimating the local learning coefficient at scale
Zach Furman and Edmund Lau. Estimating the local learning coefficient at scale. 2024. URL https://arxiv.org/abs/2402.03698
2024 arXiv
-
[20]
I. M. Gel ' fand and G. E. Shilov. Generalized functions. V ol. 1 . AMS Chelsea Publishing, Providence, RI, 2016. ISBN 978-1-4704-2658-3. doi:10.1090/chel/377. URL https://doi.org/10.1090/chel/377. Properties and operations, Translated from the 1958 Russian original [MR0097715...
2016 doi
-
[21]
On L evi's problem and the imbedding of real-analytic manifolds
Hans Grauert. On L evi's problem and the imbedding of real-analytic manifolds. Ann. of Math. (2), 68: 0 460--472, 1958. ISSN 0003-486X. doi:10.2307/1970257. URL https://doi.org/10.2307/1970257
1958 doi
-
[22]
Heat kernels on weighted manifolds and applications
Alexander Grigor ' yan. Heat kernels on weighted manifolds and applications. In The ubiquitous heat kernel, volume 398 of Contemp. Math., pages 93--191. Amer. Math. Soc., Providence, RI, 2006. ISBN 0-8218-3698-6. doi:10.1090/conm/398/07486. URL https://doi.org/10.1090/conm/398/07486
2006 doi
-
[23]
Mohammed M. Hamza. D\'etermination des formes de D irichlet sur R ^n . Th\`ese de 3\`eme cycle, Universit\'e Paris-Sud, Orsay, 1975
1975
-
[24]
HaoChen, Colin Wei, Jason D
Jeff Z. HaoChen, Colin Wei, Jason D. Lee, and Tengyu Ma. Shape matters: Understanding the implicit bias of the noise covariance. In Proceedings of the 34th Conference on Learning Theory (COLT), volume 134 of Proceedings of Machine Learning Research, pages 2315--2357. PMLR, 202...
2021
-
[25]
Resolution of singularities of an algebraic variety over a field of characteristic zero
Heisuke Hironaka. Resolution of singularities of an algebraic variety over a field of characteristic zero. I , II . Ann. of Math. (2), 79: 0 109--203; 79 (1964), 205--326, 1964. ISSN 0003-486X. doi:10.2307/1970547. URL https://doi.org/10.2307/1970547
1964 doi
-
[26]
Flat minima
Sepp Hochreiter and J \"u rgen Schmidhuber. Flat minima. Neural Computation, 9 0 (1): 0 1--42, 1997. doi:10.1162/neco.1997.9.1.1
1997 doi
-
[27]
Loss landscape degeneracy and stagewise development in transformers
Jesse Hoogland, George Wang, Matthew Farrugia-Roberts, Liam Carroll, Susan Wei, and Daniel Murfet. Loss landscape degeneracy and stagewise development in transformers. Transactions on Machine Learning Research, 2025. ISSN 2835-8856. URL https://openreview.net/forum?id=45qJyBG8Oj
2025
-
[28]
Laplace's method revisited: weak convergence of probability measures
Chii-Ruey Hwang. Laplace's method revisited: weak convergence of probability measures. Ann. Probab., 8 0 (6): 0 1177--1182, 1980. ISSN 0091-1798,2168-894X. URL http://links.jstor.org/sici?sici=0091-1798(198012)8:6<1177:LMRWCO>2.0.CO;2-1&origin=MSN
1980
-
[29]
G. S. Katzenberger. Solutions of a stochastic differential equation forced onto a manifold by a large drift. Ann. Probab., 19 0 (4): 0 1587--1628, 1991. ISSN 0091-1798,2168-894X. URL http://links.jstor.org/sici?sici=0091-1798(199110)19:4<1587:SOASDE>2.0.CO;2-H&origin=MSN
1991
-
[30]
Convergence of spectral structures: a functional analytic theory and its applications to spectral geometry
Kazuhiro Kuwae and Takashi Shioya. Convergence of spectral structures: a functional analytic theory and its applications to spectral geometry. Comm. Anal. Geom., 11 0 (4): 0 599--673, 2003. ISSN 1019-8385,1944-9992. doi:10.4310/CAG.2003.v11.n4.a1. URL https://doi.org/10.4310/C...
2003 doi
-
[31]
The local learning coefficient: A singularity-aware complexity measure, 2024
Edmund Lau, Zach Furman, George Wang, Daniel Murfet, and Susan Wei. The local learning coefficient: A singularity-aware complexity measure, 2024. URL https://arxiv.org/abs/2308.12108
2024 arXiv
-
[32]
John M. Lee. Introduction to smooth manifolds, volume 218 of Graduate Texts in Mathematics. Springer, New York, second edition, 2013. ISBN 978-1-4419-9981-8
2013
-
[33]
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E. Stochastic modified equations and adaptive stochastic gradient algorithms. In Doina Precup and Yee Whye Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Resear...
-
[34]
What happens after SGD reaches zero loss? -- a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora. What happens after SGD reaches zero loss? -- a mathematical framework. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=siCt4xZn5Ve
2022
-
[35]
Algebraic Methods for Evaluating Integrals in B ayesian Statistics
Shaowei Lin. Algebraic Methods for Evaluating Integrals in B ayesian Statistics . PhD thesis, University of California, Berkeley, 2011. Available at https://sites.google.com/site/shaoweilin/home/phd-thesis
2011
-
[36]
Hoffman, and David M
Stephan Mandt, Matthew D. Hoffman, and David M. Blei. Stochastic gradient descent as approximate bayesian inference. J. Mach. Learn. Res., 18 0 (1): 0 4873–4907, January 2017. ISSN 1532-4435
2017
-
[37]
Notes on topological stability
John Mather. Notes on topological stability. Bull. Amer. Math. Soc. (N.S.), 49 0 (4): 0 475--506, 2012. ISSN 0273-0979,1088-9485. doi:10.1090/S0273-0979-2012-01383-6. URL https://doi.org/10.1090/S0273-0979-2012-01383-6
2012 doi
-
[38]
Implicit bias of SGD for diagonal linear networks: A provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion. Implicit bias of SGD for diagonal linear networks: A provable benefit of stochasticity. In Advances in Neural Information Processing Systems, volume 34, pages 29218--29230. Curran Associates, Inc., 2021. URL https://p...
2021
-
[39]
Label noise (stochastic) gradient descent implicitly solves the L asso for quadratic parametrisation
Loucas Pillaud-Vivien, Julien Reygner, and Nicolas Flammarion. Label noise (stochastic) gradient descent implicitly solves the L asso for quadratic parametrisation. In Proceedings of the 35th Conference on Learning Theory (COLT), volume 178 of Proceedings of Machine Learning R...
2022
-
[40]
Geometry of subanalytic and semialgebraic sets, volume 150 of Progress in Mathematics
Masahiro Shiota. Geometry of subanalytic and semialgebraic sets, volume 150 of Progress in Mathematics. Birkh\"auser Boston, Inc., Boston, MA, 1997. ISBN 0-8176-4000-2. doi:10.1007/978-1-4612-2008-4. URL https://doi.org/10.1007/978-1-4612-2008-4
1997 doi
-
[41]
Differentiation and specialization of attention heads via the refined local learning coefficient
George Wang, Jesse Hoogland, Stan van Wingerden, Zach Furman, and Daniel Murfet. Differentiation and specialization of attention heads via the refined local learning coefficient. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openrevi...
2025
-
[42]
Algebraic geometry and statistical learning theory, volume 25 of Cambridge Monographs on Applied and Computational Mathematics
Sumio Watanabe. Algebraic geometry and statistical learning theory, volume 25 of Cambridge Monographs on Applied and Computational Mathematics. Cambridge University Press, Cambridge, 2009. ISBN 978-0-521-86467-1. doi:10.1017/CBO9780511800474. URL https://doi.org/10.1017/CBO978...
2009 doi
-
[43]
Mathematical theory of B ayesian statistics
Sumio Watanabe. Mathematical theory of B ayesian statistics . CRC Press, Boca Raton, FL, 2018. ISBN 978-1-482-23806-8. doi:10.1201/9781315373010. URL https://doi.org/10.1201/9781315373010
2018 doi
-
[44]
Deep learning is singular, and that's good
Susan Wei, Daniel Murfet, Mingming Gong, Hui Li, Jesse Gell-Redman, and Thomas Quella. Deep learning is singular, and that's good. IEEE Transactions on Neural Networks and Learning Systems, 34 0 (12): 0 10473--10486, 2023. doi:10.1109/TNNLS.2022.3167409
2023
-
[45]
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee Whye Teh. Bayesian learning via stochastic gradient langevin dynamics. In Proceedings of the 28th International Conference on International Conference on Machine Learning, ICML'11, page 681–688, Madison, WI, USA, 2011. Omnipress. ISBN 9781450306195
2011
-
[46]
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Zeke Xie, Issei Sato, and Masashi Sugiyama. A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima. In International Conference on Learning Representations (ICLR), 2021. URL https://openreview.net/forum?id=wXgk_iCiYGo
2021
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.