REVIEW 1 major objections 5 minor 58 references
Mathematical analysis of the gradients in deep learning
T0 review · 1 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read There is a unique generalized gradient for training ReLU networks.
desk verdict Generalizes the ReLU smoothing result to arbitrary continuous activations, but the proof of the subgradient claim has a sign error in Lemma 3.10 that must be fixed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the admissible smoothing sequence: a sequence of $C^1$ activations $A_n$ that eventually agrees pointwise with the original activation and with its generalized derivative, while staying uniformly bounded on compact intervals. Representation formulas in Proposition 2.14 write each coordinate of $G$ as an explicit integral over paths through the network, with products of weights and values of the generalized derivative at pre-activation values. Because every admissible sequence eventually evaluates the same quantities at the same finitely many pre-activation values, all smoothed-loss gradients share the same pointwise limit. A second mechanism is the one-sided continuity condition on the generalized derivative at the kink set $S$, which lets the proof approximate an arbitrary parameter by nearby differentiability points and pass $G(\theta)$ through as a limiting Fréchet subgradient.
What would settle it
Implement two different $C^1$ smoothings of ReLU whose derivatives converge pointwise to the left derivative and that are uniformly bounded on compact sets, then evaluate the smoothed-loss gradients at a parameter vector where some pre-activation equals zero. The theorem predicts that the two gradient vectors converge to the same limit; observing a nonzero difference between the limits would disprove the uniqueness claim.
Extended reading notes
Core claim
The central claim is that the gradient returned by backpropagation through a ReLU-like network is not implementation-dependent. For any continuous activation that is $C^1$ off a finite set $S$ and any generalized derivative $\psi$ that agrees with the true derivative off $S$, the paper considers sequences of $C^1$ activations $A_n$ that converge pointwise to the original activation while their derivatives converge pointwise to $\psi$, with uniform boundedness on compact intervals. Theorems 2.15 and 3.14 show that the gradients of the smoothed losses converge pointwise to a single function $G$ that depends only on the network, the loss, the data measure, and the choice of $\psi$, not on the particular smoothing sequence. Moreover, $G(\theta)$ is a limiting Fréchet subgradient of the nonsmooth loss for every $\theta$, and $G$ agrees with the classical gradient of the loss on every open set where that loss is $C^1$. For the ReLU activation with the left derivative, this is an exact description of the vector that automatic differentiation in standard deep-learning libraries produces.
Load-bearing premise
The argument rests on the generalized derivative approaching its value at each non-differentiability point from at least one side; if it oscillates on both sides of a kink, the proof that the generalized gradient is a limiting Fréchet subgradient, and with it the agreement with the true gradient on smooth regions, no longer goes through.
Editorial extensions
If this is right
- Backpropagation through a ReLU kink is deterministic: every smoothing family that respects the generalized derivative converges to the same parameter vector field, so convergence analyses can be written against a single function rather than a set-valued subgradient.
- The generalized gradient is always a limiting Fréchet subgradient, giving a nonsmooth variational certificate at every point of parameter space.
- On any open region where the loss is continuously differentiable, the generalized gradient equals the ordinary gradient, so smoothing-based training and gradient-flow dynamics coincide on smooth parts of the landscape.
- The theorems extend prior ReLU-specific analyses to arbitrary piecewise-smooth activations with finitely many kinks, general $C^1$ loss functions, and compactly supported data measures.
Reading between the lines
- A natural next question the paper leaves open is whether uniqueness survives for activations whose generalized derivative oscillates on both sides of a kink; two smoothings with different one-sided limits there would likely produce different generalized gradients.
- The single-valued uniqueness result suggests that stochastic gradient descent analyses could treat this $G$ as the canonical descent field, and it would be worth checking whether convergence guarantees known for conservative set-valued fields transfer to this smaller object.
- One could numerically probe the theorem: at parameters where a hidden pre-activation lands exactly on a kink, compare the gradients produced by two different smoothing widths; the theorem predicts identical limiting vectors.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies supervised-learning risk functionals for deep fully-connected feedforward networks with a nonsmooth activation A0 that is C1 away from a finite set S. It fixes a generalized derivative ψ of A0 and considers sequences of C1 activations A_n converging pointwise (eventually exactly) to A0 with derivatives converging to ψ. Under local boundedness and a one-sided continuity condition on ψ, it proves (Theorem 2.15 and Theorem 3.14) that there is a unique G:R^d→R^d to which the gradients ∇L_{A_n} converge pointwise regardless of the approximating sequence; that G(θ) is a limiting Fréchet subgradient of the nonsmooth risk L0 at θ; and that G coincides with ∇L0 on every open set where L0 is C1. The proof uses explicit integral representations for gradient components (Proposition 2.14), dominated convergence, Rademacher's theorem, and a limiting-subgradient density argument.
Significance. If correct, the result is a clean and fairly general mathematical description of the object that automatic differentiation libraries compute for ReLU-type networks. The main conceptual value is the uniqueness statement: unlike set-valued conservative fields, the smoothed-gradient limit is a single-valued function independent of the smoothing mechanism. The paper also extends the ReLU/MSE-specific analysis of prior work to general C1 loss functions and finite-singularity activations with a chosen generalized derivative. The explicit representation formulas and the use of standard analytic tools are strengths. However, the proof of the subgradient property currently contains a localized sign bug that must be corrected before the central claim is established.
major comments (1)
- [§3.6, Lemma 3.10, Eq. (155)–(159)] The sign used in the application of Lemma 3.8 is inconsistent with the one-sided continuity assumption. The hypothesis before Theorem 3.14 supplies z∈{−1,1} such that ψ is continuous from side z (for ReLU with ψ=1_{(0,∞)}, the good sign is z=−1). Applying Lemma 3.8 with the same z yields z(N_{k,ϑ}(x)−N_{k,θ}(x))≤0; for z=−1 this is N_{k,ϑ}(x)≥N_{k,θ}(x), so the approximating pre-activations approach any singular value from above, i.e. from the side on which ψ is discontinuous. Consequently the asserted limsup in (159) need not hold; in the scalar construction with A0=ReLU and θ=0, the sequence chosen from the U_n produced by Lemma 3.8 with z=−1 has positive pre-activations, so ψ equals 1 there while ψ(0)=0 and (153) fails for the constructed sequence. The fix is to apply Lemma 3.8 with −z, so that z(N_{k,ϑ}(x)−N_{k,θ}(x))≥0 and the pre-activations approach from the side where ψ is continuous. This repairs the proof of Lemma 3.10 and hence of Proposition 3.12 and Theorem 3.14(iii)–(iv). Because this is a proof bug in a load-bearing lemma rather than a counterexample to the theorem, it is correctable, but as written the central subgradient claim is not established.
minor comments (5)
- [§3.4, Lemma 3.7] In item (iii), the statement writes (∇L0)(x)=G(θ) but θ is not quantified in that item; it should be G(x). The proof also contains a typo in Eq. (135), where the integration domain appears as Rθ and the two functions G_k and G_k are conflated.
- [§2.2, Proposition 2.3] In Eq. (14), the definition of G_n(v) uses the variable x in the term h(y_{i_v})(x−y_{i_v}), but this should be v; as written the formula is inconsistent with the surrounding definitions.
- [§3.6, Lemma 3.10] In the sentence before Eq. (154), the proof states that there exists z∈R satisfying the displayed limit for all x∈R; the correct assertion is that z can be chosen in {−1,1} and that the limit statement holds for all x, with the singular points x∈S covered by the hypothesis and the nonsingular points covered by continuity of ψ away from S.
- [Theorem 2.15 and Theorem 3.14] The boundedness condition in items (ii) contains a malformed expression involving an indicator function of {∞} applied to supp(μ); for example, Theorem 3.14(ii) writes (z+1_{∞}(supp(μ))) without defining this notation. This should be cleaned up for readability.
- [§3.7, Proposition 3.12 and Corollary 3.13] The proof relies on [30, Lemma 3.8] and its consequence for C1 domains without restating them; since these results are central to the argument, the authors should either state the exact versions used or give a short proof, so that the paper is more self-contained.
Circularity Check
No circularity: the uniqueness theorem is proved directly via sequence-independent integral representations, and the self-citations supply independent published lemmas rather than importing the target conclusion.
full rationale
The main uniqueness claim, Theorem 2.15(ii) / Theorem 3.14(ii), is not circular. The generalized gradient G is introduced in Setting 2.2 only on the set where the smoothed gradients converge, and Proposition 2.14 then derives explicit integral representation formulas, equations (81) and (82), that depend only on the limit activation function A0 and the generalized derivative psi, not on the particular approximating sequence. Since any two admissible sequences must converge to the same explicit formula, uniqueness is proved directly from the representation, not assumed. Corollary 2.7 supplies the existence of admissible smoothings from the stated local-boundedness assumptions, and the one-sided continuity condition on psi is a genuine regularity hypothesis used only for the subgradient conclusion in Lemma 3.10, not needed for the uniqueness formula. The subgradient and C^1-agreement results in Theorem 3.14(iii)-(iv) do cite the authors' prior work [29, 30], and [30] is explicitly identified as establishing the ReLU/mean-squared-error special case. However, the paper states that it generalizes those arguments and proves new lemmas (e.g., Lemma 3.7, Lemma 3.8, Lemma 3.10) rather than merely importing the conclusion. The cited facts — that the gradient at a differentiability point is a Frechet subgradient and that a C^1 function has singleton Frechet subdifferential — are standard, published, parameter-free results whose assumptions do not include the present theorem's conclusion. Thus the self-citation is real evidence, not circular reasoning. The possible sign issue in Lemma 3.10 raised in the skeptic note is a correctness concern about the proof, not a circularity, and does not change this verdict.
Assumptions & free parameters
assumptions (5)
- standard math Rademacher's theorem and the fundamental lemma of calculus of variations (invoked via [29, Lemma 3.6 and Lemma 3.9])
- standard math [30, Lemma 3.8] relating Frechet subgradients at differentiability points and the coincidence of a locally bounded weak derivative with the gradient
- domain assumption One-sided continuity of the generalized derivative psi at the finite singular set S: min_{z in {-1,1}} sum_{x in S} limsup_{h down to 0} |psi(x+zh) - psi(x)| = 0
- domain assumption Existence of C^1 approximations A_n with eventual pointwise equality A_n = A_0 and A_n' = psi, with uniform local boundedness of |A_n| + |A_n'|
- domain assumption Growth condition on the loss gradient: for every r, sup over x in [-r,r]^ell0, y in R^ellL, theta in [-r,r]^d of (||grad_x H|| + ||grad_theta H||) / (1 + |H(X,Theta,y)|) < infinity on the support of mu
Cite this review
Pith. "Pith review of Mathematical analysis of the gradients in deep learning." pith.science (2026). https://pith.science/paper/QUAY5QWO
@misc{pith2026250115646,
author = {Pith},
title = {Pith review of: Mathematical analysis of the gradients in deep learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/QUAY5QWO}},
note = {Machine review of arXiv:2501.15646}
}
abstract
Deep learning algorithms -- typically consisting of a class of deep artificial neural networks (ANNs) trained by a stochastic gradient descent (SGD) optimization method -- are nowadays an integral part in many areas of science, industry, and also our day to day life. Roughly speaking, in their most basic form, ANNs can be regarded as functions that consist of a series of compositions of affine-linear functions with multidimensional versions of so-called activation functions. One of the most popular of such activation functions is the rectified linear unit (ReLU) function $\mathbb{R} \ni x \mapsto \max\{ x, 0 \} \in \mathbb{R}$. The ReLU function is, however, not differentiable and, typically, this lack of regularity transfers to the cost function of the supervised learning problem under consideration. Regardless of this lack of differentiability issue, deep learning practioners apply SGD methods based on suitably generalized gradients in standard deep learning libraries like {\sc TensorFlow} or {\sc Pytorch}. In this work we reveal an accurate and concise mathematical description of such generalized gradients in the training of deep fully-connected feedforward ANNs and we also study the resulting generalized gradient function analytically. Specifically, we provide an appropriate approximation procedure that uniquely describes the generalized gradient function, we prove that the generalized gradients are limiting Fr\'echet subgradients of the cost functional, and we conclude that the generalized gradients must coincide with the standard gradient of the cost functional on every open sets on which the cost functional is continuously differentiable.
Figures
Reference graph
Works this paper leans on
-
[30]
Jentzen, A., and Riekert, A. On the existence of global minima and convergence analyses for gradient descent methods in the training of dee p neural networks. J. Mach. Learn. 1, 2 (2022), 141–246
work page 2022
-
[1]
S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfell ow, I
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Co r- rado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfell ow, I. J., Harp, A., Irving, G., Isard, M., Jia, Y., J ´ozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Man ´ e, D., Monga, R., Moore, S., Murray, D. G., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutsk...
arXiv 2016
-
[2]
Learning Theory from First Principles
Bach, F. Learning Theory from First Principles . Adaptive Computation and Machine Learning series. MIT Press, 2024
2024
-
[3]
Bai, Z., Luo, T., Xu, Z.-Q. J., and Zhang, Y. Embedding principle in depth for the loss landscape analysis of deep neural networks. CSIAM Trans. Appl. Math. 5 , 2 (2024), 350–389
work page 2024
-
[4]
On the complexity of nonsmooth automatic differentiation
Bolte, J., Boustany, R., Pauwels, E., and Pesquet-Popescu, B. Nons- mooth automatic differentiation: a cheap gradient principle and other complexity results. arXiv:2206.01730v1 (2022)
work page Pith review arXiv 2022
-
[5]
Bolte, J., Daniilidis, A., and Lewis, A. The /suppress lojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems. SIAM J. Optim. 17, 4 (2006), 1205–1223. 34
work page 2006
-
[6]
A mathematical model for automatic differentiation in machine learning
Bolte, J., and Pauwels, E. A mathematical model for automatic differentiation in machine learning. arXiv:2006.02080 (2020)
work page Pith review arXiv 2020
-
[7]
Bolte, J., and Pauwels, E. Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning. Math. Program. 188 , 1 (2021), 19–51
work page 2021
Show all 58 references
-
[8]
Differentiating nonsmooth solutions to parametric monotone inclusion problems
Bolte, J., Pauwels, E., and Silveti-F alls, A. Differentiating nonsmooth solutions to parametric monotone inclusion problems. SIAM J. Optim. 34 , 1 (2024), 71–97
2024
-
[9]
Automatic differentiation of nonsmooth iter- ative algorithms
Bolte, J., Pauwels, E., and V aiter, S. Automatic differentiation of nonsmooth iter- ative algorithms. arXiv:2206.00457 (2022)
2022 arXiv
-
[10]
A proof of conver- gence for gradient descent in the training of artificial neur al networks for constant target functions
Cheridito, P., Jentzen, A., Riekert, A., and Rossmannek, F. A proof of conver- gence for gradient descent in the training of artificial neur al networks for constant target functions. J. Complexity 72 (2022), Paper No. 101646, 26
2022
-
[11]
Non-convergence of stochastic gradient descent in the training of deep neural networks
Cheridito, P., Jentzen, A., and Rossmannek, F. Non-convergence of stochastic gradient descent in the training of deep neural networks. J. Complexity 64 (2021), Paper No. 101540, 10
2021
-
[12]
Landscape analysis for shallow neural networks: complete classification of critical point s for affine target functions
Cheridito, P., Jentzen, A., and Rossmannek, F. Landscape analysis for shallow neural networks: complete classification of critical point s for affine target functions. J. Nonlinear Sci. 32 , 5 (2022), Paper No. 64, 45
2022
-
[13]
On the mathematical foundations of learning
Cucker, F., and Smale, S. On the mathematical foundations of learning. Bull. Amer. Math. Soc. (N.S.) 39 , 1 (2002), 1–49
2002
-
[14]
Conservative and semismooth derivatives are equiv- alent for semialgebraic maps
Davis, D., and Drusvyatskiy, D. Conservative and semismooth derivatives are equiv- alent for semialgebraic maps. Set-Valued Var. Anal. 30 , 2 (2022), 453–463
2022
-
[15]
Davis, D., Drusvyatskiy, D., Kakade, S., and Lee, J. D. Stochastic subgradient method converges on tame functions. Found. Comput. Math. 20 , 1 (2020), 119–154
2020
-
[16]
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
Dereich, S., Graeber, R., and Jentzen, A. Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates. arXiv:2407.08100 (2024)
2024 arXiv
-
[17]
Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes
Dereich, S., and Kassing, S. Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes. arXiv:2102.09385 (2021)
2021 arXiv
-
[18]
Adam-family Methods with Decoupled Weight Decay in Deep Learning
Ding, K., Xiao, N., and Toh, K.-C. Adam-family Methods with Decoupled Weight Decay in Deep Learning. arXiv:2310.08858 (2023)
2023 arXiv
-
[19]
Sub-Optimal Local Minima Exist for Neural Networks with Almost All Non-Linear Activations
Ding, T., Li, D., and Sun, R. Sub-Optimal Local Minima Exist for Neural Networks with Almost All Non-Linear Activations. arXiv:1911.01413 (2019)
2019 arXiv
-
[20]
Towards a Mathematical Under- standing of Neural Network-Based Machine Learning: what we know and what we don’t
E, W., Ma, C., Wojtowytsch, S., and Wu, L. Towards a Mathematical Under- standing of Neural Network-Based Machine Learning: what we know and what we don’t. arXiv:2009.10713 (2020)
2020 arXiv
-
[21]
Evans, L. C. Partial differential equations , vol. 19 of Graduate Studies in Mathematics . American Mathematical Society, Providence, RI, 1998
1998
-
[22]
Blow up phenomena for gradient descent optimization methods in the training of artificial neural networks
Gallon, D., Jentzen, A., and Lindner, F. Blow up phenomena for gradient descent optimization methods in the training of artificial neural networks. arXiv:2211.15641 (2022), 84 pages. 35
2022 arXiv
-
[23]
Garrigos, G., and Gower, R. M. Handbook of Convergence Theorems for (Stochastic) Gradient Methods. arXiv:2301.11235 (2020)
2020 arXiv
-
[24]
Approximation results for gradient descent trained shal- low neural networks in 1d
Gentile, R., and Welper, G. Approximation results for gradient descent trained shal- low neural networks in 1d. arXiv:2209.08399 (2022)
2022 arXiv
-
[25]
Hannibal, S., Jentzen, A., and Thang, D. M. Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the traini ng of deep neural networks with relu activat...
2024 arXiv
-
[26]
Conver- gence proof for stochastic gradient descent in the training of deep neural networks with ReLU activation for constant target functions
Hutzenthaler, M., Jentzen, A., Pohl, K., Riekert, A., and Scarpa , L. Conver- gence proof for stochastic gradient descent in the training of deep neural networks with ReLU activation for constant target functions. arXiv:2112.07369 (2021)
2021 arXiv
-
[27]
Ibragimov, S., Jentzen, A., and Riekert, A. Convergence to good non-optimal critical points in the training of neural networks: Gradien t descent optimization with one random initialization overcomes all bad non-global local m inima with high probability. arXiv:2212.13111 (2022)
2022 arXiv
-
[28]
Mathematical Introduction to Deep Learning: Methods, Implementations, and Theory
Jentzen, A., Kuckuck, B., and von Wurstemberger, P. Mathematical Introduction to Deep Learning: Methods, Implementations, and Theory. arXiv:2310.20360 (2024)
2024 arXiv
-
[29]
On the existence of global minima and conver- gence analyses for gradient descent methods in the training of deep neural networks
Jentzen, A., and Riekert, A. On the existence of global minima and conver- gence analyses for gradient descent methods in the training of deep neural networks. arXiv: 2112.09684v1 (2021)
2021 arXiv
-
[31]
A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions
Jentzen, A., and Riekert, A. A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions. Z. Angew. Math. Phys. 73 , 5 (2022), Paper No. 188, 30
2022
-
[32]
Convergence analysis for gradient flows in the training of artificial neural networks with ReLU activation
Jentzen, A., and Riekert, A. Convergence analysis for gradient flows in the training of artificial neural networks with ReLU activation. J. Math. Anal. Appl. 517 , 2 (2023), Paper No. 126601, 43
2023
-
[33]
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructio ns of local minimizers in the training of artificial neural networks
Jentzen, A., and Riekert, A. Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructio ns of local minimizers in the training of artificial neural networks. To appear in SIAM/ASA J. Uncertain. Quantif. arXiv:2402.05155 (2024)
2024 arXiv
-
[34]
M., and Lee, J
Kakade, S. M., and Lee, J. Provably correct automatic subdifferentiation for qualified programs. In NeurIPS (2018)
2018
-
[35]
T., Riccietti, E., and Gribonval, R
Le, Q. T., Riccietti, E., and Gribonval, R. Does a sparse ReLU network training problem always admit an optimum? arXiv:2306.02666 (2023)
2023 arXiv
-
[36]
On correctness of automatic differentiation for non-differentiable functions
Lee, W., Yu, H., Rival, X., and Yang, H. On correctness of automatic differentiation for non-differentiable functions. In Advances in Neural Information Processing Systems (2020), H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, an d H. Lin, Eds., vol. 33, Curran Associates, In...
2020
-
[37]
S., and Tian, T
Lewis, A. S., and Tian, T. The structure of conservative gradient fields. SIAM J. Optim. 31 , 3 (2021), 2080–2083. 36
2021
-
[38]
Lu, L., Shin, Y., Su, Y., and Karniadakis, G. E. Dying ReLU and initialization: theory and numerical examples. Commun. Comput. Phys. 28 , 5 (2020), 1671–1706
2020
-
[39]
Introductory lectures on convex optimization , vol
Nesterov, Y. Introductory lectures on convex optimization , vol. 87 of Applied Optimiza- tion. Kluwer Academic Publishers, Boston, MA, 2004. A basic cour se
2004
-
[40]
Automatic differentiation in PyTorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. Automatic differentiation in PyTorch. https://openreview.net/forum?id=BJJsrmfCZ (2017)
2017
-
[41]
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G ., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., K ¨opf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., F ang, L., Bai, J., and Chintala, S. PyTorch: An Im...
2019 arXiv
-
[42]
Conservative parametric optimality and the ridge method for tame min-max problems
Pauwels, E. Conservative parametric optimality and the ridge method for tame min-max problems. Set-Valued Var. Anal. 31 , 3 (2023), [Paper No. 19], 24
2023
-
[43]
Topological properties of the set of functions generated by neural networks of fixed size
Petersen, P., Raslan, M., and Voigtlaender, F. Topological properties of the set of functions generated by neural networks of fixed size. Found. Comput. Math. 21 , 2 (2021), 375–444
2021
-
[44]
Mathematical theory of deep learning
Petersen, P., and Zech, J. Mathematical theory of deep learning. arXiv:2407.18384 (2024)
2024
-
[45]
J., Kale, S., and Kumar, S
Reddi, S. J., Kale, S., and Kumar, S. On the Convergence of Adam and Beyond. arXiv:1904.09237 (2019)
2019 arXiv
-
[46]
T., and Wets, R
Rockafellar, R. T., and Wets, R. J. B. Variational Analysis, vol. 317 of Grundlehren der Mathematischen Wissenschaften. Springer-Verlag, Ber lin, 1998
1998
-
[47]
An overview of gradient descent optimization algorithms
Ruder, S. An overview of gradient descent optimization algorithms. arXiv:1609.04747 (2017)
2017 arXiv
-
[48]
Spurious Local Minima are Common in Two-Layer ReLU Neural Networks
Safran, I., and Shamir, O. Spurious Local Minima are Common in Two-Layer ReLU Neural Networks. arXiv:1712.08968 (2017)
2017 arXiv
-
[49]
The gradient’s limit of a definable family of functions is a co nservative set-valued field
Schechtman, S. The gradient’s limit of a definable family of functions is a co nservative set-valued field. arXiv:2402.08272 (2024)
2024
-
[50]
Shin, Y., and Karniadakis, G. E. Trainability of ReLU networks and Data-dependent Initialization. J. Mach. Learn. Model. Comput 1 , 1 (2020), 39–74
2020
-
[51]
Optimization for deep learning: theory and algorithms
Sun, R. Optimization for deep learning: theory and algorithms. arXiv:1912.08957 (2019)
2019 arXiv
-
[52]
M., and Pascanu, R
Swirszcz, G., Czarnecki, W. M., and Pascanu, R. Local minima in training of neural networks. arXiv:1611.06310 (2017)
2017 arXiv
-
[53]
S., and Bruna, J
Venturi, L., Bandeira, A. S., and Bruna, J. Spurious valleys in one-hidden-layer neural network optimization landscapes. J. Mach. Learn. Res. 20 (2019), Paper No. 133, 34
2019
-
[54]
Approximation and gradient descent training with neural ne tworks
Welper, G. Approximation and gradient descent training with neural ne tworks. arXiv:2405.11696 (2024)
2024 arXiv
-
[55]
Approximation results for gradient flow trained neural netw orks
Welper, G. Approximation results for gradient flow trained neural netw orks. J. Mach. Learn. 3, 2 (2024), 107–175. 37
2024
-
[56]
Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
Xiao, N., Hu, X., Liu, X., and Toh, K.-C. Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees. arXiv:2305.03938 (2023)
2023 arXiv
-
[57]
Convergence Guarantees for Stochastic Subgradient Methods in Nonsmooth Nonconvex Optimization
Xiao, N., Hu, X., and Toh, K.-C. Convergence Guarantees for Stochastic Subgradient Methods in Nonsmooth Nonconvex Optimization. arXiv:2307.10053v1 (2023)
2023 arXiv
-
[58]
Zhang, Y., Li, Y., Zhang, Z., Luo, T., and Xu, Z.-Q. J. Embedding principle: a hierarchical structure of loss landscape of deep neural ne tworks. J. Mach. Learn. 1 , 1 (2022), 60–113. 38
2022
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.