Pith. sign in

REVIEW 2 major objections 2 minor 2 cited by

Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime

T0 review · 2 major / 2 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper claims that, in a teacher-student setting, the low-regularization limit of VarPro training for two-layer mean-field networks is a weighted ultra-fast diffusion equation whose solutions converge linearly to the teacher feature…

desk verdict The VarPro/ultra-fast diffusion identification is a clean, genuinely new result, but the advertised linear-rate guarantee for small regularization does not follow from the paper's own theorems. read the letter →

arxiv 2504.18208 v2 pith:YJBKOT33 submitted 2025-04-25 cs.LG math.OC

classification cs.LGmath.OC MSC 68T0749Q2235K55
keywords two-layerneuralnetworksmean-fieldtrainingvariableprojectiontwo-timescalegradientdescentWassersteinflowultra-fastdiffusionteacher-studentfeaturelearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Variable Projection (VarPro) is the two-timescale strategy that eliminates a two-layer network's outer weights and trains only the distribution of inner features, the part responsible for feature learning. The paper shows that, when the target signal is generated by a finite teacher measure through an injective feature map, the reduced training objective is an f-divergence between student and teacher feature distributions. In the vanishing-regularization limit, its Wasserstein gradient flow is a weighted ultra-fast diffusion equation, a nonlinear diffusion with a singular negative diffusivity. Under stated regularity assumptions, the paper proves that regularized VarPro dynamics converge to this diffusion as the regularization tends to zero, and that the diffusion itself converges linearly to the teacher feature distribution. The result turns a high-dimensional, non-convex training problem into an explicitly described PDE with quantitative convergence rates, going beyond the qualitative statements common in mean-field analysis.

What carries the argument

The load-bearing object is the reduced risk $$L^\lambda_f(\mu)=\min_u \frac{1}{\$\lambda$}R^\lambda_f(\mu,u),$$ which, under the teacher-student assumption, is an infimal convolution of an f-divergence and a maximum mean discrepancy. Its dual representation, $$L^\lambda_f(\mu)=\sup_{\$\alpha$\in $L^{2}$(\rho)}\left[\int_\$\Omega$(\Phi^\top\$\alpha$)\,d\bar\nu-\int_\$\Omega$ f^*(\Phi^\top\$\alpha$)\,d\mu-\frac{\$\lambda$}{2}\|\$\alpha$\|_{$L^{2}$(\rho)}^2\right],$$ makes the envelope theorem applicable and yields the velocity field $\nabla L^\lambda_f[\mu]=-\nabla(f^*(\Phi^\top \alpha^\lambda_f[\mu]))$, whose negative is the drift of the Wasserstein gradient flow $\partial_t\mu-\operatorname{div}(\mu\nabla L^\lambda_f[\mu])=0$. For $f(t)=|t|^r/(r-1)$, the unregularized limit of this field is $-\nabla(\bar{\mu}/\mu)^r$, and the continuity equation becomes the weighted ultra-fast diffusion of Eq. (35). To pass to the limit in the regularized flows, the paper imposes a source condition — $\partial f(d\bar\nu/d\mu^\lambda_t)$ lies in the RKHS $H$, compactly embedded in $C^1(\Omega)$ — which keeps the dual variable bounded and gives compactness of the velocity fields.

What would settle it

Take a smooth teacher density on the torus, solve the weighted ultra-fast diffusion Eq. (35) numerically, run VarPro with features that make $\Phi^\star$ injective for a sequence of $\lambda$ values tending to $0$, and measure $\sup_{t\in[0,T]}W_2(\mu^\lambda_t,\mu^0_t)$; Theorem 5 predicts this tends to $0$, so a nonvanishing plateau would refute the diffusion limit. A complementary test: replace the teacher by an atomic measure, where the log-density is unbounded; the predicted linear rate should fail, exposing the boundary of the result.

Watch

Extended reading notes

Core claim

Under Assumption 1, in which $Y=\Phi^\star \bar{\nu}$ for a finite measure $\bar{\nu}$ and $\Phi^\star$ is injective, the unregularized reduced risk with $f(t)=|t|^r/(r-1)$ is $$$L^{0}$_r(\mu)=\frac{\|\bar{\nu}\|_{\mathrm{TV}}^r}{r-1}\int_\$\Omega$ \left(\frac{d\bar{\mu}}{d\mu}\right)^r d\mu ,$$ an f-divergence (for $r=2$, a $\chi^2$-divergence) between the student distribution $\mu$ and the teacher distribution $\bar{\mu}=\bar{\nu}/\|\bar{\nu}\|_{\mathrm{TV}}$. Its Wasserstein gradient flow is $$\partial_t \mu_t = -\|\bar{\nu}\|_{\mathrm{TV}}^r \operatorname{div}\left(\mu_t \nabla\left(\frac{\bar{\mu}}{\mu_t}\right)^r\right),$$ a weighted ultra-fast diffusion equation: the exponent is negative, so the diffusivity is singular where the student density vanishes. The central result is Theorem 5, which identifies the limit of the regularized dynamics: as $\lambda\to 0^+$, gradient flows of $L^r_\lambda$ converge locally uniformly in time to this diffusion. Since solutions of the diffusion converge linearly to $\bar{\mu}$, VarPro training in the low-regularization regime inherits a quantitative, essentially explicit description of feature learning.

Load-bearing premise

That the target signal is exactly generated by a finite teacher measure through an injective feature map; without exact representability or injectivity, the reduced risk is not a divergence and the ultra-fast diffusion description collapses (Theorem 5 additionally needs a source condition with the RKHS $H$ compactly embedded in $C^1(\Omega)$).

Editorial extensions

If this is right

  • In the vanishing-regularization limit, the learned feature distribution converges to the teacher's at a linear (exponential) rate in $L^2$, and the rate constant is controlled by the log-density ratio at initialization rather than by the ambient dimension.
  • At any fixed $\lambda>0$, the VarPro gradient flow converges to the minimizer of the reduced risk at an algebraic rate that, under the stated boundedness assumption, is independent of $\lambda$.
  • As $\lambda\to 0^+$, regularized dynamics approach the ultra-fast diffusion on every finite time interval, so low-regularization training runs should be accurately described by the PDE rather than by a kernel (NTK) linearization.
  • With a small enough step relative to $\lambda$, ordinary two-timescale gradient descent reproduces the VarPro dynamics; in low-regularization regimes VarPro itself keeps converging where the two-timescale scheme does not.
  • On bounded convex domains or tori, whose Poincaré constants are dimension-free, the linear rate in the diffusion limit does not degrade with dimension, indicating feature learning that avoids the curse of dimensionality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the diffusion description is exact, the proved convergence rate depends on initialization only through the log-density ratio $\|\log(\bar{\mu}/\mu_0)\|_\infty$; a testable design principle is to initialize feature distributions so that this ratio is bounded and small.
  • The infimal-convolution formula suggests a threshold transition at $\lambda$ comparable to the spectrum of the tangent kernel $K_\mu$: above it the flow behaves like an MMD gradient flow with algebraic rates, below it like an f-divergence flow with linear rates — a prediction that could be measured on other architectures.
  • Applying the same last-layer variable projection to deep networks, as the ResNet experiment does, suggests that feature learning in deep training might admit a similar effective diffusion description at the last layer; this is not covered by the paper because depth breaks the linear separability on which the proof relies.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper studies the Variable Projection (VarPro) / two-timescale gradient flow for training mean-field two-layer neural networks with square loss. Under a teacher-student assumption (Y = Phi^* nu_bar with Phi^* injective), the unregularized reduced risk L^0_f is shown to equal the f-divergence D_f(nu_bar | mu) (Eq. (17)). For the family f(t)=|t|^r/(r-1), the Wasserstein gradient flow at lambda=0 is identified as a weighted ultra-fast diffusion equation (Eq. (35)); relying on Iacobelli–Patacchini–Santambrogio, the paper states linear convergence of this diffusion to the teacher feature distribution (Theorem 3). At fixed lambda>0, an algebraic convergence rate is claimed under additional assumptions (Theorem 4). The main new result is Theorem 5, which asserts that, under a source condition and compact embedding of the RKHS H in C^1, the regularized gradient flows converge locally uniformly in time to the ultra-fast diffusion as lambda -> 0^+. Numerical experiments on S^1 and CIFAR10 are presented as supporting evidence.

Significance. If the identification is correct, the paper offers a clean, parameter-free connection between feature learning in two-layer networks and a weighted ultra-fast diffusion PDE: the lambda=0 reduced risk is exactly a scaled reverse f-divergence, and the limiting PDE is a known object with external well-posedness and convergence results. This is a valuable conceptual contribution and a rigorous stability result (Theorem 5) for the vanishing-regularization limit. The related-work discussion is careful and positions the paper well against KALE, DrMMD, and MMD flows. However, the advertised quantitative guarantee for the regularized VarPro dynamics in the low-regularization regime does not follow from the stated theorems; this gap concerns the paper's central narrative and must be addressed before publication.

major comments (2)
  1. [§5.1, Theorem 4; §5.2 setting f(t)=|t|^r/(r-1)] Theorem 4 does not apply to the regularization family used in Theorem 5. Theorem 4 requires "min f = f(\bar m) = 0" with \bar m = \bar\nu(\Omega) > 0. For f(t)=|t|^r/(r-1), the minimum is 0 at t=0 while f(\bar m) = \bar m^r/(r-1) > 0, so the hypothesis fails whenever the teacher measure has positive mass. Consequently, for every lambda > 0 and for this f, the paper establishes no convergence rate of the VarPro gradient flow to the teacher distribution; the only rigorous rate is for the lambda=0 diffusion (Theorem 3). The numerical linear rates in Figures 5 and 6 are therefore not consequences of the stated theorems, and the claims connecting them to Theorem 4 or to a transfer of the lambda=0 rate need to be corrected or replaced by a genuinely new argument.
  2. [§6.1, regularization f_u = 1/2|t-1|^2] The "unbiased" quadratic regularization f_u(t) = 1/2 |t-1|^2 used in the numerical comparison is not of the form f(t)=|t|^r/(r-1) required by Theorem 5. Since f_u differs from f_2(t)=t^2 by a constant depending on \bar\nu, the gradient-flow identification may still hold up to an additive constant, but this equivalence is not stated. The text should clarify which theorem is being tested when f_u is used, and whether Theorem 5's assumptions are intended to cover it.
minor comments (2)
  1. [Abstract and Conclusion] The phrase "provable convergence rates for the sampling of a teacher feature distribution" and the conclusion's claim that low-regularization VarPro converges "at a linear rate (Theorem 3)" overstate what is proven. These statements should be qualified to indicate that the linear rate is proven for the lambda=0 diffusion, while for lambda>0 only local-in-time approximation to that diffusion is established.
  2. [§1.1, abstract wording] The word "sampling" in the abstract is inaccurate because the dynamics is a deterministic gradient flow; "recovering" or "approximating" would be more appropriate.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the teacher-student reduction and ultra-fast-diffusion identification are self-contained; the linear rate is imported from an external theorem ([47]).

full rationale

The paper's core derivation is not circular. Eq. (17), L0_f = D_f(nu_bar | mu), follows from Assumption 1 plus injectivity of Phi*: feasibility forces nu = nu_bar in Eq. (16), so no target conclusion is hidden in the definition. Eq. (35) is the formal Wasserstein gradient flow of L0_r: the first variation of ||nu_bar||^r/(r-1) * integral (mu_bar/mu)^r dmu is -||nu_bar||^r (mu_bar/mu)^r, so the continuity equation is derived, not assumed. Theorem 3, the linear convergence of the weighted ultra-fast diffusion, is quoted from Iacobelli-Patacchini-Santambrogio [47], an external source with stated assumptions (bounded log-densities) and no overlap with the present authors; citing it as an external theorem is legitimate support. Theorem 4's algebraic rate is conditional on a uniform H^{-1} bound on mu_t - nu_bar/mbar, which the paper argues follows from bounded log-densities; this is an extra regularity hypothesis on the flow, not a restatement of convergence. Theorem 5's source condition, partial_f(dnu_bar/dmu_lambda_t) in H, is an assumption on the dynamics, and the proof identifies the cluster point by a product-limit argument and uniqueness from [47]; no equation is reused as its own conclusion. The paper itself flags in Section 6.1 that Theorem 5 'says nothing about the long time behavior of the dynamic,' so the composition with Theorem 3 is explicitly localized; any over-strong phrasing in the conclusion is a correctness or scope concern, not circularity. The self-citations [8], [9], and [88] are contextual remarks about ResNets, deep feature learning, and conditioning; they do not carry the teacher-student reduction or the ultra-fast-diffusion claim. Therefore no circular step can be exhibited, and the honest finding is a score of 0.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The ledger shows the central claim rests on a stated teacher-student modeling assumption and on standard regularity conditions; the two ad hoc assumptions (source condition, uniform dot-H^{-1} bound) are the price of the quantitative theorems. Crucially, no parameter is fitted to data to get the convergence rates.

assumptions (6)
  • domain assumption Assumption 1: Y = Phi*nu_bar for some finite measure nu_bar, and Phi* is injective.
    This is the teacher-student premise. It makes L0_f equal the f-divergence D_f(nu_bar|mu) via Eq. (17) and provides the unique teacher measure; without it the central ultra-fast diffusion identification fails. Used throughout Sections 4 and 5.
  • domain assumption Assumption 2: the regularizer f is nonnegative, strictly convex, superlinear.
    Ensures existence and uniqueness of the partial minimizer u_lambda_f[mu] (Lemma 2.1) and the duality representations used in the gradient flow derivation.
  • domain assumption Assumption 3: phi in L^2(rho, C^0), with C^1 and C^{1,1} strengthenings in Lemmas 4.1 and 4.2.
    Regularity of the feature map is needed for the continuity equation formulation, the chain rule, and the geodesic semiconvexity arguments.
  • domain assumption mu0 and mu_bar are absolutely continuous with bounded log-densities (Theorems 2 and 3, imported from [47]).
    Needed for well-posedness of the ultra-fast diffusion and for the linear-rate bound; the rate is exponentially bad in the log-density ratio, and atomic teachers are excluded.
  • ad hoc to paper Source condition for Theorem 5: partial_f(d_nu_bar/d_mu_lambda_t) in H uniformly over lambda, with H compactly embedded in C^1(Omega).
    Unverified regularity condition introduced to keep the dual variable alpha_lambda bounded as lambda->0 (Lemma 5.1) and to pass to the limit in the weak formulation. Not satisfied for generic non-smooth activations such as ReLU, which limits the scope of the lambda->0 stability theorem relative to the experiments.
  • ad hoc to paper Theorem 4: uniform boundedness of ||nu_bar/m_bar - mu_t|| in dot-H^{-1}(mu_t) along the flow.
    Assumed to obtain the algebraic decay L_lambda(mu_t) <= (L_lambda(mu_0)^{-1} + Ct)^{-1}; the paper argues it follows from bounded log-densities, but does not prove it along the lambda>0 flow.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime." pith.science (2026). https://pith.science/paper/YJBKOT33

@misc{pith2026250418208,
  author       = {Pith},
  title        = {Pith review of: Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YJBKOT33}},
  note         = {Machine review of arXiv:2504.18208}
}
read the original abstract

We study the convergence of gradient methods for the training of mean-field single-hidden-layer neural networks with square loss. For this high-dimensional and non-convex optimization problem, most known convergence results are either qualitative or rely on a neural tangent kernel analysis where nonlinear representations of the data are fixed. Using that this problem belongs to the class of separable nonlinear least squares problems, we consider here a Variable Projection (VarPro) or two-timescale learning algorithm, thereby eliminating the linear variables and reducing the learning problem to the training of nonlinear features. In a teacher-student scenario, we show such a strategy enables provable convergence rates for the sampling of a teacher feature distribution. Precisely, in the limit where the regularization strength vanishes, we show that the dynamic of the feature distribution corresponds to a weighted ultra-fast diffusion equation. Recent results on the asymptotic behavior of such PDEs then give quantitative guarantees for the convergence of the learned feature distribution.

Figures

Figures reproduced from arXiv: 2504.18208 by the authors.

Figure 1
Figure 1. Left: density of the teacher distributions [PITH_FULL_IMAGE:figures/full_fig_p029_1.png] view at source ↗
Figure 2
Figure 2. Left: Solution µt to the ultra-fast diffusion Eq. (35) equation with exponent r = 2 and weights µγ, γ = 100. Right: Evolution of the feature distribution learned by gradient descent on a SHL of width M = 1024 for the minimization the reduced risk Lˆλ f with regularization function fb : t 7→ 1 2 t 2 and λ = 10−4 (c.f. Eqs. (41) and (42)). The density is obtained by convolving the empirical feature distribution ˆµ wit… view at source ↗
Figure 3
Figure 3. Evolution of the reduced risk Lˆλ f (Eq. (42)) along iterations of gradient descent for a SHL of width M ∈ {32, 128, 512, 1024}. The regularization strength is λ = 10−3 and the regularization function is either fb : t 7→ 1 2 t 2 (left) or fu : t 7→ 1 2 |t − 1| 2 (right). Plots are averages over 6 independent runs. number of features corresponds to a better discretization. These plots also show that gradient descent … view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Evolution of the MMD distance to the teacher distribution and to the diffusion dynamic [PITH_FULL_IMAGE:figures/full_fig_p032_4.png]
Figure 5
Figure 5. Figure 5: Evolution of the MMD distance to the teacher distribution and to the diffusion dynamic [PITH_FULL_IMAGE:figures/full_fig_p032_5.png]
Figure 6
Figure 6. Figure 6: Gradient descent over the reduced risk (Eq. (42)) for a SHL of width [PITH_FULL_IMAGE:figures/full_fig_p033_6.png]
Figure 7
Figure 7. Figure 7: VarPro (Eq. (42), plain lines) and two-timescale gradient descent (Eq. (44), dashed lines) [PITH_FULL_IMAGE:figures/full_fig_p035_7.png]
Figure 8
Figure 8. Figure 8: Evolution of the training risk 1 λRˆλ B (Eq. (45)) along training for different batch sizes and different optimization methods. Plots are averages of the risk associated to each mini-batch encountered during one pass. VarPro corresponds to Eq. (46). Comparison with oth…
Figure 9
Figure 9. Figure 9: Evolution of the top-1 accuracy along training for different batch sizes and different [PITH_FULL_IMAGE:figures/full_fig_p038_9.png]
Figure 10
Figure 10. Figure 10: Left: scatter plot the empirical teacher distribution ¯µ [PITH_FULL_IMAGE:figures/full_fig_p048_10.png]
Figure 11
Figure 11. Figure 11: Evolution of the reduced risk along iterations of gradient descent for a RBF neural [PITH_FULL_IMAGE:figures/full_fig_p049_11.png]
Figure 12
Figure 12. Figure 12: Evolution of the MMD distance to the teacher distribution ¯µ [PITH_FULL_IMAGE:figures/full_fig_p049_12.png]
Figure 13
Figure 13. Figure 13: Evolution of the MMD distance to the teacher distribution ¯µ [PITH_FULL_IMAGE:figures/full_fig_p050_13.png]
Figure 14
Figure 14. Figure 14: Gradient descent over the reduced risk (Eq. (42)) for a RBF neural network (Eq. (50)) of [PITH_FULL_IMAGE:figures/full_fig_p051_14.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How are linear representations learned? Exact solutions to the dynamics of abstraction

    cs.LG 2026-07 conditional novelty 8.0 of 10

    Exact solutions show abstraction is set by input/target geometry, rises with depth, peaks under small init, and is attenuated by nonlinearities—improving LLM probes via GELU ablation.

  2. Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures

    cs.LG 2025-06 conditional novelty 8.0 of 10

    For well-separated Gaussian mixtures, over-parameterized gradient EM with n=Omega(m log m) components converges globally to the ground truth, the first such result beyond m=2.

Reference graph

Works this paper leans on

94 extracted references · 72 canonical work pages · cited by 2 Pith papers

  1. [62]

    Wasserstein Gradient Flows for Moreau Envelopes of f-Divergences in Reproducing Kernel Hilbert Spaces

    Sebastian Neumayer, Viktor Stein, and Gabriele Steidl. “Wasserstein Gradient Flows for Moreau Envelopes of f-Divergences in Reproducing Kernel Hilbert Spaces”. In: arXiv preprint arXiv:2402.04613 (2024)

  2. [1]

    A convergence theory for deep learning via over-parameterization

    Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. “A convergence theory for deep learning via over-parameterization”. In: International Conference on Machine Learning. PMLR. 2019, pp. 242–252

  3. [2]

    Gradient flows: in metric spaces and in the space of probability measures

    Luigi Ambrosio, Nicola Gigli, and Giuseppe Savar´ e. “Gradient flows: in metric spaces and in the space of probability measures”. In: Lectures in mathematics ETH Z¨ urich(2008)

  4. [3]

    C´ ecile An´ e et al.Sur les in´ egalit´ es de Sobolev logarithmiques. Vol. 10. Soci´ et´ e math´ ematique de France Paris, 2000

  5. [4]

    Maximum mean discrepancy gradient flow

    Michael Arbel et al. “Maximum mean discrepancy gradient flow”. In: Advances in Neural Information Processing Systems 32 (2019)

  6. [5]

    Breaking the curse of dimensionality with convex neural networks

    Francis Bach. “Breaking the curse of dimensionality with convex neural networks”. In: The Journal of Machine Learning Research 18.1 (2017), pp. 629–681

  7. [6]

    Gradient descent on infinitely wide neural networks: Global convergence and generalization

    Francis Bach and L´ ena¨ ıc Chizat. “Gradient descent on infinitely wide neural networks: Global convergence and generalization”. In: arXiv preprint arXiv:2110.08084 (2021)

  8. [7]

    Multiple kernel learning, conic duality, and the SMO algorithm

    Francis R Bach, Gert RG Lanckriet, and Michael I Jordan. “Multiple kernel learning, conic duality, and the SMO algorithm”. In: Proceedings of the twenty-first international conference on Machine learning . 2004, p. 6

Show all 94 references
  1. [8]

    On global convergence of ResNets: From finite to infinite width using linear parameterization

    Rapha¨ el Barboni, Gabriel Peyr´ e, and Fran¸ cois-Xavier Vialard. “On global convergence of ResNets: From finite to infinite width using linear parameterization”. In: Advances in Neural Information Processing Systems 35 (2022), pp. 16385–16397

  2. [9]

    Understanding the training of infinitely deep and wide resnets with conditional optimal transport

    Rapha¨ el Barboni, Gabriel Peyr´ e, and Fran¸ cois-Xavier Vialard. “Understanding the training of infinitely deep and wide resnets with conditional optimal transport”. In: arXiv preprint arXiv:2403.12887 (2024)

  3. [10]

    Modern regularization methods for inverse problems

    Martin Benning and Martin Burger. “Modern regularization methods for inverse problems”. In: Acta numerica 27 (2018), pp. 1–111

  4. [11]

    Learning time-scales in two-layers neural networks

    Rapha¨ el Berthier, Andrea Montanari, and Kangjie Zhou. “Learning time-scales in two-layers neural networks”. In: Foundations of Computational Mathematics (2024), pp. 1–84

  5. [12]

    On learning gaussian multi-index models with gradient flow

    Alberto Bietti, Joan Bruna, and Loucas Pillaud-Vivien. “On learning gaussian multi-index models with gradient flow”. In: arXiv preprint arXiv:2310.19793 (2023)

  6. [13]

    Stochastic approximation: a dynamical systems viewpoint

    Vivek S Borkar. Stochastic approximation: a dynamical systems viewpoint . Vol. 9. Springer, 2008

  7. [14]

    Stochastic approximation with two time scales

    Vivek S Borkar. “Stochastic approximation with two time scales”. In: Systems & Control Letters 29.5 (1997), pp. 291–294. 39

  8. [15]

    Optimization methods for large-scale machine learning

    L´ eon Bottou, Frank E Curtis, and Jorge Nocedal. “Optimization methods for large-scale machine learning”. In: SIAM review 60.2 (2018), pp. 223–311

  9. [16]

    On the global convergence of Wasserstein gradient flow of the Coulomb discrepancy

    Siwan Boufad` ene and Fran¸ cois-Xavier Vialard. “On the global convergence of Wasserstein gradient flow of the Coulomb discrepancy”. In: arXiv preprint arXiv:2312.00800 (2023)

  10. [17]

    Quantization of measures and gra- dient flows: a perturbative approach in the 2-dimensional case

    Emanuele Caglioti, Fran¸ cois Golse, and Mikaela Iacobelli. “Quantization of measures and gra- dient flows: a perturbative approach in the 2-dimensional case”. In:arXiv preprint arXiv:1607.01198 (2016)

  11. [18]

    (De)-regularized Maximum Mean Discrepancy Gradient Flow

    Zonghao Chen et al. “(De)-regularized Maximum Mean Discrepancy Gradient Flow”. In: arXiv preprint arXiv:2409.14980 (2024)

  12. [19]

    Analysis of langevin monte carlo from poincare to log-sobolev

    Sinho Chewi et al. “Analysis of langevin monte carlo from poincare to log-sobolev”. In: Foun- dations of Computational Mathematics (2024), pp. 1–51

  13. [20]

    SVGD as a kernelized Wasserstein gradient flow of the chi-squared diver- gence

    Sinho Chewi et al. “SVGD as a kernelized Wasserstein gradient flow of the chi-squared diver- gence”. In: Advances in Neural Information Processing Systems 33 (2020), pp. 2098–2109

  14. [21]

    Mean-Field Langevin Dynamics: Exponential Convergence and Annealing

    L´ ena¨ ıc Chizat. “Mean-Field Langevin Dynamics: Exponential Convergence and Annealing”. In: Transactions on Machine Learning Research (2022)

  15. [22]

    On Lazy Training in Differentiable Pro- gramming

    Lenaic Chizat, Edouard Oyallon, and Francis Bach. “On Lazy Training in Differentiable Pro- gramming”. In: NeurIPS 2019-33rd Conference on Neural Information Processing Systems . 2019, pp. 2937–2947

  16. [23]

    On the Global Convergence of Gradient Descent for Over- parameterized Models using Optimal Transport

    L´ ena¨ ıc Chizat and Francis Bach. “On the Global Convergence of Gradient Descent for Over- parameterized Models using Optimal Transport”. In: Advances in Neural Information Pro- cessing Systems 31 (2018), pp. 3036–3046

  17. [24]

    Approximation by superpositions of a sigmoidal function

    George Cybenko. “Approximation by superpositions of a sigmoidal function”. In: Mathematics of control, signals and systems 2.4 (1989), pp. 303–314

  18. [25]

    Exact reconstruction using Beurling minimal ex- trapolation

    Yohann De Castro and Fabrice Gamboa. “Exact reconstruction using Beurling minimal ex- trapolation”. In: Journal of Mathematical Analysis and applications 395.1 (2012), pp. 336– 354

  19. [26]

    High-dimensional data analysis: The curses and blessings of dimen- sionality

    David L Donoho et al. “High-dimensional data analysis: The curses and blessings of dimen- sionality”. In: AMS math challenges lecture 1.2000 (2000), p. 32

  20. [27]

    Gradient descent finds global minima of deep neural networks

    Simon Du et al. “Gradient descent finds global minima of deep neural networks”. In: Inter- national Conference on Machine Learning . PMLR. 2019, pp. 1675–1685

  21. [28]

    Exact support recovery for sparse spikes deconvolution

    Vincent Duval and Gabriel Peyr´ e. “Exact support recovery for sparse spikes deconvolution”. In: Foundations of Computational Mathematics 15.5 (2015), pp. 1315–1355

  22. [29]

    On the rate of convergence in Wasserstein distance of the empirical measure

    Nicolas Fournier and Arnaud Guillin. “On the rate of convergence in Wasserstein distance of the empirical measure”. In: Probability theory and related fields 162.3 (2015), pp. 707–738

  23. [30]

    Global convergence in training large-scale transformers

    Cheng Gao et al. “Global convergence in training large-scale transformers”. In: Advances in Neural Information Processing Systems 37 (2024), pp. 29213–29284

  24. [31]

    When do neural networks outperform kernel methods?

    Behrooz Ghorbani et al. “When do neural networks outperform kernel methods?” In: Advances in Neural Information Processing Systems 33 (2020), pp. 14820–14830

  25. [32]

    KALE flow: A relaxed KL gradient flow for probabilities with disjoint support

    Pierre Glaser, Michael Arbel, and Arthur Gretton. “KALE flow: A relaxed KL gradient flow for probabilities with disjoint support”. In: Advances in Neural Information Processing Systems 34 (2021), pp. 8018–8031

  26. [33]

    Separable nonlinear least squares: the variable projection method and its applications

    Gene H Golub and Victor Pereyra. “Separable nonlinear least squares: the variable projection method and its applications”. In: Inverse problems 19.2 (2003), R1. 40

  27. [34]

    The differentiation of pseudo-inverses and nonlinear least squares problems whose variables separate

    Gene H Golub and Victor Pereyra. “The differentiation of pseudo-inverses and nonlinear least squares problems whose variables separate”. In: SIAM Journal on numerical analysis 10.2 (1973), pp. 413–432

  28. [35]

    Deep Learning

    Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep Learning. http://www.deeplearningbook. org. MIT Press, 2016

  29. [36]

    A kernel two-sample test

    Arthur Gretton et al. “A kernel two-sample test”. In: The Journal of Machine Learning Research 13.1 (2012), pp. 723–773

  30. [37]

    Shampoo: Preconditioned stochastic tensor optimization

    Vineet Gupta, Tomer Koren, and Yoram Singer. “Shampoo: Preconditioned stochastic tensor optimization”. In: International Conference on Machine Learning . PMLR. 2018, pp. 1842– 1850

  31. [38]

    Ordinary differential equations

    Jack K Hale. Ordinary differential equations. Courier Corporation, 2009

  32. [39]

    Deep residual learning for image recognition

    Kaiming He et al. “Deep residual learning for image recognition”. In: Proceedings of the IEEE conference on computer vision and pattern recognition . 2016, pp. 770–778

  33. [40]

    Generative Sliced MMD Flows with Riesz Kernels

    Johannes Hertrich et al. “Generative Sliced MMD Flows with Riesz Kernels”. In: The Twelfth International Conference on Learning Representations

  34. [41]

    Wasserstein gradient flows of the discrepancy with distance kernel on the line

    Johannes Hertrich et al. “Wasserstein gradient flows of the discrepancy with distance kernel on the line”. In: International Conference on Scale Space and Variational Methods in Computer Vision. Springer. 2023, pp. 431–443

  35. [42]

    Wasserstein steepest descent flows of discrepancies with Riesz ker- nels

    Johannes Hertrich et al. “Wasserstein steepest descent flows of discrepancies with Riesz ker- nels”. In: Journal of Mathematical Analysis and Applications 531.1 (2024), p. 127829

  36. [43]

    ODEPACK, a systemized collection of ODE solvers

    Alan C Hindmarsh. “ODEPACK, a systemized collection of ODE solvers”. In: Scientific com- puting (1983)

  37. [44]

    Kernel methods in machine learning

    Thomas Hofmann, Bernhard Sch¨ olkopf, and Alexander J Smola. “Kernel methods in machine learning”. In: (2008)

  38. [45]

    Mean-field Langevin dynamics and energy landscape of neural networks

    Kaitong Hu et al. “Mean-field Langevin dynamics and energy landscape of neural networks”. In: Annales de l’Institut Henri Poincare (B) Probabilites et statistiques . Vol. 57. 4. Institut Henri Poincar´ e. 2021, pp. 2043–2065

  39. [46]

    Asymptotic analysis for a very fast diffusion equation arising from the 1D quantization problem

    Mikaela Iacobelli. “Asymptotic analysis for a very fast diffusion equation arising from the 1D quantization problem”. In: Discrete and Continuous Dynamical Systems 39.9 (2019), pp. 4929–4943

  40. [47]

    Weighted ultrafast diffusion equations: from well-posedness to long-time behaviour

    Mikaela Iacobelli, Francesco S Patacchini, and Filippo Santambrogio. “Weighted ultrafast diffusion equations: from well-posedness to long-time behaviour”. In: Archive for Rational Mechanics and Analysis 232 (2019), pp. 1165–1206

  41. [48]

    A note on convergence of solu- tions of total variation regularized linear inverse problems

    Jos´ e A Iglesias, Gwenael Mercier, and Otmar Scherzer. “A note on convergence of solu- tions of total variation regularized linear inverse problems”. In: Inverse Problems 34.5 (2018), p. 055011

  42. [49]

    Neural tangent kernel: Convergence and generalization in neural networks

    Arthur Jacot, Franck Gabriel, and Cl´ ement Hongler. “Neural tangent kernel: Convergence and generalization in neural networks”. In: Advances in Neural Information Processing Systems 31 (2018)

  43. [50]

    The variational formulation of the Fokker–Planck equation

    Richard Jordan, David Kinderlehrer, and Felix Otto. “The variational formulation of the Fokker–Planck equation”. In: SIAM journal on mathematical analysis 29.1 (1998), pp. 1–17. 41

  44. [51]

    Radial basis function neural network training using variable projection and fuzzy means

    Despina Karamichailidou et al. “Radial basis function neural network training using variable projection and fuzzy means”. In: Neural Computing and Applications 36.33 (2024), pp. 21137– 21151

  45. [52]

    Learning multiple layers of features from tiny im- ages

    Alex Krizhevsky, Geoffrey Hinton, et al. “Learning multiple layers of features from tiny im- ages”. In: (2009)

  46. [53]

    Learning the kernel matrix with semidefinite programming

    Gert RG Lanckriet et al. “Learning the kernel matrix with semidefinite programming”. In: Journal of Machine learning research 5.Jan (2004), pp. 27–72

  47. [54]

    Wide neural networks of any depth evolve as linear models under gradient descent

    Jaehoon Lee et al. “Wide neural networks of any depth evolve as linear models under gradient descent”. In: Advances in neural information processing systems 32 (2019), pp. 8572–8583

  48. [55]

    Optimal entropy-transport prob- lems and a new Hellinger–Kantorovich distance between positive measures

    Matthias Liero, Alexander Mielke, and Giuseppe Savar´ e. “Optimal entropy-transport prob- lems and a new Hellinger–Kantorovich distance between positive measures”. In: Inventiones mathematicae 211.3 (2018), pp. 969–1117

  49. [56]

    On the linearity of large non-linear models: when and why the tangent kernel is constant

    Chaoyue Liu, Libin Zhu, and Mikhail Belkin. “On the linearity of large non-linear models: when and why the tangent kernel is constant”. In: Advances in Neural Information Processing Systems 33 (2020)

  50. [57]

    Leveraging the two-timescale regime to demonstrate convergence of neural networks

    Pierre Marion and Rapha¨ el Berthier. “Leveraging the two-timescale regime to demonstrate convergence of neural networks”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 64996–65029

  51. [58]

    Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit

    Song Mei, Theodor Misiakiewicz, and Andrea Montanari. “Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit”. In:Conference on Learning Theory. PMLR. 2019, pp. 2388–2464

  52. [59]

    Universal Kernels

    Charles A Micchelli, Yuesheng Xu, and Haizhang Zhang. “Universal Kernels.” In: Journal of Machine Learning Research 7.12 (2006)

  53. [60]

    Envelope theorems for arbitrary choice sets

    Paul Milgrom and Ilya Segal. “Envelope theorems for arbitrary choice sets”. In: Econometrica 70.2 (2002), pp. 583–601

  54. [61]

    Kernel mean embedding of distributions: A review and beyond

    Krikamol Muandet et al. “Kernel mean embedding of distributions: A review and beyond”. In: Foundations and Trends in Machine Learning 10.1-2 (2017), pp. 1–141

  55. [63]

    Train like a (Var) Pro: Efficient training of neural networks with variable projection

    Elizabeth Newman et al. “Train like a (Var) Pro: Efficient training of neural networks with variable projection”. In: SIAM Journal on Mathematics of Data Science 3.4 (2021), pp. 1041– 1066

  56. [64]

    Convex analysis of the mean field langevin dynamics

    Atsushi Nitanda, Denny Wu, and Taiji Suzuki. “Convex analysis of the mean field langevin dynamics”. In: International Conference on Artificial Intelligence and Statistics. PMLR. 2022, pp. 9741–9757

  57. [65]

    Separable least squares, variable projection, and the Gauss-Newton algorithm

    MR Osborne. “Separable least squares, variable projection, and the Gauss-Newton algorithm”. In: Electronic Transactions on Numerical Analysis 28.2 (2007), pp. 1–15

  58. [66]

    Stochastic processes and applications

    Grigorios A Pavliotis. “Stochastic processes and applications”. In: Texts in applied mathemat- ics 60 (2014)

  59. [67]

    An optimal Poincar´ e inequality for convex do- mains

    Lawrence E Payne and Hans F Weinberger. “An optimal Poincar´ e inequality for convex do- mains”. In: Archive for Rational Mechanics and Analysis 5.1 (1960), pp. 286–292. 42

  60. [68]

    Variable projections neural network training

    V´ ıctor Pereyra, Godela Scherer, and F Wong. “Variable projections neural network training”. In: Mathematics and Computers in Simulation 73.1-4 (2006), pp. 231–243

  61. [69]

    Duality and stability in extremum problems involving convex functions

    Ralph Rockafellar. “Duality and stability in extremum problems involving convex functions”. In: Pacific Journal of Mathematics 21.1 (1967), pp. 167–187

  62. [70]

    Integrals which are convex functionals

    Ralph Rockafellar. “Integrals which are convex functionals”. In: Pacific journal of mathematics 24.3 (1968), pp. 525–539

  63. [71]

    Integrals which are convex functionals. II

    Ralph Rockafellar. “Integrals which are convex functionals. II”. In: Pacific journal of mathe- matics 39.2 (1971), pp. 439–469

  64. [72]

    Global convergence of neuron birth-death dynamics

    Grant Rotskoff et al. “Global convergence of neuron birth-death dynamics”. In: International Conference on Machine Learning. 2019

  65. [73]

    A Course in the Calculus of Variations: Optimization, Regularity, and Modeling

    Filippo Santambrogio. A Course in the Calculus of Variations: Optimization, Regularity, and Modeling. Springer Nature, 2023

  66. [74]

    {Euclidean, metric, and Wasserstein} gradient flows: an overview

    Filippo Santambrogio. “ {Euclidean, metric, and Wasserstein} gradient flows: an overview”. In: Bulletin of Mathematical Sciences 7 (2017), pp. 87–154

  67. [75]

    Optimal transport for applied mathematicians

    Filippo Santambrogio. “Optimal transport for applied mathematicians”. In: Birk¨ auser, NY 55.58-63 (2015), p. 94

  68. [76]

    Learning with kernels: support vector machines, regularization, optimization, and beyond

    Bernhard Sch¨ olkopf and Alexander J Smola. Learning with kernels: support vector machines, regularization, optimization, and beyond . 2002

  69. [77]

    Equivalence of distance-based and RKHS-based statistics in hypothesis testing

    Dino Sejdinovic et al. “Equivalence of distance-based and RKHS-based statistics in hypothesis testing”. In: The annals of statistics (2013), pp. 2263–2291

  70. [78]

    Mean field analysis of neural networks: A central limit theorem

    Justin Sirignano and Konstantinos Spiliopoulos. “Mean field analysis of neural networks: A central limit theorem”. In: Stochastic Processes and their Applications 130.3 (2020), pp. 1820– 1852

  71. [79]

    Separable non-linear least-squares minimization-possible improvements for neural net fitting

    Jonas Sjoberg and Mats Viberg. “Separable non-linear least-squares minimization-possible improvements for neural net fitting”. In:Neural networks for signal processing VII. Proceedings of the 1997 IEEE signal processing society workshop . IEEE. 1997, pp. 345–354

  72. [80]

    Universality, Char- acteristic Kernels and RKHS Embedding of Measures

    Bharath K Sriperumbudur, Kenji Fukumizu, and Gert RG Lanckriet. “Universality, Char- acteristic Kernels and RKHS Embedding of Measures.” In: Journal of Machine Learning Research 12.7 (2011)

  73. [81]

    Support vector machines

    Ingo Steinwart and Andreas Christmann. Support vector machines. Springer Science & Busi- ness Media, 2008

  74. [82]

    Mercer’s theorem on general domains: On the interaction between measures, kernels, and RKHSs

    Ingo Steinwart and Clint Scovel. “Mercer’s theorem on general domains: On the interaction between measures, kernels, and RKHSs”. In: Constructive Approximation 35 (2012), pp. 363– 417

  75. [83]

    Random Features Methods in Supervised Learning

    Yitong Sun. “Random Features Methods in Supervised Learning”. PhD thesis. 2019

  76. [84]

    Feature learning via mean-field langevin dynamics: classifying sparse parities and beyond

    Taiji Suzuki et al. “Feature learning via mean-field langevin dynamics: classifying sparse parities and beyond”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 34536–34556

  77. [85]

    Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective

    Shokichi Takakura and Taiji Suzuki. “Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective”. In: Forty-first International Conference on Machine Learning

  78. [86]

    Smoothing and decay estimates for nonlinear diffusion equations: equa- tions of porous medium type

    Juan Luis V´ azquez. Smoothing and decay estimates for nonlinear diffusion equations: equa- tions of porous medium type . Vol. 33. OUP Oxford, 2006. 43

  79. [87]

    The porous medium equation: mathematical theory

    Juan Luis V´ azquez. The porous medium equation: mathematical theory . Oxford university press, 2007

  80. [88]

    Partial optimization and Schur complement

    Fran¸ cois-Xavier Vialard. “Partial optimization and Schur complement”. In: (2019)

  81. [89]

    Optimal transport: old and new

    C´ edric Villani. Optimal transport: old and new . Vol. 338. Springer, 2009

  82. [90]

    Mean-field langevin dynam- ics for signed measures via a bilevel approach

    Guillaume Wang, Alireza Mousavi-Hosseini, and L´ ena¨ ıc Chizat. “Mean-field langevin dynam- ics for signed measures via a bilevel approach”. In:Advances in Neural Information Processing Systems 37 (2024), pp. 35165–35224

  83. [91]

    Tensor programs iv: Feature learning in infinite-width neural networks

    Greg Yang and Edward J Hu. “Tensor programs iv: Feature learning in infinite-width neural networks”. In: International Conference on Machine Learning. PMLR. 2021, pp. 11727–11737

  84. [92]

    Gradient descent optimizes over-parameterized deep ReLU networks

    Difan Zou et al. “Gradient descent optimizes over-parameterized deep ReLU networks”. In: Machine learning 109 (2020), pp. 467–492. 44 A Positive definite kernels and RKHS We recall in this section basic properties on the theory of Reproducing Kernel Hilbert Spaces and refer to...

  85. [93]

    biased” quadratic regularization fb :t7→ 1 2t2 or the “unbiased

    A scatter plot of the teacher measure ¯µγ and of the resulting teacher signal forγ = 100 are shown in Fig. 10. Finally, we consider the input data x to be distributed according to an empirical distribution ˆρ = 1 N PN i=1δxi with i.i.d. standard gaussian samples xi∼N (0, Id). ...

  86. [94]

    (50)) of width M∈{ 32, 128, 512, 1024}

    While such setting is not covered by our theory (in particular the ultra-fast diffusion equation is not necessarily well-posed), 48 100 101 102 103 104 Number of iterations 100 101 Reduced risk M=32 M=128 M=512 M=1024 0 10000 20000 30000 Number of iterations 10□4 10□3 10□2 10□...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.