{"id":"e005f1be-d6be-44ed-b3f6-2563baf265cd","arxiv_id":"2504.18208","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In the teacher-student setting, variable-projection training of two-layer networks is shown to match a weighted ultra-fast diffusion in the zero-regularization limit, giving linear convergence of the learned feature distribution.","lead":"This paper proves that a special training method for two-layer neural networks, which instantly optimizes the output weights at each step, makes the hidden feature distribution evolve like a known fast-spreading diffusion equation. It uses this link to give explicit convergence rates for recovering a teacher network's features, a gap in mean-field neural network theory.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 5 is local-in-time only; the claimed linear rate for low-regularization VarPro does not follow from the stated theorems.","rationale":"The reader's weakest assumption focused on Assumption 1 and the source condition/H-to-C^1 embedding in Theorem 5. My stress-test identifies a different, but related, gap: even granting all regularity assumptions, Theorem 5 is only a finite-time approximation result, and the paper's conclusion appears to compose it with the long-time Theorem 3 to infer a linear rate for lambda>0. This inference is not proved and is not covered by Theorem 4 for the main f(t)=|t|^r/(r-1). The paper is partially self-flagging ('says nothing about the long time behavior'), and the empirical figures are consistent with a non-uniform-in-time approximation. This reinforces the reader's CONDITIONAL verdict rather than changing it: the core local-in-time identification may be correct, but the advertised quantitative convergence rates for the actual VarPro algorithm with lambda>0 remain unproven.","tokens_in":43709,"tokens_out":36385,"duration_ms":368268,"concrete_test":"For the quadratic case r=2 on the torus with uniform teacher mu_bar, compute numerically sup_{t in [0,T]} W2(mu^lambda_t, mu^0_t) for lambda in {1e-2,1e-3,1e-4} and T up to, say, 100. If this supremum does not converge to 0 uniformly in T as lambda goes to 0 (i.e., the error grows with T at fixed lambda), the local Theorem 5 cannot justify the low-lambda linear rate. A sharper analytical check is to linearize both flows around mu_bar and compare their spectral gaps: if the lambda-flow's gap vanishes as lambda->0 while the diffusion gap is order 1, then the long-time behaviors differ and the claimed inherited linear rate is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central narrative combines Theorem 5 (lim_{lambda->0} sup_{t in [0,T]} W2(mu^lambda_t, mu^0_t)=0 for each fixed T) with Theorem 3 (linear convergence of the lambda=0 ultra-fast diffusion) to conclude that, in the low-regularization regime, the learned feature distribution converges at a linear rate. This composition requires the W2 error between the lambda-flow and the diffusion to be controlled uniformly in t, or at least to grow slowly enough with T, but Theorem 5 provides no such uniform control and the paper itself states in Section 6.1 that the result 'says nothing about the long time behavior of the dynamic.' The gap is not merely rhetorical: for the family f(t)=|t|^r/(r-1) used in Theorem 5, Theorem 4 does not apply because its hypothesis min f=f(mbar) fails whenever the teacher has positive total mass mbar>0. Hence, for every lambda>0, no convergence rate of the VarPro gradient flow to mu_bar is established for this f. The numerical linear rates in Fig. 5 are therefore not consequences of the stated theorems; indeed, the distance to the diffusion dynamic in Fig. 5 (right) grows after the initial phase, illustrating the non-uniformity in time.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the Variable Projection (VarPro) / two-timescale gradient flow for training mean-field two-layer neural networks with square loss. Under a teacher-student assumption (Y = Phi^* nu_bar with Phi^* injective), the unregularized reduced risk L^0_f is shown to equal the f-divergence D_f(nu_bar | mu) (Eq. (17)). For the family f(t)=|t|^r/(r-1), the Wasserstein gradient flow at lambda=0 is identified as a weighted ultra-fast diffusion equation (Eq. (35)); relying on Iacobelli–Patacchini–Santambrogio, the paper states linear convergence of this diffusion to the teacher feature distribution (Theorem 3). At fixed lambda>0, an algebraic convergence rate is claimed under additional assumptions (Theorem 4). The main new result is Theorem 5, which asserts that, under a source condition and compact embedding of the RKHS H in C^1, the regularized gradient flows converge locally uniformly in time to the ultra-fast diffusion as lambda -> 0^+. Numerical experiments on S^1 and CIFAR10 are presented as supporting evidence.","tokens_in":44012,"tokens_out":10009,"duration_ms":102085,"significance":"If the identification is correct, the paper offers a clean, parameter-free connection between feature learning in two-layer networks and a weighted ultra-fast diffusion PDE: the lambda=0 reduced risk is exactly a scaled reverse f-divergence, and the limiting PDE is a known object with external well-posedness and convergence results. This is a valuable conceptual contribution and a rigorous stability result (Theorem 5) for the vanishing-regularization limit. The related-work discussion is careful and positions the paper well against KALE, DrMMD, and MMD flows. However, the advertised quantitative guarantee for the regularized VarPro dynamics in the low-regularization regime does not follow from the stated theorems; this gap concerns the paper's central narrative and must be addressed before publication.","major_comments":[{"comment":"Theorem 4 does not apply to the regularization family used in Theorem 5. Theorem 4 requires \"min f = f(\\bar m) = 0\" with \\bar m = \\bar\\nu(\\Omega) > 0. For f(t)=|t|^r/(r-1), the minimum is 0 at t=0 while f(\\bar m) = \\bar m^r/(r-1) > 0, so the hypothesis fails whenever the teacher measure has positive mass. Consequently, for every lambda > 0 and for this f, the paper establishes no convergence rate of the VarPro gradient flow to the teacher distribution; the only rigorous rate is for the lambda=0 diffusion (Theorem 3). The numerical linear rates in Figures 5 and 6 are therefore not consequences of the stated theorems, and the claims connecting them to Theorem 4 or to a transfer of the lambda=0 rate need to be corrected or replaced by a genuinely new argument.","section":"§5.1, Theorem 4; §5.2 setting f(t)=|t|^r/(r-1)"},{"comment":"The \"unbiased\" quadratic regularization f_u(t) = 1/2 |t-1|^2 used in the numerical comparison is not of the form f(t)=|t|^r/(r-1) required by Theorem 5. Since f_u differs from f_2(t)=t^2 by a constant depending on \\bar\\nu, the gradient-flow identification may still hold up to an additive constant, but this equivalence is not stated. The text should clarify which theorem is being tested when f_u is used, and whether Theorem 5's assumptions are intended to cover it.","section":"§6.1, regularization f_u = 1/2|t-1|^2"}],"minor_comments":[{"comment":"The phrase \"provable convergence rates for the sampling of a teacher feature distribution\" and the conclusion's claim that low-regularization VarPro converges \"at a linear rate (Theorem 3)\" overstate what is proven. These statements should be qualified to indicate that the linear rate is proven for the lambda=0 diffusion, while for lambda>0 only local-in-time approximation to that diffusion is established.","section":"Abstract and Conclusion"},{"comment":"The word \"sampling\" in the abstract is inaccurate because the dynamics is a deterministic gradient flow; \"recovering\" or \"approximating\" would be more appropriate.","section":"§1.1, abstract wording"}],"recommendation":"major_revision","confidential_remarks":"The paper contains a genuinely interesting identification between the lambda=0 VarPro gradient flow and weighted ultra-fast diffusion, and Theorem 5 is a solid local-in-time stability result. However, the abstract and conclusion promise a quantitative convergence rate for the regularized algorithm that is not obtained by composing Theorem 5 with Theorem 3. This is not merely a presentational issue: for the f used in Theorem 5, no rate at fixed lambda>0 exists in the paper, and Theorem 4's hypotheses exclude that f. I believe the manuscript is salvageable by thoroughly rewriting the claims: state exactly what is proven (lambda=0 convergence, local-in-time approximation for lambda>0 under source conditions) and clearly separate the conjectured low-regularization rate from the proven statements. If the authors instead insist on the current abstract/conclusion as the main contribution, the paper would need a substantially stronger theorem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper: the core identification is real. Under the teacher-student assumption, the VarPro/two-timescale gradient flow of L0_r is exactly the weighted ultra-fast diffusion, and Theorem 5's local-in-time stability as lambda->0 is a clean and genuinely new result. The authors also do the right thing by importing the linear rate from Iacobelli-Patacchini-Santambrogio rather than re-deriving it. But the paper's headline claim — provable convergence rates for the low-regularization VarPro flow — is not actually established by the theorems as written.\n\nThe main problem is a composition gap. Theorem 5 gives sup_{t in [0,T]} W2(mu^lambda_t, mu^0_t) -> 0 for each fixed T, and Theorem 3 gives linear convergence of mu^0_t. These two do not yield a rate for mu^lambda_t as t->infinity. You need a control uniform in t (or at least a quantified growth in T), and the paper itself says in Section 6.1 that the result says nothing about long time behavior. The stress-test note is right to flag this. It isn't a fatal flaw in the identification itself, but it means the abstract and conclusion overstate the quantitative guarantees.\n\nThere is a second, more subtle gap: Theorem 4's algebraic rate assumes min f = f(mbar)=0, which fails for the family f(t)=|t|^r/(r-1) used in Theorem 5 whenever the teacher has positive total mass. So for the f that links to ultra-fast diffusion, there is no rate at all for lambda>0. The experiments use the 'unbiased' quadratic f(t)=1/2|t-1|^2, for which mbar=1 and the assumption holds, but that is not the f in the main theorem. That makes the theory more fragmented than the narrative suggests.\n\nThe source condition in Theorem 5 is also strong: it assumes H embeds compactly into C^1 and that partial f(dbar_nu/dmu^lambda_t) is bounded in H along the flow. For ReLU features on S^1, neither is verified, so the numerical section is an extrapolation. The paper is honest about this to a degree, but it should be explicit in the abstract or introduction that the theorems require these regularity conditions.\n\nWhat the paper does well: the duality framework, the Gamma-convergence Lemma 3.3, the careful positioning relative to [62], and the code release. The derivation of Eq. (35) is parameter-free and externally supported. It deserves a serious referee: a good referee could help the authors either prove a genuine uniform-in-time rate under stronger assumptions, or honestly reframe the contribution as a local-in-time stability result plus a conjectured rate. I would not cite the advertised linear rate in my own work, but I would cite the identification and the local stability.\n\nBottom line: send it to peer review, expect heavy revision on the rate claims, but the core mathematical observation is sound and novel.","headline":"The VarPro/ultra-fast diffusion identification is a clean, genuinely new result, but the advertised linear-rate guarantee for small regularization does not follow from the paper's own theorems.","tokens_in":44497,"tokens_out":4314,"would_cite":true,"duration_ms":43030,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T07","49Q22","35K55"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that, in a teacher-student setting, the low-regularization limit of VarPro training for two-layer mean-field networks is a weighted ultra-fast diffusion equation whose solutions converge linearly to the teacher feature…","keywords":["two-layer neural networks","mean-field training","variable projection","two-timescale gradient descent","Wasserstein gradient flow","ultra-fast diffusion","teacher-student","feature learning"],"falsifier":"Take a smooth teacher density on the torus, solve the weighted ultra-fast diffusion Eq. (35) numerically, run VarPro with features that make $\\Phi^\\star$ injective for a sequence of $\\lambda$ values tending to $0$, and measure $\\sup_{t\\in[0,T]}W_2(\\mu^\\lambda_t,\\mu^0_t)$; Theorem 5 predicts this tends to $0$, so a nonvanishing plateau would refute the diffusion limit. A complementary test: replace the teacher by an atomic measure, where the log-density is unbounded; the predicted linear rate should fail, exposing the boundary of the result.","tokens_in":43536,"feed_emoji":"⚡","tokens_out":17659,"duration_ms":149971,"temperature":0.7,"pith_summary":"Variable Projection (VarPro) is the two-timescale strategy that eliminates a two-layer network's outer weights and trains only the distribution of inner features, the part responsible for feature learning. The paper shows that, when the target signal is generated by a finite teacher measure through an injective feature map, the reduced training objective is an f-divergence between student and teacher feature distributions. In the vanishing-regularization limit, its Wasserstein gradient flow is a weighted ultra-fast diffusion equation, a nonlinear diffusion with a singular negative diffusivity. Under stated regularity assumptions, the paper proves that regularized VarPro dynamics converge to this diffusion as the regularization tends to zero, and that the diffusion itself converges linearly to the teacher feature distribution. The result turns a high-dimensional, non-convex training problem into an explicitly described PDE with quantitative convergence rates, going beyond the qualitative statements common in mean-field analysis.","feed_headline":"Eliminating outer weights turns network training into ultra-fast diffusion","feed_subtitle":"As regularization vanishes, VarPro's feature distribution follows a diffusion that converges linearly to the teacher.","key_machinery":"The load-bearing object is the reduced risk\n$$L^\\lambda_f(\\mu)=\\min_u \\frac{1}{\\$\\lambda$}R^\\lambda_f(\\mu,u),$$\nwhich, under the teacher-student assumption, is an infimal convolution of an f-divergence and a maximum mean discrepancy. Its dual representation,\n$$L^\\lambda_f(\\mu)=\\sup_{\\$\\alpha$\\in $L^{2}$(\\rho)}\\left[\\int_\\$\\Omega$(\\Phi^\\top\\$\\alpha$)\\,d\\bar\\nu-\\int_\\$\\Omega$ f^*(\\Phi^\\top\\$\\alpha$)\\,d\\mu-\\frac{\\$\\lambda$}{2}\\|\\$\\alpha$\\|_{$L^{2}$(\\rho)}^2\\right],$$\nmakes the envelope theorem applicable and yields the velocity field $\\nabla L^\\lambda_f[\\mu]=-\\nabla(f^*(\\Phi^\\top \\alpha^\\lambda_f[\\mu]))$, whose negative is the drift of the Wasserstein gradient flow $\\partial_t\\mu-\\operatorname{div}(\\mu\\nabla L^\\lambda_f[\\mu])=0$. For $f(t)=|t|^r/(r-1)$, the unregularized limit of this field is $-\\nabla(\\bar{\\mu}/\\mu)^r$, and the continuity equation becomes the weighted ultra-fast diffusion of Eq. (35). To pass to the limit in the regularized flows, the paper imposes a source condition — $\\partial f(d\\bar\\nu/d\\mu^\\lambda_t)$ lies in the RKHS $H$, compactly embedded in $C^1(\\Omega)$ — which keeps the dual variable bounded and gives compactness of the velocity fields.","core_discovery":"Under Assumption 1, in which $Y=\\Phi^\\star \\bar{\\nu}$ for a finite measure $\\bar{\\nu}$ and $\\Phi^\\star$ is injective, the unregularized reduced risk with $f(t)=|t|^r/(r-1)$ is\n$$$L^{0}$_r(\\mu)=\\frac{\\|\\bar{\\nu}\\|_{\\mathrm{TV}}^r}{r-1}\\int_\\$\\Omega$ \\left(\\frac{d\\bar{\\mu}}{d\\mu}\\right)^r d\\mu ,$$\nan f-divergence (for $r=2$, a $\\chi^2$-divergence) between the student distribution $\\mu$ and the teacher distribution $\\bar{\\mu}=\\bar{\\nu}/\\|\\bar{\\nu}\\|_{\\mathrm{TV}}$. Its Wasserstein gradient flow is\n$$\\partial_t \\mu_t = -\\|\\bar{\\nu}\\|_{\\mathrm{TV}}^r \\operatorname{div}\\left(\\mu_t \\nabla\\left(\\frac{\\bar{\\mu}}{\\mu_t}\\right)^r\\right),$$\na weighted ultra-fast diffusion equation: the exponent is negative, so the diffusivity is singular where the student density vanishes. The central result is Theorem 5, which identifies the limit of the regularized dynamics: as $\\lambda\\to 0^+$, gradient flows of $L^r_\\lambda$ converge locally uniformly in time to this diffusion. Since solutions of the diffusion converge linearly to $\\bar{\\mu}$, VarPro training in the low-regularization regime inherits a quantitative, essentially explicit description of feature learning.","pith_inferences":["If the diffusion description is exact, the proved convergence rate depends on initialization only through the log-density ratio $\\|\\log(\\bar{\\mu}/\\mu_0)\\|_\\infty$; a testable design principle is to initialize feature distributions so that this ratio is bounded and small.","The infimal-convolution formula suggests a threshold transition at $\\lambda$ comparable to the spectrum of the tangent kernel $K_\\mu$: above it the flow behaves like an MMD gradient flow with algebraic rates, below it like an f-divergence flow with linear rates — a prediction that could be measured on other architectures.","Applying the same last-layer variable projection to deep networks, as the ResNet experiment does, suggests that feature learning in deep training might admit a similar effective diffusion description at the last layer; this is not covered by the paper because depth breaks the linear separability on which the proof relies."],"forward_implications":["In the vanishing-regularization limit, the learned feature distribution converges to the teacher's at a linear (exponential) rate in $L^2$, and the rate constant is controlled by the log-density ratio at initialization rather than by the ambient dimension.","At any fixed $\\lambda>0$, the VarPro gradient flow converges to the minimizer of the reduced risk at an algebraic rate that, under the stated boundedness assumption, is independent of $\\lambda$.","As $\\lambda\\to 0^+$, regularized dynamics approach the ultra-fast diffusion on every finite time interval, so low-regularization training runs should be accurately described by the PDE rather than by a kernel (NTK) linearization.","With a small enough step relative to $\\lambda$, ordinary two-timescale gradient descent reproduces the VarPro dynamics; in low-regularization regimes VarPro itself keeps converging where the two-timescale scheme does not.","On bounded convex domains or tori, whose Poincaré constants are dimension-free, the linear rate in the diffusion limit does not degrade with dimension, indicating feature learning that avoids the curse of dimensionality."],"supporting_citations":[{"why":"Establishes existence, uniqueness, and the linear L2 convergence of weighted ultra-fast diffusion solutions, which are imported as Theorem 3.","marker":"[47]"},{"why":"Introduces variable projection for separable nonlinear least squares, the training method whose gradient flow is the object of study.","marker":"[34]"},{"why":"Provides the kernelized Moreau-envelope view of f-divergences whose Gamma-convergence structure the paper adapts (with the roles of measures reversed).","marker":"[62]"},{"why":"Supplies the variational representation of the total variation used in Lemma 3.2 to prove uniqueness of the minimizer of $L^\\lambda_r$.","marker":"[90]"},{"why":"Shows injectivity of the ReLU perceptron feature map under non-polynomial activation and full-support data, validating Assumption 1 in the numerical setting.","marker":"[83]"},{"why":"Provides the Lyapunov argument pattern used in Theorem 4 to obtain the algebraic convergence rate of the regularized gradient flow.","marker":"[32]"},{"why":"Supplies the metric gradient-flow theory, geodesic semiconvexity, and chain-rule tools used to define and prove well-posedness of the Wasserstein gradient flow.","marker":"[2]"}],"fun_headline_variants":["VarPro turns feature training into a singular ultra-fast diffusion","Teacher-student nets learn via weighted ultra-fast diffusion","Eliminating linear weights yields diffusion-like feature dynamics","As regularization vanishes, VarPro's gradient flow becomes diffusion","Feature learning as ultra-fast diffusion in low-λ regime"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That the target signal is exactly generated by a finite teacher measure through an injective feature map; without exact representability or injectivity, the reduced risk is not a divergence and the ultra-fast diffusion description collapses (Theorem 5 additionally needs a source condition with the RKHS $H$ compactly embedded in $C^1(\\Omega)$).","fun_headline_variants_meta":{"raw":{"variants":["VarPro turns feature training into a singular ultra-fast diffusion","Teacher-student nets learn via weighted ultra-fast diffusion","Eliminating linear weights yields diffusion-like feature dynamics","As regularization vanishes, VarPro's gradient flow becomes diffusion","Feature learning as ultra-fast diffusion in low-λ regime"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000219,"raw_usage":{"total_tokens":1484,"prompt_tokens":1027,"completion_tokens":457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":378}},"tokens_in":643,"tokens_out":457,"duration_ms":5270,"temperature":1.0,"reasoning_tokens":378,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:22:48.895027+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a smooth teacher density on the torus, solve the weighted ultra-fast diffusion Eq. (35) numerically, run VarPro with features that make $\\Phi^\\star$ injective for a sequence of $\\lambda$ values tending to $0$, and measure $\\sup_{t\\in[0,T]}W_2(\\mu^\\lambda_t,\\mu^0_t)$; Theorem 5 predicts this tends to $0$, so a nonvanishing plateau would refute the diffusion limit. A complementary test: replace the teacher by an atomic measure, where the log-density is unbounded; the predicted linear rate should fail, exposing the boundary of the result.","supporting_citations":[{"cited_title":"Weighted ultrafast diffusion equations: from well-posedness to long-time behaviour","cited_arxiv_id":null,"evidence_quote":"Establishes existence, uniqueness, and the linear L2 convergence of weighted ultra-fast diffusion solutions, which are imported as Theorem 3."},{"cited_title":"The differentiation of pseudo-inverses and nonlinear least squares problems whose variables separate","cited_arxiv_id":null,"evidence_quote":"Introduces variable projection for separable nonlinear least squares, the training method whose gradient flow is the object of study."},{"cited_title":"Mean-field langevin dynam- ics for signed measures via a bilevel approach","cited_arxiv_id":null,"evidence_quote":"Supplies the variational representation of the total variation used in Lemma 3.2 to prove uniqueness of the minimizer of $L^\\lambda_r$."},{"cited_title":"Random Features Methods in Supervised Learning","cited_arxiv_id":null,"evidence_quote":"Shows injectivity of the ReLU perceptron feature map under non-polynomial activation and full-support data, validating Assumption 1 in the numerical setting."},{"cited_title":"KALE flow: A relaxed KL gradient flow for probabilities with disjoint support","cited_arxiv_id":null,"evidence_quote":"Provides the Lyapunov argument pattern used in Theorem 4 to obtain the algebraic convergence rate of the regularized gradient flow."}],"review_version":1}