REVIEW 4 major objections 4 minor 1 cited by
Semiparametric M-estimation with overparameterized neural networks
T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper proves that overparameterized ReLU networks can estimate the nuisance function in semiparametric M-estimation at the minimax nonparametric rate while the parametric component stays root-n consistent and asymptotically normal.
desk verdict A plausible and novel framework for semiparametric inference with overparameterized nets, but the normality theorem has a possible gap in the least-favorable direction and the proofs are partly in a missing supplement. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the neural tangent kernel (NTK) $K^{\mathrm{NT}}(x,x')=\nabla_\theta f_\theta(x)^T\nabla_\theta f_\theta(x')$, whose RKHS is norm-equivalent to the Sobolev space $W^{(d+1)/2,2}(\Omega)$ for the ReLU network used here. Gradient flow of the neural network is shown to track the gradient flow of the same loss over this RKHS: Theorem 2.1 bounds the sup-norm gap by $o(n^{-1/2})$ once the width exceeds a polynomial in $n,\lambda^{-1},L_0,\log(1/\xi),\exp(t)$. This transfer of dynamics is what imports the rich tangent space of the RKHS into the network, so that the efficient-score remainder $P_n(S_2(\hat\beta,\hat f)[\tilde h])$ becomes negligible. The second ingredient is the Huberized margin condition (Assumption 3), a weakened margin/Bernstein-type inequality that relates excess risk to squared $L^2$ distance with a denominator allowing unbounded functions; it is what makes the peeling and entropy argument go through without boundedness assumptions.
What would settle it
Run the estimator on a partially linear model with a convex smooth loss whose pointwise risk is flat over an interval around the optimum (for example a smoothed absolute loss with large threshold), and measure coverage of 95% confidence intervals for $\beta$ as $n$ grows; if coverage remains near nominal, Assumption 3 is not necessary, while a drop in coverage would confirm that the margin condition is carrying the result.
Extended reading notes
Core claim
For the criterion $P_n l_{\beta,f} + \lambda_n \|\theta-\theta_0\|_2^2$ trained by gradient flow, Theorem 3.1 states that with probability at least $1-\xi$ over random initialization, the nuisance estimate obeys $\|\hat f_{t_s}-f_0\|^2_{L^2}=O_p(n^{-2s/(2s+d)}\log n)$, where $s=(d+1)/2$ when $f_0$ lies in the NTK reproducing kernel Hilbert space and $s>d/2$ under the relaxed assumption that $f_0$ lies in a Sobolev space; in both cases $\sqrt{n}(\hat\beta_{t_s}-\beta_0)=n^{1/2}A^{-1}P_n\tilde S(\beta_0,f_0)+o_p(1)$ converges in distribution to $N(0,\Sigma)$. In plain terms, the network learns the nuisance function at the optimal nonparametric rate while the parametric component behaves as if the nuisance were known. The result covers general convex losses, not only least squares, and requires no boundedness of the network output or of the candidate nuisance functions; it analyzes the actual gradient-flow solution rather than an idealized global minimizer.
Load-bearing premise
The load-bearing premise is the Huberized margin condition (Assumption 3), requiring excess risk to dominate the squared $L^2$ error up to a denominator that tolerates unbounded outputs; if a loss lacks sufficient curvature near the truth, the stated nonparametric rate and root-n normality collapse, and the paper verifies this condition only for specific regression and classification losses.
Editorial extensions
If this is right
- A user can fit a partially linear model or a classification model with a wide ReLU network for the nuisance term and then build confidence intervals for $\beta$ from the asymptotic normal approximation; the paper's simulations report coverage near 95% as $n$ grows.
- The same gradient-flow estimator attains the minimax rate (up to log factors) for the nonparametric component, so no separate sieve basis or kernel with closed-form expressions is needed.
- Because normality holds both when $f_0$ lies in the NTK RKHS and when it lies only in a Sobolev space with smoothness $s>d/2$, the inference on $\beta$ is robust to mis-specification of the nuisance function space.
- When the loss is the negative log-likelihood and the model contains the least-favorable submodel, the estimator is semiparametric efficient; with a misspecified loss, root-n consistency and asymptotic normality survive.
- The framework covers general convex losses with Lipschitz gradients, so beyond least squares and logistic loss it applies to other smooth robust losses that satisfy local strong curvature.
Reading between the lines
- Testable extension: replace the $\ell^2$-penalty around initialization by early stopping and check whether the same root-n normality holds with a width requirement that no longer contains $\exp(t_s)$; the paper itself notes the exponential factor is an artifact of general losses and drops out for least squares.
- By analogy with the NTK-RKHS equivalence, the same proof scheme should yield root-n normal semiparametric estimators for other kernels whose RKHS is Sobolev-norm equivalent, such as Laplace kernels, as long as the kernel's gradient-flow surrogate exists.
- The Huberized margin condition is plausible for many smooth robust losses; verifying Assumption 3 and Assumption 5 for losses such as Tukey's biweight or smooth quantile-type approximations would widen the class beyond the paper's regression and classification examples.
- A practically important question the paper leaves implicit is the finite-sample choice of $t_s$ and $\lambda_n$; the theory requires $t_s\gtrsim n^2$ for merely convex objectives, so adaptive stopping rules that track the loss decrease might yield the same guarantees at far smaller training time.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a semiparametric M-estimator in which the nuisance function is estimated by an overparameterized, randomly initialized ReLU network regularized by the squared distance of parameters from initialization and trained by gradient flow. The central device is a comparison between the network flow and the flow of a penalized empirical risk over the RKHS of the limiting neural tangent kernel; Theorem 2.1 asserts that, for sufficiently large width, the two flows stay within o(n^{-1/2}) of each other. Under a new "Huberized margin condition" (Assumption 3), assumptions placing the true f0 either in the NTK RKHS (Assumption 4) or in a Sobolev space W^{s,2} with s > d/2 (Assumption 4'), and a least-favorable direction condition (Assumption 5), Theorem 3.1 claims minimax nonparametric rates and root-n asymptotic normality of the parametric component. Sections 4 and 5 illustrate the framework on partially linear regression and classification and provide simulations.
Significance. If the claims are correct, the paper makes a substantial contribution: it studies the actual gradient-flow trajectory rather than an ideal global minimizer, avoids boundedness assumptions on the network output, and permits the true nuisance function to lie outside the NTK RKHS. The RKHS-surrogate strategy is natural, the stated nonparametric rates match known minimax benchmarks, and the numerical experiments support the qualitative conclusions. I do not see circularity: the penalty parameter is set to theoretical n-dependent rates and the claims are not fitted to the simulations. However, the main theorems are not self-contained: the proofs of Theorem 2.1 and Theorem 3.1, as well as the verification lemmas for Section 4, are relegated to a supplementary file that is not included, and one step in the normality argument under Assumption 4' appears technically problematic. The significance is therefore conditional on the missing proofs being supplied and on the tangent-space issue being resolved.
major comments (4)
- [§3, Theorem 3.1; §4.1, Assumption 6] The central results are not verifiable from the submitted manuscript. Theorem 3.1, Theorem 4.1, and Theorem 4.2 depend on technical arguments that are not present: Theorem 2.1 is stated without proof, and after Assumption 6 the text refers to "Lemma ?? in the Supplementary Material (Yan et al., 2025a)" for the verification of Assumption 3. Since the difficulty of the paper is precisely in the asymptotic normality expansion (12), the missing supplement is load-bearing. The supplement should be included with the submission, or the proofs should be written out in an appendix.
- [§3, Assumption 5(1); Proposition 2.2; Theorem 2.1] The asymptotic normality claim in Theorem 3.1 has a gap concerning the least-favorable direction when f0 is outside the NTK RKHS. Assumption 5(1) requires only that tilde h_i belongs to W^{s,2}(Ω) with s > d/2, while Proposition 2.2 identifies the NTK RKHS H_NT with W^{(d+1)/2,2}(Ω). Under Assumption 4', the relevant regime is d/2 < s < (d+1)/2, in which tilde h_i need not belong to H_NT. The proof needs P_n S_2(hat beta, hat f)[tilde h] = o_p(n^{-1/2}); Theorem 2.1 only compares the network trajectory with the RKHS flow, whose first-order conditions are taken in H_NT. For tilde h outside H_NT, neither the RKHS first-order condition nor an approximation of tilde h by the network tangent space is available from the arguments in this paper. Please either strengthen Assumption 5 to require tilde h_i in H_NT, or prove that the Sobolev-to-RKHS approximation error is negligible at the n^{-1/2} scale.
- [§3, Theorem 3.1] The boundedness condition on hat beta_{t_s} and P l_{hat beta_{t_s}, hat f_{t_s}} is a hypothesis of Theorem 3.1, but no sufficient conditions for it are stated or proved. The remark after the theorem says the condition is "often verifiable," and Section 4 does not provide a general verification. Since the paper explicitly avoids boundedness of the network output and this condition is used in the normality argument, it should either be promoted to an explicit assumption with checkable hypotheses or proved under the existing assumptions.
- [§3, Assumption 3, Eq. (8)] The "Huberized margin condition" (Assumption 3) is a global lower bound on excess risk that must hold for every beta and f, including functions far from f0 in L_infinity. For losses with flat tails or bounded loss, such as misclassification or logistic loss away from the decision boundary, the left-hand side can be much smaller than the right-hand side when ||f - f0||_{L_infinity} is large and the probability mass near the decision boundary is small. The paper verifies the condition only for the partially linear regression and classification examples in Section 4, citing a supplementary lemma, and does not prove it for the broad class of loss functions claimed in the abstract. Because the nonparametric rate in Theorem 3.1 depends on this condition through the peeling argument, a general verification or a precise statement of the loss-function restrictions is needed.
minor comments (4)
- [§2.1, Proposition 2.1] The displayed formula for K_NT contains an unresolved typographical artifact ("+/BD l≥1") that makes the term unintelligible; please re-typeset the formula.
- [§4.2, Theorem 4.2] Theorem 4.2 cites "Assumption 6(1)", but Assumption 6 is formulated for the regression loss in Section 4.1; the classification setting needs an analogous identifiability assumption stated separately.
- [§5, Tables 3 and 4] The text says simulations use 200 repetitions per setting, but the paragraph on coverage probabilities states that Tables 3 and 4 are based on 500 repeated experiments; please reconcile these numbers.
- [§3, Remark 2.2 and Theorem 2.1] The required width grows like exp(t) with t of order n^2 in the merely convex case, which makes the network width astronomically large; a short discussion of the practical interpretation of this condition would help the reader.
Circularity Check
No circularity found: the rates and normality are derived from approximation/entropy balancing and external NTK results, not from fitted inputs or self-referential definitions.
full rationale
The derivation chain is self-contained in the relevant sense. Theorem 3.1's nonparametric rates follow by balancing approximation error and entropy under Assumption 3, with lambda set to theoretical n-rates; there is no fitted parameter being renamed as a prediction. The parametric normality statement rests on standard semiparametric Taylor/entropy arguments and on the least-favorable direction in Assumption 5, neither of which is defined in terms of the conclusion. Theorem 2.1 is a mathematical comparison bound between the neural gradient flow and the RKHS flow, and Proposition 2.2 is a known NTK/Sobolev equivalence supported by the external NTK literature. The only self-references are to the authors' supplementary material (Yan et al., 2025a), which contains proofs and example verifications such as Assumption 3 for the regression model; deferring proofs to a supplement is standard practice and does not make the target theorem an input. No equation in the submitted text defines the claimed convergence rate or asymptotic normality in terms of the estimator's own fitted values, no uniqueness theorem from the same authors is used to force a choice, and no known result is merely renamed as a new contribution. The skeptical concern about the least-favorable direction lying outside the NTK tangent space is a correctness or verification gap, not a circularity: it does not reduce the theorem to its assumptions by construction. Therefore no circular step is exhibited.
Assumptions & free parameters
assumptions (5)
- domain assumption Loss l is convex, nonnegative, and has B1-Lipschitz gradient (Assumption 1)
- ad hoc to paper Huberized margin condition (Assumption 3, Eq. (8))
- domain assumption True nuisance f0 lies in NTK RKHS (Assumption 4) or in Sobolev W^{s,2}, s>d/2 (Assumption 4')
- domain assumption Existence of a least favorable direction tilde h satisfying (9), with smoothness s>d/2 and nonsingular information A (Assumption 5)
- standard math RKHS of K_NT is norm-equivalent to Sobolev W^{(d+1)/2,2}(Omega) (Prop 2.2)
Cite this review
Pith. "Pith review of Semiparametric M-estimation with overparameterized neural networks." pith.science (2026). https://pith.science/paper/QWOIWJ3N
@misc{pith2026250419089,
author = {Pith},
title = {Pith review of: Semiparametric M-estimation with overparameterized neural networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/QWOIWJ3N}},
note = {Machine review of arXiv:2504.19089}
}
abstract
We focus on semiparametric regression that has played a central role in statistics, and exploit the powerful learning ability of deep neural networks (DNNs) while enabling statistical inference on parameters of interest that offers interpretability. Despite the success of classical semiparametric method/theory, establishing the $\sqrt{n}$-consistency and asymptotic normality of the finite-dimensional parameter estimator in this context remains challenging, mainly due to nonlinearity and potential tangent space degeneration in DNNs. In this work, we introduce a foundational framework for semiparametric $M$-estimation, leveraging the approximation ability of overparameterized neural networks that circumvent tangent degeneration and align better with training practice nowadays. The optimization properties of general loss functions are analyzed, and the global convergence is guaranteed. Instead of studying the ``ideal'' solution to minimization of an objective function in most literature, we analyze the statistical properties of algorithmic estimators, and establish nonparametric convergence and parametric asymptotic normality for a broad class of loss functions. These results hold without assuming the boundedness of the network output and even when the true function lies outside the specified function space. To illustrate the applicability of the framework, we also provide examples from regression and classification, and the numerical experiments provide empirical support to the theoretical findings.
Forward citations
Cited by 1 Pith paper
-
Deep Semiparametric Partial Differential Equation Models
A semiparametric PDE model with a parametric physical part and a neural-network unknown mechanism is estimated by profiling maximum likelihood, with proofs of optimal nonparametric rates and root-n efficient inference...
Reference graph
Works this paper leans on
-
[1]
Ahmad, I., Leelahanon, S., and Li, Q. (2005). Efficient estima tion of a semiparametric partially linear varying coefficient model. Annals of statistics . Allen-Zhu, Z., Li, Y ., and Song, Z. (2019). A convergence the ory for deep learning via over- parameterization. In International conference on machine learning , pages 242–252. PMLR. Arora, S., Du, S. S., ...
work page 2005
-
[2]
in the regression model. Model The coverage rate forβ1 The coverage rate forβ2 Setting n = 500 n = 1000 n = 2000 n = 500 n = 1000 n = 2000 Case 1 0.938 0.958 0.946 0.944 0.952 0.940 Case 2 0.966 0.958 0.950 0.960 0.950 0.952 Case 3 0.946 0.970 0.948 0.968 0.952 0.950 Case 4 0.986 0.974 0.938 0.988 0.968 0.948 Table 4: The coverage probability for construc...
work page 2000
-
[3]
Chen, L. and Xu, S. (2020). Deep neural tangent kernel and lap lace kernel have the same rkhs. arXiv preprint arXiv:2009.10683. Chen, M., Jiang, H., Liao, W., and Zhao, T. (2019). Nonparametric regression on low-dimensional manifolds using deep relu networks: Function approximatio n and statistical recovery. ArXiv preprint. arXiv:1908.01842. Chen, X., Liu...
arXiv 2020
-
[4]
Uniform Generalization Bounds for Overparameterized Neural Networks
Springer. 28 Tsybakov, A. B. (2004). Optimal aggregation of classifiers i n statistical learning. The Annals of Statistics, 32(1):135–166. Vakili, S., Bromberg, M., Shiu, D., and Bernacchia, A. (2021 ). Uniform generalization bounds for overparameterized neural networks. CoRR, abs/2109.06099. Van de Geer, S. A. (2000). Empirical Processes in M-estimation, volume
work page Pith review arXiv 2004
-
[6]
Cambridge university press. Van der Vaart, A. W. (2000). Asymptotic statistics, volume
work page 2000
-
[7]
Li, Y . and Liang, Y . (2018). Learning overparameterized neural networks via stochastic gradient descent on structured data. Advances in neural information processing systems ,
work page 2018
-
[8]
Li, Y ., Yu, Z., Chen, G., and Lin, Q. (2024). On the eigenvalue decay rates of a class of neural- network related kernel functions defined on general domains . Journal of Machine Learning Research, 25(82):1–47. Liang, H., Liu, X., Li, R., and Tsai, C.-L. (2010). Estimatio n and testing for partially linear single-index models. Annals of statistics , 38(6)...
arXiv 2024
-
[12]
semipa rametric estimation with overpa- rameterized neural network
Cambridge university press. Wang, J.-L., Xue, L., Zhu, L., and Chong, Y . S. (2010). Estimation for a partial-linear single-index model. Annals of statistics . Wang, X., Zhou, L., and Lin, H. (2024). Deep regression learn ing with optimal loss function. Journal of the American Statistical Association , pages 1–20. Y an, S., Chen, Z., and Y ao, F. (2025a)....
work page 2010
Show all 15 references
-
[15]
in the classification model. Model The coverage rate forβ1 The coverage rate forβ2 Setting n = 500 n = 1000 n = 2000 n = 500 n = 1000 n = 2000 Case 1 0.964 0.968 0.950 0.966 0.954 0.938 Case 2 0.936 0.950 0.952 0.950 0.962 0.940 Case 3 0.942 0.926 0.932 0.972 0.942 0.930 Case 4...
2000
-
[31]
Jiao, Y ., Shen, G., Lin, Y ., and Huang, J. (2021). Deep nonparametric regression on approximate manifolds: Non-asymptotic error bounds with polynomial pr efactors. Kohler, M. and Langer, S. (2021). On the rate of convergence o f fully connected deep neural network regression...
2021
-
[32]
Bach, F. R. (2014). Breaking the curse of dimensionality wit h convex neural networks. CoRR, abs/1412.8690. 24 Bartlett, P . L. and Mendelson, S. (2006). Empirical minimization. Probability theory and related fields , 135(3):311–334. Bauer, B. and Kohler, M. (2019). On deep lea...
2014 arXiv
-
[61]
Lai, J., Xu, M., Chen, R., and Lin, Q
Springer. Lai, J., Xu, M., Chen, R., and Lin, Q. (2023a). Generalizatio n ability of wide neural networks on r. arXiv preprint arXiv:2302.05933. Lai, J., Yu, Z., Tian, S., and Lin, Q. (2023b). Generalizatio n ability of wide residual networks. Lecu´e, G. (2011). Interplay betw...
2023 arXiv
-
[131]
Hu, T., Wang, W., Lin, C., and Cheng, G
Springer Science & Business Media. Hu, T., Wang, W., Lin, C., and Cheng, G. (2021). Regularizati on matters: A nonparametric perspective on overparametrized neural network. In International Conference on Artificial Intelligence and Statistics, pages 829–837. PMLR. Huang, J. (19...
2021
-
[1375]
W., and Wang, J.-L
Zhong, Q., M¨ uller, J. W., and Wang, J.-L. (2021). Deep exten ded hazard models for survival analysis. In Ranzato, M., Beygelzimer, A., Dauphin, Y ., Liang, P ., and Vaughan, J. W., editors, Advances in Neural Information Processing Systems , volume 34, pages 15111–15124. Cur...
2021
-
[1897]
Shen, Z. (2020). Deep network approximation characterized by number of neurons. Communi- cations in Computational Physics , 28(5):1768–1811. Shen, Z., Y ang, H., and Zhang, S. (2022). Optimal approximat ion rate of relu networks in terms of width and depth. Journal de Math ´em...
2020 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.