Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

Semiparametric M-estimation with overparameterized neural networks

T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper proves that overparameterized ReLU networks can estimate the nuisance function in semiparametric M-estimation at the minimax nonparametric rate while the parametric component stays root-n consistent and asymptotically normal.

desk verdict A plausible and novel framework for semiparametric inference with overparameterized nets, but the normality theorem has a possible gap in the least-favorable direction and the proofs are partly in a missing supplement. read the letter →

arxiv 2504.19089 v1 pith:QWOIWJ3N submitted 2025-04-27 math.ST stat.TH

classification math.STstat.TH MSC 62G0562G0862F1268T07
keywords semiparametricM-estimationoverparameterizedneuralnetworkstangentkernelroot-nconsistencyasymptoticnormalitynonparametricminimaxrateHuberizedmarginconditiongradientflow
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether a neural network can estimate the infinite-dimensional nuisance component of a semiparametric model without destroying inference on the finite-dimensional parameter of interest. It answers yes for a penalized overparameterized ReLU network trained by gradient flow, under a new "Huberized margin condition" on the loss. If the main theorem is right, a practitioner can use a wide, overparameterized network for the nuisance function and still obtain a root-n consistent, asymptotically normal estimate of $\beta$ with valid confidence intervals; the nuisance estimate itself converges at the minimax nonparametric rate up to log factors. The known obstruction -- degenerate tangent spaces at symmetric network weights -- is shown to occur only with small probability under random initialization and to disappear in the wide-network limit.

What carries the argument

The load-bearing object is the neural tangent kernel (NTK) $K^{\mathrm{NT}}(x,x')=\nabla_\theta f_\theta(x)^T\nabla_\theta f_\theta(x')$, whose RKHS is norm-equivalent to the Sobolev space $W^{(d+1)/2,2}(\Omega)$ for the ReLU network used here. Gradient flow of the neural network is shown to track the gradient flow of the same loss over this RKHS: Theorem 2.1 bounds the sup-norm gap by $o(n^{-1/2})$ once the width exceeds a polynomial in $n,\lambda^{-1},L_0,\log(1/\xi),\exp(t)$. This transfer of dynamics is what imports the rich tangent space of the RKHS into the network, so that the efficient-score remainder $P_n(S_2(\hat\beta,\hat f)[\tilde h])$ becomes negligible. The second ingredient is the Huberized margin condition (Assumption 3), a weakened margin/Bernstein-type inequality that relates excess risk to squared $L^2$ distance with a denominator allowing unbounded functions; it is what makes the peeling and entropy argument go through without boundedness assumptions.

What would settle it

Run the estimator on a partially linear model with a convex smooth loss whose pointwise risk is flat over an interval around the optimum (for example a smoothed absolute loss with large threshold), and measure coverage of 95% confidence intervals for $\beta$ as $n$ grows; if coverage remains near nominal, Assumption 3 is not necessary, while a drop in coverage would confirm that the margin condition is carrying the result.

Watch

Extended reading notes

Core claim

For the criterion $P_n l_{\beta,f} + \lambda_n \|\theta-\theta_0\|_2^2$ trained by gradient flow, Theorem 3.1 states that with probability at least $1-\xi$ over random initialization, the nuisance estimate obeys $\|\hat f_{t_s}-f_0\|^2_{L^2}=O_p(n^{-2s/(2s+d)}\log n)$, where $s=(d+1)/2$ when $f_0$ lies in the NTK reproducing kernel Hilbert space and $s>d/2$ under the relaxed assumption that $f_0$ lies in a Sobolev space; in both cases $\sqrt{n}(\hat\beta_{t_s}-\beta_0)=n^{1/2}A^{-1}P_n\tilde S(\beta_0,f_0)+o_p(1)$ converges in distribution to $N(0,\Sigma)$. In plain terms, the network learns the nuisance function at the optimal nonparametric rate while the parametric component behaves as if the nuisance were known. The result covers general convex losses, not only least squares, and requires no boundedness of the network output or of the candidate nuisance functions; it analyzes the actual gradient-flow solution rather than an idealized global minimizer.

Load-bearing premise

The load-bearing premise is the Huberized margin condition (Assumption 3), requiring excess risk to dominate the squared $L^2$ error up to a denominator that tolerates unbounded outputs; if a loss lacks sufficient curvature near the truth, the stated nonparametric rate and root-n normality collapse, and the paper verifies this condition only for specific regression and classification losses.

Editorial extensions

If this is right

  • A user can fit a partially linear model or a classification model with a wide ReLU network for the nuisance term and then build confidence intervals for $\beta$ from the asymptotic normal approximation; the paper's simulations report coverage near 95% as $n$ grows.
  • The same gradient-flow estimator attains the minimax rate (up to log factors) for the nonparametric component, so no separate sieve basis or kernel with closed-form expressions is needed.
  • Because normality holds both when $f_0$ lies in the NTK RKHS and when it lies only in a Sobolev space with smoothness $s>d/2$, the inference on $\beta$ is robust to mis-specification of the nuisance function space.
  • When the loss is the negative log-likelihood and the model contains the least-favorable submodel, the estimator is semiparametric efficient; with a misspecified loss, root-n consistency and asymptotic normality survive.
  • The framework covers general convex losses with Lipschitz gradients, so beyond least squares and logistic loss it applies to other smooth robust losses that satisfy local strong curvature.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Testable extension: replace the $\ell^2$-penalty around initialization by early stopping and check whether the same root-n normality holds with a width requirement that no longer contains $\exp(t_s)$; the paper itself notes the exponential factor is an artifact of general losses and drops out for least squares.
  • By analogy with the NTK-RKHS equivalence, the same proof scheme should yield root-n normal semiparametric estimators for other kernels whose RKHS is Sobolev-norm equivalent, such as Laplace kernels, as long as the kernel's gradient-flow surrogate exists.
  • The Huberized margin condition is plausible for many smooth robust losses; verifying Assumption 3 and Assumption 5 for losses such as Tukey's biweight or smooth quantile-type approximations would widen the class beyond the paper's regression and classification examples.
  • A practically important question the paper leaves implicit is the finite-sample choice of $t_s$ and $\lambda_n$; the theory requires $t_s\gtrsim n^2$ for merely convex objectives, so adaptive stopping rules that track the loss decrease might yield the same guarantees at far smaller training time.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper proposes a semiparametric M-estimator in which the nuisance function is estimated by an overparameterized, randomly initialized ReLU network regularized by the squared distance of parameters from initialization and trained by gradient flow. The central device is a comparison between the network flow and the flow of a penalized empirical risk over the RKHS of the limiting neural tangent kernel; Theorem 2.1 asserts that, for sufficiently large width, the two flows stay within o(n^{-1/2}) of each other. Under a new "Huberized margin condition" (Assumption 3), assumptions placing the true f0 either in the NTK RKHS (Assumption 4) or in a Sobolev space W^{s,2} with s > d/2 (Assumption 4'), and a least-favorable direction condition (Assumption 5), Theorem 3.1 claims minimax nonparametric rates and root-n asymptotic normality of the parametric component. Sections 4 and 5 illustrate the framework on partially linear regression and classification and provide simulations.

Significance. If the claims are correct, the paper makes a substantial contribution: it studies the actual gradient-flow trajectory rather than an ideal global minimizer, avoids boundedness assumptions on the network output, and permits the true nuisance function to lie outside the NTK RKHS. The RKHS-surrogate strategy is natural, the stated nonparametric rates match known minimax benchmarks, and the numerical experiments support the qualitative conclusions. I do not see circularity: the penalty parameter is set to theoretical n-dependent rates and the claims are not fitted to the simulations. However, the main theorems are not self-contained: the proofs of Theorem 2.1 and Theorem 3.1, as well as the verification lemmas for Section 4, are relegated to a supplementary file that is not included, and one step in the normality argument under Assumption 4' appears technically problematic. The significance is therefore conditional on the missing proofs being supplied and on the tangent-space issue being resolved.

major comments (4)
  1. [§3, Theorem 3.1; §4.1, Assumption 6] The central results are not verifiable from the submitted manuscript. Theorem 3.1, Theorem 4.1, and Theorem 4.2 depend on technical arguments that are not present: Theorem 2.1 is stated without proof, and after Assumption 6 the text refers to "Lemma ?? in the Supplementary Material (Yan et al., 2025a)" for the verification of Assumption 3. Since the difficulty of the paper is precisely in the asymptotic normality expansion (12), the missing supplement is load-bearing. The supplement should be included with the submission, or the proofs should be written out in an appendix.
  2. [§3, Assumption 5(1); Proposition 2.2; Theorem 2.1] The asymptotic normality claim in Theorem 3.1 has a gap concerning the least-favorable direction when f0 is outside the NTK RKHS. Assumption 5(1) requires only that tilde h_i belongs to W^{s,2}(Ω) with s > d/2, while Proposition 2.2 identifies the NTK RKHS H_NT with W^{(d+1)/2,2}(Ω). Under Assumption 4', the relevant regime is d/2 < s < (d+1)/2, in which tilde h_i need not belong to H_NT. The proof needs P_n S_2(hat beta, hat f)[tilde h] = o_p(n^{-1/2}); Theorem 2.1 only compares the network trajectory with the RKHS flow, whose first-order conditions are taken in H_NT. For tilde h outside H_NT, neither the RKHS first-order condition nor an approximation of tilde h by the network tangent space is available from the arguments in this paper. Please either strengthen Assumption 5 to require tilde h_i in H_NT, or prove that the Sobolev-to-RKHS approximation error is negligible at the n^{-1/2} scale.
  3. [§3, Theorem 3.1] The boundedness condition on hat beta_{t_s} and P l_{hat beta_{t_s}, hat f_{t_s}} is a hypothesis of Theorem 3.1, but no sufficient conditions for it are stated or proved. The remark after the theorem says the condition is "often verifiable," and Section 4 does not provide a general verification. Since the paper explicitly avoids boundedness of the network output and this condition is used in the normality argument, it should either be promoted to an explicit assumption with checkable hypotheses or proved under the existing assumptions.
  4. [§3, Assumption 3, Eq. (8)] The "Huberized margin condition" (Assumption 3) is a global lower bound on excess risk that must hold for every beta and f, including functions far from f0 in L_infinity. For losses with flat tails or bounded loss, such as misclassification or logistic loss away from the decision boundary, the left-hand side can be much smaller than the right-hand side when ||f - f0||_{L_infinity} is large and the probability mass near the decision boundary is small. The paper verifies the condition only for the partially linear regression and classification examples in Section 4, citing a supplementary lemma, and does not prove it for the broad class of loss functions claimed in the abstract. Because the nonparametric rate in Theorem 3.1 depends on this condition through the peeling argument, a general verification or a precise statement of the loss-function restrictions is needed.
minor comments (4)
  1. [§2.1, Proposition 2.1] The displayed formula for K_NT contains an unresolved typographical artifact ("+/BD l≥1") that makes the term unintelligible; please re-typeset the formula.
  2. [§4.2, Theorem 4.2] Theorem 4.2 cites "Assumption 6(1)", but Assumption 6 is formulated for the regression loss in Section 4.1; the classification setting needs an analogous identifiability assumption stated separately.
  3. [§5, Tables 3 and 4] The text says simulations use 200 repetitions per setting, but the paragraph on coverage probabilities states that Tables 3 and 4 are based on 500 repeated experiments; please reconcile these numbers.
  4. [§3, Remark 2.2 and Theorem 2.1] The required width grows like exp(t) with t of order n^2 in the merely convex case, which makes the network width astronomically large; a short discussion of the practical interpretation of this condition would help the reader.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the rates and normality are derived from approximation/entropy balancing and external NTK results, not from fitted inputs or self-referential definitions.

full rationale

The derivation chain is self-contained in the relevant sense. Theorem 3.1's nonparametric rates follow by balancing approximation error and entropy under Assumption 3, with lambda set to theoretical n-rates; there is no fitted parameter being renamed as a prediction. The parametric normality statement rests on standard semiparametric Taylor/entropy arguments and on the least-favorable direction in Assumption 5, neither of which is defined in terms of the conclusion. Theorem 2.1 is a mathematical comparison bound between the neural gradient flow and the RKHS flow, and Proposition 2.2 is a known NTK/Sobolev equivalence supported by the external NTK literature. The only self-references are to the authors' supplementary material (Yan et al., 2025a), which contains proofs and example verifications such as Assumption 3 for the regression model; deferring proofs to a supplement is standard practice and does not make the target theorem an input. No equation in the submitted text defines the claimed convergence rate or asymptotic normality in terms of the estimator's own fitted values, no uniqueness theorem from the same authors is used to force a choice, and no known result is merely renamed as a new contribution. The skeptical concern about the least-favorable direction lying outside the NTK tangent space is a correctness or verification gap, not a circularity: it does not reduce the theorem to its assumptions by construction. Therefore no circular step is exhibited.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The theory rests on a mix of standard semiparametric conditions and one new assumption (Huberized margin). The least favorable direction and RKHS equivalence are classical or derived, while the margin condition is the main novel premise. No invented physical entities appear.

assumptions (5)
  • domain assumption Loss l is convex, nonnegative, and has B1-Lipschitz gradient (Assumption 1)
    Guarantees gradient flow convergence (Prop 2.3) and the closeness of neural and RKHS flows (Thm 2.1).
  • ad hoc to paper Huberized margin condition (Assumption 3, Eq. (8))
    New condition: P(lβ,f - lβ0,f0) ≥ B2 (||β-β0||^2 + ||f-f0||^2_L2)/(1 + ||β-β0|| + ||f-f0||_L∞). Replaces margin/Bernstein conditions and is load-bearing for the peeling/entropy argument.
  • domain assumption True nuisance f0 lies in NTK RKHS (Assumption 4) or in Sobolev W^{s,2}, s>d/2 (Assumption 4')
    Ensures the regularized problem is well-posed and yields the stated n^{-2s/(2s+d)} rate; not checkable from data.
  • domain assumption Existence of a least favorable direction tilde h satisfying (9), with smoothness s>d/2 and nonsingular information A (Assumption 5)
    Standard efficient-score condition; needed to make the parametric component root-n estimable once the nuisance rate is faster than n^{-1/4}.
  • standard math RKHS of K_NT is norm-equivalent to Sobolev W^{(d+1)/2,2}(Omega) (Prop 2.2)
    Bridges NTK to Sobolev entropy bounds; proof deferred to the supplement.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Semiparametric M-estimation with overparameterized neural networks." pith.science (2026). https://pith.science/paper/QWOIWJ3N

@misc{pith2026250419089,
  author       = {Pith},
  title        = {Pith review of: Semiparametric M-estimation with overparameterized neural networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QWOIWJ3N}},
  note         = {Machine review of arXiv:2504.19089}
}
abstract

We focus on semiparametric regression that has played a central role in statistics, and exploit the powerful learning ability of deep neural networks (DNNs) while enabling statistical inference on parameters of interest that offers interpretability. Despite the success of classical semiparametric method/theory, establishing the $\sqrt{n}$-consistency and asymptotic normality of the finite-dimensional parameter estimator in this context remains challenging, mainly due to nonlinearity and potential tangent space degeneration in DNNs. In this work, we introduce a foundational framework for semiparametric $M$-estimation, leveraging the approximation ability of overparameterized neural networks that circumvent tangent degeneration and align better with training practice nowadays. The optimization properties of general loss functions are analyzed, and the global convergence is guaranteed. Instead of studying the ``ideal'' solution to minimization of an objective function in most literature, we analyze the statistical properties of algorithmic estimators, and establish nonparametric convergence and parametric asymptotic normality for a broad class of loss functions. These results hold without assuming the boundedness of the network output and even when the true function lies outside the specified function space. To illustrate the applicability of the framework, we also provide examples from regression and classification, and the numerical experiments provide empirical support to the theoretical findings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Semiparametric Partial Differential Equation Models

    stat.ME 2025-06 conditional novelty 6.0 of 10

    A semiparametric PDE model with a parametric physical part and a neural-network unknown mechanism is estimated by profiling maximum likelihood, with proofs of optimal nonparametric rates and root-n efficient inference...

Reference graph

Works this paper leans on

15 extracted references · 10 canonical work pages · cited by 1 Pith paper

  1. [1]

    Ahmad, I., Leelahanon, S., and Li, Q. (2005). Efficient estima tion of a semiparametric partially linear varying coefficient model. Annals of statistics . Allen-Zhu, Z., Li, Y ., and Song, Z. (2019). A convergence the ory for deep learning via over- parameterization. In International conference on machine learning , pages 242–252. PMLR. Arora, S., Du, S. S., ...

  2. [2]

    in the regression model. Model The coverage rate forβ1 The coverage rate forβ2 Setting n = 500 n = 1000 n = 2000 n = 500 n = 1000 n = 2000 Case 1 0.938 0.958 0.946 0.944 0.952 0.940 Case 2 0.966 0.958 0.950 0.960 0.950 0.952 Case 3 0.946 0.970 0.948 0.968 0.952 0.950 Case 4 0.986 0.974 0.938 0.988 0.968 0.948 Table 4: The coverage probability for construc...

  3. [3]

    and Xu, S

    Chen, L. and Xu, S. (2020). Deep neural tangent kernel and lap lace kernel have the same rkhs. arXiv preprint arXiv:2009.10683. Chen, M., Jiang, H., Liao, W., and Zhao, T. (2019). Nonparametric regression on low-dimensional manifolds using deep relu networks: Function approximatio n and statistical recovery. ArXiv preprint. arXiv:1908.01842. Chen, X., Liu...

  4. [4]

    Uniform Generalization Bounds for Overparameterized Neural Networks

    Springer. 28 Tsybakov, A. B. (2004). Optimal aggregation of classifiers i n statistical learning. The Annals of Statistics, 32(1):135–166. Vakili, S., Bromberg, M., Shiu, D., and Bernacchia, A. (2021 ). Uniform generalization bounds for overparameterized neural networks. CoRR, abs/2109.06099. Van de Geer, S. A. (2000). Empirical Processes in M-estimation, volume

  5. [6]

    Van der Vaart, A

    Cambridge university press. Van der Vaart, A. W. (2000). Asymptotic statistics, volume

  6. [7]

    and Liang, Y

    Li, Y . and Liang, Y . (2018). Learning overparameterized neural networks via stochastic gradient descent on structured data. Advances in neural information processing systems ,

  7. [8]

    Li, Y ., Yu, Z., Chen, G., and Lin, Q. (2024). On the eigenvalue decay rates of a class of neural- network related kernel functions defined on general domains . Journal of Machine Learning Research, 25(82):1–47. Liang, H., Liu, X., Li, R., and Tsai, C.-L. (2010). Estimatio n and testing for partially linear single-index models. Annals of statistics , 38(6)...

  8. [12]

    semipa rametric estimation with overpa- rameterized neural network

    Cambridge university press. Wang, J.-L., Xue, L., Zhu, L., and Chong, Y . S. (2010). Estimation for a partial-linear single-index model. Annals of statistics . Wang, X., Zhou, L., and Lin, H. (2024). Deep regression learn ing with optimal loss function. Journal of the American Statistical Association , pages 1–20. Y an, S., Chen, Z., and Y ao, F. (2025a)....

Show all 15 references
  1. [15]

    in the classification model. Model The coverage rate forβ1 The coverage rate forβ2 Setting n = 500 n = 1000 n = 2000 n = 500 n = 1000 n = 2000 Case 1 0.964 0.968 0.950 0.966 0.954 0.938 Case 2 0.936 0.950 0.952 0.950 0.962 0.940 Case 3 0.942 0.926 0.932 0.972 0.942 0.930 Case 4...

  2. [31]

    Jiao, Y ., Shen, G., Lin, Y ., and Huang, J. (2021). Deep nonparametric regression on approximate manifolds: Non-asymptotic error bounds with polynomial pr efactors. Kohler, M. and Langer, S. (2021). On the rate of convergence o f fully connected deep neural network regression...

  3. [32]

    Bach, F. R. (2014). Breaking the curse of dimensionality wit h convex neural networks. CoRR, abs/1412.8690. 24 Bartlett, P . L. and Mendelson, S. (2006). Empirical minimization. Probability theory and related fields , 135(3):311–334. Bauer, B. and Kohler, M. (2019). On deep lea...

  4. [61]

    Lai, J., Xu, M., Chen, R., and Lin, Q

    Springer. Lai, J., Xu, M., Chen, R., and Lin, Q. (2023a). Generalizatio n ability of wide neural networks on r. arXiv preprint arXiv:2302.05933. Lai, J., Yu, Z., Tian, S., and Lin, Q. (2023b). Generalizatio n ability of wide residual networks. Lecu´e, G. (2011). Interplay betw...

  5. [131]

    Hu, T., Wang, W., Lin, C., and Cheng, G

    Springer Science & Business Media. Hu, T., Wang, W., Lin, C., and Cheng, G. (2021). Regularizati on matters: A nonparametric perspective on overparametrized neural network. In International Conference on Artificial Intelligence and Statistics, pages 829–837. PMLR. Huang, J. (19...

  6. [1375]

    W., and Wang, J.-L

    Zhong, Q., M¨ uller, J. W., and Wang, J.-L. (2021). Deep exten ded hazard models for survival analysis. In Ranzato, M., Beygelzimer, A., Dauphin, Y ., Liang, P ., and Vaughan, J. W., editors, Advances in Neural Information Processing Systems , volume 34, pages 15111–15124. Cur...

  7. [1897]

    Shen, Z. (2020). Deep network approximation characterized by number of neurons. Communi- cations in Computational Physics , 28(5):1768–1811. Shen, Z., Y ang, H., and Zhang, S. (2022). Optimal approximat ion rate of relu networks in terms of width and depth. Journal de Math ´em...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.