{"id":"a1757ea0-e010-4a59-b51a-54b4c0d29aa5","arxiv_id":"2504.19089","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For semiparametric M-estimation with overparameterized ReLU networks, the gradient-flow estimator achieves minimax nonparametric rates and root-n-consistent, asymptotically normal estimates of the finite-dimensional parameter.","lead":"This paper proves that semiparametric models with the nuisance function estimated by a very wide neural network can still give the usual square-root-n-accurate confidence intervals for the parameter of interest. It works in the neural tangent kernel regime, where wide networks behave like kernel smoothers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Root-n normality needs the least-favorable direction in the NTK tangent space; Assumption 5 only gives W^{s,2}, which can lie outside H_NT when f0 is outside the RKHS.","rationale":"The reader's chosen weakest assumption, Assumption 3, is a reasonable place to look: the nonparametric rate and the peeling argument both pass through it, and it is only verified in examples via references to the unavailable supplement. However, the central new claim is not only the L2 rate but the parametric asymptotic normality, and the step that makes or breaks that claim is the control of the efficient-score contribution P_nS_2[\\tilde h]. The paper itself frames the tangent-space approximation of the least-favorable direction as the core difficulty, and its resolution is supposed to come from random initialization plus overparameterization in Theorem 2.1. The manuscript's Assumption 5, however, only asks for \\tilde h \\in W^{s,2}, s > d/2, which is strictly weaker than membership in H_NT = W^{(d+1)/2,2} whenever Assumption 4' is active. If the proof requires \\tilde h to be (approximately) in the tangent space, this assumption is insufficient as stated. Because the supplementary proof is absent, I cannot confirm whether there is a hidden argument using the enormous width to approximate \\tilde h in L2 or a dual norm; but the missing display is precisely the load-bearing step, and the reader's Assumption 3 concern, while valid, does not address it. My verdict remains CONDITIONAL: the claimed result may be true, but its key normality step is not verifiable from the submitted text, and the stated assumptions may need strengthening to \\tilde h \\in H_NT or a separate approximation condition.","tokens_in":22331,"tokens_out":23363,"duration_ms":249653,"concrete_test":"Obtain the supplementary file and locate the bound for P_nS_2(\\hat\\beta_{t_s},\\hat f_{t_s})[\\tilde h]. If it is of the form O(\\|\\nabla_f \\tilde L_n\\|_{H_NT} \\|\\tilde h\\|_{H_NT}), check whether Assumption 5 plus s > d/2 actually guarantees \\tilde h \\in H_NT; for d = 5 and s = 2.8 it does not. If the proof instead approximates \\tilde h by an element of H_NT, verify that the width m(t_s,\\xi) = poly(exp(t_s)) from Theorem 2.1 makes the resulting score error o_p(n^{-1/2}) in the norm that appears; otherwise Theorem 3.1's normality claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 3.1's asymptotic-normality conclusion depends on showing P_nS_2(\\hat\\beta_{t_s},\\hat f_{t_s})[\\tilde h] = o_p(n^{-1/2}) for the least-favorable direction \\tilde h introduced in Assumption 5. The only control the paper proves is the gradient-flow comparison Theorem 2.1 between neural training and the RKHS flow of (7), i.e. derivatives along functions in the NTK RKHS H_NT. Proposition 2.2 identifies H_NT with W^{(d+1)/2,2}(\\Omega). Assumption 5(1) only assumes \\tilde h \\in W^{s,2} with s > d/2; in the regime where Assumption 4' is genuinely needed (f0 \\in W^{s,2} \\setminus H_NT) one must have d/2 < s < (d+1)/2, so W^{s,2} contains functions not in H_NT. For such \\tilde h, the RKHS first-order condition P_n l'_2 \\tilde h = -2\\lambda\\langle \\tilde f, \\tilde h\\rangle_{H_NT} is not available, and it is not clear that the tangent-space approximation error—exactly the issue the paper says it overcomes—is o_p(n^{-1/2}). The proof is relegated to the missing supplement (Lemma ?? in Yan et al., 2025a), so this step cannot be checked from the submitted manuscript. The boundedness condition on \\hat\\beta_{t_s} and P l is also stated without proof, but the tangent-space/least-favorable-direction gap is the more essential one for the normality claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a semiparametric M-estimator in which the nuisance function is estimated by an overparameterized, randomly initialized ReLU network regularized by the squared distance of parameters from initialization and trained by gradient flow. The central device is a comparison between the network flow and the flow of a penalized empirical risk over the RKHS of the limiting neural tangent kernel; Theorem 2.1 asserts that, for sufficiently large width, the two flows stay within o(n^{-1/2}) of each other. Under a new \"Huberized margin condition\" (Assumption 3), assumptions placing the true f0 either in the NTK RKHS (Assumption 4) or in a Sobolev space W^{s,2} with s > d/2 (Assumption 4'), and a least-favorable direction condition (Assumption 5), Theorem 3.1 claims minimax nonparametric rates and root-n asymptotic normality of the parametric component. Sections 4 and 5 illustrate the framework on partially linear regression and classification and provide simulations.","tokens_in":15,"tokens_out":7015,"duration_ms":155930,"significance":"If the claims are correct, the paper makes a substantial contribution: it studies the actual gradient-flow trajectory rather than an ideal global minimizer, avoids boundedness assumptions on the network output, and permits the true nuisance function to lie outside the NTK RKHS. The RKHS-surrogate strategy is natural, the stated nonparametric rates match known minimax benchmarks, and the numerical experiments support the qualitative conclusions. I do not see circularity: the penalty parameter is set to theoretical n-dependent rates and the claims are not fitted to the simulations. However, the main theorems are not self-contained: the proofs of Theorem 2.1 and Theorem 3.1, as well as the verification lemmas for Section 4, are relegated to a supplementary file that is not included, and one step in the normality argument under Assumption 4' appears technically problematic. The significance is therefore conditional on the missing proofs being supplied and on the tangent-space issue being resolved.","major_comments":[{"comment":"The central results are not verifiable from the submitted manuscript. Theorem 3.1, Theorem 4.1, and Theorem 4.2 depend on technical arguments that are not present: Theorem 2.1 is stated without proof, and after Assumption 6 the text refers to \"Lemma ?? in the Supplementary Material (Yan et al., 2025a)\" for the verification of Assumption 3. Since the difficulty of the paper is precisely in the asymptotic normality expansion (12), the missing supplement is load-bearing. The supplement should be included with the submission, or the proofs should be written out in an appendix.","section":"§3, Theorem 3.1; §4.1, Assumption 6"},{"comment":"The asymptotic normality claim in Theorem 3.1 has a gap concerning the least-favorable direction when f0 is outside the NTK RKHS. Assumption 5(1) requires only that tilde h_i belongs to W^{s,2}(Ω) with s > d/2, while Proposition 2.2 identifies the NTK RKHS H_NT with W^{(d+1)/2,2}(Ω). Under Assumption 4', the relevant regime is d/2 < s < (d+1)/2, in which tilde h_i need not belong to H_NT. The proof needs P_n S_2(hat beta, hat f)[tilde h] = o_p(n^{-1/2}); Theorem 2.1 only compares the network trajectory with the RKHS flow, whose first-order conditions are taken in H_NT. For tilde h outside H_NT, neither the RKHS first-order condition nor an approximation of tilde h by the network tangent space is available from the arguments in this paper. Please either strengthen Assumption 5 to require tilde h_i in H_NT, or prove that the Sobolev-to-RKHS approximation error is negligible at the n^{-1/2} scale.","section":"§3, Assumption 5(1); Proposition 2.2; Theorem 2.1"},{"comment":"The boundedness condition on hat beta_{t_s} and P l_{hat beta_{t_s}, hat f_{t_s}} is a hypothesis of Theorem 3.1, but no sufficient conditions for it are stated or proved. The remark after the theorem says the condition is \"often verifiable,\" and Section 4 does not provide a general verification. Since the paper explicitly avoids boundedness of the network output and this condition is used in the normality argument, it should either be promoted to an explicit assumption with checkable hypotheses or proved under the existing assumptions.","section":"§3, Theorem 3.1"},{"comment":"The \"Huberized margin condition\" (Assumption 3) is a global lower bound on excess risk that must hold for every beta and f, including functions far from f0 in L_infinity. For losses with flat tails or bounded loss, such as misclassification or logistic loss away from the decision boundary, the left-hand side can be much smaller than the right-hand side when ||f - f0||_{L_infinity} is large and the probability mass near the decision boundary is small. The paper verifies the condition only for the partially linear regression and classification examples in Section 4, citing a supplementary lemma, and does not prove it for the broad class of loss functions claimed in the abstract. Because the nonparametric rate in Theorem 3.1 depends on this condition through the peeling argument, a general verification or a precise statement of the loss-function restrictions is needed.","section":"§3, Assumption 3, Eq. (8)"}],"minor_comments":[{"comment":"The displayed formula for K_NT contains an unresolved typographical artifact (\"+/BD l≥1\") that makes the term unintelligible; please re-typeset the formula.","section":"§2.1, Proposition 2.1"},{"comment":"Theorem 4.2 cites \"Assumption 6(1)\", but Assumption 6 is formulated for the regression loss in Section 4.1; the classification setting needs an analogous identifiability assumption stated separately.","section":"§4.2, Theorem 4.2"},{"comment":"The text says simulations use 200 repetitions per setting, but the paragraph on coverage probabilities states that Tables 3 and 4 are based on 500 repeated experiments; please reconcile these numbers.","section":"§5, Tables 3 and 4"},{"comment":"The required width grows like exp(t) with t of order n^2 in the merely convex case, which makes the network width astronomically large; a short discussion of the practical interpretation of this condition would help the reader.","section":"§3, Remark 2.2 and Theorem 2.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript repeatedly refers to its own supplement (Yan et al., 2025a) for the proofs of the main theorems and for key lemma verifications, but that supplement is not part of this submission. For a statistics journal, the theorem-proof chain cannot be checked without it. I recommend requesting the supplement before any further decision and, in particular, asking the authors to address the tilde h / H_NT gap under Assumption 4' explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Mark,\n\nQuick take: this is a real attempt to get root-n normality for the parametric component while the nuisance is an overparameterized ReLU network trained by gradient flow. That's genuinely new. Earlier DNN semiparametric work (Zhong et al., Chen et al.) either assumed tangent-space approximation or used debiased two-step procedures, and mostly analyzed ideal minimizers. Here they analyze the algorithmic solution, allow general losses, and even let the true f live outside the NTK RKHS. The Huberized margin condition is a reasonable device to handle unbounded network outputs. The nonparametric rates match the known minimax rates, and the simulations look supportive. Credit where it's due: the paper is honestly written and the proof skeleton is standard.\n\nThe soft spots are real though. The full proofs live in a supplement that isn't attached to this arXiv submission; many lemmas are referenced as 'Lemma ??'. That alone would make me want a revised version before acceptance. More substantively, Theorem 3.1 carries a boundedness assumption on \\hat\\beta and P l that is only argued for the specific examples, not for the general M-estimation setting. That's probably fixable, but it's not in the paper.\n\nThe bigger issue is the one the stress-test flagged. The normality argument needs the efficient score remainder P_n S_2(\\hat\\beta,\\hat f)[\\tilde h] to be o_p(n^{-1/2}). The paper controls this via the RKHS first-order condition, which holds for directions in the NTK RKHS H_NT = W^{(d+1)/2,2}. But Assumption 5 only assumes the least-favorable direction \\tilde h is in W^{s,2} with s > d/2. When d/2 < s < (d+1)/2, \\tilde h need not belong to H_NT, so the RKHS condition isn't available. The paper says it overcame the tangent-space approximation issue, but this step is exactly where such an issue would resurface. The proof is in the missing supplement, so I can't check whether they have a workaround. This is not a fatal flaw on its face; they might approximate \\tilde h by H_NT functions with acceptable error. But as submitted, the central normality claim rests on an uncheckable step.\n\nVerdict: conditional acceptance is right. The paper deserves serious refereeing—the question is important and the approach is promising. I'd send it to referees, but with the explicit request that the supplement be included and that the least-favorable-direction step be spelled out.\n\nMe","headline":"A plausible and novel framework for semiparametric inference with overparameterized nets, but the normality theorem has a possible gap in the least-favorable direction and the proofs are partly in a missing supplement.","tokens_in":23172,"tokens_out":3591,"would_cite":false,"duration_ms":34702,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G08","62F12","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that overparameterized ReLU networks can estimate the nuisance function in semiparametric M-estimation at the minimax nonparametric rate while the parametric component stays root-n consistent and asymptotically normal.","keywords":["semiparametric M-estimation","overparameterized neural networks","neural tangent kernel","root-n consistency","asymptotic normality","nonparametric minimax rate","Huberized margin condition","gradient flow"],"falsifier":"Run the estimator on a partially linear model with a convex smooth loss whose pointwise risk is flat over an interval around the optimum (for example a smoothed absolute loss with large threshold), and measure coverage of 95% confidence intervals for $\\beta$ as $n$ grows; if coverage remains near nominal, Assumption 3 is not necessary, while a drop in coverage would confirm that the margin condition is carrying the result.","tokens_in":22130,"feed_emoji":"🧠","tokens_out":10300,"duration_ms":87449,"temperature":0.7,"pith_summary":"The paper asks whether a neural network can estimate the infinite-dimensional nuisance component of a semiparametric model without destroying inference on the finite-dimensional parameter of interest. It answers yes for a penalized overparameterized ReLU network trained by gradient flow, under a new \"Huberized margin condition\" on the loss. If the main theorem is right, a practitioner can use a wide, overparameterized network for the nuisance function and still obtain a root-n consistent, asymptotically normal estimate of $\\beta$ with valid confidence intervals; the nuisance estimate itself converges at the minimax nonparametric rate up to log factors. The known obstruction -- degenerate tangent spaces at symmetric network weights -- is shown to occur only with small probability under random initialization and to disappear in the wide-network limit.","feed_headline":"Wide neural nets give semiparametric beta root-n inference","feed_subtitle":"Overparameterized network learns the nuisance function at minimax rate while beta stays asymptotically normal, enabling confidence…","key_machinery":"The load-bearing object is the neural tangent kernel (NTK) $K^{\\mathrm{NT}}(x,x')=\\nabla_\\theta f_\\theta(x)^T\\nabla_\\theta f_\\theta(x')$, whose RKHS is norm-equivalent to the Sobolev space $W^{(d+1)/2,2}(\\Omega)$ for the ReLU network used here. Gradient flow of the neural network is shown to track the gradient flow of the same loss over this RKHS: Theorem 2.1 bounds the sup-norm gap by $o(n^{-1/2})$ once the width exceeds a polynomial in $n,\\lambda^{-1},L_0,\\log(1/\\xi),\\exp(t)$. This transfer of dynamics is what imports the rich tangent space of the RKHS into the network, so that the efficient-score remainder $P_n(S_2(\\hat\\beta,\\hat f)[\\tilde h])$ becomes negligible. The second ingredient is the Huberized margin condition (Assumption 3), a weakened margin/Bernstein-type inequality that relates excess risk to squared $L^2$ distance with a denominator allowing unbounded functions; it is what makes the peeling and entropy argument go through without boundedness assumptions.","core_discovery":"For the criterion $P_n l_{\\beta,f} + \\lambda_n \\|\\theta-\\theta_0\\|_2^2$ trained by gradient flow, Theorem 3.1 states that with probability at least $1-\\xi$ over random initialization, the nuisance estimate obeys $\\|\\hat f_{t_s}-f_0\\|^2_{L^2}=O_p(n^{-2s/(2s+d)}\\log n)$, where $s=(d+1)/2$ when $f_0$ lies in the NTK reproducing kernel Hilbert space and $s>d/2$ under the relaxed assumption that $f_0$ lies in a Sobolev space; in both cases $\\sqrt{n}(\\hat\\beta_{t_s}-\\beta_0)=n^{1/2}A^{-1}P_n\\tilde S(\\beta_0,f_0)+o_p(1)$ converges in distribution to $N(0,\\Sigma)$. In plain terms, the network learns the nuisance function at the optimal nonparametric rate while the parametric component behaves as if the nuisance were known. The result covers general convex losses, not only least squares, and requires no boundedness of the network output or of the candidate nuisance functions; it analyzes the actual gradient-flow solution rather than an idealized global minimizer.","pith_inferences":["Testable extension: replace the $\\ell^2$-penalty around initialization by early stopping and check whether the same root-n normality holds with a width requirement that no longer contains $\\exp(t_s)$; the paper itself notes the exponential factor is an artifact of general losses and drops out for least squares.","By analogy with the NTK-RKHS equivalence, the same proof scheme should yield root-n normal semiparametric estimators for other kernels whose RKHS is Sobolev-norm equivalent, such as Laplace kernels, as long as the kernel's gradient-flow surrogate exists.","The Huberized margin condition is plausible for many smooth robust losses; verifying Assumption 3 and Assumption 5 for losses such as Tukey's biweight or smooth quantile-type approximations would widen the class beyond the paper's regression and classification examples.","A practically important question the paper leaves implicit is the finite-sample choice of $t_s$ and $\\lambda_n$; the theory requires $t_s\\gtrsim n^2$ for merely convex objectives, so adaptive stopping rules that track the loss decrease might yield the same guarantees at far smaller training time."],"forward_implications":["A user can fit a partially linear model or a classification model with a wide ReLU network for the nuisance term and then build confidence intervals for $\\beta$ from the asymptotic normal approximation; the paper's simulations report coverage near 95% as $n$ grows.","The same gradient-flow estimator attains the minimax rate (up to log factors) for the nonparametric component, so no separate sieve basis or kernel with closed-form expressions is needed.","Because normality holds both when $f_0$ lies in the NTK RKHS and when it lies only in a Sobolev space with smoothness $s>d/2$, the inference on $\\beta$ is robust to mis-specification of the nuisance function space.","When the loss is the negative log-likelihood and the model contains the least-favorable submodel, the estimator is semiparametric efficient; with a misspecified loss, root-n consistency and asymptotic normality survive.","The framework covers general convex losses with Lipschitz gradients, so beyond least squares and logistic loss it applies to other smooth robust losses that satisfy local strong curvature."],"supporting_citations":[{"why":"introduces the neural tangent kernel whose wide-network limit defines the RKHS surrogate the proof tracks.","marker":"Jacot et al., 2018"},{"why":"gives exact infinite-width kernel computation, supporting the interpretation of the limit kernel.","marker":"Arora et al., 2019"},{"why":"establishes the uniform convergence of finite-width NTK to its limit used in the network-to-RKHS approximation bound.","marker":"Lai et al., 2023a"},{"why":"supplies the semiparametric efficient-score and asymptotic-normality framework that Theorem 3.1 exports to the neural estimator.","marker":"Van der Vaart, 2000"},{"why":"provides the empirical-process and entropy background for the peeling argument and for semiparametric M-estimation theory.","marker":"Kosorok, 2008"},{"why":"contributes the margin condition that Assumption 3 generalizes to unbounded nuisance functions.","marker":"Tsybakov, 2004"},{"why":"contributes the Bernstein-type condition that the new Huberized margin condition relaxes.","marker":"Bartlett and Mendelson, 2006"},{"why":"earlier neural-network semiparametric estimator whose tangent-space assumptions the paper shows can be avoided.","marker":"Zhong et al., 2022"}],"fun_headline_variants":["Wide nets achieve root-n beta and smooth nuisance learning","Neural semiparametrics: root-n beta with overparameterized nets","Gradient flow nets give root-n inference in semiparametrics","Overparameterized nets: beta root-n, nuisance optimal rate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is the Huberized margin condition (Assumption 3), requiring excess risk to dominate the squared $L^2$ error up to a denominator that tolerates unbounded outputs; if a loss lacks sufficient curvature near the truth, the stated nonparametric rate and root-n normality collapse, and the paper verifies this condition only for specific regression and classification losses.","fun_headline_variants_meta":{"raw":{"variants":["Wide nets achieve root-n beta and smooth nuisance learning","Neural semiparametrics: root-n beta with overparameterized nets","Gradient flow nets give root-n inference in semiparametrics","Overparameterized nets: beta root-n, nuisance optimal rate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000655,"raw_usage":{"total_tokens":3043,"prompt_tokens":1033,"completion_tokens":2010,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":1936}},"tokens_in":649,"tokens_out":2010,"duration_ms":14203,"temperature":1.0,"reasoning_tokens":1936,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T06:02:15.526434+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the estimator on a partially linear model with a convex smooth loss whose pointwise risk is flat over an interval around the optimum (for example a smoothed absolute loss with large threshold), and measure coverage of 95% confidence intervals for $\\beta$ as $n$ grows; if coverage remains near nominal, Assumption 3 is not necessary, while a drop in coverage would confirm that the margin condition is carrying the result.","supporting_citations":[],"review_version":1}