{"id":"cad87b2f-05aa-4155-a288-e8d589193668","arxiv_id":"1909.02088","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Under finite moments and a local envelope growth condition, the least squares estimator in nonparametric regression can achieve minimax rates with heavy-tailed, covariate-dependent errors.","lead":"This paper derives finite-sample upper bounds on how fast the least squares estimator converges in nonparametric regression when the noise can be heavy-tailed and depend on the covariates. It shows the estimator can still reach the best possible rate if the noise has enough moments, and it gives explicit moment thresholds based on the complexity and local shape of the function class.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed q* threshold hinges entirely on the sup-norm local envelope condition (13), which is not implied by (L∞) and can fail for natural classes; the scope of the minimax claim is narrower than stated.","rationale":"The reader identified the local envelope growth condition as the weakest assumption, and I agree. The paper's main theorems are conditional on this condition, and the claimed rate threshold is a direct function of s. The concern is therefore about the scope and verification of the envelope condition: it is not implied by the entropy assumptions, it can fail for natural shape-constrained classes, and even for the multiple-index example the stated verification (Proposition A.3) relies on a density assumption that is not satisfied in common regimes. This is a real load-bearing limitation, but it is not an internal inconsistency in the conditional theorem. The central theorems appear internally consistent, and the paper does flag that necessity and general verification of s are open. Given the reader's CONDITIONAL verdict, my read does not change the verdict; the concern reinforces the need for explicit checking of (13) or (19) before applying the rate conclusions.","tokens_in":47793,"tokens_out":46546,"duration_ms":395310,"concrete_test":"Fix a class F satisfying (L∞) but not known to have a shrinking sup-norm envelope, e.g., the two-index model M_{gamma,3,2} with a truth f0(x)=m0(B0x), or the class of bounded monotone functions. For a sequence delta -> 0, compute or numerically estimate F_delta(x)=sup_{f:||f-f0||<=delta}|(f-f0)(x)| and check whether sup_x F_delta(x) <= C delta^s for the claimed s (for multiple index, s=2gamma/(2gamma+2); for monotone, s=0). If this inequality fails, Theorem 4.1 does not apply and the slower Corollary 4.1 rate (or Theorem 5.1 if VC-type conditions hold) is the actual guarantee.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central rate in Theorem 4.1 reduces, for constant A and Phi, to max{n^{-1/(2+alpha)}, n^{-(q-1)/(q(2-s)+alpha s(q-1))}}, and the threshold q*=(2+alpha(1-s))/(s+alpha(1-s)) is exactly a function of s from the sup-norm envelope condition (13): ||F_delta||_infty <= C Phi^{1-s} delta^s. If this condition fails, the claimed n^{-1/(2+alpha)} rate is not guaranteed. Condition (13) is not a consequence of the L∞-entropy assumption; there are classes satisfying (L∞) for which the local envelope does not shrink in sup norm (e.g., monotone or convex functions, where ||F_delta||_infty is of constant order), giving s=0 and the slower rate of Corollary 4.1. The paper verifies (13) only for specific smooth classes via interpolation inequalities. A particularly concrete gap is the multiple-index example: Proposition A.3 requires ((BX),(B0X)) to have a density on R^{2p} lower bounded away from zero, which is impossible when 2p>d (e.g., a two-index model in R^3) and degenerates as B approaches B0. Thus the envelope growth condition is the load-bearing assumption, and its verification is not automatic for every class to which the theorem might be applied.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the least squares estimator (LSE) in nonparametric regression with errors that are heteroscedastic and have only finitely many moments. The authors provide finite-sample tail bounds for ||hat f - f0|| under three global complexity conditions: bracketing L2 entropy (Theorem 3.1), L-infinity entropy (Theorem 4.1), and uniform VC-type entropy (Theorem 5.1). The rates are expressed in terms of the entropy parameter alpha, the moment index q of the errors, and a local-envelope growth parameter s that controls how the sup-norm or L2-norm of the local envelope F_delta(x)=sup_{f:||f-f0||<=delta}|(f-f0)(x)| shrinks as delta goes to zero. The headline conclusion is that, under the sup-norm envelope condition (13), the LSE attains the minimax rate n^{-1/(2+alpha)} once q is at least (2+alpha(1-s))/(s+alpha(1-s)), with slower rates otherwise; under the bracketing condition the analogous threshold is q>=2/s. Applications include Holder, Sobolev, additive, multiple-index, convex, and isotonic regression, with new interpolation inequalities in Appendix A and a new truncated peeling theorem in Appendix C. Section 5.2 contains a lower bound showing the role of s in the worst-case LSE rate. The authors correctly stress that their conditions are sufficient and that some optimality questions remain open.","tokens_in":48048,"tokens_out":16881,"duration_ms":158018,"significance":"If the main theorems are correct, this is a useful contribution: it gives explicit finite-sample rates for a widely used estimator under realistic noise assumptions, allows arbitrary dependence between errors and covariates, and identifies the local envelope as the quantity that controls the moment threshold. The new truncation-peeling result (Theorem C.1) and the finite-maximum maximal inequality (Proposition B.1) are potentially reusable tools. The paper is also honest about what is not proved: it states open questions about sharpness in Corollary 3.1 and says that the necessity of the envelope-growth conditions is under investigation. Two load-bearing points, however, need attention before the results can be taken as established: the proof of Theorem 4.1 omits the condition needed for the sup-norm chaining bound, and the multiple-index verification of the envelope condition rests on an assumption that fails for a natural part of the parameter space. These are fixable in a revision, in my view.","major_comments":[{"comment":"The verification of condition (13) for the multiple-index model is not valid as stated. Proposition A.3 requires ((BX)^T,(B0X)^T) to have a density on R^{2p} that is bounded below by C>0. This is impossible when 2p>d (for instance, a two-index model in d=3 gives a random vector in R^4 supported on a lower-dimensional subspace), and even when 2p<=d the lower bound cannot hold uniformly as B approaches B0 because the joint distribution degenerates onto the diagonal subspace. Consequently the sentence \"By Proposition A.3, M_{gamma,d,d1} satisfies (13) with s=2gamma/(2gamma+d1)\" is not supported for the full parameter space. The authors should either add an explicit design condition (for example, d>=2p and a uniform lower bound on the density in a neighborhood of the parameter space) or replace the multiple-index example with one for which the envelope growth can be verified.","section":"Section 4.1 and Proposition A.3"},{"comment":"The proof of the maximal inequality used for Theorem 4.1 requires the sup-norm entropy integral sum_t 2^{t(1-1/q)} epsilon_{infinity,t} to converge, which is true only if alpha(1-1/q)<1. This condition is not stated in Theorem 4.1. For alpha in (1,2) and q>alpha/(alpha-1) the displayed bound is not available, so the rates in (14) and (17) are not established in that regime. This is not merely a technical nuisance: the threshold q* displayed after (17) is always smaller than alpha/(alpha-1) for alpha>1, so the paper's stated moment threshold lies below the range covered by the proof, and for larger q the theorem is simply unproved. Please add the condition alpha(1-1/q)<1 to the theorem (possibly with a remark on how larger q can be handled by using a smaller q0 in Proposition B.1) or supply the missing truncated-chaining argument.","section":"Supplement S.6, Eq. (S.17), and Theorem 4.1"},{"comment":"The sup-norm envelope condition (13) is not implied by the L-infinity entropy assumption, and for natural non-smooth classes such as monotone or convex functions the sup-norm local envelope is of constant order, so only s=0 is available in Theorem 4.1, leading to the slower rates of Corollary 4.1. The paper does verify the L2-envelope analogue for convex and isotonic regression, but those verifications are used with Theorems 3.1 and 5.1, not with Theorem 4.1. The main text should state more explicitly that the improved moment threshold in Theorem 4.1 applies only to classes for which (13) has been established, and that for shape-constrained classes the corresponding claims live in Sections 3 and 5. As written, the abstract and Section 2.4 could be read as promising the n^{-1/(2+alpha)} rate for a broad family of heavy-tailed heteroscedastic problems without this caveat.","section":"Section 2.3 and Theorem 4.1, condition (13)"}],"minor_comments":[{"comment":"The entropy exponent is defined as alpha in [0,2) for (L2) and (L-infinity), but Sections 2.4 and 2.5 use alpha in [0,2]; the range should be made consistent.","section":"Section 2.2"},{"comment":"The text says Theorem 5.1 improves \"Theorem 2 of [36]\" but then compares with \"[36, Theorem 1]\"; clarify which result is meant.","section":"Remark 5.2"},{"comment":"The misspecification discussion states that the proofs go through when F is convex by replacing epsilon with xi=Y-bar f(X), but E(xi|X) is not zero; because Theorem C.1 relies on the identity E[M_n(f)-M_n(f0)]=||f-f0||^2, the statement needs either a proof or an explicit \"sketch only\" caveat.","section":"Section 6"},{"comment":"The version I received contains pervasive typographical artifacts such as \"/T_he\" and \"/f_ind\" in displayed text; the final version should be typeset cleanly.","section":"General typography"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid, well-written contribution, but the proof gap in Theorem 4.1 and the unsupported multiple-index envelope verification are central enough that I cannot recommend acceptance before they are resolved. I would welcome a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read of 1909.02088. The paper does something real: it gives finite-sample tail bounds for the least squares estimator when errors have q moments and can depend on the covariates, under bracketing, L∞-entropy, and VC-type conditions. That genuinely extends Han-Wellner (error independent of X) and Mendelson (function class sub-Gaussian). Theorems 3.1, 4.1, and 5.1 are new, and Theorem C.1's peeling-with-truncation plus Proposition B.1's maximal inequality are useful tools in their own right. The examples—convex, Hölder, multiple-index, isotonic—are concrete, and the s=0 corollaries are honest about what happens when the local envelope doesn't shrink. I could not machine-check the supplement, but the proof structure is standard empirical process chaining/symmetrization/peeling and the rate formulas are internally consistent. No fitted constants; comparisons to external minimax bounds are clean.\n\nThe soft spots, in order. First, Proposition A.3: the density assumption on ((BX),(B0X)) on R^{2p} cannot hold when 2p > d—the vectors live in a subspace of dimension at most d—and it degenerates as B→B0. So the multiple-index example's verification of the envelope condition is incomplete as stated. That's localized to the example; the core theorems don't depend on it, but the claim for M_{γ,d,d1} needs a fix (restrict to 2p≤d or use a different argument). Second, Section 6's misspecification claim is explicitly 'we leave the details to the reader'—fine as a sketch, but it shouldn't read like a proved result. Third, Theorem 3.1's tail exponent 1{s=1}/10 is needlessly cryptic; Remark 3.2 clears it up, so I call that minor presentation.\n\nOn the stress-test note: the worry that (13) is load-bearing is basically right but off-target as a criticism—the paper states (13) as an assumption and is explicit about the s=0 fallbacks. The honest summary is that the minimax-rate claims are proven for classes where the sup-norm local envelope shrinks; for monotone/convex classes that envelope is constant order and the L2 envelope results (Theorem 5.1) are what carry the adaptation claims. The reader's CONDITIONAL verdict is fair. I'd send this to a serious referee, with the appendix bug flagged. Worth citing for Theorem 4.1 and the peeling lemma.","headline":"Solid new rates for the LSE under heavy-tailed heteroscedastic errors, worth refereeing; the multiple-index example has a dimensional bug in a density assumption, and the local envelope condition does more work than the packaging suggests.","tokens_in":48624,"tokens_out":4428,"would_cite":true,"duration_ms":43179,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G08","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that the least squares estimator attains its usual minimax rate in nonparametric regression even when errors are heavy-tailed and heteroscedastic, provided the noise has enough finite moments and the function class…","keywords":["least squares estimator","heavy-tailed errors","heteroscedastic errors","nonparametric regression","rate of convergence","local envelope","metric entropy","finite moments"],"falsifier":"Produce a uniformly bounded function class satisfying the paper's entropy condition and a point $f_0$ where the local envelope satisfies the stated growth with parameter $s$, yet the least squares estimator with errors having the theorem's required number $q$ of moments provably fails to converge at the claimed rate, e.g. empirically the $L_2$ error decays strictly slower than $n^{-1/(2+\\alpha)}$. A concrete starting point is the class of 1-Lipschitz functions on $[0,1]$ with $f_0(x)=x$, where $\\alpha=1$ and $s=2/3$: the theorem predicts the $n^{-1/3}$ rate when $q\\ge 7/3$, so a simulation or matching lower bound showing that the LSE with exactly three conditional moments cannot reach $O_p(n^{-1/3})$ would contradict the central claim.","tokens_in":47578,"feed_emoji":"📊","tokens_out":7541,"duration_ms":79034,"temperature":0.7,"pith_summary":"Nonparametric least squares is known to be rate-optimal under sub-Gaussian errors, but real noise is often heavier-tailed and dependent on the covariates. This paper shows that the sub-Gaussian assumption can be replaced by a finite-moment condition together with a geometric condition on the function class near the true regression function. The resulting rates interpolate between the familiar $n^{-1/(2+\\alpha)}$ sub-Gaussian rate and slower rates, with the required number of error moments given by an explicit formula in terms of the entropy parameter $\\alpha$ and a local envelope parameter $s$. The bounds are finite-sample, allow heteroscedastic errors, and come with polynomial tail bounds. A matching lower bound shows that the local envelope parameter genuinely controls the worst-case rate, so the new condition is not a proof artifact.","feed_headline":"Least squares keeps its rate with heavy-tailed errors","feed_subtitle":"Finite moments plus local structure replace sub-Gaussian noise assumptions in nonparametric regression.","key_machinery":"The load-bearing object is the local envelope $F_\\delta(x)=\\sup_{f:\\|f-f_0\\|\\le\\delta}|(f-f_0)(x)|$, and its assumed growth $\\|F_\\delta\\|_\\infty\\le C\\Phi^{1-s}\\delta^s$ or the corresponding $L_2$/$L_q$ versions. This envelope converts the local geometry of the function class around the truth into a clean scaling $\\delta^s$, which combines with metric or bracketing entropy conditions to control the empirical process. The other key mechanism is a new peeling theorem with truncation (Theorem C.1) that handles an unbounded empirical process under finite moments by splitting the tail into a bounded, truncatable part and a remainder controlled by a $q$th-moment Markov bound; a new maximal inequality for maxima over finite sets (Proposition B.1) supplies the control needed in the $L_\\infty$-entropy case.","core_discovery":"The central claim is a set of rate identities: under $L_2$-bracketing entropy, the least squares estimator converges at the sub-Gaussian rate $n^{-1/(2+\\alpha)}$ once the error has at least $2/s$ conditional moments; under $L_\\infty$-entropy and a sup-norm local envelope bound $\\|F_\\delta\\|_\\infty \\le C\\Phi^{1-s}\\delta^s$, the rate is at most $\\max\\{n^{-1/(2+\\alpha)}, n^{-(q-1)/(q(2-s)+\\alpha s(q-1))}\\}$, collapsing to $n^{-1/(2+\\alpha)}$ when $q\\ge (2+\\alpha(1-s))/(s+\\alpha(1-s))$; and under a VC-type entropy condition, only two moments suffice for the rate $n^{-1/(2(2-s))}$. These rates hold for heteroscedastic errors that may depend on the covariates, and the tail probability of $\\delta_n^{-1}\\|\\hat f-f_0\\|$ decays polynomially at degree close to $q$, not exponentially as in the sub-Gaussian case.","pith_inferences":["The explicit threshold formula could be used as a practical diagnostic: for a given model class, estimate the entropy exponent $\\alpha$ and the local envelope exponent $s$, then decide whether the error's moment budget is large enough to trust ordinary least squares or whether a robust estimator is needed.","Because the local envelope can depend on the true function $f_0$, the same class can exhibit different rates at different truths; this suggests an adaptive story where shape-constrained least squares automatically speeds up when the truth lies in a smoother submodel, without any model-selection step.","The truncation-based peeling argument is not tied to squared error, so the same technique likely transfers to other smooth losses with heavy-tailed inputs, giving analogous moment-threshold formulas for robust regression with non-quadratic losses.","If one could estimate the envelope scaling $s$ from data, the theorem suggests a testable extension: compare the empirical scaling of $F_\\delta$ near a candidate truth against the moment threshold, and choose the estimator class accordingly."],"forward_implications":["Sub-Gaussianity is not necessary for minimax-rate least squares in standard nonparametric classes: Hölder, Sobolev, Lipschitz, convex, and isotonic regression all satisfy the envelope growth condition with explicit $s$, leading to explicit moment thresholds.","For univariate convex regression, the paper shows the convex LSE converges at $n^{-2/5}$ up to logarithmic factors when $E(|\\epsilon|^3\\mid X)$ is bounded, a setting where only sub-Gaussian-rate results were previously known.","For isotonic regression at a constant truth, the LSE achieves a near-parametric $n^{-1/2}$ rate, up to log factors, using only a bounded conditional second moment and allowing heteroscedastic errors dependent on $X$.","The paper's lower bound shows that for a constructed VC-type class, the LSE cannot beat $n^{-1/(2(2-s))}$ up to log factors even with finite-variance errors, so the envelope parameter $s$ is a genuine driver of worst-case rates rather than a technical convenience."],"supporting_citations":[{"why":"Supplies the standard peeling theorem and empirical process machinery that the paper refines into its truncation-based Theorem C.1.","marker":"[76]"},{"why":"Establishes heavy-tailed LSE rates under the assumption that errors are independent of covariates, the baseline that this paper relaxes.","marker":"[37]"},{"why":"Provides the envelope perspective and the lower-bound construction for VC-type classes that the paper adapts in Theorem 5.2.","marker":"[36]"},{"why":"Gives the interpolation inequalities and local envelope calculations used to compute the envelope growth parameter $s$ for Hölder, Sobolev, and Lipschitz classes.","marker":"[14]"},{"why":"Supplies symmetrization, contraction, and entropy facts used to bound the empirical process in the proofs of Theorems 3.1 and 4.1.","marker":"[27]"},{"why":"Provides the empirical process theory and entropy-integral framework used to translate local complexity into rate bounds.","marker":"[72]"},{"why":"Gives maximal inequalities for multiplier empirical processes with $q$ moments; the paper compares its new maximal inequality against this approach.","marker":"[55]"},{"why":"Establishes sub-Gaussian-class learning rates and highlights the gap between the empirical process bound and the local envelope that motivates the paper's setup.","marker":"[47]"}],"fun_headline_variants":["Heavy-tailed errors don't slow least squares convergence","LSE matches sub-Gaussian rate with just finite moments","Heavy tails, same speed: least squares rate holds","Least squares wants moments, not Gaussian tails"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The local envelope of the function class around the true regression function must shrink at least like a power of the radius $\\delta$ in the relevant norm; if functions very close to $f_0$ still differ from it at many points, the claimed rates and moment thresholds are not guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["Heavy-tailed errors don't slow least squares convergence","LSE matches sub-Gaussian rate with just finite moments","Heavy tails, same speed: least squares rate holds","Least squares wants moments, not Gaussian tails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000356,"raw_usage":{"total_tokens":1916,"prompt_tokens":910,"completion_tokens":1006,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":943}},"tokens_in":526,"tokens_out":1006,"duration_ms":8799,"temperature":1.0,"reasoning_tokens":943,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:00:26.970735+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Produce a uniformly bounded function class satisfying the paper's entropy condition and a point $f_0$ where the local envelope satisfies the stated growth with parameter $s$, yet the least squares estimator with errors having the theorem's required number $q$ of moments provably fails to converge at the claimed rate, e.g. empirically the $L_2$ error decays strictly slower than $n^{-1/(2+\\alpha)}$. A concrete starting point is the class of 1-Lipschitz functions on $[0,1]$ with $f_0(x)=x$, where $\\alpha=1$ and $s=2/3$: the theorem predicts the $n^{-1/3}$ rate when $q\\ge 7/3$, so a simulation or matching lower bound showing that the LSE with exactly three conditional moments cannot reach $O_p(n^{-1/3})$ would contradict the central claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the standard peeling theorem and empirical process machinery that the paper refines into its truncation-based Theorem C.1."},{"cited_title":"and Wellner, J","cited_arxiv_id":null,"evidence_quote":"Establishes heavy-tailed LSE rates under the assumption that errors are independent of covariates, the baseline that this paper relaxes."},{"cited_title":"and Shen, X","cited_arxiv_id":null,"evidence_quote":"Gives the interpolation inequalities and local envelope calculations used to compute the envelope growth parameter $s$ for Hölder, Sobolev, and Lipschitz classes."},{"cited_title":"and Nickl, R","cited_arxiv_id":null,"evidence_quote":"Supplies symmetrization, contraction, and entropy facts used to bound the empirical process in the proofs of Theorems 3.1 and 4.1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the empirical process theory and entropy-integral framework used to translate local complexity into rate bounds."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives maximal inequalities for multiplier empirical processes with $q$ moments; the paper compares its new maximal inequality against this approach."}],"review_version":1}