Pith. sign in

REVIEW 3 major objections 4 minor 82 references

On Least Squares Estimation under Heteroscedastic and Heavy-Tailed Errors

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This paper establishes that the least squares estimator attains its usual minimax rate in nonparametric regression even when errors are heavy-tailed and heteroscedastic, provided the noise has enough finite moments and the function class…

desk verdict Solid new rates for the LSE under heavy-tailed heteroscedastic errors, worth refereeing; the multiple-index example has a dimensional bug in a density assumption, and the local envelope condition does more work than the packaging suggests. read the letter →

arxiv 1909.02088 v3 pith:CLHUP6QY submitted 2019-09-04 math.ST cs.LGstat.MLstat.TH

classification math.STcs.LGstat.MLstat.TH MSC 62G0862G20
keywords leastsquaresestimatorheavy-tailederrorsheteroscedasticnonparametricregressionrateofconvergencelocalenvelopemetricentropyfinitemoments
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Nonparametric least squares is known to be rate-optimal under sub-Gaussian errors, but real noise is often heavier-tailed and dependent on the covariates. This paper shows that the sub-Gaussian assumption can be replaced by a finite-moment condition together with a geometric condition on the function class near the true regression function. The resulting rates interpolate between the familiar $n^{-1/(2+\alpha)}$ sub-Gaussian rate and slower rates, with the required number of error moments given by an explicit formula in terms of the entropy parameter $\alpha$ and a local envelope parameter $s$. The bounds are finite-sample, allow heteroscedastic errors, and come with polynomial tail bounds. A matching lower bound shows that the local envelope parameter genuinely controls the worst-case rate, so the new condition is not a proof artifact.

What carries the argument

The load-bearing object is the local envelope $F_\delta(x)=\sup_{f:\|f-f_0\|\le\delta}|(f-f_0)(x)|$, and its assumed growth $\|F_\delta\|_\infty\le C\Phi^{1-s}\delta^s$ or the corresponding $L_2$/$L_q$ versions. This envelope converts the local geometry of the function class around the truth into a clean scaling $\delta^s$, which combines with metric or bracketing entropy conditions to control the empirical process. The other key mechanism is a new peeling theorem with truncation (Theorem C.1) that handles an unbounded empirical process under finite moments by splitting the tail into a bounded, truncatable part and a remainder controlled by a $q$th-moment Markov bound; a new maximal inequality for maxima over finite sets (Proposition B.1) supplies the control needed in the $L_\infty$-entropy case.

What would settle it

Produce a uniformly bounded function class satisfying the paper's entropy condition and a point $f_0$ where the local envelope satisfies the stated growth with parameter $s$, yet the least squares estimator with errors having the theorem's required number $q$ of moments provably fails to converge at the claimed rate, e.g. empirically the $L_2$ error decays strictly slower than $n^{-1/(2+\alpha)}$. A concrete starting point is the class of 1-Lipschitz functions on $[0,1]$ with $f_0(x)=x$, where $\alpha=1$ and $s=2/3$: the theorem predicts the $n^{-1/3}$ rate when $q\ge 7/3$, so a simulation or matching lower bound showing that the LSE with exactly three conditional moments cannot reach $O_p(n^{-1/3})$ would contradict the central claim.

Watch

Extended reading notes

Core claim

The central claim is a set of rate identities: under $L_2$-bracketing entropy, the least squares estimator converges at the sub-Gaussian rate $n^{-1/(2+\alpha)}$ once the error has at least $2/s$ conditional moments; under $L_\infty$-entropy and a sup-norm local envelope bound $\|F_\delta\|_\infty \le C\Phi^{1-s}\delta^s$, the rate is at most $\max\{n^{-1/(2+\alpha)}, n^{-(q-1)/(q(2-s)+\alpha s(q-1))}\}$, collapsing to $n^{-1/(2+\alpha)}$ when $q\ge (2+\alpha(1-s))/(s+\alpha(1-s))$; and under a VC-type entropy condition, only two moments suffice for the rate $n^{-1/(2(2-s))}$. These rates hold for heteroscedastic errors that may depend on the covariates, and the tail probability of $\delta_n^{-1}\|\hat f-f_0\|$ decays polynomially at degree close to $q$, not exponentially as in the sub-Gaussian case.

Load-bearing premise

The local envelope of the function class around the true regression function must shrink at least like a power of the radius $\delta$ in the relevant norm; if functions very close to $f_0$ still differ from it at many points, the claimed rates and moment thresholds are not guaranteed.

Editorial extensions

If this is right

  • Sub-Gaussianity is not necessary for minimax-rate least squares in standard nonparametric classes: Hölder, Sobolev, Lipschitz, convex, and isotonic regression all satisfy the envelope growth condition with explicit $s$, leading to explicit moment thresholds.
  • For univariate convex regression, the paper shows the convex LSE converges at $n^{-2/5}$ up to logarithmic factors when $E(|\epsilon|^3\mid X)$ is bounded, a setting where only sub-Gaussian-rate results were previously known.
  • For isotonic regression at a constant truth, the LSE achieves a near-parametric $n^{-1/2}$ rate, up to log factors, using only a bounded conditional second moment and allowing heteroscedastic errors dependent on $X$.
  • The paper's lower bound shows that for a constructed VC-type class, the LSE cannot beat $n^{-1/(2(2-s))}$ up to log factors even with finite-variance errors, so the envelope parameter $s$ is a genuine driver of worst-case rates rather than a technical convenience.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The explicit threshold formula could be used as a practical diagnostic: for a given model class, estimate the entropy exponent $\alpha$ and the local envelope exponent $s$, then decide whether the error's moment budget is large enough to trust ordinary least squares or whether a robust estimator is needed.
  • Because the local envelope can depend on the true function $f_0$, the same class can exhibit different rates at different truths; this suggests an adaptive story where shape-constrained least squares automatically speeds up when the truth lies in a smoother submodel, without any model-selection step.
  • The truncation-based peeling argument is not tied to squared error, so the same technique likely transfers to other smooth losses with heavy-tailed inputs, giving analogous moment-threshold formulas for robust regression with non-quadratic losses.
  • If one could estimate the envelope scaling $s$ from data, the theorem suggests a testable extension: compare the empirical scaling of $F_\delta$ near a candidate truth against the moment threshold, and choose the estimator class accordingly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper studies the least squares estimator (LSE) in nonparametric regression with errors that are heteroscedastic and have only finitely many moments. The authors provide finite-sample tail bounds for ||hat f - f0|| under three global complexity conditions: bracketing L2 entropy (Theorem 3.1), L-infinity entropy (Theorem 4.1), and uniform VC-type entropy (Theorem 5.1). The rates are expressed in terms of the entropy parameter alpha, the moment index q of the errors, and a local-envelope growth parameter s that controls how the sup-norm or L2-norm of the local envelope F_delta(x)=sup_{f:||f-f0||<=delta}|(f-f0)(x)| shrinks as delta goes to zero. The headline conclusion is that, under the sup-norm envelope condition (13), the LSE attains the minimax rate n^{-1/(2+alpha)} once q is at least (2+alpha(1-s))/(s+alpha(1-s)), with slower rates otherwise; under the bracketing condition the analogous threshold is q>=2/s. Applications include Holder, Sobolev, additive, multiple-index, convex, and isotonic regression, with new interpolation inequalities in Appendix A and a new truncated peeling theorem in Appendix C. Section 5.2 contains a lower bound showing the role of s in the worst-case LSE rate. The authors correctly stress that their conditions are sufficient and that some optimality questions remain open.

Significance. If the main theorems are correct, this is a useful contribution: it gives explicit finite-sample rates for a widely used estimator under realistic noise assumptions, allows arbitrary dependence between errors and covariates, and identifies the local envelope as the quantity that controls the moment threshold. The new truncation-peeling result (Theorem C.1) and the finite-maximum maximal inequality (Proposition B.1) are potentially reusable tools. The paper is also honest about what is not proved: it states open questions about sharpness in Corollary 3.1 and says that the necessity of the envelope-growth conditions is under investigation. Two load-bearing points, however, need attention before the results can be taken as established: the proof of Theorem 4.1 omits the condition needed for the sup-norm chaining bound, and the multiple-index verification of the envelope condition rests on an assumption that fails for a natural part of the parameter space. These are fixable in a revision, in my view.

major comments (3)
  1. [Section 4.1 and Proposition A.3] The verification of condition (13) for the multiple-index model is not valid as stated. Proposition A.3 requires ((BX)^T,(B0X)^T) to have a density on R^{2p} that is bounded below by C>0. This is impossible when 2p>d (for instance, a two-index model in d=3 gives a random vector in R^4 supported on a lower-dimensional subspace), and even when 2p<=d the lower bound cannot hold uniformly as B approaches B0 because the joint distribution degenerates onto the diagonal subspace. Consequently the sentence "By Proposition A.3, M_{gamma,d,d1} satisfies (13) with s=2gamma/(2gamma+d1)" is not supported for the full parameter space. The authors should either add an explicit design condition (for example, d>=2p and a uniform lower bound on the density in a neighborhood of the parameter space) or replace the multiple-index example with one for which the envelope growth can be verified.
  2. [Supplement S.6, Eq. (S.17), and Theorem 4.1] The proof of the maximal inequality used for Theorem 4.1 requires the sup-norm entropy integral sum_t 2^{t(1-1/q)} epsilon_{infinity,t} to converge, which is true only if alpha(1-1/q)<1. This condition is not stated in Theorem 4.1. For alpha in (1,2) and q>alpha/(alpha-1) the displayed bound is not available, so the rates in (14) and (17) are not established in that regime. This is not merely a technical nuisance: the threshold q* displayed after (17) is always smaller than alpha/(alpha-1) for alpha>1, so the paper's stated moment threshold lies below the range covered by the proof, and for larger q the theorem is simply unproved. Please add the condition alpha(1-1/q)<1 to the theorem (possibly with a remark on how larger q can be handled by using a smaller q0 in Proposition B.1) or supply the missing truncated-chaining argument.
  3. [Section 2.3 and Theorem 4.1, condition (13)] The sup-norm envelope condition (13) is not implied by the L-infinity entropy assumption, and for natural non-smooth classes such as monotone or convex functions the sup-norm local envelope is of constant order, so only s=0 is available in Theorem 4.1, leading to the slower rates of Corollary 4.1. The paper does verify the L2-envelope analogue for convex and isotonic regression, but those verifications are used with Theorems 3.1 and 5.1, not with Theorem 4.1. The main text should state more explicitly that the improved moment threshold in Theorem 4.1 applies only to classes for which (13) has been established, and that for shape-constrained classes the corresponding claims live in Sections 3 and 5. As written, the abstract and Section 2.4 could be read as promising the n^{-1/(2+alpha)} rate for a broad family of heavy-tailed heteroscedastic problems without this caveat.
minor comments (4)
  1. [Section 2.2] The entropy exponent is defined as alpha in [0,2) for (L2) and (L-infinity), but Sections 2.4 and 2.5 use alpha in [0,2]; the range should be made consistent.
  2. [Remark 5.2] The text says Theorem 5.1 improves "Theorem 2 of [36]" but then compares with "[36, Theorem 1]"; clarify which result is meant.
  3. [Section 6] The misspecification discussion states that the proofs go through when F is convex by replacing epsilon with xi=Y-bar f(X), but E(xi|X) is not zero; because Theorem C.1 relies on the identity E[M_n(f)-M_n(f0)]=||f-f0||^2, the statement needs either a proof or an explicit "sketch only" caveat.
  4. [General typography] The version I received contains pervasive typographical artifacts such as "/T_he" and "/f_ind" in displayed text; the final version should be typeset cleanly.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the rate and moment-threshold results are derived from stated entropy and envelope assumptions, with external minimax benchmarks.

full rationale

The derivation chain is self-contained in the sense that matters for circularity: Theorem 4.1 takes (L∞), (CVar), (Eq), and the sup-norm local envelope bound (13) as assumptions and derives the rate (17) and the moment threshold q* by algebra and chaining/peeling arguments; the envelope growth parameter s is a geometric property of the class around f0, not a parameter fitted to the LSE's rate or to any data. Theorem 3.1 similarly uses the Lq-envelope condition (9) as a sufficient condition, and the lower bound Theorem 5.2 constructs a class satisfying the assumptions rather than importing the upper bound's conclusion. The minimax comparisons are to external lower bounds (Birgé–Massart, Györfi et al., Yang–Barron) and to sub-Gaussian benchmarks [76]; no benchmark is defined in terms of the paper's own rate. The self-citations ([43], [44]) are applications or peripheral references and are not load-bearing for the main rate theorems. The local envelope condition can indeed fail for some natural classes, and Proposition A.3's density requirement is restrictive, but those are scope/correctness concerns about the assumptions' range of applicability, not circularity: the theorems explicitly condition on those assumptions. Thus no step reduces by construction to its inputs.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The central claims rest on structural conditions on the error distribution and the function class. No free parameters are fit to data; alpha, s, Phi, sigma, A are properties of the problem or constants in the bounds. The envelope growth condition is the assumption whose failure would break the rate claims. Standard empirical process inequalities are invoked as background.

assumptions (6)
  • domain assumption Uniform boundedness of the function class: sup_{f in F} ||f||_infty <= Phi
    Used throughout (e.g., Section 2.3, Theorem 3.1) to bound the empirical process via contraction and to construct the local envelope. The paper notes it can be relaxed to ||f_hat||_infty = O_p(1) in some examples.
  • domain assumption Bounded conditional variance (CVar): E(epsilon^2|X) <= sigma^2 a.e.
    Invoked in all main theorems; used to bound bracketing entropy of {epsilon(f-f0)} and the variance term in the peeling inequality. The paper restricts attention to sigma bounded away from zero.
  • domain assumption Finite q-th moment (Eq): E(|epsilon|^q) <= K_q^q for some q >= 2
    Required in Theorems 3.1 and 4.1; Theorem 5.1 needs only q=2 through (CVar). The tail decay order D^{-q} depends on this.
  • domain assumption Entropy conditions (L2), (L-infinity), or (VC(f0)) on the function class.
    Each theorem assumes one of these complexity measures; they control the metric entropy integrals used in maximal inequalities.
  • domain assumption Local envelope growth: ||(|epsilon|+Phi)F_delta(X)||_q <= C Phi^2 delta^s (or L-infinity analogues ||F_delta||_infty <= C Phi^{1-s} delta^s, ||F_delta|| <= C Phi^{1-s} delta^s).
    This is the key local structural assumption. It is verified for Holder, Sobolev, convex, and isotonic classes via interpolation inequalities, but is not automatic.
  • standard math Background empirical process inequalities: Lemma 3.4.2 of van der Vaart and Wellner, Proposition 3.1 of Gine et al., Nagaev's inequality.
    Used in proofs without reproving; these are standard results in the empirical process literature.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On Least Squares Estimation under Heteroscedastic and Heavy-Tailed Errors." pith.science (2026). https://pith.science/paper/CLHUP6QY

@misc{pith2026190902088,
  author       = {Pith},
  title        = {Pith review of: On Least Squares Estimation under Heteroscedastic and Heavy-Tailed Errors},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CLHUP6QY}},
  note         = {Machine review of arXiv:1909.02088}
}
read the original abstract

We consider least squares estimation in a general nonparametric regression model. The rate of convergence of the least squares estimator (LSE) for the unknown regression function is well studied when the errors are sub-Gaussian. We find upper bounds on the rates of convergence of the LSE when the errors have uniformly bounded conditional variance and have only finitely many moments. We show that the interplay between the moment assumptions on the error, the metric entropy of the class of functions involved, and the "local" structure of the function class around the truth drives the rate of convergence of the LSE. We find sufficient conditions on the errors under which the rate of the LSE matches the rate of the LSE under sub-Gaussian error. Our results are finite sample and allow for heteroscedastic and heavy-tailed errors.

Figures

Figures reproduced from arXiv: 1909.02088 by the authors.

Figure 1
Figure 1. Illustration of Fδ (le‰ panel) and Fδ (right panel) when F := {f : [0, 1] → R||f(x) − f(y)| ≤ |x − y|} and f0(x) = x for δ = .2 (solid gray) and δ = .05 (solid black). Any 1-Lipschitz function f that satis€es kf − f0k ≤ .2 lies in the “band” created by the two solid gray lines. Here s = 2/3. ‘e dashed line in the le‰ panel is f0. In general, if supf∈F kfk∞ ≤ Φ, then one can invoke the rich theory of interpolation in… view at source ↗
Figure 2
Figure 2. Illustration of Fδ (le‰ panel) and Fδ (right panel) when f0(x) = x 2 and F := {f : [0, 1] → R | kfk∞ ≤ 2 and f is convex} for δ = .2 (solid black) and δ = .05 (solid gray). Any convex function f that is uniformly bounded by 2 and satis€es kf − f0k ≤ .2 lies in the band created by the solid gray lines. ‘e dashed line in the le‰ panel is f0. for some s ∈ [0, 1], and let rn := min  (nA−1 ) 1/(2+α) (σ + Φ)2/(2+α) , n (… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

82 extracted references · 71 canonical work pages

  1. [1]

    Agmon, S. (2010). Lectures on elliptic boundary value problems . AMS Chelsea Publishing, Providence, RI. Prepared for publication by B. Frank Jones, Jr. with the assistance of George W. Ba/t_ten, Jr., Revised edition of the 1965 original

  2. [2]

    and Catoni, O

    Audibert, J.-Y. and Catoni, O. (2011). Robust linear least squares regression. /T_he Annals of Statistics, 39(5):2766–2794

  3. [3]

    Bartle/t_t, P. L. (2008). Fast rates for estimation error and oracle inequalities for model selection.Econo- metric /T_heory, 24(2):545–552. 22

  4. [4]

    Bellec, P. C. (2018). Sharp oracle inequalities for least squares estimators in shape restricted regression. /T_he Annals of Statistics, 46(2):745–780

  5. [5]

    Birge, L. (1989). /T_he grenader estimator: A nonasymptotic approach. /T_he Annals of Statistics, pages 1532–1549

  6. [6]

    and Massart, P

    Birg ´e, L. and Massart, P. (1993). Rates of convergence for minimum contrast estimators. Probability /T_heory and Related Fields, 97(1-2):113–150

  7. [7]

    and Massart, P

    Birg ´e, L. and Massart, P. (1998). Minimum contrast estimators on sieves: exponential bounds and rates of convergence. Bernoulli, 4(3):329–375

  8. [8]

    Boucheron, S., Lugosi, G., and Massart, P. (2013). Concentration inequalities: A nonasymptotic theory of independence. Oxford university press

Show all 82 references
  1. [9]

    Brownlees, C., Joly, E., and Lugosi, G. (2015). Empirical risk minimization for heavy-tailed losses. /T_he Annals of Statistics, 43(6):2507–2536

  2. [10]

    Buja, A., Hastie, T., and Tibshirani, R. (1989). Linear smoothers and additive models. Ann. Statist., 17(2):453–555

  3. [11]

    Cha/t_terjee, S., Guntuboyina, A., and Sen, B. (2015). On risk bounds in isotonic shape restricted re- gression problems. /T_he Annals of Statistics, 43(4):1774–1800

  4. [12]

    Cha/t_terjee, S., Guntuboyina, A., and Sen, B. (2018). On matrix estimation under monotonicity con- straints. Bernoulli, 24(2):1072–1100

  5. [13]

    and Lafferty, J

    Cha/t_terjee, S. and Lafferty, J. (2015). Adaptive risk bounds in unimodal regression. arXiv preprint arXiv:1512.02956

  6. [14]

    and Shen, X

    Chen, X. and Shen, X. (1998). Sieve extremum estimates for weakly dependent data. Econometrica, pages 289–314

  7. [15]

    Chernozhukov, V., Chetverikov, D., and Kato, K. (2015). Comparison and anti-concentration bounds for maxima of Gaussian random vectors. Probab. /T_heory Related Fields, 162(1-2):47–70

  8. [16]

    de la Pe ˜na, V. H. and Gin ´e, E. (1999). Decoupling. Probability and its Applications (New York). Springer-Verlag, New York. From dependence to independence, Randomly stopped processes. U- statistics and processes. Martingales and beyond

  9. [17]

    Dirksen, S. (2015). Tail bounds via generic chaining. Electron. J. Probab., 20:no. 53, 1–29

  10. [18]

    Doss, C. R. (2015). Bracketing Numbers of Convex Functions on Polytopes. arXiv preprint arXiv:1506.00034

  11. [19]

    A., Veraar, M

    D ¨umbgen, L., Van De Geer, S. A., Veraar, M. C., and Wellner, J. A. (2010). Nemirovski’s inequalities revisited. /T_he American Mathematical Monthly, 117(2):138–160

  12. [20]

    Friedman, J. H. and Stuetzle, W. (1981). Projection pursuit regression. Journal of the American statis- tical Association, 76(376):817–823

  13. [21]

    and Gerchinovitz, S

    Gaillard, P. and Gerchinovitz, S. (2015). A chaining algorithm for online nonparametric regression. In Conference on Learning /T_heory, pages 764–796. 23

  14. [22]

    Gao, C., Han, F., and Zhang, C.-H. (2017). On Estimation of Isotonic Piecewise Constant Signals. arXiv preprint arXiv:1705.06386

  15. [23]

    Gao, C., Han, F., and Zhang, C.-H. (2020). On estimation of isotonic piecewise constant signals.Annals of Statistics, 48(2):629–654

  16. [24]

    and Sen, B

    Ghosal, P. and Sen, B. (2017). On univariate convex regression. Sankhya A, 79(2):215–253

  17. [25]

    and Koltchinskii, V

    Gin ´e, E. and Koltchinskii, V. (2006). Concentration inequalities and asymptotic results for ratio type empirical processes. Ann. Probab., 34(3):1143–1216

  18. [26]

    Gin ´e, E., Latała, R., and Zinn, J. (2000). Exponential and moment inequalities for U-statistics. In High dimensional probability, II (Sea/t_tle, W A, 1999), volume 47 of Progr. Probab., pages 13–38. Birkh¨auser Boston, Boston, MA

  19. [27]

    and Nickl, R

    Gin ´e, E. and Nickl, R. (2016).Mathematical foundations of in/f_inite-dimensional statistical models. Cam- bridge Series in Statistical and Probabilistic Mathematics, [40]. Cambridge University Press, New York

  20. [28]

    and Zinn, J

    Gin ´e, E. and Zinn, J. (1983). Central limit theorems and weak laws of large numbers in certain banach spaces. Zeitschri/f_t f¨ur Wahrscheinlichkeitstheorie und Verwandte Gebiete, 62(3):323–354

  21. [29]

    and Lepski, O

    Goldenshluger, A. and Lepski, O. (2020). Minimax estimation of norms of a probability density: Ii. rate-optimal estimation procedures. arXiv:2008.10987

  22. [30]

    and Sen, B

    Guntuboyina, A. and Sen, B. (2015a). Global risk bounds and adaptation in univariate convex regres- sion. Probability /T_heory and Related Fields, 163(1-2):379–411

  23. [31]

    and Sen, B

    Guntuboyina, A. and Sen, B. (2015b). Global risk bounds and adaptation in univariate convex regres- sion. Probab. /T_heory Related Fields, 163(1-2):379–411

  24. [32]

    and Sen, B

    Guntuboyina, A. and Sen, B. (2018). Nonparametric shape-restricted regression. Statistical Science, 33(4):568–594

  25. [33]

    Gy ¨or/f_i, L., Kohler, M., Krzy˙zak, A., and Walk, H. (2002). A distribution-free theory of nonparametric regression. Springer Series in Statistics. Springer-Verlag, New York

  26. [34]

    Han, Q., Wang, T., Cha/t_terjee, S., and Samworth, R. J. (2017). Isotonic regression in general dimen- sions. arXiv preprint arXiv:1708.09468

  27. [35]

    and Wellner, J

    Han, Q. and Wellner, J. A. (2017). A sharp multiplier inequality with applications to heavy-tailed regression problems. arXiv preprint arXiv:1706.02410v1

  28. [36]

    and Wellner, J

    Han, Q. and Wellner, J. A. (2018). Robustness of shape-restricted regression estimators: an envelope perspective. ArXiv preprints arxiv:1805.02542

  29. [37]

    and Wellner, J

    Han, Q. and Wellner, J. A. (2019). Convergence rates of least squares regression estimators with heavy-tailed errors. Ann. Statist., 47:2286 – 2319

  30. [38]

    Hastie, T., Tibshirani, R., and Wainwright, M. (2015). Statistical learning with sparsity: the lasso and generalizations. Chapman and Hall/CRC

  31. [39]

    Hristache, M., Juditsky, A., Polzehl, J., and Spokoiny, V. (2001). Structure adaptive approach for dimension reduction. Ann. Statist., 29(6):1537–1566. 24

  32. [40]

    Kolmogorov, A. N. (1949). On inequalities between the upper bounds of the successive derivatives of an arbitrary function on an in/f_inite interval.American Mathematical Society Translations, (1-2):233–243

  33. [41]

    Koltchinskii, V. (2006). Local rademacher complexities and oracle inequalities in risk minimization. /T_he Annals of Statistics, 34(6):2593–2656

  34. [42]

    Koltchinskii, V. (2011). Oracle inequalities in empirical risk minimization and sparse recovery problems, volume 2033 of Lecture Notes in Mathematics . Springer, Heidelberg. Lectures from the 38th Probability Summer School held in Saint-Flour, 2008

  35. [43]

    Kuchibhotla, A. K. and Chakrabor/t_ty, A. (2018). Moving beyond sub-gaussianity in high-dimensional statistics: Applications in covariance estimation and linear regression. arXiv preprint arXiv:1804.02605

  36. [44]

    K., Patra, R

    Kuchibhotla, A. K., Patra, R. K., and Sen, B. (2021). Semiparametric Efficiency in Convex- ity Constrained Single Index Model. Journal of the American Statistical Association (to appear). arXiv:1708.00145v3

  37. [45]

    Kur, G., Dagan, Y., and Rakhlin, A. (2019). Optimality of maximum likelihood for log-concave density estimation and bounded convex regression. arXiv:1903.05315

  38. [46]

    and Lerasle, M

    Lecu ´e, G. and Lerasle, M. (2020). Robust machine learning by median-of-means: theory and practice. Annals of Statistics, 48(2):906–931

  39. [47]

    and Mendelson, S

    Lecu ´e, G. and Mendelson, S. (2013). Learning subgaussian classes: Upper and minimax bounds.arXiv preprint arXiv:1305.4825

  40. [48]

    and Mendelson, S

    Lecu ´e, G. and Mendelson, S. (2016). Performance of empirical risk minimization in linear aggregation. Bernoulli, 22(3):1520–1534

  41. [49]

    and Mendelson, S

    Lugosi, G. and Mendelson, S. (2016). Risk minimization by median-of-means tournaments. arXiv preprint arXiv:1608.00757

  42. [50]

    and N ´ed´elec, ´E

    Massart, P. and N ´ed´elec, ´E. (2006). Risk bounds for statistical learning. /T_he Annals of Statistics, 34(5):2326–2366

  43. [51]

    and Rossignol, R

    Massart, P. and Rossignol, R. (2013). Around nemirovski’s inequality. In From Probability to Statistics and Back: High-Dimensional Models and Processes–A Festschri/f_t in Honor of Jon A. Wellner , pages 254–

  44. [52]

    Mendelson, S. (2008). On weakly bounded empirical processes. Mathematische Annalen, 340(2):293– 314

  45. [53]

    Mendelson, S. (2014). Learning without concentration. In Conference on Learning /T_heory, pages 25–39

  46. [54]

    Mendelson, S. (2015). ‘Local’ vs. ‘global’ parameters – breaking the gaussian complexity barrier. ArXiv preprints arXiv:1504.02191

  47. [55]

    Mendelson, S. (2016). Upper bounds on product and multiplier empirical processes. Stochastic Pro- cesses and their Applications, 126(12):3652–3680

  48. [56]

    Mendelson, S. (2019). An unrestricted learning procedure. Journal of the ACM (JACM) , 66(6):1–42

  49. [57]

    Nagaev, S. V. (1979). Large deviations of sums of independent random variables. /T_he Annals of Probability, pages 745–789. 25

  50. [58]

    and P ¨otscher, B

    Nickl, R. and P ¨otscher, B. M. (2007). Bracketing metric entropy rates and empirical central limit theorems for function classes of besov-and sobolev-type. Journal of /T_heoretical Probability, 20(2):177– 199

  51. [59]

    Nirenberg, L. (2011). On elliptic partial differential equations. In /T_he principle of minimum and its applications to functional equations , pages 1–48. Springer

  52. [60]

    Pollard, D. (1984). Convergence of stochastic processes . Springer Series in Statistics. Springer-Verlag, New York

  53. [61]

    Rakhlin, A., Sridharan, K., and Tsybakov, A. B. (2017). Empirical entropy, minimax regret and minimax risk. Bernoulli, 23(2):789–824

  54. [62]

    and Wong, W

    Shen, X. and Wong, W. H. (1994). Convergence rate of sieve estimates. /T_he Annals of Statistics, pages 580–615

  55. [63]

    Srebro, N., Sridharan, K., and Tewari, A. (2010). Optimistic rates for learning with a smooth loss. arXiv preprint arXiv:1009.3896

  56. [64]

    Stone, C. J. (1994). /T_he use of polynomial splines and their tensor products in multivariate function estimation. Ann. Statist., 22(1):118–184. With discussion by Andreas Buja and Trevor Hastie and a rejoinder by the author

  57. [65]

    Talagrand, M. (1996). Majorizing measures: the generic chaining. Ann. Probab., 24(3):1049–1103

  58. [66]

    Talagrand, M. (2014). Upper and lower bounds for stochastic processes , volume 60 of A Series of Modern Surveys in Mathematics. Springer, Heidelberg. Modern methods and classical problems

  59. [67]

    van de Geer, S. (1990). Estimating a regression function. /T_he Annals of Statistics, pages 907–924

  60. [68]

    and Lederer, J

    van de Geer, S. and Lederer, J. (2013). /T_he Bernstein-Orlicz norm and deviation inequalities.Probab. /T_heory Related Fields, 157(1-2):225–250

  61. [69]

    and Muro, A

    van de Geer, S. and Muro, A. (2014). On higher order isotropy conditions and lower bounds for sparse quadratic forms. Electronic Journal of Statistics , 8(2):3031–3061

  62. [70]

    and Wainwright, M

    van de Geer, S. and Wainwright, M. J. (2017). On concentration for (regularized) empirical risk mini- mization. Sankhya A, 79(2):159–200

  63. [71]

    and Wegkamp, M

    van de Geer, S. and Wegkamp, M. (1996). Consistency for the least squares estimator in nonparametric regression. /T_he Annals of Statistics, pages 2513–2523

  64. [72]

    van de Geer, S. A. (2000). Applications of empirical process theory , volume 6 of Cambridge Series in Statistical and Probabilistic Mathematics . Cambridge University Press, Cambridge

  65. [73]

    van der Vaart, A. (2002). Semiparametric statistics. In Lectures on probability theory and statistics (Saint-Flour, 1999), volume 1781 of Lecture Notes in Math. , pages 331–457. Springer, Berlin

  66. [74]

    and Wellner, J

    van der Vaart, A. and Wellner, J. A. (2011). A local maximal inequality under uniform entropy. Elec- tronic Journal of Statistics , 5(2011):192

  67. [75]

    van der Vaart, A. W. (1998). Asymptotic statistics , volume 3 of Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, Cambridge. 26

  68. [76]

    van der Vaart, A. W. and Wellner, J. A. (1996). Weak convergence and empirical processes . Springer Series in Statistics. Springer-Verlag, New York

  69. [77]

    and Barron, A

    Yang, Y. and Barron, A. (1999). Information-theoretic determination of minimax rates of convergence. Annals of Statistics, pages 1564–1599

  70. [78]

    On Least Squares Estimation Under Heteroscedastic and Heavy-Tailed Errors

    Zhang, C.-H. (2002). Risk bounds in isotonic regression. /T_he Annals of Statistics, 30(2):528–555. 27 Supplement to “On Least Squares Estimation Under Heteroscedastic and Heavy-Tailed Errors” S.1 Discussion on the local envelope In this section, we will provide a heuristic ar...

  71. [80]

    7Note that, we can chooseB to be∞

    If α∈ [0, 2) andβ≥ 0, then φn(δ,B )≤CA1/2(σ + Φ)Φδs, for everyδ,B >0.7 (S.21) 6Hereα andβ are as described in (VC(f0)). 7Note that, we can chooseB to be∞. 42

  72. [81]

    If α = 2 andβ = 0, then φn(δ,B )≤CA1/2(σ + Φ)Φδs log ( A−1n ) , for alln≥A and everyδ,B >0. (S.22)

  73. [82]

    (S.23) Application of /T_heorem C.1:To apply /T_heorem C.1, we needτ, γ, andsγ(δ)

    If α> 2 andβ = 0, then φn(δ,B )≤CA1/αn1/2−1/α(σ + Φ)Φδs, for alln≥A and everyδ,B >0. (S.23) Application of /T_heorem C.1:To apply /T_heorem C.1, we needτ, γ, andsγ(δ). If we choose τ =γ = 2 and s2(δ) = E ( |U(ϵ,X,δ )|2) ≡ 22σ2Φ2δ2s, thensq(δ) satis/f_iesE [Uγ(ϵ,X ;δ)]≤sγ(δ). /...

  74. [265]

    Institute of Mathematical Statistics

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.