REVIEW 2 cited by
Exponential Tail Local Rademacher Complexity Risk Bounds Without the Bernstein Condition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The local Rademacher complexity framework is one of the most successful general-purpose toolboxes for establishing sharp excess risk bounds for statistical estimators based on the framework of empirical risk minimization. Applying this toolbox typically requires using the Bernstein condition, which often restricts applicability to convex and proper settings. Recent years have witnessed several examples of problems where optimal statistical performance is only achievable via non-convex and improper estimators originating from aggregation theory, including the fundamental problem of model selection. These examples are currently outside of the reach of the classical localization theory. In this work, we build upon the recent approach to localization via offset Rademacher complexities, for which a general high-probability theory has yet to be established. Our main result is an exponential-tail excess risk bound expressed in terms of the offset Rademacher complexity that yields results at least as sharp as those obtainable via the classical theory. However, our bound applies under an estimator-dependent geometric condition (the "offset condition") instead of the estimator-independent (but, in general, distribution-dependent) Bernstein condition on which the classical theory relies. Our results apply to improper prediction regimes not directly covered by the classical theory.
Forward citations
Cited by 2 Pith papers
-
Reweighting Improves Conditional Risk Bounds
Weighted ERM with margin or inverse-variance weights tightens conditional risk bounds on high-confidence subregions by a constant factor, and yields an O(1/n) rate for estimating heteroscedastic variance.
-
On the Efficiency of ERM in Feature Learning
ERM over a union of linear feature classes achieves excess risk within a factor of two of the oracle that knows the optimal feature map, asymptotically, under size conditions on the feature set.
Discussion (0). Continue with ORCID to comment.