REVIEW 5 cited by
Optimistic Rates for Learning with a Smooth Loss
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We establish an excess risk bound of O(H R_n^2 + R_n \sqrt{H L*}) for empirical risk minimization with an H-smooth loss function and a hypothesis class with Rademacher complexity R_n, where L* is the best risk achievable by the hypothesis class. For typical hypothesis classes where R_n = \sqrt{R/n}, this translates to a learning rate of O(RH/n) in the separable (L*=0) case and O(RH/n + \sqrt{L^* RH/n}) more generally. We also provide similar guarantees for online and stochastic convex optimization with a smooth non-negative objective.
Forward citations
Cited by 5 Pith papers
-
Optimistic Rates for Multiclass PAC Learning
For multiclass PAC learning, the optimal excess risk at any fixed oracle error L* equals the square root of L* times the Natarajan dimension over n, plus the realizable DS-dimension rate, with matching upper and lower bounds.
-
Online Learning and Unlearning
Two OGD-based algorithms provably make deleted points statistically invisible in future outputs while adding only modest regret overhead.
-
On Least Squares Estimation under Heteroscedastic and Heavy-Tailed Errors
Under finite moments and a local envelope growth condition, the least squares estimator in nonparametric regression can achieve minimax rates with heavy-tailed, covariate-dependent errors.
-
Tight Generalization Bound for AdaBoost
AdaBoost’s generalization error is Θ(d ln(nγ²/d)/(nγ²) + ln(1/δ)/n), via a new zero-margin-loss bound for voting classifiers.
-
What Makes Local Updates Effective: The Role of Data Heterogeneity and Smoothness
Under bounded second-order heterogeneity, local updates are shown to achieve faster convergence than mini-batch SGD in several convex and non-convex regimes, with matching lower bounds.
Discussion (0). Continue with ORCID to comment.