REVIEW 3 cited by
High-Dimensional $L_2$Boosting: Rate of Convergence
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Boosting is one of the most significant developments in machine learning. This paper studies the rate of convergence of $L_2$Boosting, which is tailored for regression, in a high-dimensional setting. Moreover, we introduce so-called \textquotedblleft post-Boosting\textquotedblright. This is a post-selection estimator which applies ordinary least squares to the variables selected in the first stage by $L_2$Boosting. Another variant is \textquotedblleft Orthogonal Boosting\textquotedblright\ where after each step an orthogonal projection is conducted. We show that both post-$L_2$Boosting and the orthogonal boosting achieve the same rate of convergence as LASSO in a sparse, high-dimensional setting. We show that the rate of convergence of the classical $L_2$Boosting depends on the design matrix described by a sparse eigenvalue constant. To show the latter results, we derive new approximation results for the pure greedy algorithm, based on analyzing the revisiting behavior of $L_2$Boosting. We also introduce feasible rules for early stopping, which can be easily implemented and used in applied work. Our results also allow a direct comparison between LASSO and boosting which has been missing from the literature. Finally, we present simulation studies and applications to illustrate the relevance of our theoretical results and to provide insights into the practical aspects of boosting. In these simulation studies, post-$L_2$Boosting clearly outperforms LASSO.
Forward citations
Cited by 3 Pith papers
-
Forward-Selected Panel Data Approach for Program Evaluation
Forward selection of control units in the panel data approach yields valid normal inference for average treatment effects even when the number of controls grows much faster than the time dimension and the true model is dense.
-
Nonparametric estimation of causal heterogeneity under high-dimensional confounding
The paper derives coupled convergence conditions under which a two-step estimator with machine-learned nuisance parameters consistently estimates group average treatment effects in high-dimensional settings, and shows...
-
Double Machine Learning for Conditional Moment Restrictions: IV Regression, Proximal Causal Learning and Beyond
A DML estimator for conditional moment restrictions is proposed, but its central N^{-1/2} rate theorem is broken because the selected score is degenerate at the truth.
Discussion (0). Continue with ORCID to comment.