REVIEW 7 cited by
Deep Equals Shallow for ReLU Networks in Kernel Regimes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep networks are often considered to be more expressive than shallow ones in terms of approximation. Indeed, certain functions can be approximated by deep networks provably more efficiently than by shallow ones, however, no tractable algorithms are known for learning such deep models. Separately, a recent line of work has shown that deep networks trained with gradient descent may behave like (tractable) kernel methods in a certain over-parameterized regime, where the kernel is determined by the architecture and initialization, and this paper focuses on approximation for such kernels. We show that for ReLU activations, the kernels derived from deep fully-connected networks have essentially the same approximation properties as their shallow two-layer counterpart, namely the same eigenvalue decay for the corresponding integral operator. This highlights the limitations of the kernel framework for understanding the benefits of such deep architectures. Our main theoretical result relies on characterizing such eigenvalue decays through differentiability properties of the kernel function, which also easily applies to the study of other kernels defined on the sphere.
Forward citations
Cited by 7 Pith papers
-
The Cost of Discretization in Functional Linear Regression: Minimax Rates and Adaptation
Matching minimax prediction rates for discretely observed functional linear regression are n^{-ν/(ν+1)}+(nm)^{-ν/κ} under independent design, and those two terms plus m^{-ν}+m^{-4α} under common design.
-
Large Dimensional Kernel Ridge Regression: Extending to Product Kernels
Extends high-dimensional KRR to product kernels, proving convergence rates that recover minimax optimality for source condition s ≤ 1, saturation for s > 1, and multiple-descent phenomena with respect to sample size n.
-
The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently
Temporal correlations from lazy random walks enable efficient SGD learning of k-juntas via temporal-difference loss on ReLU networks, achieving linear sample complexity in d.
-
Variable Importance Identification Through Lazy Training for Binary Classification
LazyVI tests feature importance in binary classifiers by fitting a linearized logistic model on neural tangent features and claims op(n^{-1/2}) asymptotic normality.
-
Revisiting the Neural Tangent Kernel: the role of large width and depth
For infinite-width ReLU nets, the normalized neural tangent kernel provably collapses toward the all-ones matrix with depth, while the claimed convergence of the kernel-regression solution to a nontrivial limit is onl...
-
On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains
A method is given to determine eigenvalue decay rates of NTK and related kernels on general domains, leading to minimax optimality results for wide neural networks under smoothness assumptions on the target function.
-
On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations
The paper introduces EF and ΔEF as spectral metrics, but ΔEF is derived from EF, making the complexity-faithfulness trade-off partly tautological.
Discussion (0). Sign in to comment.