REVIEW 3 cited by
A Universal Law of Robustness via Isoperimetry
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Classically, data interpolation with a parametrized model class is possible as long as the number of parameters is larger than the number of equations to be satisfied. A puzzling phenomenon in deep learning is that models are trained with many more parameters than what this classical theory would suggest. We propose a partial theoretical explanation for this phenomenon. We prove that for a broad class of data distributions and model classes, overparametrization is necessary if one wants to interpolate the data smoothly. Namely we show that smooth interpolation requires $d$ times more parameters than mere interpolation, where $d$ is the ambient data dimension. We prove this universal law of robustness for any smoothly parametrized function class with polynomial size weights, and any covariate distribution verifying isoperimetry. In the case of two-layers neural networks and Gaussian covariates, this law was conjectured in prior work by Bubeck, Li and Nagaraj. We also give an interpretation of our result as an improved generalization bound for model classes consisting of smooth functions.
Forward citations
Cited by 3 Pith papers
-
A law of robustness for two-layer neural networks with arbitrary weights
Any width-m two-layer piecewise-linear network with arbitrary weights that fits n noisy labels below the noise floor has Lip ≳ ε sqrt(n/(m log(m n d/ε))) with high probability on the sphere or Gaussian.
-
Does Order Matter : Connecting The Law of Robustness to Robust Generalization
The paper proves R(ℓρ∘B_L∘S) ≤ 8R(B_L∘S) but does not derive the advertised Ω(n^{1/d}) recovery or the missing local-scale result.
-
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
Wavelet-based positional encodings are claimed to improve how transformers extrapolate to longer sequences, with a toy experiment supporting the claim but with weak theory.
Discussion (0). Continue with ORCID to comment.