Pith. sign in

REVIEW 3 cited by

Representing smooth functions as compositions of near-identity functions with implications for deep network optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.05012 v2 pith:SP7AUMLS submitted 2018-04-13 cs.LG cs.AIcs.NEmath.STstat.MLstat.TH

classification cs.LGcs.AIcs.NEmath.STstat.MLstat.TH
keywords functionsnear-identitycriticalresiduallipschitznetworknonlinearpoints
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We show that any smooth bi-Lipschitz $h$ can be represented exactly as a composition $h_m \circ ... \circ h_1$ of functions $h_1,...,h_m$ that are close to the identity in the sense that each $\left(h_i-\mathrm{Id}\right)$ is Lipschitz, and the Lipschitz constant decreases inversely with the number $m$ of functions composed. This implies that $h$ can be represented to any accuracy by a deep residual network whose nonlinear layers compute functions with a small Lipschitz constant. Next, we consider nonlinear regression with a composition of near-identity nonlinear maps. We show that, regarding Fr\'echet derivatives with respect to the $h_1,...,h_m$, any critical point of a quadratic criterion in this near-identity region must be a global minimizer. In contrast, if we consider derivatives with respect to parameters of a fixed-size residual network with sigmoid activation functions, we show that there are near-identity critical points that are suboptimal, even in the realizable case. Informally, this means that functional gradient methods for residual networks cannot get stuck at suboptimal critical points corresponding to near-identity layers, whereas parametric gradient methods for sigmoidal residual networks suffer from suboptimal critical points in the near-identity region.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Chaining Meets Chain Rule: Multilevel Entropic Regularization and Training of Neural Nets

    cs.LG 2019-06 unverdicted novelty 6.0 of 10

    Derives algorithm-dependent generalization bounds for neural nets using multilevel entropic regularization and proposes a Metropolis-simulated multi-scale Gibbs training procedure tested on a two-layer net for MNIST.

  2. Smooth Model Compression without Fine-Tuning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Training a ResNet-18 with smoothness penalties on weights, then compressing via truncated SVD, keeps 91% CIFAR-10 accuracy at 70% sparsity with no post-compression fine-tuning.

  3. Prior-aware and Context-guided Group Sampling for Active Probabilistic Subsampling

    eess.SY 2026-07 conditional novelty 4.0 of 10

    Combining dataset-prior sampling with top-k group active sampling yields more stable optimization and better downstream task performance than prior active subsampling methods across classification, reconstruction, and...

Pith tools