Pith. sign in

REVIEW 4 cited by

Deep Network Approximation Characterized by Number of Neurons

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.05497 v5 pith:OALZLWN3 submitted 2019-06-13 math.NA cs.LGcs.NA

classification math.NAcs.LGcs.NA
keywords deltamathcalsqrtapproximationomegatfracvarepsilonalpha
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

This paper quantitatively characterizes the approximation power of deep feed-forward neural networks (FNNs) in terms of the number of neurons. It is shown by construction that ReLU FNNs with width $\mathcal{O}\big(\max\{d\lfloor N^{1/d}\rfloor,\, N+1\}\big)$ and depth $\mathcal{O}(L)$ can approximate an arbitrary H\"older continuous function of order $\alpha\in (0,1]$ on $[0,1]^d$ with a nearly tight approximation rate $\mathcal{O}\big(\sqrt{d} N^{-2\alpha/d}L^{-2\alpha/d}\big)$ measured in $L^p$-norm for any $N,L\in \mathbb{N}^+$ and $p\in[1,\infty]$. More generally for an arbitrary continuous function $f$ on $[0,1]^d$ with a modulus of continuity $\omega_f(\cdot)$, the constructive approximation rate is $\mathcal{O}\big(\sqrt{d}\,\omega_f( N^{-2/d}L^{-2/d})\big)$. We also extend our analysis to $f$ on irregular domains or those localized in an $\varepsilon$-neighborhood of a $d_{\mathcal{M}}$-dimensional smooth manifold $\mathcal{M}\subseteq [0,1]^d$ with $d_{\mathcal{M}}\ll d$. Especially, in the case of an essentially low-dimensional domain, we show an approximation rate $\mathcal{O}\big(\omega_f(\tfrac{\varepsilon}{1-\delta}\sqrt{\tfrac{d}{d_\delta}}+\varepsilon)+\sqrt{d}\,\omega_f(\tfrac{\sqrt{d}}{(1-\delta)\sqrt{d_\delta}}N^{-2/d_\delta}L^{-2/d_\delta})\big)$ for ReLU FNNs to approximate $f$ in the $\varepsilon$-neighborhood, where $d_\delta=\mathcal{O}\big(d_{\mathcal{M}}\tfrac{\ln (d/\delta)}{\delta^2}\big)$ for any $\delta\in(0,1)$ as a relative error for a projection to approximate an isometry when projecting $\mathcal{M}$ to a $d_{\delta}$-dimensional domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Boosting Statistic Learning with Synthetic Data from Pretrained Large Models

    stat.ML 2025-05 reject novelty 6.0 of 10

    The paper claims synthetic tabular data generated by pass-through Stable Diffusion, filtered by Wasserstein distance or hypothesis tests, improves predictive accuracy, but the evidence is weakened by missing baselines...

  2. Do Neural Networks Really Beat the Curse of Dimensionality? A Bit-Complexity View

    cs.LG 2026-08 conditional novelty 5.0 of 10

    When approximation quality is measured per bit instead of per parameter, neural networks do not fundamentally beat classical methods; the real limit is the metric entropy of the target function class.

  3. Calibration Prediction Interval for Non-parametric Regression and Neural Networks

    stat.ME 2025-09 conditional novelty 5.0 of 10

    The authors construct prediction intervals by calibrating estimated conditional CDF values on a grid, and show asymptotic validity for DNN and kernel estimators, with a finite-sample coverage guarantee only under an o...

  4. Nonparametric Regression on Low-Dimensional Manifolds using Deep ReLU Networks : Function Approximation and Statistical Recovery

    cs.LG 2019-08 conditional novelty 5.0 of 10

    When the regression function is (s+α)-Hölder on a d-dimensional manifold, the empirical risk minimizer over deep ReLU networks achieves mean squared error n^{-2(s+α)/(2(s+α)+d)} log^3 n.

Pith tools