Pith. sign in

REVIEW 4 cited by

On approximating nabla f with neural networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.12744 v2 pith:2JZRKJAI submitted 2019-10-28 stat.ML cs.LG

On approximating nabla f with neural networks

classification stat.ML cs.LG
keywords nablamathbbapproxhiddenlayernetworkneuralrightarrow
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Consider a feedforward neural network $\psi: \mathbb{R}^d\rightarrow \mathbb{R}^d$ such that $\psi\approx \nabla f$, where $f:\mathbb{R}^d \rightarrow \mathbb{R}$ is a smooth function, therefore $\psi$ must satisfy $\partial_j \psi_i = \partial_i \psi_j$ pointwise. We prove a theorem that a $\psi$ network with more than one hidden layer can only represent one feature in its first hidden layer; this is a dramatic departure from the well-known results for one hidden layer. The proof of the theorem is straightforward, where two backward paths and a weight-tying matrix play the key roles. We then present the alternative, the implicit parametrization, where the neural network is $\phi: \mathbb{R}^d \rightarrow \mathbb{R}$ and $\nabla \phi \approx \nabla f$; in addition, a "soft analysis" of $\nabla \phi$ gives a dual perspective on the theorem. Throughout, we come back to recent probabilistic models that are formulated as $\nabla \phi \approx \nabla f$, and conclude with a critique of denoising autoencoders.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Diffusion Models Observe Only Gradients: A Geometric Perspective on Score Matching Errors

    stat.ML 2026-06 unverdicted novelty 7.0

    Helmholtz-Hodge decomposition of score errors shows only the gradient component affects marginal Fokker-Planck dynamics in diffusion models, yielding an impossibility result for L2 error bounding divergences and a tra...

  2. Diffusion Models Observe Only Gradients: A Geometric Perspective on Score Matching Errors

    stat.ML 2026-06 unverdicted novelty 7.0

    Only the gradient component of score errors affects marginal distributions in diffusion models, so L2 error can be arbitrarily large with perfect match; this yields an impossibility result, a gradient-only KL bound, a...

  3. Tessellations of Semi-Discrete Flow Matching

    cs.LG 2026-05 unverdicted novelty 7.0

    Semi-discrete Flow Matching produces terminal assignment regions that are topologically simple (open, simply connected, homeomorphic to the ball under assumption) yet geometrically distinct from optimal transport Lagu...

  4. Universal Representation of Generalized Convex Functions and their Gradients

    math.OC 2025-08 unverdicted novelty 6.0

    A new differentiable layer with convex parameter space universally approximates generalized convex functions and their gradients, enabling single-level reformulations of bilevel problems in optimal transport and multi...