REVIEW 2 cited by
Your diffusion model secretly knows the dimension of the data manifold
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this work, we propose a novel framework for estimating the dimension of the data manifold using a trained diffusion model. A diffusion model approximates the score function i.e. the gradient of the log density of a noise-corrupted version of the target distribution for varying levels of corruption. We prove that, if the data concentrates around a manifold embedded in the high-dimensional ambient space, then as the level of corruption decreases, the score function points towards the manifold, as this direction becomes the direction of maximal likelihood increase. Therefore, for small levels of corruption, the diffusion model provides us with access to an approximation of the normal bundle of the data manifold. This allows us to estimate the dimension of the tangent space, thus, the intrinsic dimension of the data manifold. To the best of our knowledge, our method is the first estimator of the data manifold dimension based on diffusion models and it outperforms well established statistical estimators in controlled experiments on both Euclidean and image data.
Forward citations
Cited by 2 Pith papers
-
Diffusion models recover accurate mixture weights despite score function insensitivity
Mixture-weight recovery errors in diffusion models are controlled by the curvature of the diffusion score-matching loss (the DSSI), not by the target score's sensitivity.
-
On the Local Complexity of Linear Regions in Deep ReLU Networks
The paper introduces a noise-regularized measure of linear-region density and derives inequalities relating it to representation rank, total variation, and representation cost, but key inequalities contain a false nor...
Discussion (0). Continue with ORCID to comment.