Pith. sign in

REVIEW 7 cited by

The Intrinsic Dimension of Images and Its Impact on Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08894 v1 pith:2I2UMCCD submitted 2021-04-18 cs.CV cs.LGstat.ML

classification cs.CVcs.LGstat.ML
keywords dimensiondatadatasetsimageintrinsiclearningcommondeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

It is widely believed that natural image data exhibits low-dimensional structure despite the high dimensionality of conventional pixel representations. This idea underlies a common intuition for the remarkable success of deep learning in computer vision. In this work, we apply dimension estimation tools to popular datasets and investigate the role of low-dimensional structure in deep learning. We find that common natural image datasets indeed have very low intrinsic dimension relative to the high number of pixels in the images. Additionally, we find that low dimensional datasets are easier for neural networks to learn, and models solving these tasks generalize better from training to test data. Along the way, we develop a technique for validating our dimension estimation tools on synthetic data generated by GANs allowing us to actively manipulate the intrinsic dimension by controlling the image generation process. Code for our experiments may be found here https://github.com/ppope/dimensions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CASC: Causal Adversarial Subspace Clustering for Multivariate Spatiotemporal Data

    cs.LG 2026-07 reject novelty 6.0 of 10

    CASC uses a U-Net-style adversarial autoencoder with attention and causal-regularized self-expression to cluster multivariate spatiotemporal series into evolving regimes, validated only on internal cluster metrics.

  2. Sparse Autoencoders, Again?

    cs.LG 2025-06 conditional novelty 6.0 of 10

    VAEase gates the VAE decoder input by the encoder's variance, combining sparse-autoencoder adaptive sparsity with a hyperparameter-free loss; a global-minimizer theorem says active latent dimensions recover per-manifo...

  3. Diffusion Sampling Path Tells More: An Efficient Plug-and-Play Strategy for Sample Filtering

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CFG-Rejection filters low-quality diffusion samples early using the accumulated norm of the classifier-free guidance vector, improving quality scores without external reward models.

  4. Computing Optimal Transport Maps and Wasserstein Barycenters Using Conditional Normalizing Flows

    stat.ML 2025-05 conditional novelty 6.0 of 10

    A conditional normalizing flow method that solves the primal optimal transport problem and computes Wasserstein-2 barycenters as weighted averages of maps from a shared latent distribution.

  5. A Multiclass Quantum Aligned Centroid Kernel

    quant-ph 2026-07 conditional novelty 5.0 of 10

    A sample-to-centroid fidelity kernel enables linear-scaling multiclass quantum classification; in simulation it beats pure quantum baselines, and untrained 124-qubit hardware results match an RBF kernel.

  6. Provable diffusion-based posterior sampling for linear inverse problems via DDIM

    cs.LG 2026-07 reject novelty 5.0 of 10

    A SVD-based, coordinate-wise DDIM sampler is claimed to asymptotically sample from the posterior for noisy linear inverse problems, but the proof's posterior identification step does not follow from the stated updates.

  7. Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Training flow matching along sphere geodesics with a curvature-aware loss weight lets standard DiT-B converge on DINOv2 features (FID 3.37 with guidance), contradicting the need for width scaling.

Pith tools