REVIEW 7 cited by
The Intrinsic Dimension of Images and Its Impact on Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
It is widely believed that natural image data exhibits low-dimensional structure despite the high dimensionality of conventional pixel representations. This idea underlies a common intuition for the remarkable success of deep learning in computer vision. In this work, we apply dimension estimation tools to popular datasets and investigate the role of low-dimensional structure in deep learning. We find that common natural image datasets indeed have very low intrinsic dimension relative to the high number of pixels in the images. Additionally, we find that low dimensional datasets are easier for neural networks to learn, and models solving these tasks generalize better from training to test data. Along the way, we develop a technique for validating our dimension estimation tools on synthetic data generated by GANs allowing us to actively manipulate the intrinsic dimension by controlling the image generation process. Code for our experiments may be found here https://github.com/ppope/dimensions.
Forward citations
Cited by 7 Pith papers
-
CASC: Causal Adversarial Subspace Clustering for Multivariate Spatiotemporal Data
CASC uses a U-Net-style adversarial autoencoder with attention and causal-regularized self-expression to cluster multivariate spatiotemporal series into evolving regimes, validated only on internal cluster metrics.
-
Sparse Autoencoders, Again?
VAEase gates the VAE decoder input by the encoder's variance, combining sparse-autoencoder adaptive sparsity with a hyperparameter-free loss; a global-minimizer theorem says active latent dimensions recover per-manifo...
-
Diffusion Sampling Path Tells More: An Efficient Plug-and-Play Strategy for Sample Filtering
CFG-Rejection filters low-quality diffusion samples early using the accumulated norm of the classifier-free guidance vector, improving quality scores without external reward models.
-
Computing Optimal Transport Maps and Wasserstein Barycenters Using Conditional Normalizing Flows
A conditional normalizing flow method that solves the primal optimal transport problem and computes Wasserstein-2 barycenters as weighted averages of maps from a shared latent distribution.
-
A Multiclass Quantum Aligned Centroid Kernel
A sample-to-centroid fidelity kernel enables linear-scaling multiclass quantum classification; in simulation it beats pure quantum baselines, and untrained 124-qubit hardware results match an RBF kernel.
-
Provable diffusion-based posterior sampling for linear inverse problems via DDIM
A SVD-based, coordinate-wise DDIM sampler is claimed to asymptotically sample from the posterior for noisy linear inverse problems, but the proof's posterior identification step does not follow from the stated updates.
-
Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders
Training flow matching along sphere geodesics with a curvature-aware loss weight lets standard DiT-B converge on DINOv2 features (FID 3.37 with guidance), contradicting the need for width scaling.
Discussion (0). Continue with ORCID to comment.