Pith. sign in

REVIEW 3 cited by

Interpreting the Curse of Dimensionality from Distance Concentration and Manifold Effect

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.00422 v3 pith:AIPTODFY submitted 2023-12-31 cs.LG cs.DS

classification cs.LGcs.DS
keywords dimensionalitydistancecursecausesdataclassificationclusteringconcentration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The characteristics of data like distribution and heterogeneity, become more complex and counterintuitive as dimensionality increases. This phenomenon is known as curse of dimensionality, where common patterns and relationships (e.g., internal pattern and boundary pattern) that hold in low-dimensional space may be invalid in higher-dimensional space. It leads to a decreasing performance for the regression, classification, or clustering models or algorithms. Curse of dimensionality can be attributed to many causes. In this paper, we first summarize the potential challenges associated with manipulating high-dimensional data, and explains the possible causes for the failure of regression, classification, or clustering tasks. Subsequently, we delve into two major causes of the curse of dimensionality, distance concentration, and manifold effect, by performing theoretical and empirical analyses. The results demonstrate that, as the dimensionality increases, nearest neighbor search (NNS) using three classical distance measurements, Minkowski distance, Chebyshev distance, and cosine distance, becomes meaningless. Meanwhile, the data incorporates more redundant features, and the variance contribution of principal component analysis (PCA) is skewed towards a few dimensions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 5 citations worldwide. Full citation record

  1. Measuring Distortion in the Empty Regions of Dimensionality Reduction Scatterplots with the Gap Index

    cs.LG 2026-07 conditional novelty 5.0 of 10

    The Gap Index quantifies visual distortion in empty regions of DR scatterplots via Delaunay triangle area deformation and is more sensitive to salient gap artifacts than stress or trustworthiness.

  2. Hyperellipsoid Density Sampling: Exploitative Sequences to Accelerate High-Dimensional Numerical Optimization

    math.NA 2025-11 conditional novelty 4.0 of 10

    A new sampling scheme that clusters an initial uniform sequence into hyperellipsoids is claimed to improve differential evolution results by 3–37% on CEC2017 benchmarks, but the evidence is internally inconsistent.

  3. Towards High Supervised Learning Utility Training Data Generation: Data Pruning and Column Reordering

    cs.LG 2025-07 reject novelty 4.0 of 10

    PRRO combines signal-based data pruning and column reordering to improve the supervised learning utility of synthetic tabular data, but its evaluation is undermined by data manipulation and an ill-defined correlation measure.

Pith tools