REVIEW 3 cited by
Projected support points: a new method for high-dimensional data reduction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In an era where big and high-dimensional data is readily available, data scientists are inevitably faced with the challenge of reducing this data for expensive downstream computation or analysis. To this end, we present here a new method for reducing high-dimensional big data into a representative point set, called projected support points (PSPs). A key ingredient in our method is the so-called sparsity-inducing (SpIn) kernel, which encourages the preservation of low-dimensional features when reducing high-dimensional data. We begin by introducing a unifying theoretical framework for data reduction, connecting PSPs with fundamental sampling principles from experimental design and Quasi-Monte Carlo. Through this framework, we then derive sparsity conditions under which the curse-of-dimensionality in data reduction can be lifted for our method. Next, we propose two algorithms for one-shot and sequential reduction via PSPs, both of which exploit big data subsampling and majorization-minimization for efficient optimization. Finally, we demonstrate the practical usefulness of PSPs in two real-world applications, the first for data reduction in kernel learning, and the second for reducing Markov Chain Monte Carlo (MCMC) chains.
Forward citations
Cited by 3 Pith papers
-
Stein Kernelized Molecular Dynamics for Active Learning of Interatomic Potentials
SKMD adapts Stein variational gradient descent into molecular dynamics with asynchronous updates and global atomic descriptor kernels to acquire non-redundant training configurations while preserving the Boltzmann dis...
-
Weighted Support Points from Random Measures: An Interpretable Alternative for Generative Modeling
Randomly reweighting a dataset and then optimizing a set of support points to match the weighted data produces diverse, interpretable sample sets at low cost, according to visual results on MNIST and CelebA.
-
Robust designs for Gaussian process emulation of computer experiments
Energy-distance-minimizing support points and projected support points are shown to be robust Gaussian process emulation designs, with a theory linking them to maximum-entropy, minimax, and maximin designs.
Discussion (0). Continue with ORCID to comment.