REVIEW 4 cited by
Hyperparameter Optimization: A Spectral Approach
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We give a simple, fast algorithm for hyperparameter optimization inspired by techniques from the analysis of Boolean functions. We focus on the high-dimensional regime where the canonical example is training a neural network with a large number of hyperparameters. The algorithm --- an iterative application of compressed sensing techniques for orthogonal polynomials --- requires only uniform sampling of the hyperparameters and is thus easily parallelizable. Experiments for training deep neural networks on Cifar-10 show that compared to state-of-the-art tools (e.g., Hyperband and Spearmint), our algorithm finds significantly improved solutions, in some cases better than what is attainable by hand-tuning. In terms of overall running time (i.e., time required to sample various settings of hyperparameters plus additional computation time), we are at least an order of magnitude faster than Hyperband and Bayesian Optimization. We also outperform Random Search 8x. Additionally, our method comes with provable guarantees and yields the first improvements on the sample complexity of learning decision trees in over two decades. In particular, we obtain the first quasi-polynomial time algorithm for learning noisy decision trees with polynomial sample complexity.
Forward citations
Cited by 4 Pith papers
-
ID3 Learns Juntas for Smoothed Product Distributions
ID3 learns log n-juntas in polynomial time under the smoothed analysis model for product distributions.
-
Efficient Automatic Meta Optimization Search for Few-Shot Learning
A NAS controller and Reptile meta-learning are jointly optimized to automatically search few-shot learner architectures, reaching 74.2% on Mini-ImageNet 5-shot 5-way transductive classification in 1 to 2 GPU days.
-
Which Hyperparameters Matter? A Game-Theoretic Framework for Interpretable Hyperparameter Sensitivity Analysis
A framework using Shapley Effects and Pareto fronts ranks hyperparameter influence per objective from a coarse grid-search lookup table, without proposing a new optimizer.
-
Adaptive Parameter Optimization in Gaussian Processes: A Comprehensive Study of Uncertainty Quantification and Dimensional Scaling
An adaptive-kappa, uncertainty-penalized GP-UCB is claimed to outperform fixed-parameter baselines, but the supporting theory is sketched and the empirical evidence is not shipped.
Discussion (0). Continue with ORCID to comment.