Pith. sign in

REVIEW 1 cited by

Hyperparameter Optimization Is Deceiving Us, and How to Stop It

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.03034 v4 pith:JAHTEC2M submitted 2021-02-05 cs.LG cs.LO

classification cs.LGcs.LO
keywords frameworkhyperparameteroptimizationconclusionsdeceptiondefendedehpoinconsistent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent empirical work shows that inconsistent results based on choice of hyperparameter optimization (HPO) configuration are a widespread problem in ML research. When comparing two algorithms J and K searching one subspace can yield the conclusion that J outperforms K, whereas searching another can entail the opposite. In short, the way we choose hyperparameters can deceive us. We provide a theoretical complement to this prior work, arguing that, to avoid such deception, the process of drawing conclusions from HPO should be made more rigorous. We call this process epistemic hyperparameter optimization (EHPO), and put forth a logical framework to capture its semantics and how it can lead to inconsistent conclusions about performance. Our framework enables us to prove EHPO methods that are guaranteed to be defended against deception, given bounded compute time budget t. We demonstrate our framework's utility by proving and empirically validating a defended variant of random search.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What Do Machine Learning Researchers Mean by "Reproducible"?

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A survey-based taxonomy that splits reproducibility in AI/ML into eight rigor aspects (repeatability, reproducibility, replicability, adaptability, model selection, label/data quality, meta/incentive, maintainability)...

Pith tools