Pith. sign in

REVIEW 2 cited by

Descending through a Crowded Valley - Benchmarking Deep Learning Optimizers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.01547 v6 pith:DFLTS3PJ submitted 2020-07-03 cs.LG stat.ML

classification cs.LGstat.ML
keywords optimizersdeeplearningmethodsoptimizationoptimizeracrossanecdotes
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Choosing the optimizer is considered to be among the most crucial design decisions in deep learning, and it is not an easy one. The growing literature now lists hundreds of optimization methods. In the absence of clear theoretical guidance and conclusive empirical evidence, the decision is often made based on anecdotes. In this work, we aim to replace these anecdotes, if not with a conclusive ranking, then at least with evidence-backed heuristics. To do so, we perform an extensive, standardized benchmark of fifteen particularly popular deep learning optimizers while giving a concise overview of the wide range of possible choices. Analyzing more than $50,000$ individual runs, we contribute the following three points: (i) Optimizer performance varies greatly across tasks. (ii) We observe that evaluating multiple optimizers with default parameters works approximately as well as tuning the hyperparameters of a single, fixed optimizer. (iii) While we cannot discern an optimization method clearly dominating across all tested tasks, we identify a significantly reduced subset of specific optimizers and parameter choices that generally lead to competitive results in our experiments: Adam remains a strong contender, with newer methods failing to significantly and consistently outperform it. Our open-sourced results are available as challenging and well-tuned baselines for more meaningful evaluations of novel optimization methods without requiring any further computational efforts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 34 citations worldwide. Full citation record

  1. Jess+: designing embodied AI for interactive music-making

    cs.HC 2024-12 conditional novelty 6.0 of 10

    A proof-of-concept embodied AI system using a trained neural-network factory and a robotic arm that draws or gestures to co-improvise with a mixed ensemble, with qualitative reports of transformed and inclusive music-making.

  2. Consistency of Feature Attribution in Deep Learning Architectures for Multi-Omics

    stat.ML 2025-07 conditional novelty 5.0 of 10

    SHAP feature rankings in multi-view deep learning models for multi-omics data are sensitive to architecture and random weight initialization, so important-biomolecule lists should be inspected across repeated runs.

Pith tools