Pith. sign in

REVIEW 1 cited by

Tune As You Scale: Hyperparameter Optimization For Compute Efficient Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.08055 v1 pith:ZODWH2LK submitted 2023-06-13 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelstuningsearchcomputehyperparameterscalingbayesiancarbs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hyperparameter tuning of deep learning models can lead to order-of-magnitude performance gains for the same amount of compute. Despite this, systematic tuning is uncommon, particularly for large models, which are expensive to evaluate and tend to have many hyperparameters, necessitating difficult judgment calls about tradeoffs, budgets, and search bounds. To address these issues and propose a practical method for robustly tuning large models, we present Cost-Aware Pareto Region Bayesian Search (CARBS), a Bayesian optimization algorithm that performs local search around the performance-cost Pareto frontier. CARBS does well even in unbounded search spaces with many hyperparameters, learns scaling relationships so that it can tune models even as they are scaled up, and automates much of the "black magic" of tuning. Among our results, we effectively solve the entire ProcGen benchmark just by tuning a simple baseline (PPO, as provided in the original ProcGen paper). We also reproduce the model size vs. training tokens scaling result from the Chinchilla project (Hoffmann et al. 2022), while simultaneously discovering scaling laws for every other hyperparameter, via an easy automated process that uses significantly less compute and is applicable to any deep learning problem (not just language models).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Combining Pre-Trained Models for Enhanced Feature Representation in Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A lightweight attention module that weights embeddings from multiple pre-trained models achieves comparable Atari RL performance to end-to-end training, with improved robustness to visual changes.

Pith tools