Pith. sign in

REVIEW 2 cited by

AutoRL Hyperparameter Landscapes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.02396 v4 pith:2DID3WR7 submitted 2023-04-05 cs.LG cs.AIcs.ROcs.SYeess.SY

classification cs.LGcs.AIcs.ROcs.SYeess.SY
keywords hyperparameterautorllandscapestimeapproachesconfigurationsdynamicallyhyperparameters
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although Reinforcement Learning (RL) has shown to be capable of producing impressive results, its use is limited by the impact of its hyperparameters on performance. This often makes it difficult to achieve good results in practice. Automated RL (AutoRL) addresses this difficulty, yet little is known about the dynamics of the hyperparameter landscapes that hyperparameter optimization (HPO) methods traverse in search of optimal configurations. In view of existing AutoRL approaches dynamically adjusting hyperparameter configurations, we propose an approach to build and analyze these hyperparameter landscapes not just for one point in time but at multiple points in time throughout training. Addressing an important open question on the legitimacy of such dynamic AutoRL approaches, we provide thorough empirical evidence that the hyperparameter landscapes strongly vary over time across representative algorithms from RL literature (DQN, PPO, and SAC) in different kinds of environments (Cartpole, Bipedal Walker, and Hopper) This supports the theory that hyperparameters should be dynamically adjusted during training and shows the potential for more insights on AutoRL problems that can be gained through landscape analyses. Our code can be found at https://github.com/automl/AutoRL-Landscape

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

    cs.AI 2024-12 conditional novelty 6.0 of 10

    B-STaR dynamically tunes sampling temperature and reward thresholds during iterative self-training, improving Pass@1 on math, code, and commonsense reasoning benchmarks versus static self-improvement baselines.

  2. Rethinking Performance Analysis for Configurable Software Systems: A Case Study from a Fitness Landscape Perspective

    cs.PF 2024-12 conditional novelty 6.0 of 10

    Modeling configurable software performance as a spatial fitness landscape reveals highly rugged terrain, many scattered local optima, few consistently important options, and prevalent high-order interactions.

Pith tools