Pith. sign in

REVIEW 2 cited by

Hyperparameter Tuning for Deep Reinforcement Learning Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.11182 v1 pith:IWNJHZ4S submitted 2022-01-26 cs.LG cs.NE

classification cs.LGcs.NE
keywords applicationsdeephyperparameterlearningreinforcementalgorithmapproacharchitectures
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Reinforcement learning (RL) applications, where an agent can simply learn optimal behaviors by interacting with the environment, are quickly gaining tremendous success in a wide variety of applications from controlling simple pendulums to complex data centers. However, setting the right hyperparameters can have a huge impact on the deployed solution performance and reliability in the inference models, produced via RL, used for decision-making. Hyperparameter search itself is a laborious process that requires many iterations and computationally expensive to find the best settings that produce the best neural network architectures. In comparison to other neural network architectures, deep RL has not witnessed much hyperparameter tuning, due to its algorithm complexity and simulation platforms needed. In this paper, we propose a distributed variable-length genetic algorithm framework to systematically tune hyperparameters for various RL applications, improving training time and robustness of the architecture, via evolution. We demonstrate the scalability of our approach on many RL problems (from simple gyms to complex applications) and compared with Bayesian approach. Our results show that with more generations, optimal solutions that require fewer training episodes and are computationally cheap while being more robust for deployment. Our results are imperative to advance deep reinforcement learning controllers for real-world problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

    cs.AI 2024-12 conditional novelty 6.0 of 10

    B-STaR dynamically tunes sampling temperature and reward thresholds during iterative self-training, improving Pass@1 on math, code, and commonsense reasoning benchmarks versus static self-improvement baselines.

  2. Parameter Estimation using Reinforcement Learning Causal Curiosity: Limits and Challenges

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Systematic analysis of Causal Curiosity in a simulated robotic manipulator shows high accuracy in single-factor and high-granularity settings, but frequent failures when multiple causal factors vary simultaneously.

Pith tools