Pith. sign in

REVIEW 1 cited by

Online Hyper-parameter Tuning in Off-policy Learning via Evolutionary Strategies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.07554 v1 pith:ANZ7I2DE submitted 2020-06-13 cs.LG cs.NEstat.ML

classification cs.LGcs.NEstat.ML
keywords learningoff-policyhyper-parametersalgorithmsevolutionaryhyper-parametermeta-gradientsonline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Off-policy learning algorithms have been known to be sensitive to the choice of hyper-parameters. However, unlike near on-policy algorithms for which hyper-parameters could be optimized via e.g. meta-gradients, similar techniques could not be straightforwardly applied to off-policy learning. In this work, we propose a framework which entails the application of Evolutionary Strategies to online hyper-parameter tuning in off-policy learning. Our formulation draws close connections to meta-gradients and leverages the strengths of black-box optimization with relatively low-dimensional search spaces. We show that our method outperforms state-of-the-art off-policy learning baselines with static hyper-parameters and recent prior work over a wide range of continuous control benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning

    cs.NE 2025-01 conditional novelty 4.0 of 10

    A JAX-based framework runs evolutionary reinforcement learning end-to-end on GPUs and reports large training speed-ups over CPU-based libraries on Brax locomotion tasks.

Pith tools