Learning to Optimize for Reinforcement Learning

A. Rupam Mahmood; Qingfeng Lan; Shuicheng Yan; Zhongwen Xu

arxiv: 2302.01470 · v3 · pith:VIHQXEGBnew · submitted 2023-02-03 · 💻 cs.LG · cs.AI

Learning to Optimize for Reinforcement Learning

Qingfeng Lan , A. Rupam Mahmood , Shuicheng Yan , Zhongwen Xu This is my paper

classification 💻 cs.LG cs.AI

keywords learningoptimizertaskslearnedoptimizersreinforcementbiasissues

0 comments

read the original abstract

In recent years, by leveraging more data, computation, and diverse tasks, learned optimizers have achieved remarkable success in supervised learning, outperforming classical hand-designed optimizers. Reinforcement learning (RL) is essentially different from supervised learning, and in practice, these learned optimizers do not work well even in simple RL tasks. We investigate this phenomenon and identify two issues. First, the agent-gradient distribution is non-independent and identically distributed, leading to inefficient meta-training. Moreover, due to highly stochastic agent-environment interactions, the agent-gradients have high bias and variance, which increases the difficulty of learning an optimizer for RL. We propose pipeline training and a novel optimizer structure with a good inductive bias to address these issues, making it possible to learn an optimizer for reinforcement learning from scratch. We show that, although only trained in toy tasks, our learned optimizer can generalize to unseen complex tasks in Brax.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

AI Training Manager: Bounded Closed-Loop Control of Adaptive Training Recipes
cs.AI 2026-06 unverdicted novelty 4.0

An LLM-based bounded controller adapts ML training parameters from structured telemetry to correct overfitting and exploration issues, shown on TinyStories and robotic RL tasks.