REVIEW 4 cited by
Maximum Entropy RL (Provably) Solves Some Robust RL Problems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Many potential applications of reinforcement learning (RL) require guarantees that the agent will perform well in the face of disturbances to the dynamics or reward function. In this paper, we prove theoretically that maximum entropy (MaxEnt) RL maximizes a lower bound on a robust RL objective, and thus can be used to learn policies that are robust to some disturbances in the dynamics and the reward function. While this capability of MaxEnt RL has been observed empirically in prior work, to the best of our knowledge our work provides the first rigorous proof and theoretical characterization of the MaxEnt RL robust set. While a number of prior robust RL algorithms have been designed to handle similar disturbances to the reward function or dynamics, these methods typically require additional moving parts and hyperparameters on top of a base RL algorithm. In contrast, our results suggest that MaxEnt RL by itself is robust to certain disturbances, without requiring any additional modifications. While this does not imply that MaxEnt RL is the best available robust RL method, MaxEnt RL is a simple robust RL method with appealing formal guarantees.
Forward citations
Cited by 4 Pith papers
-
RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
RobustVLA benchmarks VLA robot policies under 17 multi-modal perturbations and uses adversarial flow-matching plus UCB-based noise selection to raise robustness by up to 12.6 absolute points on LIBERO.
-
When Maximum Entropy Misleads Policy Optimization
Maximum entropy RL can be formally steered into arbitrary suboptimal policies at convergence by adding entropy trap states, while standard RL is unaffected.
-
Reinforcement Learning on Reconfigurable Hardware: Overcoming Material Variability in Laser Material Processing
A real-time FPGA-based reinforcement learning controller, trained on the fly with the optical reflection signal as reward, adapts laser power to surface roughness and reports reward gains of up to 23% over constant-po...
-
Distributionally Robust Deep Q-Learning
Sinkhorn-ball robust Bellman equation for continuous-state MDPs is implemented as a Robust DQN that learns policies robust to transition-model misspecification.
Discussion (0). Continue with ORCID to comment.