Pith. sign in

REVIEW 4 cited by

Maximum Entropy RL (Provably) Solves Some Robust RL Problems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.06257 v2 pith:SMXVIWIH submitted 2021-03-10 cs.LG cs.RO

classification cs.LGcs.RO
keywords robustmaxentdisturbancesdynamicsfunctionrewardwhileadditional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many potential applications of reinforcement learning (RL) require guarantees that the agent will perform well in the face of disturbances to the dynamics or reward function. In this paper, we prove theoretically that maximum entropy (MaxEnt) RL maximizes a lower bound on a robust RL objective, and thus can be used to learn policies that are robust to some disturbances in the dynamics and the reward function. While this capability of MaxEnt RL has been observed empirically in prior work, to the best of our knowledge our work provides the first rigorous proof and theoretical characterization of the MaxEnt RL robust set. While a number of prior robust RL algorithms have been designed to handle similar disturbances to the reward function or dynamics, these methods typically require additional moving parts and hyperparameters on top of a base RL algorithm. In contrast, our results suggest that MaxEnt RL by itself is robust to certain disturbances, without requiring any additional modifications. While this does not imply that MaxEnt RL is the best available robust RL method, MaxEnt RL is a simple robust RL method with appealing formal guarantees.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

    cs.CV 2025-09 conditional novelty 6.0 of 10

    RobustVLA benchmarks VLA robot policies under 17 multi-modal perturbations and uses adversarial flow-matching plus UCB-based noise selection to raise robustness by up to 12.6 absolute points on LIBERO.

  2. When Maximum Entropy Misleads Policy Optimization

    cs.LG 2025-06 reject novelty 6.0 of 10

    Maximum entropy RL can be formally steered into arbitrary suboptimal policies at convergence by adding entropy trap states, while standard RL is unaffected.

  3. Reinforcement Learning on Reconfigurable Hardware: Overcoming Material Variability in Laser Material Processing

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A real-time FPGA-based reinforcement learning controller, trained on the fly with the optical reflection signal as reward, adapts laser power to surface roughness and reports reward gains of up to 23% over constant-po...

  4. Distributionally Robust Deep Q-Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Sinkhorn-ball robust Bellman equation for continuous-state MDPs is implemented as a Robust DQN that learns policies robust to transition-model misspecification.

Pith tools