Pith. sign in

REVIEW 1 cited by

Robust Reinforcement Learning for Continuous Control with Model Misspecification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.07516 v2 pith:KPBIPVNP submitted 2019-06-18 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords continuouscontrollearningrobustrobustnessadditionalgorithmbellman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We provide a framework for incorporating robustness -- to perturbations in the transition dynamics which we refer to as model misspecification -- into continuous control Reinforcement Learning (RL) algorithms. We specifically focus on incorporating robustness into a state-of-the-art continuous control RL algorithm called Maximum a-posteriori Policy Optimization (MPO). We achieve this by learning a policy that optimizes for a worst case expected return objective and derive a corresponding robust entropy-regularized Bellman contraction operator. In addition, we introduce a less conservative, soft-robust, entropy-regularized objective with a corresponding Bellman operator. We show that both, robust and soft-robust policies, outperform their non-robust counterparts in nine Mujoco domains with environment perturbations. In addition, we show improved robust performance on a high-dimensional, simulated, dexterous robotic hand. Finally, we present multiple investigative experiments that provide a deeper insight into the robustness framework. This includes an adaptation to another continuous control RL algorithm as well as learning the uncertainty set from offline data. Performance videos can be found online at https://sites.google.com/view/robust-rl.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

    cs.CV 2025-09 conditional novelty 6.0 of 10

    RobustVLA benchmarks VLA robot policies under 17 multi-modal perturbations and uses adversarial flow-matching plus UCB-based noise selection to raise robustness by up to 12.6 absolute points on LIBERO.

Pith tools