Pith. sign in

REVIEW 2 cited by

COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RL

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07220 v2 pith:MEUYH63V submitted 2023-10-11 cs.LG

classification cs.LG
keywords modelcoplannerenvironmentmodel-basedtextttexplorationlearningrollouts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Dyna-style model-based reinforcement learning contains two phases: model rollouts to generate sample for policy learning and real environment exploration using current policy for dynamics model learning. However, due to the complex real-world environment, it is inevitable to learn an imperfect dynamics model with model prediction error, which can further mislead policy learning and result in sub-optimal solutions. In this paper, we propose $\texttt{COPlanner}$, a planning-driven framework for model-based methods to address the inaccurately learned dynamics model problem with conservative model rollouts and optimistic environment exploration. $\texttt{COPlanner}$ leverages an uncertainty-aware policy-guided model predictive control (UP-MPC) component to plan for multi-step uncertainty estimation. This estimated uncertainty then serves as a penalty during model rollouts and as a bonus during real environment exploration respectively, to choose actions. Consequently, $\texttt{COPlanner}$ can avoid model uncertain regions through conservative model rollouts, thereby alleviating the influence of model error. Simultaneously, it explores high-reward model uncertain regions to reduce model error actively through optimistic real environment exploration. $\texttt{COPlanner}$ is a plug-and-play framework that can be applied to any dyna-style model-based methods. Experimental results on a series of proprioceptive and visual continuous control tasks demonstrate that both sample efficiency and asymptotic performance of strong model-based methods are significantly improved combined with $\texttt{COPlanner}$.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control

    cs.LG 2026-08 conditional novelty 6.0 of 10

    V-Simba, a visual RL architecture combining layer normalization, weight decay, and a distributional critic, matches or outperforms complex baselines on 29 continuous control tasks while using less compute.

  2. Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A TD-trained vision value model guides sentence-level inference-time search in VLMs, cutting hallucination and improving caption quality, with self-training gains on nine benchmarks.

Pith tools