Pith. sign in

REVIEW 2 cited by

Constrained Policy Optimization via Bayesian World Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.09802 v4 pith:PSDKYXBU submitted 2022-01-24 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords safetyworldapproachbayesianboundsconstrainedlambdamodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Improving sample-efficiency and safety are crucial challenges when deploying reinforcement learning in high-stakes real world applications. We propose LAMBDA, a novel model-based approach for policy optimization in safety critical tasks modeled via constrained Markov decision processes. Our approach utilizes Bayesian world models, and harnesses the resulting uncertainty to maximize optimistic upper bounds on the task objective, as well as pessimistic upper bounds on the safety constraints. We demonstrate LAMBDA's state of the art performance on the Safety-Gym benchmark suite in terms of sample efficiency and constraint violation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    DROPJ trains a world-model-based MPC agent from one-shot human preferences plus safety justifications, cutting training cost and improving deployment safety in car-racing simulations.

  2. Safe Planning and Policy Optimization via World Model Learning

    cs.AI 2025-06 conditional novelty 6.0 of 10

    SPOWL is a model-based safe RL method that uses a value-equivalent world model, a Lagrangian-trained safe policy, and adaptive planning thresholds to achieve low-cost, high-reward control on SafetyGymnasium tasks.

Pith tools