Pith. sign in

REVIEW 1 cited by

Safe Exploration Using Bayesian World Models and Log-Barrier Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.05890 v1 pith:ZUNI4JCF submitted 2024-05-09 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningcerlsafebayesianduringexplorationmethodmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A major challenge in deploying reinforcement learning in online tasks is ensuring that safety is maintained throughout the learning process. In this work, we propose CERL, a new method for solving constrained Markov decision processes while keeping the policy safe during learning. Our method leverages Bayesian world models and suggests policies that are pessimistic w.r.t. the model's epistemic uncertainty. This makes CERL robust towards model inaccuracies and leads to safe exploration during learning. In our experiments, we demonstrate that CERL outperforms the current state-of-the-art in terms of safety and optimality in solving CMDPs from image observations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Koopman-Equivariant Gaussian Processes

    cs.LG 2025-02 reject novelty 6.0 of 10

    Koopman-equivariant Gaussian processes give a new kernel family for forecasting nonlinear dynamics with closed-form multi-step uncertainty and a claimed sample-complexity reduction.

Pith tools