REVIEW 3 cited by
Value Functions are Control Barrier Functions: Verification of Safe Policies using Control Theory
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Guaranteeing safe behaviour of reinforcement learning (RL) policies poses significant challenges for safety-critical applications, despite RL's generality and scalability. To address this, we propose a new approach to apply verification methods from control theory to learned value functions. By analyzing task structures for safety preservation, we formalize original theorems that establish links between value functions and control barrier functions. Further, we propose novel metrics for verifying value functions in safe control tasks and practical implementation details to improve learning. Our work presents a novel method for certificate learning, which unlocks a diversity of verification techniques from control theory for RL policies, and marks a significant step towards a formal framework for the general, scalable, and verifiable design of RL-based control systems. Code and videos are available at this https url: https://rl-cbf.github.io/
Forward citations
Cited by 3 Pith papers
-
CBF-RL: Safety Filtering Reinforcement Learning in Training with Control Barrier Functions
Training RL policies with a closed-form CBF safety filter plus CBF reward lets a Unitree G1 humanoid avoid obstacles and climb stairs without a runtime safety filter.
-
Towards Safe Robot Foundation Models Using Inductive Biases
A modular ATACOM safety layer is added to robot foundation models pi0 and OCTO, yielding provably safe actions with minimal performance loss.
-
Learning Ensembles of Vision-based Safety Control Filters
Ensembles of vision-based safety filters with diverse backbones and aggregation methods improve safe/unsafe classification accuracy over individual models on the DeepAccident dataset.
Discussion (0). Continue with ORCID to comment.