Pith. sign in

REVIEW 3 cited by

Reinforcement Learning for Safety-Critical Control under Model Uncertainty, using Control Lyapunov Functions and Control Barrier Functions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.07584 v2 pith:VHBUE2GI submitted 2020-04-16 eess.SY cs.LGcs.ROcs.SY

classification eess.SYcs.LGcs.ROcs.SY
keywords controlmodeluncertaintycbf-clf-qpconstraintsreinforcementbarrierfunction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, the issue of model uncertainty in safety-critical control is addressed with a data-driven approach. For this purpose, we utilize the structure of an input-ouput linearization controller based on a nominal model along with a Control Barrier Function and Control Lyapunov Function based Quadratic Program (CBF-CLF-QP). Specifically, we propose a novel reinforcement learning framework which learns the model uncertainty present in the CBF and CLF constraints, as well as other control-affine dynamic constraints in the quadratic program. The trained policy is combined with the nominal model-based CBF-CLF-QP, resulting in the Reinforcement Learning-based CBF-CLF-QP (RL-CBF-CLF-QP), which addresses the problem of model uncertainty in the safety constraints. The performance of the proposed method is validated by testing it on an underactuated nonlinear bipedal robot walking on randomly spaced stepping stones with one step preview, obtaining stable and safe walking under model uncertainty.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A scalable end-to-end training method for neural controllers with embedded control-barrier-function safety filters, demonstrated up to 1200 state dimensions and 400 control dimensions, with convergence guarantees unde...

  2. Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Forecasting context shifts and treating adaptation demand against calibrated recovery capacity as a safety gate reduces transient violations in a nonstationary highway-driving simulator.

  3. End-to-End Humanoid Robot Safe and Comfortable Locomotion Policy

    cs.RO 2025-08 conditional novelty 5.0 of 10

    An end-to-end humanoid locomotion policy maps raw LiDAR point clouds to motor commands using P3O with CBF-inspired safety costs and comfort rewards, with sim-to-real tests on a Unitree G1.

Pith tools