Pith. sign in

REVIEW 1 cited by

Learning Safety Constraints from Demonstrations with Unknown Rewards

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16147 v2 pith:VCIQ54OL submitted 2023-05-25 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords demonstrationscocorlconstraintssafelearningdifferentrewardsconvex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose Convex Constraint Learning for Reinforcement Learning (CoCoRL), a novel approach for inferring shared constraints in a Constrained Markov Decision Process (CMDP) from a set of safe demonstrations with possibly different reward functions. While previous work is limited to demonstrations with known rewards or fully known environment dynamics, CoCoRL can learn constraints from demonstrations with different unknown rewards without knowledge of the environment dynamics. CoCoRL constructs a convex safe set based on demonstrations, which provably guarantees safety even for potentially sub-optimal (but safe) demonstrations. For near-optimal demonstrations, CoCoRL converges to the true safe set with no policy regret. We evaluate CoCoRL in gridworld environments and a driving simulation with multiple constraints. CoCoRL learns constraints that lead to safe driving behavior. Importantly, we can safely transfer the learned constraints to different tasks and environments. In contrast, alternative methods based on Inverse Reinforcement Learning (IRL) often exhibit poor performance and learn unsafe policies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Soft Driving Constraints from Vectorized Scene Embeddings while Imitating Expert Trajectories

    cs.RO 2024-12 conditional novelty 4.0 of 10

    A motion planner that learns soft driving constraints from vectorized scene embeddings improves closed-loop safety and interpretability over a reward-only imitation-learning baseline.

Pith tools