REVIEW 1 cited by
Constrained episodic reinforcement learning in concave-convex and knapsack settings
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose an algorithm for tabular episodic reinforcement learning with constraints. We provide a modular analysis with strong theoretical guarantees for settings with concave rewards and convex constraints, and for settings with hard constraints (knapsacks). Most of the previous work in constrained reinforcement learning is limited to linear constraints, and the remaining work focuses on either the feasibility question or settings with a single episode. Our experiments demonstrate that the proposed algorithm significantly outperforms these approaches in existing constrained episodic environments.
Forward citations
Cited by 1 Pith paper
-
Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment
Evaluation blindness, silent measurement failure where monitoring looks healthy while systems fail, is formalized, classified into six production classes, and found in 53% of 36 verifiable public incidents.
Discussion (0). Continue with ORCID to comment.