Pith. sign in

REVIEW 1 cited by

Constrained episodic reinforcement learning in concave-convex and knapsack settings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.05051 v2 pith:IJZCKUTI submitted 2020-06-09 cs.LG cs.AIcs.DSstat.ML

classification cs.LGcs.AIcs.DSstat.ML
keywords constraintssettingsconstrainedepisodiclearningreinforcementalgorithmwork
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose an algorithm for tabular episodic reinforcement learning with constraints. We provide a modular analysis with strong theoretical guarantees for settings with concave rewards and convex constraints, and for settings with hard constraints (knapsacks). Most of the previous work in constrained reinforcement learning is limited to linear constraints, and the remaining work focuses on either the feasibility question or settings with a single episode. Our experiments demonstrate that the proposed algorithm significantly outperforms these approaches in existing constrained episodic environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment

    cs.LG 2026-08 conditional novelty 5.0 of 10

    Evaluation blindness, silent measurement failure where monitoring looks healthy while systems fail, is formalized, classified into six production classes, and found in 53% of 36 verifiable public incidents.

Pith tools