REVIEW 6 cited by
Exploration-Exploitation in Constrained MDPs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In many sequential decision-making problems, the goal is to optimize a utility function while satisfying a set of constraints on different utilities. This learning problem is formalized through Constrained Markov Decision Processes (CMDPs). In this paper, we investigate the exploration-exploitation dilemma in CMDPs. While learning in an unknown CMDP, an agent should trade-off exploration to discover new information about the MDP, and exploitation of the current knowledge to maximize the reward while satisfying the constraints. While the agent will eventually learn a good or optimal policy, we do not want the agent to violate the constraints too often during the learning process. In this work, we analyze two approaches for learning in CMDPs. The first approach leverages the linear formulation of CMDP to perform optimistic planning at each episode. The second approach leverages the dual formulation (or saddle-point formulation) of CMDP to perform incremental, optimistic updates of the primal and dual variables. We show that both achieves sublinear regret w.r.t.\ the main utility while having a sublinear regret on the constraint violations. That being said, we highlight a crucial difference between the two approaches; the linear programming approach results in stronger guarantees than in the dual formulation based approach.
Forward citations
Cited by 6 Pith papers
-
Decoupling Corruption and Horizon in Robust Contextual Pricing
Robust contextual pricing admits regret O(Cd + d² log T), the first bound that additively separates corruption budget C from horizon T.
-
Beyond Slater's Condition in Online CMDPs with Stochastic and Adversarial Constraints
A new algorithm achieves Õ(√T) regret and constraint violation in online CMDPs without Slater's condition, plus sublinear α-regret against the unconstrained optimum under adversarial constraints.
-
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
A new RL reward-shaping method that rewards low confidence on hard problems and high confidence on easy ones improves LLM math reasoning over a GRPO baseline.
-
No-Regret Learning Under Adversarial Resource Constraints: A Spending Plan Is All You Need!
Given a spending plan, primal-dual no-regret algorithms achieve tilde-O(sqrt T) dynamic or static regret under adversarially changing reward and cost distributions, and tilde-O(T^{3/4}) when the plan is highly imbalanced.
-
Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees
Two new Monte Carlo tree search algorithms, CVaR-MCTS and W-MCTS, give provable PAC-level tail-risk controls and regret bounds for worst-case outcome scenarios.
-
An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints
A primal-dual algorithm with optimistic mirror descent is claimed to achieve O~(sqrt K) regret and O~(sqrt K) strong constraint violation in episodic CMDPs with anytime adversarial constraints, without Slater's condition.
Discussion (0). Continue with ORCID to comment.