Pith. sign in

REVIEW 2 cited by

Span-Based Optimal Sample Complexity for Average Reward MDPs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.13469 v2 pith:XSJ47OKA submitted 2023-11-22 cs.LG cs.ITmath.ITmath.OCstat.ML

classification cs.LGcs.ITmath.ITmath.OCstat.ML
keywords varepsilongammaoptimalfracmdpsboundscomplexitydiscounted
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We study the sample complexity of learning an $\varepsilon$-optimal policy in an average-reward Markov decision process (MDP) under a generative model. We establish the complexity bound $\widetilde{O}\left(SA\frac{H}{\varepsilon^2} \right)$, where $H$ is the span of the bias function of the optimal policy and $SA$ is the cardinality of the state-action space. Our result is the first that is minimax optimal (up to log factors) in all parameters $S,A,H$ and $\varepsilon$, improving on existing work that either assumes uniformly bounded mixing times for all policies or has suboptimal dependence on the parameters. Our result is based on reducing the average-reward MDP to a discounted MDP. To establish the optimality of this reduction, we develop improved bounds for $\gamma$-discounted MDPs, showing that $\widetilde{O}\left(SA\frac{H}{(1-\gamma)^2\varepsilon^2} \right)$ samples suffice to learn a $\varepsilon$-optimal policy in weakly communicating MDPs under the regime that $\gamma \geq 1 - \frac{1}{H}$, circumventing the well-known lower bound of $\widetilde{\Omega}\left(SA\frac{1}{(1-\gamma)^3\varepsilon^2} \right)$ for general $\gamma$-discounted MDPs. Our analysis develops upper bounds on certain instance-dependent variance parameters in terms of the span parameter. These bounds are tighter than those based on the mixing time or diameter of the MDP and may be of broader use.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

    cs.LG 2025-05 reject novelty 7.0 of 10

    RHI is claimed to find an epsilon-optimal robust policy under the average-reward criterion with about SAH^2/epsilon^2 samples under the communicating assumption, with a parameter-free variant that avoids knowing H.

  2. A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs

    cs.LG 2025-04 conditional novelty 6.0 of 10

    A discounted value-iteration algorithm with visited-state clipping and deviation-controlled updates achieves ~O(sp(v*) sqrt(d^3 T)) regret for infinite-horizon average-reward linear MDPs with computational cost indepe...

Pith tools