Pith. sign in

REVIEW 1 cited by

Average Cost Optimality of Partially Observed MDPS: Contraction of Non-linear Filters, Optimal Solutions and Approximations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.14111 v3 pith:MMVJJ2VO submitted 2023-12-21 math.OC

classification math.OC
keywords optimalityaveragecostanalysisapproximationsavailablebeyondconditions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The average cost optimality is known to be a challenging problem for partially observable stochastic control, with few results available beyond the finite state, action, and measurement setup, for which somewhat restrictive conditions are available. In this paper, we present explicit and easily testable conditions for the existence of solutions to the average cost optimality equation where the state space is compact. In particular, we present a new contraction based analysis, which is new to the literature to our knowledge, building on recent regularity results for non-linear filters. Beyond establishing existence, we also present several implications of our analysis that are new to the literature: (i) robustness to incorrect priors (ii) near optimality of policies based on quantized approximations, (iii) near optimality of policies with finite memory, and (iv) convergence in Q-learning. In addition to our main theorem, each of these represents a novel contribution for average cost criteria.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning

    math.OC 2024-12 conditional novelty 2.0 of 10

    A survey of regularity, approximation, and reinforcement learning guarantees for partially observed Markov decision processes, drawing mostly on the authors' earlier work.

Pith tools