Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper establishes that, under regularity of the belief filter (weak Feller or Wasserstein continuity plus filter stability), finite-model and finite-window approximations are near-optimal for partially observed MDPs, and that…

desk verdict A useful consolidation of the authors' own POMDP results, but the one new-looking theorem (Thm 5.3) that anchors the learning claim is stated without proof or citation, so read the survey as a map, not as a new result. read the letter →

arxiv 2412.06735 v2 pith:KZXQGT4O submitted 2024-12-09 math.OC cs.SYeess.SY

classification math.OCcs.SYeess.SY MSC 93E2090C4093E35
keywords partiallyobservedMarkovdecisionprocessesbelief-MDPfilterstabilityWassersteincontractionquantizationapproximationQ-learningnear-optimalityaveragecost
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper surveys and consolidates a program of results whose common claim is that partially observed Markov decision processes can be controlled and learned near-optimally whenever their nonlinear filter is regular: weakly Feller or Wasserstein-contractive, and stable with respect to initialization errors. Under those properties, optimal policies exist for discounted and average cost, quantized finite models and finite-window memory policies approximate the true problem with explicit error bounds, and Q-learning on either quantized beliefs or finite observation windows converges almost surely with a stated suboptimality gap. The paper's headline learning result is Theorem 5.3: under asymptotic unique ergodicity of the belief-action process with positive invariant mass on every quantization bin, the policy built from the limit Q-values is within $2\alpha_c/((1-\beta)^2(1-\beta\alpha_\eta)) \bar L$ of optimal for discounted cost. The same set of ideas transfers the guarantee to average cost through the vanishing discount method, so the paper's overall thesis is that near-optimal learning for POMDPs is a theorem under checkable conditions rather than a heuristic.

What carries the argument

The machinery is the belief-MDP reduction together with three regularity doctrines. The first is weak Feller continuity of the filter kernel $\eta(\cdot|\pi,u)$, established under two alternative assumption sets (Theorems 2.1 and 2.2); it makes the belief space a well-behaved MDP and supports existence and asymptotic optimality. The second is Wasserstein contraction, $W_1(\eta(\cdot|z,u), \eta(\cdot|z',u)) \le K_2 W_1(z,z')$ with $K_2 = \alpha D(3-2\delta(Q))/2$; this gives Lipschitz value functions, the average-cost optimality equation under $K_2 < 1$, and the explicit constants in the quantization error bound. The third is filter stability: the total-variation distance between filters started from different priors, either in expectation ($L_N^t$, bounded via the Dobrushin coefficient by $2\alpha^N$) or uniformly ($\bar L_{TV}^N$, bounded geometrically via Hilbert-metric contraction under Assumption 4.2). Filter stability converts finite-window memory into a near-optimal state, and supplies the empirical-averages-approach-expectations that Q-learning needs. For the learning theorems specifically, the load-bearing identity is the fixed-point equation $Q^*(s,u) = C^*(s,u) + \beta \sum_{s_1} V^*(s_1)P^*(s_1|s,u)$, which is the Bellman equation of a limiting approximate MDP; asymptotic ergodicity (Assumptions 5.1, 5.3, 5.4) guarantees that the running averages in the Q-learning update converge to these $C^*$ and $P^*$, so the iterates converge to $Q^*$ almost surely.

What would settle it

A concrete calculation that would test the reach of the claim: construct a POMDP with finite state and output spaces whose filter under a memoryless randomized exploration policy has an absorbing belief or a periodic cycle, so that the invariant measure has zero mass on some quantization bin. In that case the Q-learning iterates for that bin never update, and no near-optimality bound of the form in Theorem 5.3 can be obtained; such an example would demarcate exactly which problems the learning theorem covers.

Watch

Extended reading notes

Core claim

The paper's central claim is that the obstacles to rigorous POMDP control—infinite-dimensional belief space, non-Markovian information, and unknown models—can each be overcome by quantitative regularity of the filter. On the existence side, weak Feller continuity (Theorems 2.1 and 2.2) guarantees an optimal policy for discounted cost, and the Wasserstein contraction with $K_2 = \alpha D(3-2\delta(Q))/2 < 1$ yields a solution to the average-cost optimality equation and a constant optimal cost for every initial belief (Theorem 3.2(i)). On the approximation side, quantizing the belief space gives finite models whose value functions deviate from the true one by at most $2K_1/((1-\beta)^2(1-\beta K_2)) \bar L$, while finite-window reductions, under controlled filter stability, give memory-$N$ policies whose error is controlled by the filter-error term $L_N^t$ (Theorems 4.2 and 4.3). On the learning side, the paper claims that Q-learning is not just an algorithm but a theorem: if the exploration policy makes $\{\pi_t, U_t\}$ asymptotically uniquely ergodic with positive mass on every bin, the iterates converge almost surely, and the learned policy satisfies the bound in Theorem 5.3; for the model-free finite-window variant, Theorem 5.2 gives the same type of guarantee using $L_N^t$ under ergodicity of the hidden state. In the average-cost case, these results transfer through the vanishing discount method, so both criteria are covered.

Load-bearing premise

The load-bearing premise is that during exploration the observer's posterior-probability state visits every cell of the quantized state space infinitely often with positive long-run frequency; the paper itself notes this must be verified problem by problem.

Editorial extensions

If this is right

  • Under the Wasserstein contraction condition $K_2 < 1$, any quantized approximation of the belief MDP is provably near-optimal, with an error proportional to the largest diameter of the quantization bins; refining the partition shrinks the gap at a known rate.
  • Finite-window policies are near-optimal whenever the filter is stable: the expected loss is bounded by a discounted sum of the filter-error terms $L_N^t$, and under geometric filter stability the bound decays geometrically in the window length $N$.
  • Q-learning with quantized beliefs converges almost surely and the learned policy's suboptimality is bounded by an explicit function of the quantization level, provided the exploration policy makes the belief process visit every bin infinitely often.
  • Learning under the average-cost criterion is reduced to learning under discounted cost: with $K_2 < 1$, a near-optimal discounted policy is near-optimal for the average cost, so the same Q-learning guarantees carry over.
  • When only weak Feller continuity holds, asymptotic (rate-free) near-optimality of the learned policies still follows as the quantization becomes infinitely fine, even without a uniform filter stability estimate.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the paper's learning theorem suggests a practical design rule—choose exploration policies that are randomized with full-support observation channels so that all bins in the trained set are visited; checking Assumption 5.4 for a specific channel is the main bottleneck to turning the theorem into an algorithm.
  • Editorial inference: the same proof architecture should extend to sample-based quantization of beliefs (for example, particle-filter outputs), but the convergence would then require controlling both quantization error and particle approximation error simultaneously, which the paper does not do.
  • Editorial inference: the average-cost transfer via discounted near-optimality implies that any finite-time regret bound for the discounted Q-learning analysis would automatically yield an average-cost regret bound up to a $\beta$-dependent factor; the paper does not compute such finite-time rates.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper is a survey/tutorial on optimal control of partially observed Markov decision processes (POMDPs), organized around the belief-MDP reduction. It reviews weak Feller and Wasserstein regularity of the controlled filter, filter stability, existence of optimal policies for discounted and average cost criteria, quantization and finite-window approximations with explicit error bounds, robustness of optimal costs to model and prior errors, and reinforcement learning results. The central advertised claim is that recent RL results rigorously establish convergence to near optimality under both discounted and average cost criteria; for discounted cost, the paper presents two routes: finite-window memory under uniform geometric filter stability, and quantized belief states under asymptotic unique ergodicity, culminating in Theorem 5.3.

Significance. If fully supported, the paper would be a useful unifying review: it collects regularity conditions, existence theorems, and approximation bounds with explicit constants, and it identifies which filter-stability assumptions are needed for learning. The paper is generally careful to attribute results and to flag limitations, e.g., footnote 1 on Theorem 3.2(ii) and Remark 5.1. However, the strongest advertised contribution — RL convergence to near optimality for belief-quantized POMDPs — is not yet supported because Theorem 5.3 is stated without proof or attribution and its key Assumption 5.4 is not established or supplied with verifiable sufficient conditions. The finite-window route (Theorems 5.1 and 5.2) is properly anchored in the cited literature, but the belief-quantization route needs substantial additional support before the abstract's claim can be accepted.

major comments (3)
  1. [Section V-C, Theorem 5.3] The abstract's central claim — that the paper presents RL results that rigorously establish convergence to near optimality — rests on Theorem 5.3, but this theorem is stated without proof, without a citation to a published theorem, and without a derivation. The preceding text only says that the Section V-A and IV-A setup applies 'with a significantly more tedious analysis involving ergodicity requirements.' Every other main theorem in the manuscript is a restatement of a named reference (e.g., Theorem 4.1 from [37], Theorem 5.1 from [42], Theorem 5.2 from [41]), so the status of Theorem 5.3 is unclear. The authors should either supply a complete proof, identify the exact published statement it restates, or explicitly say that it is new; if it is new, the proof must be included.
  2. [Section V-C, Assumption 5.4] Theorem 5.3(a)–(b) requires the controlled belief process {pi_t, U_t} to be asymptotically uniquely ergodic with eta_gamma(B) > 0 for every quantization bin B. This is not a consequence of Assumption 4.1, weak Feller regularity, or filter stability; the paper itself concedes that 'the condition that eta_gamma(B) > 0 requires an analysis tailored for each problem.' The suggested sufficient routes (quantizing via finite past windows, or initializing from the invariant measure of the hidden Markov source) do not cover the general quantized belief-MDP used in Theorem 5.3, and the invariant measure of the hidden source is an object on X, not on P(X). Without a verifiable sufficient condition for eta_gamma(B) > 0, the claimed almost-sure convergence and the bound J*_beta(pi_0, gamma_hat) - J*_beta(pi_0) <= 2 alpha_c / ((1-beta)^2 (1-beta alpha_eta)) bar_L are not established.
  3. [Section V-C, Remark 5.1(i) and Theorem 5.3(b)] The near-optimality bound for the learned policy assumes not only convergence of the Q-iterates, but also that the policy constructed from the limit Q-values, when applied to the true model, keeps the closed-loop process inside the trained set P_eta for all time. The manuscript only references [42] and [12, Lemma 6 and Corollary 2] for this, and Remark 5.1(i) notes that 'one needs to ensure' the closed-loop condition without stating conditions under which it follows. Since Theorem 5.3(b) is the statement that connects the learned Q-values to near optimality in the original POMDP, this closed-loop condition is load-bearing and should be made an explicit hypothesis or derived from the assumptions.
minor comments (5)
  1. [Section V-A, after Eq. (27)] The definition 'V*(u) := min_u Q*(s, u)' has a variable mismatch; it should read 'V*(s) := min_u Q*(s, u)'.
  2. [Section V-B, Assumption 5.3(ii)] Assumption 5.3(ii) refers to 'Assumption 5.1(i)', but Assumption 5.1 has no numbered parts; please renumber or cross-reference precisely.
  3. [Section IV-A, quantization construction] The construction requires a weight measure pi* with pi*(Z_i) > 0 for every quantization bin; the paper does not discuss how such a measure is chosen or why it exists uniformly as the number of bins M grows. A brief justification would help.
  4. [Section III-B, footnote 1] Footnote 1 states that the optimality result in Theorem 3.2(ii) 'may only hold for a restrictive class of initial conditions or initializations'; this is too vague. Please specify the class or point to the exact statement in [72].
  5. [Section I-B, Eq. (8)] The bounded Lipschitz metric is defined using a distance d on the underlying space, but the same symbol d is used later for the metric on P(X) without explicit comment; a sentence noting which metric is used for probability measures would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a review whose learning results are cited from the authors' prior published work, and no claimed bound reduces to its inputs by construction.

full rationale

This is a review/tutorial, not a paper fitting parameters to data. The main theorems are explicitly attributed: Theorem 2.1 to [23], Theorem 4.1 to [37], Theorem 5.1 to [42], Theorem 5.2 to [41], and Theorem 4.4 to [15]. These prior results are peer-reviewed, state explicit assumptions (e.g., Assumptions 4.1 and 5.2), and do not build the target near-optimality claim into their hypotheses, so under the review rule they are real evidence rather than circular self-citation. The strongest new-looking claim is Theorem 5.3 in Section V-C, which is stated without proof: the text says the quantized Q-learning setup applies 'with a significantly more tedious analysis involving ergodicity requirements,' and Assumption 5.4, requiring eta_gamma(B) > 0 on every quantization bin, is conceded to 'require an analysis tailored for each problem.' Remark 5.1 also defers to [42] and [12] for the condition that the closed loop under the learned policy remains in the trained set. These are omitted-proof and unverified-hypothesis concerns about the abstract's 'rigorously establish' wording; they are not circularity. The Theorem 5.3 bound 2 alpha_c / ((1-beta)^2 (1-beta alpha_eta)) bar_L has the same form as the quantization bound in Theorem 4.1, but that is a mathematical implication, not a definitional identity or a fitted parameter disguised as a prediction. No quantity is defined in terms of the quantity it is claimed to predict. Therefore no circular step can be exhibited, and the circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper is a review, so it does not fit free parameters. The central claims depend on a chain of regularity and ergodicity assumptions inherited from prior papers; most are stated explicitly as Assumptions 2.1-2.3, 4.1-4.2, and 5.1-5.5. The strongest additional burden is the unique ergodicity assumption on the belief process in Section V-C.

assumptions (5)
  • domain assumption POMDP model: standard Borel spaces, transition kernel T, observation kernel Q, bounded continuous cost, and admissible measurable policies.
    This defines the problem class; Section I states the model and invokes the Ionescu Tulcea theorem for the strategic measure.
  • domain assumption Either Assumption 2.1 or 2.2, i.e. weak continuity or total-variation continuity of the transition and observation kernels.
    The weak Feller property of the filter (Theorems 2.1-2.2) is conditional on these, and the existence results in Section III rely on it.
  • domain assumption Assumption 2.3 with K2 = alpha D (3 - 2 delta(Q)) / 2 < 1 for Wasserstein contraction of the belief kernel.
    Used in Theorem 2.3, the Lipschitz value function result, Theorem 3.2(i) for the average cost optimality equation, and explicit quantization bounds.
  • domain assumption Controlled filter stability: Dobrushin coefficient contraction, one-step stochastic observability, or mixing kernel/Hilbert metric contraction (Assumptions 2.3, 2.7 and 4.2).
    Finite-window approximations and finite-window Q-learning bounds in Theorems 4.2-4.4 and 5.2 are obtained only under one of these stability regimes.
  • domain assumption Asymptotic unique ergodicity of the relevant state or belief process under exploration, as in Assumptions 5.1-5.4.
    The Q-learning convergence theorem (Theorem 5.1 from [42]) and Theorem 5.3 require time averages to converge and every relevant state-action pair to be visited infinitely often.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning." pith.science (2026). https://pith.science/paper/KZXQGT4O

@misc{pith2026241206735,
  author       = {Pith},
  title        = {Pith review of: Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KZXQGT4O}},
  note         = {Machine review of arXiv:2412.06735}
}
read the original abstract

In this review/tutorial article, we present recent progress on optimal control of partially observed Markov Decision Processes (POMDPs). We first present regularity and continuity conditions for POMDPs and their belief-MDP reductions, where these constitute weak Feller and Wasserstein regularity and controlled filter stability. These are then utilized to arrive at existence results on optimal policies for both discounted and average cost problems, and regularity of value functions. Then, we study rigorous approximation results involving quantization based finite model approximations as well as finite window approximations under controlled filter stability. Finally, we present several recent reinforcement learning theoretic results which rigorously establish convergence to near optimality under both criteria.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sensitivity of Filter Kernels and Robustness Bounds to Transition and Measurement Kernel Perturbations in Partially Observable Stochastic Control

    math.OC 2025-08 conditional novelty 6.0 of 10

    Explicit upper bounds on filter-kernel sensitivity and POMDP value differences in terms of Wasserstein and total-variation distances between transition and observation kernels, with decaying error bounds for joint sta...

Reference graph

Works this paper leans on

76 extracted references · 71 canonical work pages · cited by 1 Pith paper

  1. [37]

    Saldi, and S

    A.D Kara, N. Saldi, and S. Y¨ uksel. Q-learning for MDPs w ith general spaces: Convergence and near optimality via quantization u nder weak continuity. Journal of Machine Learning Research , pages 1–34, 2023

  2. [42]

    Q-Learning for Stochastic Control under General Information Structures and Non-Markovian Environments

    A.D. Kara and S. Y¨ uksel. Q-learning for stochastic con trol under gen- eral information structures and non-Markovian environmen ts. Trans- actions on Machine Learning Research (arXiv:2311.00123) , 2024

  3. [41]

    Y¨ uksel

    A.D Kara and S. Y¨ uksel. Convergence of finite memory Q-l earning for POMDPs and near optimality of learned policies under filter s tability. Mathematics of Operations Research , 48(4):2066–2093, 2023

  4. [1]

    R. J. Aumann. Mixed and behavior strategies in infinite ex tensive games. Technical report, Princeton University NJ, 1961

  5. [2]

    Reinforcement learning of pomdps using spectr al methods

    Kamyar Azizzadenesheli, Alessandro Lazaric, and Anima shree Anandkumar. Reinforcement learning of pomdps using spectr al methods. In Conference on Learning Theory , pages 193–256. PMLR, 2016

  6. [3]

    Billingsley

    P . Billingsley. Convergence of probability measures. New Y ork: Wiley, 2nd edition, 1999

  7. [4]

    Blackwell

    D. Blackwell. Memoryless strategies in finite-stage dyn amic program- ming. Annals of Mathematical Statistics , 35:863–865, 1964

  8. [5]

    V . S. Borkar. White-noise representations in stochasti c realization theory. SIAM J. on Control and Optimization , 31:1093–1102, 1993

Show all 76 references
  1. [6]

    V . S. Borkar. Average cost dynamic programming equation s for controlled Markov chains with partial observations. SIAM J. Control Optim., 39(3):673–681, 2000

  2. [7]

    V . S. Borkar. Convex analytic methods in Markov decision processes. In Handbook of Markov Decision Processes, E. A. Feinberg, A. Shwartz (Eds.) , pages 347–375. Kluwer, Boston, MA, 2001

  3. [8]

    Budhiraja

    A. Budhiraja. On invariant measures of discrete time filt ers in the correlated signal-noise case. The Annals of Applied Probability , 12(3):1096–1113, 2002

  4. [9]

    Cayci, N

    S. Cayci, N. He, and R. Srikant. Finite-time analysis of e ntropy- regularized neural natural actor-critic algorithm. arXiv preprint arXiv:2206.00833, 2022

  5. [10]

    Chandak, V .S

    S. Chandak, V .S. Borkar, and P . Dodhia. Reinforcement l earning in non-markovian environments. Systems & Control Letters , 185:105751, 2024

  6. [11]

    Chigansky and R

    P . Chigansky and R. V an Handel. A complete solution to Bl ackwell’s unique ergodicity problem for hidden Markov chains. The Annals of Applied Probability, 20(6):2318–2345, 2010

  7. [12]

    Cregg, T

    L. Cregg, T. Linder, and S. Y¨ uksel. Reinforcement lear ning for near-optimal design of zero-delay codes for markov sources . IEEE Transactions on Information Theory, arXiv:2311.12609 , 2024

  8. [13]

    Crisan and A

    D. Crisan and A. Doucet. A survey of convergence results on particle filtering methods for practitioners. IEEE Transactions on Signal Processing, 50(3):736–746, 2002

  9. [14]

    Demirci, A.D

    Y .E. Demirci, A.D. Kara, and S. Y¨ uksel. Average cost op timality of partially observed mdps: Contraction of non-linear filters and existence of optimal solutions. SIAM Journal on Control and Optimization, arXiv:2312.14111, 2024

  10. [15]

    Demirci, A.D

    Y .E. Demirci, A.D. Kara, and S. Y¨ uksel. Refined bounds o n near optimality finite window policies in pomdps and their reinfo rcement learning. arXiv, 2024

  11. [16]

    Demirci and S

    Y .E. Demirci and S. Y¨ uksel. Geometric ergodicity and w asserstein continuity of non-linear filters. arXiv preprint arXiv:2307.15764 , 2023

  12. [17]

    Devroye and L

    L. Devroye and L. Gy¨ orfi. Non-parametric Density Estimation: The L1 View. John Wiley, New Y ork, 1985

  13. [18]

    Dobrushin

    R.L. Dobrushin. Central limit theorem for nonstationa ry Markov chains. i. Theory of Probability & Its Applications , 1(1):65–80, 1956

  14. [19]

    S. Dong, B. van Roy, and Z. Zhou. Simple agent, complex en viron- ment: Efficient reinforcement learning with agent states. The Journal of Machine Learning Research , 23(1):11627–11680, 2022

  15. [20]

    R. M. Dudley. Real Analysis and Probability . Cambridge University Press, Cambridge, 2nd edition, 2002

  16. [21]

    Reinf orcement learning in pomdps without resets

    Eyal Even-Dar, Sham M Kakade, and Yishay Mansour. Reinf orcement learning in pomdps without resets. In IJCAI, pages 690–695, 2005

  17. [22]

    Feinberg and P .O

    E.A. Feinberg and P .O. Kasyanov. Equivalent condition s for weak continuity of nonlinear filters. Systems & Control Letters , 173:105458, 2023

  18. [23]

    Feinberg, P .O

    E.A. Feinberg, P .O. Kasyanov, and M.Z. Zgurovsky. Part ially observ- able total-cost Markov decision process with weakly contin uous tran- sition probabilities. Mathematics of Operations Research , 41(2):656– 681, 2016

  19. [24]

    Feinberg, P .O

    E.A. Feinberg, P .O. Kasyanov, and M.Z. Zgurovsky. Mark ov deci- sion processes with incomplete information and semiunifor m feller transition probabilities. SIAM Journal on Control and Optimization , 60(4):2488–2513, 2022

  20. [25]

    Fern´ andez-Gaucherand, A

    E. Fern´ andez-Gaucherand, A. Arapostathis, and S. I. M arcus. Remarks on the existence of solutions to the average cost optimality equation in markov decision processes. Systems & control letters , 15(5):425–432, 1990

  21. [26]

    F.Kochman and J. Reeds. A simple proof of kaijser’s uniq ue ergodicity result for hidden markov α -chains. The Annals of Applied Probability , 16(4):1805–1815, 2006

  22. [27]

    I. I. Gihman and A. V . Skorohod. Controlled Stochastic Processes . Springer Science & Business Media, 2012

  23. [28]

    Le Gland and N

    F. Le Gland and N. Oudjane. Stability and uniform approx imation of nonlinear filters using the Hilbert metric and application t o particle filters. The Annals of Applied Probability , 14(1):144–187, 2004

  24. [29]

    Hern´ andez-Lerma

    O. Hern´ andez-Lerma. Adaptive Markov Control Processes . Springer- V erlag, 1989

  25. [30]

    Hern´ andez-Lerma and J

    O. Hern´ andez-Lerma and J. B. Lasserre. Discrete-Time Markov Control Processes: Basic Optimality Criteria . Springer, 1996

  26. [31]

    Hsu, D.-M Chuang, and A

    S.-P . Hsu, D.-M Chuang, and A. Arapostathis. On the exis tence of stationary optimal policies for partially observed mdps un der the long- run average cost criterion. Systems & control letters , 55(2):165–173, 2006

  27. [32]

    Sample-efficient reinforcement learning of undercomplete pomdps

    Chi Jin, Sham Kakade, Akshay Krishnamurthy, and Qinghu a Liu. Sample-efficient reinforcement learning of undercomplete pomdps. Advances in Neural Information Processing Systems , 33:18530–18539, 2020

  28. [33]

    T. Kaijser. A limit theorem for partially observed Mark ov chains. Annals of Probability , 4(677), 1975

  29. [34]

    T. Kaijser. On markov chains induced by partitioned tra nsition proba- bility matrices. Acta Mathematica Sinica, English Series , 27(3):441– 476, 2011

  30. [35]

    A. D. Kara, M. Raginsky, and S. Y¨ uksel. Robustness to in correct models and data-driven learning in average-cost optimal st ochastic control. Automatica, 139:110179, 2022

  31. [36]

    Saldi, and S

    A.D Kara, N. Saldi, and S. Y¨ uksel. Weak Feller property of non-linear filters. Systems & Control Letters , 134:104–512, 2019

  32. [38]

    Y¨ uksel

    A.D Kara and S. Y¨ uksel. Robustness to incorrect priors in partially ob- served stochastic control. SIAM Journal on Control and Optimization , 57(3):1929–1964, 2019

  33. [39]

    Y¨ uksel

    A.D Kara and S. Y¨ uksel. Robustness to incorrect system models in stochastic control. SIAM Journal on Control and Optimization , 58(2):1144–1182, 2020

  34. [40]

    Y¨ uksel

    A.D Kara and S. Y¨ uksel. Near optimality of finite memory feedback policies in partially observed markov decision processes. Journal of Machine Learning Research , 23(11):1–46, 2022

  35. [43]

    Kara and S

    Ali D. Kara and S. Y¨ uksel. Robustness to approximation s and model learning in MDPs and POMDPs. In A. B. Piunovskiy and Y . Zhang, editors, Modern Trends in Controlled Stochastic Processes: Theory and Applications, V olume III . Luniver Press, 2021

  36. [44]

    Rl for latent mdps: Regret guarantees and a lower bou nd

    Jeongyeol Kwon, Y onathan Efroni, Constantine Caraman is, and Shie Mannor. Rl for latent mdps: Regret guarantees and a lower bou nd. Advances in Neural Information Processing Systems , 34:24523–24534, 2021

  37. [45]

    H.J. Langen. Convergence of dynamic programming model s. Mathe- matics of Operations Research , 6(4):493–512, Nov. 1981

  38. [46]

    Di Masi and L

    G. Di Masi and L. Stettner. Ergodicity of hidden markov m odels. Mathematics of Control, Signals and Systems , 17(4):269–296, 2005

  39. [47]

    McDonald and S

    C. McDonald and S. Y¨ uksel. Exponential filter stabilit y via Do- brushin’s coefficient. Electronic Communications in Probability , 25, 2020

  40. [48]

    McDonald and S

    C. McDonald and S. Y¨ uksel. Robustness to incorrect pri ors and controlled filter stability in partially observed stochast ic control. SIAM Journal on Control and Optimization , 60(2):842–870, 2022

  41. [49]

    McDonald and S

    C. McDonald and S. Y¨ uksel. Stochastic observability a nd filter stability under several criteria. IEEE Transactions on Automatic Control, 69(5):2931–2946, 2024

  42. [50]

    Parthasarathy

    K.R. Parthasarathy. Probability Measures on Metric Spaces . AMS Bookstore, 1967

  43. [51]

    L. K. Platzman. Optimal infinite-horizon undiscounted control of finite probabilistic systems. SIAM Journal on Control and Optimization , 18(4):362–380, 1980

  44. [52]

    D. Rhenius. Incomplete information in Markovian decis ion models. Ann. Statist. , 2:1327–1334, 1974

  45. [53]

    Approximations of discrete time partially observed control problems

    Wolfgang J Runggaldier and Lukasz Stettner. Approximations of discrete time partially observed control problems . Giardini Pisa, 1994

  46. [54]

    Saldi, T

    N. Saldi, T. Linder, and S. Y¨ uksel. Finite Approximations in Discrete- Time Stochastic Control: Quantized Models and Asymptotic O ptimal- ity. Springer, Cham, 2018

  47. [55]

    Saldi, S

    N. Saldi, S. Y¨ uksel, and T. Linder. Near optimality of q uantized poli- cies in stochastic control under weak continuity condition s. Journal of Mathematical Analysis and Applications , 435(1):321–337, 2016

  48. [56]

    Saldi, S

    N. Saldi, S. Y¨ uksel, and T. Linder. On the asymptotic op timality of finite approximations to Markov decision processes with Bor el spaces. Mathematics of Operations Research , 42(4):945–978, 2017

  49. [57]

    Saldi, S

    N. Saldi, S. Y¨ uksel, and T. Linder. Finite model approx imations for partially observed Markov decision processes with discoun ted cost. IEEE Transactions on Automatic Control , 65, 2020

  50. [58]

    R. Serfozo. Convergence of Lebesgue integrals with var ying measures. Sankhy¯ a: The Indian Journal of Statistics, Series A , pages 380–402, 1982

  51. [59]

    Seyedsalehi, N

    E. Seyedsalehi, N. Akbarzadeh, A. Sinha, and A. Mahajan . Approx- imate information state based convergence analysis of recu rrent q- learning. arXiv preprint arXiv:2306.05991 , 2023

  52. [60]

    Singh, T

    S.P . Singh, T. Jaakkola, and M.I. Jordan. Learning with out state- estimation in partially observable markovian decision pro cesses. In Machine Learning Proceedings 1994 , pages 284–292. Elsevier, 1994

  53. [61]

    Sinha and A

    A. Sinha and A. Mahajan. Agent-state based policies in p omdps: Beyond belief-state mdps. In IEEE Conference on Decision and Control, Tutorial Paper , 2024

  54. [62]

    Stettner

    L. Stettner. Long run control with degenerate observat ion. SIAM Journal on Control and Optimization , 57(2):880–899, 2019

  55. [63]

    Subramanian, A

    J. Subramanian, A. Sinha, R. Seraj, and A. Mahajan. Appr oximate information state for approximate planning and reinforcem ent learning in partially observed systems. The Journal of Machine Learning Research, 23(1):483–565, 2022

  56. [64]

    T. Szarek. Feller processes on nonlocally compact spac es. Annals of Probability, pages 1849–1863, 2006

  57. [65]

    Budhiraja V

    A. Budhiraja V . S. Borkar. A further remark on dynamic pr ogramming for partially observed markov processes. Stochastic processes and their applications, 112(1):79–93, 2004

  58. [66]

    van Handel

    R. van Handel. Observability and nonlinear filtering. Probability theory and related fields , 145(1-2):35–74, 2009

  59. [67]

    van Handel

    R. van Handel. Uniform time average consistency of Mont e Carlo particle filters. Stochastic Processes and their Applications , 119(11):3835–3861, 2009

  60. [68]

    van Handel

    R. van Handel. Nonlinear filtering and systems theory. I n Proceedings of the 19th International Symposium on Mathematical Theory of Networks and Systems (MTNS semi-plenary paper) , 2010

  61. [69]

    C. Villani. Optimal Transport: Old and New . Springer, 2008

  62. [70]

    Su blinear regret for learning pomdps

    Yi Xiong, Ningyuan Chen, Xuefeng Gao, and Xiang Zhou. Su blinear regret for learning pomdps. Production and Operations Management , 31(9):3491–3504, 2022

  63. [71]

    Y u and D

    H. Y u and D. P . Bertsekas. On near optimality of the set of finite- state controllers for average cost pomdp. Mathematics of Operations Research, 33(1):1–11, 2008

  64. [72]

    Y¨ uksel

    S. Y¨ uksel. Another look at partially observed optimal stochastic control: Existence, ergodicity, and approximations witho ut belief- reduction. arXiv, 2023

  65. [73]

    Y¨ uksel

    S. Y¨ uksel. Optimization and Control of Stochastic Systems . Queen’s University, Lecture Notes, available online, Queen’s Univ ersity, Lec- ture notes

  66. [74]

    Y¨ uksel and T

    S. Y¨ uksel and T. Bas ¸ar. Stochastic Teams, Games, and Control under Information Constraints . Springer, Cham, 2024

  67. [75]

    Y¨ uksel and T

    S. Y¨ uksel and T. Linder. Optimization and convergence of observation channels in stochastic control. SIAM J. on Control and Optimization , 50:864–887, 2012

  68. [76]

    Y ushkevich

    A.A. Y ushkevich. Reduction of a controlled Markov mode l with incomplete data to a problem with complete information in th e case of Borel state and control spaces. Theory Prob. Appl. , 21:153–158, 1976

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.