REVIEW 3 major objections 5 minor 1 cited by
Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper establishes that, under regularity of the belief filter (weak Feller or Wasserstein continuity plus filter stability), finite-model and finite-window approximations are near-optimal for partially observed MDPs, and that…
desk verdict A useful consolidation of the authors' own POMDP results, but the one new-looking theorem (Thm 5.3) that anchors the learning claim is stated without proof or citation, so read the survey as a map, not as a new result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the belief-MDP reduction together with three regularity doctrines. The first is weak Feller continuity of the filter kernel $\eta(\cdot|\pi,u)$, established under two alternative assumption sets (Theorems 2.1 and 2.2); it makes the belief space a well-behaved MDP and supports existence and asymptotic optimality. The second is Wasserstein contraction, $W_1(\eta(\cdot|z,u), \eta(\cdot|z',u)) \le K_2 W_1(z,z')$ with $K_2 = \alpha D(3-2\delta(Q))/2$; this gives Lipschitz value functions, the average-cost optimality equation under $K_2 < 1$, and the explicit constants in the quantization error bound. The third is filter stability: the total-variation distance between filters started from different priors, either in expectation ($L_N^t$, bounded via the Dobrushin coefficient by $2\alpha^N$) or uniformly ($\bar L_{TV}^N$, bounded geometrically via Hilbert-metric contraction under Assumption 4.2). Filter stability converts finite-window memory into a near-optimal state, and supplies the empirical-averages-approach-expectations that Q-learning needs. For the learning theorems specifically, the load-bearing identity is the fixed-point equation $Q^*(s,u) = C^*(s,u) + \beta \sum_{s_1} V^*(s_1)P^*(s_1|s,u)$, which is the Bellman equation of a limiting approximate MDP; asymptotic ergodicity (Assumptions 5.1, 5.3, 5.4) guarantees that the running averages in the Q-learning update converge to these $C^*$ and $P^*$, so the iterates converge to $Q^*$ almost surely.
What would settle it
A concrete calculation that would test the reach of the claim: construct a POMDP with finite state and output spaces whose filter under a memoryless randomized exploration policy has an absorbing belief or a periodic cycle, so that the invariant measure has zero mass on some quantization bin. In that case the Q-learning iterates for that bin never update, and no near-optimality bound of the form in Theorem 5.3 can be obtained; such an example would demarcate exactly which problems the learning theorem covers.
Extended reading notes
Core claim
The paper's central claim is that the obstacles to rigorous POMDP control—infinite-dimensional belief space, non-Markovian information, and unknown models—can each be overcome by quantitative regularity of the filter. On the existence side, weak Feller continuity (Theorems 2.1 and 2.2) guarantees an optimal policy for discounted cost, and the Wasserstein contraction with $K_2 = \alpha D(3-2\delta(Q))/2 < 1$ yields a solution to the average-cost optimality equation and a constant optimal cost for every initial belief (Theorem 3.2(i)). On the approximation side, quantizing the belief space gives finite models whose value functions deviate from the true one by at most $2K_1/((1-\beta)^2(1-\beta K_2)) \bar L$, while finite-window reductions, under controlled filter stability, give memory-$N$ policies whose error is controlled by the filter-error term $L_N^t$ (Theorems 4.2 and 4.3). On the learning side, the paper claims that Q-learning is not just an algorithm but a theorem: if the exploration policy makes $\{\pi_t, U_t\}$ asymptotically uniquely ergodic with positive mass on every bin, the iterates converge almost surely, and the learned policy satisfies the bound in Theorem 5.3; for the model-free finite-window variant, Theorem 5.2 gives the same type of guarantee using $L_N^t$ under ergodicity of the hidden state. In the average-cost case, these results transfer through the vanishing discount method, so both criteria are covered.
Load-bearing premise
The load-bearing premise is that during exploration the observer's posterior-probability state visits every cell of the quantized state space infinitely often with positive long-run frequency; the paper itself notes this must be verified problem by problem.
Editorial extensions
If this is right
- Under the Wasserstein contraction condition $K_2 < 1$, any quantized approximation of the belief MDP is provably near-optimal, with an error proportional to the largest diameter of the quantization bins; refining the partition shrinks the gap at a known rate.
- Finite-window policies are near-optimal whenever the filter is stable: the expected loss is bounded by a discounted sum of the filter-error terms $L_N^t$, and under geometric filter stability the bound decays geometrically in the window length $N$.
- Q-learning with quantized beliefs converges almost surely and the learned policy's suboptimality is bounded by an explicit function of the quantization level, provided the exploration policy makes the belief process visit every bin infinitely often.
- Learning under the average-cost criterion is reduced to learning under discounted cost: with $K_2 < 1$, a near-optimal discounted policy is near-optimal for the average cost, so the same Q-learning guarantees carry over.
- When only weak Feller continuity holds, asymptotic (rate-free) near-optimality of the learned policies still follows as the quantization becomes infinitely fine, even without a uniform filter stability estimate.
Reading between the lines
- Editorial inference: the paper's learning theorem suggests a practical design rule—choose exploration policies that are randomized with full-support observation channels so that all bins in the trained set are visited; checking Assumption 5.4 for a specific channel is the main bottleneck to turning the theorem into an algorithm.
- Editorial inference: the same proof architecture should extend to sample-based quantization of beliefs (for example, particle-filter outputs), but the convergence would then require controlling both quantization error and particle approximation error simultaneously, which the paper does not do.
- Editorial inference: the average-cost transfer via discounted near-optimality implies that any finite-time regret bound for the discounted Q-learning analysis would automatically yield an average-cost regret bound up to a $\beta$-dependent factor; the paper does not compute such finite-time rates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a survey/tutorial on optimal control of partially observed Markov decision processes (POMDPs), organized around the belief-MDP reduction. It reviews weak Feller and Wasserstein regularity of the controlled filter, filter stability, existence of optimal policies for discounted and average cost criteria, quantization and finite-window approximations with explicit error bounds, robustness of optimal costs to model and prior errors, and reinforcement learning results. The central advertised claim is that recent RL results rigorously establish convergence to near optimality under both discounted and average cost criteria; for discounted cost, the paper presents two routes: finite-window memory under uniform geometric filter stability, and quantized belief states under asymptotic unique ergodicity, culminating in Theorem 5.3.
Significance. If fully supported, the paper would be a useful unifying review: it collects regularity conditions, existence theorems, and approximation bounds with explicit constants, and it identifies which filter-stability assumptions are needed for learning. The paper is generally careful to attribute results and to flag limitations, e.g., footnote 1 on Theorem 3.2(ii) and Remark 5.1. However, the strongest advertised contribution — RL convergence to near optimality for belief-quantized POMDPs — is not yet supported because Theorem 5.3 is stated without proof or attribution and its key Assumption 5.4 is not established or supplied with verifiable sufficient conditions. The finite-window route (Theorems 5.1 and 5.2) is properly anchored in the cited literature, but the belief-quantization route needs substantial additional support before the abstract's claim can be accepted.
major comments (3)
- [Section V-C, Theorem 5.3] The abstract's central claim — that the paper presents RL results that rigorously establish convergence to near optimality — rests on Theorem 5.3, but this theorem is stated without proof, without a citation to a published theorem, and without a derivation. The preceding text only says that the Section V-A and IV-A setup applies 'with a significantly more tedious analysis involving ergodicity requirements.' Every other main theorem in the manuscript is a restatement of a named reference (e.g., Theorem 4.1 from [37], Theorem 5.1 from [42], Theorem 5.2 from [41]), so the status of Theorem 5.3 is unclear. The authors should either supply a complete proof, identify the exact published statement it restates, or explicitly say that it is new; if it is new, the proof must be included.
- [Section V-C, Assumption 5.4] Theorem 5.3(a)–(b) requires the controlled belief process {pi_t, U_t} to be asymptotically uniquely ergodic with eta_gamma(B) > 0 for every quantization bin B. This is not a consequence of Assumption 4.1, weak Feller regularity, or filter stability; the paper itself concedes that 'the condition that eta_gamma(B) > 0 requires an analysis tailored for each problem.' The suggested sufficient routes (quantizing via finite past windows, or initializing from the invariant measure of the hidden Markov source) do not cover the general quantized belief-MDP used in Theorem 5.3, and the invariant measure of the hidden source is an object on X, not on P(X). Without a verifiable sufficient condition for eta_gamma(B) > 0, the claimed almost-sure convergence and the bound J*_beta(pi_0, gamma_hat) - J*_beta(pi_0) <= 2 alpha_c / ((1-beta)^2 (1-beta alpha_eta)) bar_L are not established.
- [Section V-C, Remark 5.1(i) and Theorem 5.3(b)] The near-optimality bound for the learned policy assumes not only convergence of the Q-iterates, but also that the policy constructed from the limit Q-values, when applied to the true model, keeps the closed-loop process inside the trained set P_eta for all time. The manuscript only references [42] and [12, Lemma 6 and Corollary 2] for this, and Remark 5.1(i) notes that 'one needs to ensure' the closed-loop condition without stating conditions under which it follows. Since Theorem 5.3(b) is the statement that connects the learned Q-values to near optimality in the original POMDP, this closed-loop condition is load-bearing and should be made an explicit hypothesis or derived from the assumptions.
minor comments (5)
- [Section V-A, after Eq. (27)] The definition 'V*(u) := min_u Q*(s, u)' has a variable mismatch; it should read 'V*(s) := min_u Q*(s, u)'.
- [Section V-B, Assumption 5.3(ii)] Assumption 5.3(ii) refers to 'Assumption 5.1(i)', but Assumption 5.1 has no numbered parts; please renumber or cross-reference precisely.
- [Section IV-A, quantization construction] The construction requires a weight measure pi* with pi*(Z_i) > 0 for every quantization bin; the paper does not discuss how such a measure is chosen or why it exists uniformly as the number of bins M grows. A brief justification would help.
- [Section III-B, footnote 1] Footnote 1 states that the optimality result in Theorem 3.2(ii) 'may only hold for a restrictive class of initial conditions or initializations'; this is too vague. Please specify the class or point to the exact statement in [72].
- [Section I-B, Eq. (8)] The bounded Lipschitz metric is defined using a distance d on the underlying space, but the same symbol d is used later for the metric on P(X) without explicit comment; a sentence noting which metric is used for probability measures would improve readability.
Circularity Check
No circularity: the paper is a review whose learning results are cited from the authors' prior published work, and no claimed bound reduces to its inputs by construction.
full rationale
This is a review/tutorial, not a paper fitting parameters to data. The main theorems are explicitly attributed: Theorem 2.1 to [23], Theorem 4.1 to [37], Theorem 5.1 to [42], Theorem 5.2 to [41], and Theorem 4.4 to [15]. These prior results are peer-reviewed, state explicit assumptions (e.g., Assumptions 4.1 and 5.2), and do not build the target near-optimality claim into their hypotheses, so under the review rule they are real evidence rather than circular self-citation. The strongest new-looking claim is Theorem 5.3 in Section V-C, which is stated without proof: the text says the quantized Q-learning setup applies 'with a significantly more tedious analysis involving ergodicity requirements,' and Assumption 5.4, requiring eta_gamma(B) > 0 on every quantization bin, is conceded to 'require an analysis tailored for each problem.' Remark 5.1 also defers to [42] and [12] for the condition that the closed loop under the learned policy remains in the trained set. These are omitted-proof and unverified-hypothesis concerns about the abstract's 'rigorously establish' wording; they are not circularity. The Theorem 5.3 bound 2 alpha_c / ((1-beta)^2 (1-beta alpha_eta)) bar_L has the same form as the quantization bound in Theorem 4.1, but that is a mathematical implication, not a definitional identity or a fitted parameter disguised as a prediction. No quantity is defined in terms of the quantity it is claimed to predict. Therefore no circular step can be exhibited, and the circularity score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption POMDP model: standard Borel spaces, transition kernel T, observation kernel Q, bounded continuous cost, and admissible measurable policies.
- domain assumption Either Assumption 2.1 or 2.2, i.e. weak continuity or total-variation continuity of the transition and observation kernels.
- domain assumption Assumption 2.3 with K2 = alpha D (3 - 2 delta(Q)) / 2 < 1 for Wasserstein contraction of the belief kernel.
- domain assumption Controlled filter stability: Dobrushin coefficient contraction, one-step stochastic observability, or mixing kernel/Hilbert metric contraction (Assumptions 2.3, 2.7 and 4.2).
- domain assumption Asymptotic unique ergodicity of the relevant state or belief process under exploration, as in Assumptions 5.1-5.4.
Cite this review
Pith. "Pith review of Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning." pith.science (2026). https://pith.science/paper/KZXQGT4O
@misc{pith2026241206735,
author = {Pith},
title = {Pith review of: Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/KZXQGT4O}},
note = {Machine review of arXiv:2412.06735}
}
read the original abstract
In this review/tutorial article, we present recent progress on optimal control of partially observed Markov Decision Processes (POMDPs). We first present regularity and continuity conditions for POMDPs and their belief-MDP reductions, where these constitute weak Feller and Wasserstein regularity and controlled filter stability. These are then utilized to arrive at existence results on optimal policies for both discounted and average cost problems, and regularity of value functions. Then, we study rigorous approximation results involving quantization based finite model approximations as well as finite window approximations under controlled filter stability. Finally, we present several recent reinforcement learning theoretic results which rigorously establish convergence to near optimality under both criteria.
Forward citations
Cited by 1 Pith paper
-
Sensitivity of Filter Kernels and Robustness Bounds to Transition and Measurement Kernel Perturbations in Partially Observable Stochastic Control
Explicit upper bounds on filter-kernel sensitivity and POMDP value differences in terms of Wasserstein and total-variation distances between transition and observation kernels, with decaying error bounds for joint sta...
Reference graph
Works this paper leans on
-
[37]
A.D Kara, N. Saldi, and S. Y¨ uksel. Q-learning for MDPs w ith general spaces: Convergence and near optimality via quantization u nder weak continuity. Journal of Machine Learning Research , pages 1–34, 2023
work page 2023
-
[42]
A.D. Kara and S. Y¨ uksel. Q-learning for stochastic con trol under gen- eral information structures and non-Markovian environmen ts. Trans- actions on Machine Learning Research (arXiv:2311.00123) , 2024
work page Pith review arXiv 2024
- [41]
-
[1]
R. J. Aumann. Mixed and behavior strategies in infinite ex tensive games. Technical report, Princeton University NJ, 1961
work page 1961
-
[2]
Reinforcement learning of pomdps using spectr al methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Anima shree Anandkumar. Reinforcement learning of pomdps using spectr al methods. In Conference on Learning Theory , pages 193–256. PMLR, 2016
work page 2016
-
[3]
P . Billingsley. Convergence of probability measures. New Y ork: Wiley, 2nd edition, 1999
work page 1999
- [4]
-
[5]
V . S. Borkar. White-noise representations in stochasti c realization theory. SIAM J. on Control and Optimization , 31:1093–1102, 1993
work page 1993
Show all 76 references
-
[6]
V . S. Borkar. Average cost dynamic programming equation s for controlled Markov chains with partial observations. SIAM J. Control Optim., 39(3):673–681, 2000
2000
-
[7]
V . S. Borkar. Convex analytic methods in Markov decision processes. In Handbook of Markov Decision Processes, E. A. Feinberg, A. Shwartz (Eds.) , pages 347–375. Kluwer, Boston, MA, 2001
2001
-
[8]
Budhiraja
A. Budhiraja. On invariant measures of discrete time filt ers in the correlated signal-noise case. The Annals of Applied Probability , 12(3):1096–1113, 2002
2002
-
[9]
Cayci, N
S. Cayci, N. He, and R. Srikant. Finite-time analysis of e ntropy- regularized neural natural actor-critic algorithm. arXiv preprint arXiv:2206.00833, 2022
2022 arXiv
-
[10]
Chandak, V .S
S. Chandak, V .S. Borkar, and P . Dodhia. Reinforcement l earning in non-markovian environments. Systems & Control Letters , 185:105751, 2024
2024
-
[11]
Chigansky and R
P . Chigansky and R. V an Handel. A complete solution to Bl ackwell’s unique ergodicity problem for hidden Markov chains. The Annals of Applied Probability, 20(6):2318–2345, 2010
2010
-
[12]
Cregg, T
L. Cregg, T. Linder, and S. Y¨ uksel. Reinforcement lear ning for near-optimal design of zero-delay codes for markov sources . IEEE Transactions on Information Theory, arXiv:2311.12609 , 2024
2024 arXiv
-
[13]
Crisan and A
D. Crisan and A. Doucet. A survey of convergence results on particle filtering methods for practitioners. IEEE Transactions on Signal Processing, 50(3):736–746, 2002
2002
-
[14]
Demirci, A.D
Y .E. Demirci, A.D. Kara, and S. Y¨ uksel. Average cost op timality of partially observed mdps: Contraction of non-linear filters and existence of optimal solutions. SIAM Journal on Control and Optimization, arXiv:2312.14111, 2024
2024 arXiv
-
[15]
Demirci, A.D
Y .E. Demirci, A.D. Kara, and S. Y¨ uksel. Refined bounds o n near optimality finite window policies in pomdps and their reinfo rcement learning. arXiv, 2024
2024
-
[16]
Demirci and S
Y .E. Demirci and S. Y¨ uksel. Geometric ergodicity and w asserstein continuity of non-linear filters. arXiv preprint arXiv:2307.15764 , 2023
2023 arXiv
-
[17]
Devroye and L
L. Devroye and L. Gy¨ orfi. Non-parametric Density Estimation: The L1 View. John Wiley, New Y ork, 1985
1985
-
[18]
Dobrushin
R.L. Dobrushin. Central limit theorem for nonstationa ry Markov chains. i. Theory of Probability & Its Applications , 1(1):65–80, 1956
1956
-
[19]
S. Dong, B. van Roy, and Z. Zhou. Simple agent, complex en viron- ment: Efficient reinforcement learning with agent states. The Journal of Machine Learning Research , 23(1):11627–11680, 2022
2022
-
[20]
R. M. Dudley. Real Analysis and Probability . Cambridge University Press, Cambridge, 2nd edition, 2002
2002
-
[21]
Reinf orcement learning in pomdps without resets
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour. Reinf orcement learning in pomdps without resets. In IJCAI, pages 690–695, 2005
2005
-
[22]
Feinberg and P .O
E.A. Feinberg and P .O. Kasyanov. Equivalent condition s for weak continuity of nonlinear filters. Systems & Control Letters , 173:105458, 2023
2023
-
[23]
Feinberg, P .O
E.A. Feinberg, P .O. Kasyanov, and M.Z. Zgurovsky. Part ially observ- able total-cost Markov decision process with weakly contin uous tran- sition probabilities. Mathematics of Operations Research , 41(2):656– 681, 2016
2016
-
[24]
Feinberg, P .O
E.A. Feinberg, P .O. Kasyanov, and M.Z. Zgurovsky. Mark ov deci- sion processes with incomplete information and semiunifor m feller transition probabilities. SIAM Journal on Control and Optimization , 60(4):2488–2513, 2022
2022
-
[25]
Fern´ andez-Gaucherand, A
E. Fern´ andez-Gaucherand, A. Arapostathis, and S. I. M arcus. Remarks on the existence of solutions to the average cost optimality equation in markov decision processes. Systems & control letters , 15(5):425–432, 1990
1990
-
[26]
F.Kochman and J. Reeds. A simple proof of kaijser’s uniq ue ergodicity result for hidden markov α -chains. The Annals of Applied Probability , 16(4):1805–1815, 2006
2006
-
[27]
I. I. Gihman and A. V . Skorohod. Controlled Stochastic Processes . Springer Science & Business Media, 2012
2012
-
[28]
Le Gland and N
F. Le Gland and N. Oudjane. Stability and uniform approx imation of nonlinear filters using the Hilbert metric and application t o particle filters. The Annals of Applied Probability , 14(1):144–187, 2004
2004
-
[29]
Hern´ andez-Lerma
O. Hern´ andez-Lerma. Adaptive Markov Control Processes . Springer- V erlag, 1989
1989
-
[30]
Hern´ andez-Lerma and J
O. Hern´ andez-Lerma and J. B. Lasserre. Discrete-Time Markov Control Processes: Basic Optimality Criteria . Springer, 1996
1996
-
[31]
Hsu, D.-M Chuang, and A
S.-P . Hsu, D.-M Chuang, and A. Arapostathis. On the exis tence of stationary optimal policies for partially observed mdps un der the long- run average cost criterion. Systems & control letters , 55(2):165–173, 2006
2006
-
[32]
Sample-efficient reinforcement learning of undercomplete pomdps
Chi Jin, Sham Kakade, Akshay Krishnamurthy, and Qinghu a Liu. Sample-efficient reinforcement learning of undercomplete pomdps. Advances in Neural Information Processing Systems , 33:18530–18539, 2020
2020
-
[33]
T. Kaijser. A limit theorem for partially observed Mark ov chains. Annals of Probability , 4(677), 1975
1975
-
[34]
T. Kaijser. On markov chains induced by partitioned tra nsition proba- bility matrices. Acta Mathematica Sinica, English Series , 27(3):441– 476, 2011
2011
-
[35]
A. D. Kara, M. Raginsky, and S. Y¨ uksel. Robustness to in correct models and data-driven learning in average-cost optimal st ochastic control. Automatica, 139:110179, 2022
2022
-
[36]
Saldi, and S
A.D Kara, N. Saldi, and S. Y¨ uksel. Weak Feller property of non-linear filters. Systems & Control Letters , 134:104–512, 2019
2019
-
[38]
Y¨ uksel
A.D Kara and S. Y¨ uksel. Robustness to incorrect priors in partially ob- served stochastic control. SIAM Journal on Control and Optimization , 57(3):1929–1964, 2019
1929
-
[39]
Y¨ uksel
A.D Kara and S. Y¨ uksel. Robustness to incorrect system models in stochastic control. SIAM Journal on Control and Optimization , 58(2):1144–1182, 2020
2020
-
[40]
Y¨ uksel
A.D Kara and S. Y¨ uksel. Near optimality of finite memory feedback policies in partially observed markov decision processes. Journal of Machine Learning Research , 23(11):1–46, 2022
2022
-
[43]
Kara and S
Ali D. Kara and S. Y¨ uksel. Robustness to approximation s and model learning in MDPs and POMDPs. In A. B. Piunovskiy and Y . Zhang, editors, Modern Trends in Controlled Stochastic Processes: Theory and Applications, V olume III . Luniver Press, 2021
2021
-
[44]
Rl for latent mdps: Regret guarantees and a lower bou nd
Jeongyeol Kwon, Y onathan Efroni, Constantine Caraman is, and Shie Mannor. Rl for latent mdps: Regret guarantees and a lower bou nd. Advances in Neural Information Processing Systems , 34:24523–24534, 2021
2021
-
[45]
H.J. Langen. Convergence of dynamic programming model s. Mathe- matics of Operations Research , 6(4):493–512, Nov. 1981
1981
-
[46]
Di Masi and L
G. Di Masi and L. Stettner. Ergodicity of hidden markov m odels. Mathematics of Control, Signals and Systems , 17(4):269–296, 2005
2005
-
[47]
McDonald and S
C. McDonald and S. Y¨ uksel. Exponential filter stabilit y via Do- brushin’s coefficient. Electronic Communications in Probability , 25, 2020
2020
-
[48]
McDonald and S
C. McDonald and S. Y¨ uksel. Robustness to incorrect pri ors and controlled filter stability in partially observed stochast ic control. SIAM Journal on Control and Optimization , 60(2):842–870, 2022
2022
-
[49]
McDonald and S
C. McDonald and S. Y¨ uksel. Stochastic observability a nd filter stability under several criteria. IEEE Transactions on Automatic Control, 69(5):2931–2946, 2024
2024
-
[50]
Parthasarathy
K.R. Parthasarathy. Probability Measures on Metric Spaces . AMS Bookstore, 1967
1967
-
[51]
L. K. Platzman. Optimal infinite-horizon undiscounted control of finite probabilistic systems. SIAM Journal on Control and Optimization , 18(4):362–380, 1980
1980
-
[52]
D. Rhenius. Incomplete information in Markovian decis ion models. Ann. Statist. , 2:1327–1334, 1974
1974
-
[53]
Approximations of discrete time partially observed control problems
Wolfgang J Runggaldier and Lukasz Stettner. Approximations of discrete time partially observed control problems . Giardini Pisa, 1994
1994
-
[54]
Saldi, T
N. Saldi, T. Linder, and S. Y¨ uksel. Finite Approximations in Discrete- Time Stochastic Control: Quantized Models and Asymptotic O ptimal- ity. Springer, Cham, 2018
2018
-
[55]
Saldi, S
N. Saldi, S. Y¨ uksel, and T. Linder. Near optimality of q uantized poli- cies in stochastic control under weak continuity condition s. Journal of Mathematical Analysis and Applications , 435(1):321–337, 2016
2016
-
[56]
Saldi, S
N. Saldi, S. Y¨ uksel, and T. Linder. On the asymptotic op timality of finite approximations to Markov decision processes with Bor el spaces. Mathematics of Operations Research , 42(4):945–978, 2017
2017
-
[57]
Saldi, S
N. Saldi, S. Y¨ uksel, and T. Linder. Finite model approx imations for partially observed Markov decision processes with discoun ted cost. IEEE Transactions on Automatic Control , 65, 2020
2020
-
[58]
R. Serfozo. Convergence of Lebesgue integrals with var ying measures. Sankhy¯ a: The Indian Journal of Statistics, Series A , pages 380–402, 1982
1982
-
[59]
Seyedsalehi, N
E. Seyedsalehi, N. Akbarzadeh, A. Sinha, and A. Mahajan . Approx- imate information state based convergence analysis of recu rrent q- learning. arXiv preprint arXiv:2306.05991 , 2023
2023 arXiv
-
[60]
Singh, T
S.P . Singh, T. Jaakkola, and M.I. Jordan. Learning with out state- estimation in partially observable markovian decision pro cesses. In Machine Learning Proceedings 1994 , pages 284–292. Elsevier, 1994
1994
-
[61]
Sinha and A
A. Sinha and A. Mahajan. Agent-state based policies in p omdps: Beyond belief-state mdps. In IEEE Conference on Decision and Control, Tutorial Paper , 2024
2024
-
[62]
Stettner
L. Stettner. Long run control with degenerate observat ion. SIAM Journal on Control and Optimization , 57(2):880–899, 2019
2019
-
[63]
Subramanian, A
J. Subramanian, A. Sinha, R. Seraj, and A. Mahajan. Appr oximate information state for approximate planning and reinforcem ent learning in partially observed systems. The Journal of Machine Learning Research, 23(1):483–565, 2022
2022
-
[64]
T. Szarek. Feller processes on nonlocally compact spac es. Annals of Probability, pages 1849–1863, 2006
2006
-
[65]
Budhiraja V
A. Budhiraja V . S. Borkar. A further remark on dynamic pr ogramming for partially observed markov processes. Stochastic processes and their applications, 112(1):79–93, 2004
2004
-
[66]
van Handel
R. van Handel. Observability and nonlinear filtering. Probability theory and related fields , 145(1-2):35–74, 2009
2009
-
[67]
van Handel
R. van Handel. Uniform time average consistency of Mont e Carlo particle filters. Stochastic Processes and their Applications , 119(11):3835–3861, 2009
2009
-
[68]
van Handel
R. van Handel. Nonlinear filtering and systems theory. I n Proceedings of the 19th International Symposium on Mathematical Theory of Networks and Systems (MTNS semi-plenary paper) , 2010
2010
-
[69]
C. Villani. Optimal Transport: Old and New . Springer, 2008
2008
-
[70]
Su blinear regret for learning pomdps
Yi Xiong, Ningyuan Chen, Xuefeng Gao, and Xiang Zhou. Su blinear regret for learning pomdps. Production and Operations Management , 31(9):3491–3504, 2022
2022
-
[71]
Y u and D
H. Y u and D. P . Bertsekas. On near optimality of the set of finite- state controllers for average cost pomdp. Mathematics of Operations Research, 33(1):1–11, 2008
2008
-
[72]
Y¨ uksel
S. Y¨ uksel. Another look at partially observed optimal stochastic control: Existence, ergodicity, and approximations witho ut belief- reduction. arXiv, 2023
2023
-
[73]
Y¨ uksel
S. Y¨ uksel. Optimization and Control of Stochastic Systems . Queen’s University, Lecture Notes, available online, Queen’s Univ ersity, Lec- ture notes
-
[74]
Y¨ uksel and T
S. Y¨ uksel and T. Bas ¸ar. Stochastic Teams, Games, and Control under Information Constraints . Springer, Cham, 2024
2024
-
[75]
Y¨ uksel and T
S. Y¨ uksel and T. Linder. Optimization and convergence of observation channels in stochastic control. SIAM J. on Control and Optimization , 50:864–887, 2012
2012
-
[76]
Y ushkevich
A.A. Y ushkevich. Reduction of a controlled Markov mode l with incomplete data to a problem with complete information in th e case of Borel state and control spaces. Theory Prob. Appl. , 21:153–158, 1976
1976
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.