{"id":"f1902457-0275-41ce-abab-d28e5727a2ae","arxiv_id":"2412.06735","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of regularity, approximation, and reinforcement learning guarantees for partially observed Markov decision processes, drawing mostly on the authors' earlier work.","lead":"This review consolidates recent mathematical results on controlling systems whose true state is hidden and only noisy observations are available. It maps the conditions under which optimal policies exist, finite approximations are accurate, and Q-learning converges to near-optimal performance.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 5.3 is stated without proof or citation; its key hypothesis Assumption 5.4 (positive mass on every bin under the exploration invariant measure) is not established for the quantized belief process, so the claimed Q-learning convergence to near optimality is not yet supported.","rationale":"The reader's verdict is CONDITIONAL with Assumption 5.4 identified as the weakest assumption. I agree: that is the single most load-bearing concern because Theorem 5.3 is the quantitative justification for the abstract's claim of rigorous RL convergence. The concern is not that unique ergodicity is 'hard to verify' in an aesthetic sense; it is that the paper does not supply a proof or an explicit verifiable condition connecting Assumption 5.4 to the stated hypotheses of Theorem 5.3. The theorem is presented without proof and without a citation for the result itself; the text only gestures at 'a significantly more tedious analysis.' The listed sufficient conditions (Assumption 5.5, initialization from the invariant measure, positive Harris recurrence) are either themselves strong, or not established for the quantized belief state with the quantizer used in the theorem. Additional internal evidence supports this: Remark 5.1(i) admits that 'the initialization cannot be arbitrary' and that one must ensure the closed-loop process remains in the trained set, referencing other papers for this. The paper's own limitation statements therefore corroborate the concern. This is not an outside-consensus disagreement; it is a supportability gap relative to the paper's own strongest claim. I do not see a reason to move the verdict to REJECT, because the surrounding survey content is well-structured and much of the learning analysis is built on published work (Theorems 5.1, 5.2 are attributed), but the unproved Theorem 5.3 warrants the CONDITIONAL status. The concrete test proposed—instantiating the theorem in a minimal binary POMDP and checking whether Assumption 5.4 holds for a natural quantizer—would directly determine whether the missing analysis is a genuine gap or merely an omitted detail.","tokens_in":22417,"tokens_out":2120,"duration_ms":18940,"concrete_test":"Attempt to instantiate Theorem 5.3 in a minimal POMDP that satisfies the paper's weak Feller and contraction assumptions: e.g., binary state X={0,1}, binary observation Y with Q(y|x) continuous in total variation, and a control-dependent transition T(·|x,u) satisfying Assumption 2.3 with K2<1. Choose a quantizer g of P({0,1}) with a bin around a belief that is not reachable in the support of the exploration invariant measure. Check numerically and analytically whether Assumption 5.4 holds (i.e., whether eta_gamma(B)>0 for every bin) under (i) i.i.d. exploration and (ii) epsilon-greedy exploration. If some bin has zero invariant mass, then Theorem 5.3(b) cannot apply for that quantizer, and the theorem needs either a proof that Assumption 5.4 is implied by the stated assumptions or an explicit construction of a quantizer for which it holds.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract's central claim is that the paper presents RL results rigorously establishing convergence to near optimality. For the belief-quantization route, Theorem 5.3 states a.s. convergence of Q-iterates and the near-optimality bound 2 alpha_c /((1-beta)^2 (1-beta alpha_eta)) bar_L, but the statement is asserted without a proof or a pointer to a specific published theorem; Section V-C only says the setup of Sections V-A and IV-A applies 'with a significantly more tedious analysis involving ergodicity requirements.' The most fragile premise is Assumption 5.4: under the exploration policy, the joint process {pi_t, U_t} is asymptotically uniquely ergodic with eta_gamma(B) > 0 for every quantization bin B. This is a substantive property of the controlled filter dynamics, not a consequence of the weak Feller regularity or of Assumption 4.1, and the paper itself concedes the eta_gamma(B) > 0 condition 'requires an analysis tailored for each problem.' The suggested routes (quantizing via finite past windows, or initializing from the invariant measure of the HMM) do not cover the general belief-MDP reduction used in Theorem 5.3; moreover part (b) requires the learned policy's closed loop to remain inside the trained set, a condition only referenced to other work. Since the theorem is the basis for the abstract's strongest claim, the absence of a proof or a verifiable sufficient condition for Assumption 5.4 is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a survey/tutorial on optimal control of partially observed Markov decision processes (POMDPs), organized around the belief-MDP reduction. It reviews weak Feller and Wasserstein regularity of the controlled filter, filter stability, existence of optimal policies for discounted and average cost criteria, quantization and finite-window approximations with explicit error bounds, robustness of optimal costs to model and prior errors, and reinforcement learning results. The central advertised claim is that recent RL results rigorously establish convergence to near optimality under both discounted and average cost criteria; for discounted cost, the paper presents two routes: finite-window memory under uniform geometric filter stability, and quantized belief states under asymptotic unique ergodicity, culminating in Theorem 5.3.","tokens_in":22682,"tokens_out":5575,"duration_ms":56157,"significance":"If fully supported, the paper would be a useful unifying review: it collects regularity conditions, existence theorems, and approximation bounds with explicit constants, and it identifies which filter-stability assumptions are needed for learning. The paper is generally careful to attribute results and to flag limitations, e.g., footnote 1 on Theorem 3.2(ii) and Remark 5.1. However, the strongest advertised contribution — RL convergence to near optimality for belief-quantized POMDPs — is not yet supported because Theorem 5.3 is stated without proof or attribution and its key Assumption 5.4 is not established or supplied with verifiable sufficient conditions. The finite-window route (Theorems 5.1 and 5.2) is properly anchored in the cited literature, but the belief-quantization route needs substantial additional support before the abstract's claim can be accepted.","major_comments":[{"comment":"The abstract's central claim — that the paper presents RL results that rigorously establish convergence to near optimality — rests on Theorem 5.3, but this theorem is stated without proof, without a citation to a published theorem, and without a derivation. The preceding text only says that the Section V-A and IV-A setup applies 'with a significantly more tedious analysis involving ergodicity requirements.' Every other main theorem in the manuscript is a restatement of a named reference (e.g., Theorem 4.1 from [37], Theorem 5.1 from [42], Theorem 5.2 from [41]), so the status of Theorem 5.3 is unclear. The authors should either supply a complete proof, identify the exact published statement it restates, or explicitly say that it is new; if it is new, the proof must be included.","section":"Section V-C, Theorem 5.3"},{"comment":"Theorem 5.3(a)–(b) requires the controlled belief process {pi_t, U_t} to be asymptotically uniquely ergodic with eta_gamma(B) > 0 for every quantization bin B. This is not a consequence of Assumption 4.1, weak Feller regularity, or filter stability; the paper itself concedes that 'the condition that eta_gamma(B) > 0 requires an analysis tailored for each problem.' The suggested sufficient routes (quantizing via finite past windows, or initializing from the invariant measure of the hidden Markov source) do not cover the general quantized belief-MDP used in Theorem 5.3, and the invariant measure of the hidden source is an object on X, not on P(X). Without a verifiable sufficient condition for eta_gamma(B) > 0, the claimed almost-sure convergence and the bound J*_beta(pi_0, gamma_hat) - J*_beta(pi_0) <= 2 alpha_c / ((1-beta)^2 (1-beta alpha_eta)) bar_L are not established.","section":"Section V-C, Assumption 5.4"},{"comment":"The near-optimality bound for the learned policy assumes not only convergence of the Q-iterates, but also that the policy constructed from the limit Q-values, when applied to the true model, keeps the closed-loop process inside the trained set P_eta for all time. The manuscript only references [42] and [12, Lemma 6 and Corollary 2] for this, and Remark 5.1(i) notes that 'one needs to ensure' the closed-loop condition without stating conditions under which it follows. Since Theorem 5.3(b) is the statement that connects the learned Q-values to near optimality in the original POMDP, this closed-loop condition is load-bearing and should be made an explicit hypothesis or derived from the assumptions.","section":"Section V-C, Remark 5.1(i) and Theorem 5.3(b)"}],"minor_comments":[{"comment":"The definition 'V*(u) := min_u Q*(s, u)' has a variable mismatch; it should read 'V*(s) := min_u Q*(s, u)'.","section":"Section V-A, after Eq. (27)"},{"comment":"Assumption 5.3(ii) refers to 'Assumption 5.1(i)', but Assumption 5.1 has no numbered parts; please renumber or cross-reference precisely.","section":"Section V-B, Assumption 5.3(ii)"},{"comment":"The construction requires a weight measure pi* with pi*(Z_i) > 0 for every quantization bin; the paper does not discuss how such a measure is chosen or why it exists uniformly as the number of bins M grows. A brief justification would help.","section":"Section IV-A, quantization construction"},{"comment":"Footnote 1 states that the optimality result in Theorem 3.2(ii) 'may only hold for a restrictive class of initial conditions or initializations'; this is too vague. Please specify the class or point to the exact statement in [72].","section":"Section III-B, footnote 1"},{"comment":"The bounded Lipschitz metric is defined using a distance d on the underlying space, but the same symbol d is used later for the metric on P(X) without explicit comment; a sentence noting which metric is used for probability measures would improve readability.","section":"Section I-B, Eq. (8)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a review/tutorial, and most technical theorems are restatements of prior published papers, including many by the authors. The self-citation rate is high, but there are independent anchors (Feinberg et al., Le Gland and Oudjane, van Handel, Stettner, Borkar), so I do not see a circularity problem. The more serious issue is that the one potentially new-looking theorem (5.3) is not anchored; if the authors can either prove it or map it to a published theorem, the paper is likely acceptable. If not, the abstract's RL claim should be weakened to match what is actually proved. Given the journal's scope, a tutorial with an unsupported central theorem is not acceptable in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on 2412.06735. It's a review/tutorial of the Yuksel–Kara line on POMDPs: regularity of the belief MDP, existence results, quantization and finite-window approximations, and Q-learning with near-optimality guarantees. The consolidation is genuinely useful: the theorems are stated cleanly, the conditions are laid out, and the paper is honest about gaps in the cited results—see the footnote on Theorem 3.2(ii) and Remark 5.1. If you want the statements of the main results on weak Feller continuity, Wasserstein contraction, and controlled filter stability in one place, this does that.\n\nThe soft spot is Theorem 5.3. In a survey that otherwise attributes every numbered theorem to a prior publication, 5.3 is stated without proof or citation, and it's the load-bearing result for the abstract's claim about RL convergence to near optimality. The key assumption, Assumption 5.4, asks that the controlled belief process under exploration be asymptotically uniquely ergodic with positive invariant mass on every quantization bin. That's a real condition on the filter dynamics, not a consequence of weak Feller regularity. The paper concedes that eta_gamma(B)>0 'requires an analysis tailored for each problem,' and the suggested routes (finite past windows, initializing from the HMM invariant measure) do not cover the general belief-MDP reduction used in the theorem. Part (b) also requires the learned policy to keep the closed loop inside the trained set, which is only referenced elsewhere. So the central learning guarantee is not yet supported as stated.\n\nThe self-citation density is high, but it's largely legitimate: these are published results, many with independent anchors (Feinberg, Le Gland, van Handel). The issue is not provenance, it's that one new-looking statement sits in a review without the same rigorous treatment.\n\nWho should read it: someone who wants a structured map of the authors' POMDP regularity/approximation results and the conditions under which they hold. It's less useful as a source for the learning claims until 5.3 gets fixed.\n\nRecommendation: send to review as a survey, but the referee should require the authors to either prove Theorem 5.3, cite a published proof, or clearly demote it to a conjecture with a precise statement of Assumption 5.4's status. That seems like a standard, fair bar for a tutorial.","headline":"A useful consolidation of the authors' own POMDP results, but the one new-looking theorem (Thm 5.3) that anchors the learning claim is stated without proof or citation, so read the survey as a map, not as a new result.","tokens_in":23258,"tokens_out":2338,"would_cite":false,"duration_ms":22618,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93E20","90C40","93E35"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that, under regularity of the belief filter (weak Feller or Wasserstein continuity plus filter stability), finite-model and finite-window approximations are near-optimal for partially observed MDPs, and that…","keywords":["partially observed Markov decision processes","belief-MDP","filter stability","Wasserstein contraction","quantization approximation","Q-learning","near-optimality","average cost"],"falsifier":"A concrete calculation that would test the reach of the claim: construct a POMDP with finite state and output spaces whose filter under a memoryless randomized exploration policy has an absorbing belief or a periodic cycle, so that the invariant measure has zero mass on some quantization bin. In that case the Q-learning iterates for that bin never update, and no near-optimality bound of the form in Theorem 5.3 can be obtained; such an example would demarcate exactly which problems the learning theorem covers.","tokens_in":22162,"feed_emoji":"","tokens_out":9784,"duration_ms":97498,"temperature":0.7,"pith_summary":"This paper surveys and consolidates a program of results whose common claim is that partially observed Markov decision processes can be controlled and learned near-optimally whenever their nonlinear filter is regular: weakly Feller or Wasserstein-contractive, and stable with respect to initialization errors. Under those properties, optimal policies exist for discounted and average cost, quantized finite models and finite-window memory policies approximate the true problem with explicit error bounds, and Q-learning on either quantized beliefs or finite observation windows converges almost surely with a stated suboptimality gap. The paper's headline learning result is Theorem 5.3: under asymptotic unique ergodicity of the belief-action process with positive invariant mass on every quantization bin, the policy built from the limit Q-values is within $2\\alpha_c/((1-\\beta)^2(1-\\beta\\alpha_\\eta)) \\bar L$ of optimal for discounted cost. The same set of ideas transfers the guarantee to average cost through the vanishing discount method, so the paper's overall thesis is that near-optimal learning for POMDPs is a theorem under checkable conditions rather than a heuristic.","feed_headline":"Proofs now cover near-optimal Q-learning for POMDPs","feed_subtitle":"Regularity of the nonlinear filter plus quantization yields explicit error bounds for discounted and average cost.","key_machinery":"The machinery is the belief-MDP reduction together with three regularity doctrines. The first is weak Feller continuity of the filter kernel $\\eta(\\cdot|\\pi,u)$, established under two alternative assumption sets (Theorems 2.1 and 2.2); it makes the belief space a well-behaved MDP and supports existence and asymptotic optimality. The second is Wasserstein contraction, $W_1(\\eta(\\cdot|z,u), \\eta(\\cdot|z',u)) \\le K_2 W_1(z,z')$ with $K_2 = \\alpha D(3-2\\delta(Q))/2$; this gives Lipschitz value functions, the average-cost optimality equation under $K_2 < 1$, and the explicit constants in the quantization error bound. The third is filter stability: the total-variation distance between filters started from different priors, either in expectation ($L_N^t$, bounded via the Dobrushin coefficient by $2\\alpha^N$) or uniformly ($\\bar L_{TV}^N$, bounded geometrically via Hilbert-metric contraction under Assumption 4.2). Filter stability converts finite-window memory into a near-optimal state, and supplies the empirical-averages-approach-expectations that Q-learning needs. For the learning theorems specifically, the load-bearing identity is the fixed-point equation $Q^*(s,u) = C^*(s,u) + \\beta \\sum_{s_1} V^*(s_1)P^*(s_1|s,u)$, which is the Bellman equation of a limiting approximate MDP; asymptotic ergodicity (Assumptions 5.1, 5.3, 5.4) guarantees that the running averages in the Q-learning update converge to these $C^*$ and $P^*$, so the iterates converge to $Q^*$ almost surely.","core_discovery":"The paper's central claim is that the obstacles to rigorous POMDP control—infinite-dimensional belief space, non-Markovian information, and unknown models—can each be overcome by quantitative regularity of the filter. On the existence side, weak Feller continuity (Theorems 2.1 and 2.2) guarantees an optimal policy for discounted cost, and the Wasserstein contraction with $K_2 = \\alpha D(3-2\\delta(Q))/2 < 1$ yields a solution to the average-cost optimality equation and a constant optimal cost for every initial belief (Theorem 3.2(i)). On the approximation side, quantizing the belief space gives finite models whose value functions deviate from the true one by at most $2K_1/((1-\\beta)^2(1-\\beta K_2)) \\bar L$, while finite-window reductions, under controlled filter stability, give memory-$N$ policies whose error is controlled by the filter-error term $L_N^t$ (Theorems 4.2 and 4.3). On the learning side, the paper claims that Q-learning is not just an algorithm but a theorem: if the exploration policy makes $\\{\\pi_t, U_t\\}$ asymptotically uniquely ergodic with positive mass on every bin, the iterates converge almost surely, and the learned policy satisfies the bound in Theorem 5.3; for the model-free finite-window variant, Theorem 5.2 gives the same type of guarantee using $L_N^t$ under ergodicity of the hidden state. In the average-cost case, these results transfer through the vanishing discount method, so both criteria are covered.","pith_inferences":["Editorial inference: the paper's learning theorem suggests a practical design rule—choose exploration policies that are randomized with full-support observation channels so that all bins in the trained set are visited; checking Assumption 5.4 for a specific channel is the main bottleneck to turning the theorem into an algorithm.","Editorial inference: the same proof architecture should extend to sample-based quantization of beliefs (for example, particle-filter outputs), but the convergence would then require controlling both quantization error and particle approximation error simultaneously, which the paper does not do.","Editorial inference: the average-cost transfer via discounted near-optimality implies that any finite-time regret bound for the discounted Q-learning analysis would automatically yield an average-cost regret bound up to a $\\beta$-dependent factor; the paper does not compute such finite-time rates."],"forward_implications":["Under the Wasserstein contraction condition $K_2 < 1$, any quantized approximation of the belief MDP is provably near-optimal, with an error proportional to the largest diameter of the quantization bins; refining the partition shrinks the gap at a known rate.","Finite-window policies are near-optimal whenever the filter is stable: the expected loss is bounded by a discounted sum of the filter-error terms $L_N^t$, and under geometric filter stability the bound decays geometrically in the window length $N$.","Q-learning with quantized beliefs converges almost surely and the learned policy's suboptimality is bounded by an explicit function of the quantization level, provided the exploration policy makes the belief process visit every bin infinitely often.","Learning under the average-cost criterion is reduced to learning under discounted cost: with $K_2 < 1$, a near-optimal discounted policy is near-optimal for the average cost, so the same Q-learning guarantees carry over.","When only weak Feller continuity holds, asymptotic (rate-free) near-optimality of the learned policies still follows as the quantization becomes infinitely fine, even without a uniform filter stability estimate."],"supporting_citations":[{"why":"Establishes weak Feller continuity of the filter kernel under weak continuity of the transition and total-variation continuity of the channel; the basis for Theorem 2.1 and for existence of optimal policies.","marker":"[23]"},{"why":"Provides the complementary weak Feller result (Theorem 2.2) under total-variation continuity of the transition and a control-independent channel.","marker":"[36]"},{"why":"Derives the Wasserstein contraction and near-optimality of finite-memory policies; supplies Assumption 2.3/4.1 constants and the uniform bound in (19).","marker":"[40]"},{"why":"Solves the average-cost optimality equation under the contraction condition and gives the vanishing-discount transfer used in Section VI.","marker":"[14]"},{"why":"Proves finite-memory Q-learning convergence and the near-optimality bounds Theorems 4.2, 4.3, and 5.2 in terms of the filter-error term $L_N^t$.","marker":"[41]"},{"why":"Establishes Q-learning convergence for non-Markov environments generally, giving Theorem 5.1 and the asymptotic ergodicity framework of Assumptions 5.1-5.2.","marker":"[42]"},{"why":"Proves Q-learning convergence and near-optimality via quantization for weakly continuous MDPs, yielding Assumption 4.1 and Theorem 4.1.","marker":"[37]"},{"why":"Develops quantization-based finite model approximations for POMDPs under discounted cost, supplying the quantization scheme and asymptotic results.","marker":"[57]"},{"why":"Gives controlled filter stability and robustness-to-prior results, supporting stochastic observability-based conditions such as Assumption 5.5 and Theorem 2.5.","marker":"[48]"},{"why":"Provides the exponential filter stability bound via the Dobrushin coefficient that controls $L_N^t$ in (20) and Theorem 2.4.","marker":"[47]"}],"fun_headline_variants":["Rigorous Q-learning guarantees for partially observed MDPs","Near-optimality proven for POMDP control and learning","Rigorous proofs for near-optimal policies in POMDPs","Error bounds and Q-learning convergence for POMDPs","Theoretical guarantees for POMDP learning and control"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that during exploration the observer's posterior-probability state visits every cell of the quantized state space infinitely often with positive long-run frequency; the paper itself notes this must be verified problem by problem.","fun_headline_variants_meta":{"raw":{"variants":["Rigorous Q-learning guarantees for partially observed MDPs","Near-optimality proven for POMDP control and learning","Rigorous proofs for near-optimal policies in POMDPs","Error bounds and Q-learning convergence for POMDPs","Theoretical guarantees for POMDP learning and control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3288,"prompt_tokens":995,"completion_tokens":2293,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":2208}},"tokens_in":611,"tokens_out":2293,"duration_ms":16263,"temperature":1.0,"reasoning_tokens":2208,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:19:39.509844+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete calculation that would test the reach of the claim: construct a POMDP with finite state and output spaces whose filter under a memoryless randomized exploration policy has an absorbing belief or a periodic cycle, so that the invariant measure has zero mass on some quantization bin. In that case the Q-learning iterates for that bin never update, and no near-optimality bound of the form in Theorem 5.3 can be obtained; such an example would demarcate exactly which problems the learning theorem covers.","supporting_citations":[{"cited_title":"Feinberg, P .O","cited_arxiv_id":null,"evidence_quote":"Establishes weak Feller continuity of the filter kernel under weak continuity of the transition and total-variation continuity of the channel; the basis for Theorem 2.1 and for existence of optimal policies."},{"cited_title":"Saldi, and S","cited_arxiv_id":null,"evidence_quote":"Provides the complementary weak Feller result (Theorem 2.2) under total-variation continuity of the transition and a control-independent channel."},{"cited_title":"Y¨ uksel","cited_arxiv_id":null,"evidence_quote":"Derives the Wasserstein contraction and near-optimality of finite-memory policies; supplies Assumption 2.3/4.1 constants and the uniform bound in (19)."},{"cited_title":"Average Cost Optimality of Partially Observed MDPS: Contraction of Non-linear Filters, Optimal Solutions and Approximations","cited_arxiv_id":"2312.14111","evidence_quote":"Solves the average-cost optimality equation under the contraction condition and gives the vanishing-discount transfer used in Section VI."},{"cited_title":"Y¨ uksel","cited_arxiv_id":null,"evidence_quote":"Proves finite-memory Q-learning convergence and the near-optimality bounds Theorems 4.2, 4.3, and 5.2 in terms of the filter-error term $L_N^t$."},{"cited_title":"Q-Learning for Stochastic Control under General Information Structures and Non-Markovian Environments","cited_arxiv_id":"2311.00123","evidence_quote":"Establishes Q-learning convergence for non-Markov environments generally, giving Theorem 5.1 and the asymptotic ergodicity framework of Assumptions 5.1-5.2."},{"cited_title":"Saldi, and S","cited_arxiv_id":null,"evidence_quote":"Proves Q-learning convergence and near-optimality via quantization for weakly continuous MDPs, yielding Assumption 4.1 and Theorem 4.1."},{"cited_title":"Saldi, S","cited_arxiv_id":null,"evidence_quote":"Develops quantization-based finite model approximations for POMDPs under discounted cost, supplying the quantization scheme and asymptotic results."},{"cited_title":"McDonald and S","cited_arxiv_id":null,"evidence_quote":"Gives controlled filter stability and robustness-to-prior results, supporting stochastic observability-based conditions such as Assumption 5.5 and Theorem 2.5."}],"review_version":1}