REVIEW 4 major objections 3 minor 24 references
Risk-averse formulations of Stochastic Optimal Control and Markov Decision Processes
T0 review · 4 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper proves a finite-sample guarantee for Value-at-Risk-averse Bellman equations: with finite-support noise, empirical fixed points are ε-accurate with sample size scaling like κ_α^{-2}ε^{-2}(1-β)^{-1} up to logarithms.
desk verdict A useful synthesis of risk-averse SOC/MDP with a genuinely new static sample-complexity bound, but the main infinite-horizon theorem is stated without the βL>1 condition its proof requires. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two pieces carry the argument. First, the nested conditional risk-functional construction: for a law-invariant risk functional $\rho$, one forms the conditional functional $R^P_{|\xi_{[t-1]}}(Z_t)=\rho\big(F^P_{Z_t|\xi_{[t-1]}}\big)$ and composes these maps across stages; in the rectangular case this reduces to applying a per-stage risk functional $R^{P_t}$ to the stage cost plus continuation value, yielding dynamic equations of the form (3.8) for stochastic optimal control and (4.5) for Markov decision processes. Second, the sample-complexity engine: the empirical Bellman operator $T_N$ is compared with $T$ through $\|V-V_N\|_\infty \le (1-\beta)^{-1}\big(\|(T-T_N)\tilde V\|_\infty + 2\|V-\tilde V\|_\infty\big)$, and the term $\|(T-T_N)\tilde V\|_\infty$ is controlled by the 1-Lipschitz property of $\mathrm{V@R}$ in the sup-norm, a covering $\eta$-net of $X\times U$, and the DKW inequality applied at the quantile gap $\kappa_\alpha$.
What would settle it
Run the empirical Bellman fixed point on a finite-support Lipschitz example with $\beta L\le1$ and $\kappa_\alpha>0$; if $\|V-V_N\|_\infty$ exceeds the bound in (3.20) at the prescribed $N$ with probability larger than $\delta$, then Theorem 3.1 as stated is false, since the proof requires $\beta L>1$ through the denominator $\beta L-1$.
Extended reading notes
Core claim
The paper's central discovery is that the risk-averse Bellman fixed point is learnable from data at a concrete rate. For $V(x)=\inf_{u\in U}\mathrm{V@R}^P_\alpha\big(c(x,u,\xi)+\beta V(\Phi(x,u,\xi))\big)$, Theorem 3.1 says the fixed point $V_N$ of the empirical Bellman operator $T_N$ satisfies $\|V-V_N\|_\infty\le\varepsilon$ with probability at least $1-\delta$ whenever $N$ exceeds the bound in (3.20), which is $\kappa_\alpha^{-2}\varepsilon^{-2}$ times a factor linear in $(1-\beta)^{-1}$ up to logarithms, under compactness, Lipschitz continuity, and finite support of $P$. The proof works by approximating $V$ with a Lipschitz iterate $\tilde V$ and bounding the empirical error $\|(T-T_N)\tilde V\|_\infty$ over an $\eta$-net of $X\times U$; the finite-support quantile gap $\kappa_\alpha$ makes each sample quantile exact with high probability. On the policy side, Proposition 3.1 shows that under Assumption 3.2 (convex ambiguity sets, compact actions, risk functionals concave in the distribution), non-randomized optimal policies exist if and only if saddle points (3.30)--(3.31) exist. Thus randomization is never strictly needed in the robust rectangular setting under these conditions.
Load-bearing premise
The load-bearing premise is that the Lipschitz constants of the cost and dynamics, the discount factor, and the finite-support quantile gap $\kappa_\alpha$ line up so that the proof's Lipschitz-approximation step yields a finite Lipschitz constant and exact empirical quantiles; in particular the theorem as stated does not make the $\beta L>1$ condition explicit that its own denominator $\beta L-1$ requires.
Editorial extensions
If this is right
- For infinite-horizon VaR-averse stochastic optimal control with finite-support noise, sample average approximation of the Bellman equation is justified: with $N$ of order $\kappa_\alpha^{-2}\varepsilon^{-2}(1-\beta)^{-1}\log(1/\delta)$ times dimension factors, the empirical value function is $\varepsilon$-accurate uniformly with probability at least $1-\delta$.
- When the ambiguity sets are convex, the action spaces compact, and the risk functional is concave in the distribution (Value-at-Risk and Average Value-at-Risk qualify), the robust risk-averse MDP has a non-randomized optimal policy if and only if a saddle point exists, so randomization is unnecessary.
- Dynamic-programming optimality conditions (3.7) are sufficient but, without strict monotonicity, need not be necessary: an optimal policy can fail to be dynamically consistent.
- In the continuous-distribution case, the same polynomial sample complexity would follow if the uniform slope condition (3.24) held; the paper leaves the verification of that condition as an open problem.
Reading between the lines
- The linear dependence on $(1-\beta)^{-1}$ in Theorem 3.1 suggests VaR-averse infinite-horizon problems are no harder than risk-neutral ones as the discount factor approaches one; comparing this rate with the risk-neutral sample complexity discussed in the paper's Remark 3.2 would test whether the phenomenon is general.
- The saddle-point equivalence suggests a recipe for other non-convex or non-coherent risk measures: whenever the functional is concave in the distribution, deterministic optimal policies should exist, and for measures that fail this concavity one could try to construct a two-stage MDP where randomized policies strictly outperform deterministic ones.
- The Lipschitz-approximation step in Lemma 3.1 is modular: any risk functional with a finite Lipschitz modulus in the sup-norm fits the same sample-complexity machinery, so the theorem should extend to entropic or expectile-based risk measures with only the constants changed.
- The open density question in Remark 3.2 is the main obstacle to a fully unconditional continuous-distribution theorem; a coarea-type argument under smoothness of the cost and dynamics is a plausible route toward removing the 'difficult to verify' caveat.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a framework for risk-averse stochastic optimal control and Markov decision processes based on nested conditional risk functionals, with special attention to Value-at-Risk. It presents dynamic-programming equations for both finite and infinite horizons, derives finite-sample guarantees for empirical Value-at-Risk and for the empirical Bellman fixed point in the finite-support case (Theorem 3.1), and gives conditions under which randomized policies can be replaced by non-randomized optimal policies (Proposition 3.1). The proofs use contraction arguments, DKW-type uniform deviation bounds, and net arguments.
Significance. If the main results are repaired as indicated below, the paper makes a useful contribution: it gives an explicit, dimension-dependent sample-size bound for VaR-averse infinite-horizon SOC, a regime where quantitative guarantees are scarce, and it cleanly separates the rectangular and non-rectangular constructions of nested risk functionals. The DKW-based arguments are transparent and the contraction framework is standard. The randomized-policy equivalence is also of interest, but the current statement and proof need additional hypotheses. Overall, the central ideas are defensible, but the main theorem is incomplete as stated and one supporting claim about VaR is incorrect.
major comments (4)
- [§3.1.1, Theorem 3.1 and Lemma 3.1] Theorem 3.1 is stated for arbitrary β∈(0,1) and L>0, but its proof and the displayed bound (3.20) are valid only when βL>1. Lemma 3.1 explicitly assumes βL_RL>1 (with L_R=1 for VaR), and the bound (3.20) contains 1/(βL−1) and log(βL); for βL≤1 the denominator vanishes or the logarithm becomes negative, so the expression is not a meaningful sample size. The parenthetical 'without loss of generality' in Lemma 3.1 does not supply the missing condition, and Remark 3.2 itself notes that the βL<1 case requires a different argument. Please add βL>1 to the theorem statement or give a separate statement and proof for βL≤1.
- [§3.1.1, Lemma 3.1 and Theorem 3.1] The choice k=1/(1−β) log(1/ε) in Lemma 3.1 only ensures ||V−V^(k)||∞≤ε when ||V||∞≤1. The contraction property gives ||V−V^(k)||∞≤β^k||V||∞, and Assumption 3.1 imposes no normalization on the cost function. If the cost is bounded by C, the value function is only known to be bounded by C/(1−β), and k should involve an additional log(C/[ε(1−β)])-type factor. This missing constant also propagates into the sample bound (3.20). Please add a boundedness or normalization assumption and track the resulting constant through the bound.
- [§3.2, paragraph after Assumption 3.2] The statement that V@R_α is concave in P is false. For a random variable Z taking values 0 and 1, let P_1=0.4δ_0+0.6δ_1, P_2=0.6δ_0+0.4δ_1, and α=0.5. Then V@R_{0.5}^{P_1}(Z)=1, V@R_{0.5}^{P_2}(Z)=0, and V@R_{0.5}^{(P_1+P_2)/2}(Z)=0, violating the concavity inequality R_{τP+(1−τ)P'} ≥ τR_P+(1−τ)R_{P'}. If the intended use is only Sion's theorem, quasi-concavity in P would suffice and does hold for VaR; the assumption and the surrounding discussion should be corrected accordingly.
- [§3.2, Proposition 3.1] The converse direction of Proposition 3.1 invokes Sion's theorem, but Assumption 3.2 does not include compactness of M_t, continuity or semicontinuity of R_P in P, or existence of optimal solutions for the dual problem (3.29). The proof's sentence 'Because both respective problems have optimal solutions' introduces an unstated hypothesis. As written, the claimed necessary-and-sufficient characterization is not established. Please add the missing compactness, continuity, and solution-existence hypotheses, or weaken the proposition accordingly.
minor comments (3)
- [Proof of Theorem 3.1, eq. (3.22)] The symbol L is reused for L(1+β\tilde L) in (3.22), while L already denotes the Assumption 3.1 Lipschitz constant in the theorem statement. Please use a distinct symbol, such as L_{\tilde Z}, to avoid ambiguity.
- [§2.1, Proposition 2.2 proof] The displayed expression involving the minimum in the proof appears to have a typesetting error; the intended quantity should be \sqrt{\log(2/\delta)/(2N c^2)} rather than a misplaced square-root sign. Please correct the formula.
- [Remark 3.2] The remark says the sample size is 'not sensitive to the discount factor' while displaying O((1−β)^{-2}c^{-2}ε^{-2}); it would be clearer to state that the dependence is polynomial rather than exponential.
Circularity Check
No circularity found: the sample-complexity theorems follow from DKW inequalities and Lipschitz net arguments, and the self-citations are to independent published theorems, not to the paper's own conclusions.
full rationale
The derivation chain is not circular. In Theorem 3.1 the empirical Bellman fixed point V_N is defined by the empirical Bellman equation, not in terms of the true value function V or the quantile gap κ_α; the bound (3.20) is a consequence of Proposition 2.1, Lemma 3.1, and the ε-net argument, so no prediction is a renamed fitted parameter. The condition κ_α>0 is a hypothesis about the fixed reference distribution P and is not fitted from the data or from V_N. The cited prior work [17], [18], [20], and [21] supplies the interchangeability principle, the conditional functional construction, and comparison results; each is a parameter-free published theorem with stated assumptions that do not include the finite-sample bound being derived, so these are independent support rather than circular self-citation. The paper itself flags a genuine limitation in Remark 3.2 ('it remains open to establish the existence of density function for Z_{x,u}'), and there is a correctness gap in Theorem 3.1: the bound (3.20) contains the factors (βL−1)^{-1} and log(βL), while the theorem statement omits the βL>1 condition that Lemma 3.1 invokes 'without loss of generality'; for βL≤1 the displayed bound is undefined. This is an incompleteness or correctness risk, not a circularity, because no equation in the paper is equivalent to its own input by construction and no fitted value is presented as a prediction. Accordingly no circular step is identified and the score is 0.
Assumptions & free parameters
assumptions (6)
- domain assumption Risk functionals satisfy monotonicity (A1) and translation equivariance (A3), with RP(0)=0 in the infinite-horizon sections.
- domain assumption In non-rectangular settings the paper postulates law-invariant risk functionals representable as ρ(F^P_Z), so conditional counterparts are well-defined versions.
- ad hoc to paper Uniform growth condition (2.17): the cdfs of Z_x have slope at least c on [ν_x-b, ν_x+b] uniformly in x.
- ad hoc to paper Finite support of the base measure P and the quantile gap κ_α>0 in (3.19) for Theorem 3.1.
- domain assumption Rectangularity and stagewise independence: the distribution of (ξ_1,...,ξ_T) does not depend on states or controls, and in rectangular robust settings M = {P_1×...×P_T}.
- standard math Sion's minimax theorem is applicable to the inf-sup problems (3.25) and (3.28), requiring convexity and compactness or continuity conditions not fully stated in Assumption 3.2.
Cite this review
Pith. "Pith review of Risk-averse formulations of Stochastic Optimal Control and Markov Decision Processes." pith.science (2026). https://pith.science/paper/Z6WM6S46
@misc{pith2026250516651,
author = {Pith},
title = {Pith review of: Risk-averse formulations of Stochastic Optimal Control and Markov Decision Processes},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z6WM6S46}},
note = {Machine review of arXiv:2505.16651}
}
read the original abstract
The aim of this paper is to investigate risk-averse and distributionally robust modeling of Stochastic Optimal Control (SOC) and Markov Decision Process (MDP). We discuss construction of conditional nested risk functionals, a particular attention is given to the Value-at-Risk measure. Necessary and sufficient conditions for existence of non-randomized optimal policies in the framework of robust SOC and MDP are derived. We also investigate sample complexity of optimization problems involving the Value-at-Risk measure.
Reference graph
Works this paper leans on
-
[20]
A. Shapiro and Yan Li. Distributionally robust stochastic optimal control. Operations Research Letters, 2025
work page 2025
-
[17]
A. Shapiro. Interchangeability principle and dynamic equations in risk averse stochastic pro- gramming. Operations Research Letters, 45:377–381, 2017
work page 2017
-
[18]
A. Shapiro and Y. Cheng. Central limit theorem and sample complexity of stationary stochastic programs. Operations Research Letters, 49:676–681, 2021
work page 2021
-
[21]
A. Shapiro and A. Pichler. Conditional distributionally robust functionals. Operations Re- search, 72:2745 – 2757, 2024
work page 2024
-
[1]
P. Artzner, F. Delbaen, J.-M. Eber, and D. Heath. Coherent measures of risk. Mathematical Finance, 9:203–228, 1999
work page 1999
-
[2]
D.P. Bertsekas and S.E. Shreve. Stochastic Optimal Control, The Discrete Time Case . Aca- demic Press, New York, 1978
work page 1978
-
[3]
Claude Dellacherie and Paul-Andr´ e Meyer. Probabilities and Potential . North-Holland Pub- lishing Co., Amsterdam, The Netherlands., 1988
work page 1988
-
[4]
P. M. Esfahani and D. Kuhn. Data-driven distributionally robust optimization using the Wasserstein metric: performance guarantees and tractable reformulations. Mathematical Pro- gramming, pages 115–166, 2018
work page 2018
Show all 24 references
-
[5]
H. Federer. Curvature measures. Transactions of the American Mathematical Society , 93(3):418–491, 1959
1959
-
[6]
G.N. Iyengar. Robust Dynamic Programming. Mathematics of Operations Research , 30:257– 280, 2005
2005
-
[7]
M. R. Kosorok. Introduction to Empirical Processes and Semiparametric Inference . Springer, 2008
2008
-
[8]
D. Kuhn, S. Shafiee, and W. Wiesemann. Distributionally robust optimization. https://arxiv.org/abs/2411.02549, 2024. 21
2024 arXiv
-
[9]
Yan Li and A. Shapiro. Rectangularity and duality of distributionally robust markov decision processes. https://arxiv.org/abs/2308.11139, 2023
2023 arXiv
-
[10]
Nilim and L
A. Nilim and L. El Ghaoui. Robust control of Markov decision processes with uncertain transition probabilities. Operations Research, 53:780–798, 2005
2005
-
[11]
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman. Markov decision processes: discrete stochastic dynamic programming . John Wiley & Sons, 2014
2014
-
[12]
R.T Rockafellar and S. Uryasev. Conditional value-at-risk for general loss distributions. J. of Banking and Finance , 26(7):1443–1471, 2002
2002
-
[13]
Ruszczy´ nski
A. Ruszczy´ nski. Risk-averse dynamic programming for Markov decision processes.Mathemat- ical Programming, 125:235–261, 2010
2010
-
[14]
Ruszczy´ nski and A
A. Ruszczy´ nski and A. Shapiro. Conditional risk mappings. Mathematics of Operations Re- search, 31:544–561, 2006
2006
-
[15]
Ruszczy´ nski and A
A. Ruszczy´ nski and A. Shapiro. Optimization of convex risk functions. Mathematics of Oper- ations Research, 31:433–452, 2006
2006
-
[16]
H. Scarf. A min-max solution of an inventory problem. In Studies in the Mathematical Theory of Inventory and Production , pages 201–209. Stanford University Press, 1958
1958
-
[19]
Shapiro, D
A. Shapiro, D. Dentcheva, and A. Ruszczy´ nski.Lectures on Stochastic Programming: Modeling and Theory. SIAM, Philadelphia, third edition, 2021
2021
-
[22]
M. Sion. On general minimax theorems. Pacific Journal of Mathematics , 8:171–176, 1958
1958
-
[23]
Mesures dans les espaces produits
Ionescu Tulcea. Mesures dans les espaces produits. Atti Accad. Naz. Lincei Rend, 7:208 – 211, 1949
1949
-
[24]
Wiesemann, D
W. Wiesemann, D. Kuhn, and B. Rustem. Robust Markov decision processes. Mathematics of Operations Research, 38(1):153 – 183, 2013. 22
2013
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.