REVIEW 2 major objections 4 minor 2 cited by
Turnpike Property of Stochastic Linear-Quadratic Optimal Control Problems in Large Horizons with Regime Switching I: Homogeneous Cases
T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper proves a strong turnpike property for stochastic linear-quadratic optimal control with regime switching: on a long finite horizon, the optimal state–control pair is exponentially close, throughout the middle of the interval, to…
desk verdict Genuinely new regime-switching turnpike result with a real but repairable gap in the key Riccati convergence proof; worth refereeing after the authors patch Step 4. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the family of Riccati equations attached to the value function. For horizon $T$ the value is quadratic, $V_T(t,x,i)=\tfrac12\langle P_T(t,i)x,x\rangle$, with $P_T$ solving a differential Riccati equation (DRE) that runs backward to the terminal condition $P_T(T)=0$; for the infinite horizon the value is $\tfrac12\langle P_\infty(i)x,x\rangle$ with $P_\infty$ solving an algebraic Riccati equation (ARE). The central step is Lemma 4.1(iii): under uniform strong convexity of the running cost (assumption (A3)), the DRE solution increases monotonically to the ARE solution and the gap decays exponentially, $0\le P_\infty(i)-P_T(t,i)\le K e^{-\delta(T-t)}I$, with the closed-loop gains $\Theta_T(t,i)$ and $\Theta_\infty(i)$ (defined through $P_T$ and $P_\infty$) converging at the same rate. The proof feeds this gain convergence into a dissipation inequality: because the limiting gain $\Theta_\infty$ stabilizes the system, the closed-loop dynamics of the difference process $\bar X_T-\bar X_\infty$ are exponentially stable up to forcing terms carrying the gain error and the decaying state bound (4.26), and Gronwall's inequality converts this into the estimate (4.27) of Theorem 4.2.
What would settle it
A concrete check: take a two-regime example with $n=m=1$ (one stable and one unstable regime, unit state and control costs, positive jump rates), solve the differential Riccati equation (4.5) backward from $P_T(T)=0$ and the algebraic Riccati equation (4.10), and plot $\log\|P_\infty-P_T(t)\|$ against $T-t$. The theorem predicts a straight line whose slope does not depend on $T$; a gap that decays only polynomially, or a slope that changes as $T$ grows, would refute Lemma 4.1(iii) and with it Theorem 4.2. A direct simulation of the two closed-loop optimal systems could then verify whether $E|\bar X_T(s)-\bar X_\infty(s)|^2$ obeys (4.27) uniformly in $s\in[t,T]$.
Extended reading notes
Core claim
The central result, Theorem 4.2, is the exponential turnpike estimate: under assumptions (A1), (A2)$'$, and (A3) there are constants $\delta>0$ and $K>0$ such that for any two initial triples $(t,x_T,i)$ and $(t,x_\infty,i)$ and every $s\in[t,T]$, $$E\big(|\bar X_T(s)-\bar X_\infty(s)|^2+|\bar u_T(s)-\bar u_\infty(s)|^2\big)\le K $e^{{-\delta(s-t)}}$|x_T-x_\infty|^2+K $e^{{-\delta(s-t)}}$$e^{{-2\delta(T-s)}}$|x_T|^2,$$ where $(\bar X_T,\bar u_T)$ is the optimal pair of the finite-horizon problem $(LQ)_{t,T}$ and $(\bar X_\infty,\bar u_\infty)$ is the optimal pair of the infinite-horizon problem $(LQ)_{t,\infty}$. The first summand is the forgetting of initial conditions: whatever the starting difference $x_T-x_\infty$, the two optimal processes merge exponentially fast. The second summand is the finite-horizon policy's memory of its terminal boundary, which dies out at rate $\delta$ as one moves into the interior. On the middle of a long horizon both terms are negligible, so the finite-horizon optimal pair is essentially the $T$-independent pair of the infinite-horizon problem — the strong turnpike property. The mechanism is the exponential convergence of the differential Riccati solutions $P_T$ to the algebraic Riccati solution $P_\infty$, with $0\le P_\infty(i)-P_T(t,i)\le K e^{-\delta(T-t)}I$ and the same exponential rate for the closed-loop gains $\Theta_T\to\Theta_\infty$; the difference of the two closed-loop systems then decays by dissipativity of the limiting loop (Lemma 4.1 and the proof of Theorem 4.2).
Load-bearing premise
The load-bearing premise is that the running cost is uniformly strongly convex — the same positive curvature constant across every regime and every horizon length — because when that fails a simple example in Section 3 already drives the optimal value to $-\infty$, and every exponential estimate in Section 4 collapses; a second load-bearing premise is the time-homogeneity of the Markov chain's generator, without which the time-shift identity (4.1) anchoring the Riccati convergence breaks.
Editorial extensions
If this is right
- For any long horizon, the infinite-horizon optimal pair is a provably close substitute for the finite-horizon one: in the middle of the interval the mean-square gap is bounded by the two exponential terms of (4.27), and both are negligible.
- The closed-loop gain $\Theta_\infty$ computed once from the algebraic Riccati equation serves as a stationary feedback law for every long horizon, with the same uniform error bound $K e^{-\delta(T-s)}$ near the terminal boundary.
- Even with a single regime (no switching at all), the estimate refines the earlier stochastic turnpike bounds, and the proof supplies a new route to the exponential convergence of the differential Riccati solution to the algebraic one.
- The Riccati convergence lemma is the announced tool for the sequel on non-homogeneous problems, where the bound (1.7) anticipates an additional, horizon-independent function $h(s-t)$.
- When both problems start from the same state at time zero, the whole-horizon mean-square gap $E\int_0^T(|\bar X_T-\bar X_\infty|^2+|\bar u_T-\bar u_\infty|^2)\,ds$ converges to $0$ as $T\to\infty$, so the finite-horizon solution converges to the infinite-horizon one in the time-integrated sense (eq. (4.28)).
Reading between the lines
- Because the proof uses only uniform strong convexity of the cost and dissipativity of the limiting closed loop, the same two-step pattern — exponential Riccati convergence, then a dissipation inequality — should extend to mean-field or partial-information LQ problems with regime switching, as long as an analogue of (A3) holds.
- Read as a statement about receding-horizon control, the result implies that a controller solving the differential Riccati equation on a rolling window of length $L$ behaves like the infinite-horizon law except on a final stretch of width roughly $1/\delta$, so the window length need not grow with the overall planning horizon.
- The two-term structure of (4.27) suggests what the non-homogeneous sequel's extra term $h(s-t)$ should look like: a transient contributed by the linear terms in the cost, decaying in $s-t$, so the interior turnpike survives and only the two ends of the horizon carry additional error.
- A quantitative prediction a practitioner could test: the rate $\delta$ should be governed by the stability margin of the infinite-horizon closed-loop system, so weakly stabilizable regimes should slow the turnpike convergence; the paper does not compute this dependence explicitly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies homogeneous stochastic linear-quadratic optimal control problems with Markov regime switching, treating finite horizons [t,T] and the infinite horizon as a limit. Under assumptions (A1), (A2)', and uniform strong convexity (A3), the authors prove that the solution P_T of the differential Riccati equation converges exponentially to the solution P_infinity of the algebraic Riccati equation, that the associated feedback gains converge at the same rate, and that the finite-horizon optimal pairs satisfy the strong turnpike estimate (4.27). The proof follows the architecture of Sun-Wang-Yong [29] but uses a Gronwall/dissipativity argument for the exponential Riccati convergence. The paper closes with a brief discussion of future non-homogeneous cases.
Significance. If valid, the result is a useful extension of turnpike theory to regime-switching stochastic LQ problems; even in the no-switching case the rate in (4.27) is sharper than the previous two-exponential bound (1.6). The clean formulation of the Riccati convergence, the uniform-in-regime strong convexity condition, and the explicit use of a dissipativity inequality are assets. Most of the estimates in Theorem 4.2 are algebraically sound; the main weakness is a gap in the proof of Lemma 4.1(iii), the key exponential Riccati convergence. That gap appears repairable by a standard dynamic-programming argument, but the printed proof is incomplete.
major comments (2)
- [Section 4, Lemma 4.1, Step 4] The displayed chain after "By the dynamic program principle" is not a valid identity. Since Problem (LQ)_{T,T} has empty horizon, V_T(T,z,j)=0 and the DRE terminal condition gives P_T(T,.)=0; thus the term E< P_T(T,alpha(T))X(T), X(T)> vanishes and the chain asserts V_infinity(x,i)=inf_{u in U[0,T]} J_T(0,x,i;u)=V_T(0,x,i), which is false under (A3) whenever the infinite-horizon tail cost is positive. The shift from U[t,T] to U[0,T] also changes the horizon length from T-t to T, so even the intended comparison is misindexed. This step is precisely what yields 0 <= P_infinity - P_T(t,i) <= K e^{-delta(T-t)} I in (4.18), and (4.18) drives (4.19) and Theorem 4.2. The argument can be repaired by applying the dynamic programming principle with tail cost V_infinity (equivalently 1/2 < P_infinity(alpha(T))X(T),X(T)>), inserting the finite-horizon optimal pair, subtracting V_T(t,x,i), and using (4.26) with horizon length T-t; the missing factor 1/2 in the P_T(T) term must also be restored in the repaired version.
- [Section 4, Lemma 4.1, Step 2] The proof of eTheta in S[A,C;B,D] uses the uniform convergence (4.23) before it has been established; Step 1 only proves pointwise monotone convergence P_T(t,i) -> P_infinity(i). The missing justification is Dini's theorem, since P_T(s,i)=P_{T-s}(0,i) is monotone in T and continuous in s and the limit is continuous. In the same step, the nonexistent reference (A6) should be (A3), the displayed "J_T = epsilon E int ..." should read "J_T >= epsilon E int ...", and the two limits "T -> 0" should be "T -> infinity". With these corrections the Fatou argument is sound.
minor comments (4)
- [Section 4, Lemma 4.1, Step 3] The displayed identity for J_infinity is missing the factors 1/2 present in the earlier derivation: it should be J_infinity = (1/2) < eP(i)x,x> + (1/2) E int_0^infty <(R+D^T eP D)(u-eTheta X), u-eTheta X> ds; also "Let X(.) be the solution to (2.3) under u(.)" should refer to the controlled equation (1.1).
- [Section 4, Theorem 4.2 proof] In the Ito computation, the term C_{Theta_infinity}(\bar X_T - \bar X_T) should read C_{Theta_infinity}(\bar X_T - \bar X_infinity).
- [Introduction, estimate (1.6)] The quantifier "for all t in [0,T]" in (1.6) should be "for all s in [0,T]", since s is the running time in the displayed expectation.
- [References] Reference [21] has a typographical error: "L. Mou and J. YongTwo-person" should be "L. Mou and J. Yong, Two-person".
Circularity Check
No significant circularity: the Riccati-convergence engine is proved from the assumptions and external lemmas; the DPP flaw in Lemma 4.1 is a proof gap, not a circular reduction.
full rationale
The derivation chain from (A1)-(A3) to the turnpike estimate (4.27) is not circular. DRE existence and δ-strong regularity are imported from Zhang-Li-Xiong [39], an external source with no author overlap; the stabilizability/dissipating-strategy equivalence from [20] is a supporting theorem used to set up the infinite-horizon problem, not the target claim. Lemma 4.1 proves monotone convergence P_T(t,i) -> P_inf(i) using time homogeneity (4.20)-(4.22), identifies the limit with the ARE solution via Fatou and uniqueness, and derives the exponential rate (4.18) from the Gronwall estimate (4.26) together with a value-function comparison. Theorem 4.2 then combines (4.18)-(4.19) with dissipativity of [A_{Theta_inf}, C_{Theta_inf}] to bound the state and control differences. No parameter is fitted and no estimate is assumed from the conclusion. The proof as printed contains a defect in Lemma 4.1 Step 4 where the dynamic-programming tail uses V_T / P_T(T)=0; taken literally this would assert V_inf = V_T, which is generally false. This is a correctness gap in the written proof, not a circularity: replacing the tail value by V_inf and using (4.26) repairs the argument without turning the target estimate into an input.
Assumptions & free parameters
assumptions (6)
- domain assumption Coefficient measurability and boundedness (A1)
- domain assumption Stabilizability of [A,C;B,D], reduced to stability of [A,C] as (A2)'
- domain assumption Uniform strong convexity (A3): Q(i) - S(i)^T R(i)^{-1} S(i) in S^n_++ and R(i) in S^m_++
- standard math Extended Ito formula for Markov-modulated SDEs (Proposition 2.5)
- domain assumption Time-homogeneous Markov chain with constant generator (2.1)
- standard math Existence and uniqueness of delta-strong regular DRE solution (Lemma 4.1(i))
Cite this review
Pith. "Pith review of Turnpike Property of Stochastic Linear-Quadratic Optimal Control Problems in Large Horizons with Regime Switching I: Homogeneous Cases." pith.science (2026). https://pith.science/paper/CQBCSAIJ
@misc{pith2026250609337,
author = {Pith},
title = {Pith review of: Turnpike Property of Stochastic Linear-Quadratic Optimal Control Problems in Large Horizons with Regime Switching I: Homogeneous Cases},
year = {2026},
howpublished = {\url{https://pith.science/paper/CQBCSAIJ}},
note = {Machine review of arXiv:2506.09337}
}
read the original abstract
This paper is concerned with optimal control problems for a linear homogeneous stochastic differential equation having regime switching with purely quadratic functional in the large time horizons. We establish the so-called turnpike properties for the optimal pairs. The key is to prove a proper convergence of the solutions to the differential Riccati equations to the algebraic Riccati equation. Even for the problems without regime switchings, our result provides a refined estimate compared to those in the previous literature, which also provides a new tool for further research.
Forward citations
Cited by 2 Pith papers
-
Long-time behavior and turnpike properties of linear-quadratic graphon mean field control problems
Finite-horizon optimal pairs in linear-quadratic graphon mean field control converge exponentially to the ergodic optimal pair away from time boundaries.
-
Turnpike properties for zero-sum stochastic linear quadratic differential games of Markovian regime switching system
Finite-horizon optimal feedback gains in zero-sum stochastic linear-quadratic games with regime switching converge exponentially to infinite-horizon gains, yielding a turnpike theorem for the optimal triple.
Reference graph
Works this paper leans on
-
[29]
J. Sun, H. Wang, and J. Yong,Turnpike properties for stochastic linear-quadratic optimal control problems,Chin. Ann. Math.,Ser B, 43 (2022), 999–1022
work page 2022
- [31]
- [32]
-
[13]
J. Jian, S. Jin, Q. Song, and J. Yong,Long-Time Behaviors of Stochastic Linear- Quadratic Optimal Control Problems,arXiv preprint arXiv:2409.11633
-
[1]
M. Basei and H. Pham,A weak martingale approach to linear-quadratic McKean- Vlasov stochastic control problems,J. Optim. Theory Appl.,181 (2019), 347–382
work page 2019
-
[2]
E. Bayraktar and J. Jian,Ergodicity and turnpike properties of linear-quadratic mean field control problems,arXiv preprint arXiv:2502.08935
-
[3]
T. Breiten and L. Pfeiffer,On the turnpike property and the receding-horizon method for linear-quadratic optimal control problems,SIAM J. Control Optim.,58 (2020), 1077–1102
work page 2020
-
[4]
D. A. Carlson, A. B. Haurie, and A. Leizarowitz,Infinite Horizon Optimal Control — Deterministic and Stochastic Systems, 2nd ed.,Springer-Verlag, Berlin, 1991
work page 1991
Show all 39 references
-
[5]
Chen and P
Y. Chen and P. Luo,Turnpike properties for stochastic backward linear-quadratic optimal problems,arXiv preprint arXiv:2309.03456
-
[6]
Conforti,Coupling by reflection for controlled diffusion processes: Turnpike prop- erty and large time behavior of Hamilton-Jacobi-Bellman equations,Ann
G. Conforti,Coupling by reflection for controlled diffusion processes: Turnpike prop- erty and large time behavior of Hamilton-Jacobi-Bellman equations,Ann. Appl. Probab.,33 (2023), 4608–4644
2023
-
[7]
T. Damm, L. Gr¨ une, M. Stieler, and K. Worthmann,An exponential turnpike theorem for dissipative discrete time optimal control problems,SIAM J. Control Optim.,52 (2014), 1935–1957
2014
-
[8]
Dorfman, P
R. Dorfman, P. A. Samuelson, and R. M. Solow,Linear Programming and Economics Analysis,McGraw-Hill, New York, 1958
1958
-
[9]
Faulwasser and L
T. Faulwasser and L. Gr¨ une,Turnpike properties in optimal control: An overview of discrete-time and continuous-time results,Numerical Control: Part A,23 (2022), 367–400
2022
-
[10]
Gr¨ une and R
L. Gr¨ une and R. Guglielmi,Turnpike properties and strict dissipativity for discrete time linear quadratic optimal control problems,SIAM J. Control Optim.,56 (2018), 1282–1302
2018
-
[11]
Gr¨ une and R
L. Gr¨ une and R. Guglielmi,On the relation between turnpike properties and dissipa- tivity for continuous time linear quadratic optimal control problems,Math. Control Relat. Fields,11 (2021), 169–188
2021
-
[12]
Huang, X
J. Huang, X. Li, and J. Yong,A linear-quadratic optimal control problem for mean- field stochastic differential equations in infinite horizon,Math. Control Relat. Fields, 5 (2015), 97–139. 21
2015
-
[14]
L¨ u,Stochastic linear quadratic optimal control problems for mean-field stochastic evolution equation,ESAIM Control Optim
Q. L¨ u,Stochastic linear quadratic optimal control problems for mean-field stochastic evolution equation,ESAIM Control Optim. Calc. Var.,26 (2020), 127
2020
-
[15]
Lou and W
H. Lou and W. Wang,Turnpike properties of optimal relaxed control problems,ESAIM Control Optim. Calc. Var.,25 (2019), 74
2019
-
[16]
Mao,Stability of stochastic differential equations with Markovian switching,Stoch
X. Mao,Stability of stochastic differential equations with Markovian switching,Stoch. Proc. Appl.,79 (1999) 45–67
1999
-
[17]
Marimon,Stochastic turnpike property and stationary equilibrium,J
R. Marimon,Stochastic turnpike property and stationary equilibrium,J. Economic Theory,47 (1989), 282–306
1989
-
[18]
L. W. McKenzie,Turnpike theory,Econometrica,44 (1976), 841–865
1976
-
[19]
H. Mei, Q. Wei, and J. Yong,Linear-quadratic optimal control problem for mean-field stochastic differential equations with a type of random coefficients,Numer. Algebra, Control Optim., 14 (2024), 813–852
2024
-
[20]
H. Mei, Q. Wei, and J. Yong,Linear-quadrartic optimal control for mean-field stochas- tic differential equations in infinite-horizon with regime switching,Chin. Ann. Math. to appear
-
[21]
Mou and J
L. Mou and J. YongTwo-person zero-sum linear quadratic stochastic differential games by a Hilbert space method,J. Industrial & Management Optim.,2 (2006), 95–117
2006
-
[22]
von Neumann,A model of general economic equilibrium,Rev
J. von Neumann,A model of general economic equilibrium,Rev. Econ. Stud.,13 (1945), 1–9
1945
-
[23]
Porretta and E
A. Porretta and E. Zuazua,Long time versus steady state optimal control,SIAM J. Control Optim.,51 (2013), 4242–4273
2013
-
[24]
F. P. Ramsey,A mathematical theory of saving,The economic journal,38 (1928), 543–559
1928
-
[25]
Sakamoto and E
N. Sakamoto and E. Zuazua,The turnpike property in nonlinear optimal control—A geometric approach,Automatica,134 (2021), 109939
2021
-
[26]
Schiessl, M
J. Schiessl, M. H., Baumann, T. Faulwasser, and L. Gr¨ une, L.,On the relationship between stochastic turnpike and dissipativity notions,arXiv:2311.07281v2 [math.OC] 21 Aug 2024
2024 arXiv
-
[27]
A. V. Skorokhod,Asymptotic Methods in the Theory of Stochastic Differential Equa- tions,AMS, Providence, RI., 1989
1989
-
[28]
Sun,Mean-field stochastic linear quadratic optimal control problems: Open-loop solvabilities,ESAIM Control Optim
J. Sun,Mean-field stochastic linear quadratic optimal control problems: Open-loop solvabilities,ESAIM Control Optim. Calc. Var.,23 (2017), 1099–1127
2017
-
[30]
Sun and J
J. Sun and J. Yong,Stochastic Linear-Quadratic Optimal Control Theory: Open- Loop and Closed-Loop Solutions,Springer Briefs Math., Springer, 2020. 22
2020
-
[33]
Tr´ elat and E
E. Tr´ elat and E. Zuazua,The turnpike property in finite-dimensional nonlinear opti- mal control,J. Diff. Equ.,258 (2015), 81–114
2015
-
[34]
Yong,Linear-quadratic optimal control problems for mean-field stochastic differen- tial equations,SIAM J
J. Yong,Linear-quadratic optimal control problems for mean-field stochastic differen- tial equations,SIAM J. Control Optim. 51(2013), 2809–2838
2013
-
[35]
Yong,Optimization Theory: A Concise Introduction,World Scientific, Singapore, 2018
J. Yong,Optimization Theory: A Concise Introduction,World Scientific, Singapore, 2018
2018
-
[36]
A. J. Zaslavski,Turnpike Properties in the Calculus of Variations and Optimal Con- trol,Nonconvex Optim. Appl. 80, Springer, New York, 2006
2006
-
[37]
A. J. Zaslavski,Turnpike Theory of Continuous-Time Linear Optimal Control Prob- lems,Springer Optim. Appl. 148, Springer, 2019
2019
-
[38]
Zuazua,Large time control and turnpike properties for wave equations,Ann
E. Zuazua,Large time control and turnpike properties for wave equations,Ann. Rev. Control, 44 (2017), 199–210
2017
-
[39]
Zhang, X
X. Zhang, X. Li, and J. Xiong,Open-loop and closed-loop solvabilities for stochas- tic linear quadratic optimal control problems of Markovian regime switching system, ESAIM: Control, Optim. Calc. Var.,27(2021), 69. 23
2021
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.