Online policy-gradient updates for unknown LQR systems are shown to be sequentially stable and convergent to the optimal gain, for indirect, direct, natural-gradient, Gauss-Newton and regularized versions.
Adaptive control and intersections with reinforce- ment learning,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.OC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Policy Gradient Adaptive Control for the LQR: Indirect and Direct Approaches
Online policy-gradient updates for unknown LQR systems are shown to be sequentially stable and convergent to the optimal gain, for indirect, direct, natural-gradient, Gauss-Newton and regularized versions.