Pith. sign in

REVIEW 1 cited by

A Q-learning algorithm for discrete-time linear-quadratic control with random parameters of unknown distribution: convergence and stabilization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.04970 v1 pith:G6E5LYFN submitted 2020-11-10 math.OC math.PR

classification math.OCmath.PR
keywords controlparametersproblemrandomalgebraicalgorithmconvergencediscrete-time
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper studies an infinite horizon optimal control problem for discrete-time linear systems and quadratic criteria, both with random parameters which are independent and identically distributed with respect to time. A classical approach is to solve an algebraic Riccati equation that involves mathematical expectations and requires certain statistical information of the parameters. In this paper, we propose an online iterative algorithm in the spirit of Q-learning for the situation where only one random sample of parameters emerges at each time step. The first theorem proves the equivalence of three properties: the convergence of the learning sequence, the well-posedness of the control problem, and the solvability of the algebraic Riccati equation. The second theorem shows that the adaptive feedback control in terms of the learning sequence stabilizes the system as long as the control problem is well-posed. Numerical examples are presented to illustrate our results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Automatic Differentiation-based Full Waveform Inversion with Flexible Workflows

    cs.LG 2024-11 conditional novelty 4.0 of 10

    ADFWI is an open-source PyTorch framework that uses automatic differentiation to replace hand-derived adjoint-state gradients in full waveform inversion across acoustic, elastic, and anisotropic media.

Pith tools