Pith. sign in

REVIEW 1 cited by

A Moreau Envelope Approach for LQR Meta-Policy Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.17364 v2 pith:J3Y2Y2XX submitted 2024-03-26 math.OC cs.LGcs.SYeess.SY

classification math.OCcs.LGcs.SYeess.SY
keywords linearrealizationsapproachcostestimationmeta-policymoreausystem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the problem of policy estimation for the Linear Quadratic Regulator (LQR) in discrete-time linear time-invariant uncertain dynamical systems. We propose a Moreau Envelope-based surrogate LQR cost, built from a finite set of realizations of the uncertain system, to define a meta-policy efficiently adjustable to new realizations. Moreover, we design an algorithm to find an approximate first-order stationary point of the meta-LQR cost function. Numerical results show that the proposed approach outperforms naive averaging of controllers on new realizations of the linear system. We also provide empirical evidence that our method has better sample complexity than Model-Agnostic Meta-Learning (MAML) approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Coreset-Based Task Selection for Sample-Efficient Meta-Reinforcement Learning

    math.OC 2025-02 conditional novelty 6.0 of 10

    A derivative-free coreset task-selection algorithm for MAML-RL trains on a small weighted task subset and provably reduces sample complexity by O(1/epsilon), provided the task-selection bias is small.

Pith tools