Pith. sign in

REVIEW 1 cited by

A Pontryagin Perspective on Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.18100 v3 pith:D7M55WX6 submitted 2024-05-28 cs.LG math.OC

classification cs.LGmath.OC
keywords learningreinforcementalgorithmscontrolmethodsopen-loopoptimalpontryagin
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Reinforcement learning has traditionally focused on learning state-dependent policies to solve optimal control problems in a closed-loop fashion. In this work, we introduce the paradigm of open-loop reinforcement learning where a fixed action sequence is learned instead. We present three new algorithms: one robust model-based method and two sample-efficient model-free methods. Rather than basing our algorithms on Bellman's equation from dynamic programming, our work builds on Pontryagin's principle from the theory of open-loop optimal control. We provide convergence guarantees and evaluate all methods empirically on a pendulum swing-up task, as well as on two high-dimensional MuJoCo tasks, significantly outperforming existing baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DATA-DRIVEN PRONTO: a Model-free Solution for Numerical Optimal Control

    eess.SY 2025-06 conditional novelty 6.0 of 10

    DATA-DRIVEN PRONTO iteratively estimates local linearized dynamics from perturbed closed-loop experiments, solves an LQR subproblem with those estimates, and provably converges to a neighborhood of the optimal solutio...

Pith tools