Pith. sign in

REVIEW 1 cited by

Duality between density function and value function with applications in constrained optimal control and Markov Decision Process

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1902.09583 v3 pith:5S3XC4WP submitted 2019-02-25 eess.SY cs.SY

classification eess.SYcs.SY
keywords controlfunctiondensityproblemoptimalconstraintsdualoptimization
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Density function describes the density of states in the state space of a dynamic system or a Markov Decision Process (MDP). Its evolution follows the Liouville equation. We show that the density function is the dual of the value function in the optimal control problems. By utilizing the duality, constraints that are hard to enforce in the primal value function optimization such as safety constraints in robot navigation, traffic capacity constraints in traffic flow control can be posed on the density function, and the constrained optimal control problem can be solved with a primal-dual algorithm that alternates between the primal and dual optimization. The primal optimization follows the standard optimal control algorithm with a perturbation term generated by the density constraint, and the dual problem solves the Liouville equation to get the density function under a fixed control strategy and updates the perturbation term. Moreover, the proposed method can be extended to the case with exogenous disturbance, and guarantee robust safety under the worst-case disturbance. We apply the proposed method to three examples, a robot navigation problem and a traffic control problem in sim, and a segway control problem with experiment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Situational-Constrained Sequential Resources Allocation via Reinforcement Learning

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A constrained-reinforcement-learning algorithm encodes situational if-then allocation rules as disjunctive penalties and shows lower violations on simulated medical and agricultural allocation tasks.

Pith tools