Pith. sign in

REVIEW 2 cited by

V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.12238 v1 pith:WYF4VB3M submitted 2019-09-26 cs.AI cs.LG

classification cs.AIcs.LG
keywords policyv-mpocontinuouscontrolon-policypreviouslyreportedscores
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Some of the most successful applications of deep reinforcement learning to challenging domains in discrete and continuous control have used policy gradient methods in the on-policy setting. However, policy gradients can suffer from large variance that may limit performance, and in practice require carefully tuned entropy regularization to prevent policy collapse. As an alternative to policy gradient algorithms, we introduce V-MPO, an on-policy adaptation of Maximum a Posteriori Policy Optimization (MPO) that performs policy iteration based on a learned state-value function. We show that V-MPO surpasses previously reported scores for both the Atari-57 and DMLab-30 benchmark suites in the multi-task setting, and does so reliably without importance weighting, entropy regularization, or population-based tuning of hyperparameters. On individual DMLab and Atari levels, the proposed algorithm can achieve scores that are substantially higher than has previously been reported. V-MPO is also applicable to problems with high-dimensional, continuous action spaces, which we demonstrate in the context of learning to control simulated humanoids with 22 degrees of freedom from full state observations and 56 degrees of freedom from pixel observations, as well as example OpenAI Gym tasks where V-MPO achieves substantially higher asymptotic scores than previously reported.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 39 citations worldwide. Full citation record

  1. Multiple-Frequencies Population-Based Training

    cs.LG 2025-06 conditional novelty 7.0 of 10

    MF-PBT combines sub-populations that evolve at different frequencies with asymmetric migration to reduce PBT's short-sightedness and improve long-term RL rewards.

  2. A Minimax Approach to Ad Hoc Teamwork

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A minimax-Bayes choice of teammate-training distribution improves worst-case robustness in ad hoc teamwork, with the strongest empirical gains on Overcooked and Melting Pot.

Pith tools