Pith. sign in

REVIEW 2 cited by

On Wasserstein Reinforcement Learning and the Fokker-Planck equation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1712.07185 v1 pith:5T4KTDRH submitted 2017-12-19 cs.LG

classification cs.LG
keywords policysmallwassersteinchangedistanceequationfokker-planckgradients
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Policy gradients methods often achieve better performance when the change in policy is limited to a small Kullback-Leibler divergence. We derive policy gradients where the change in policy is limited to a small Wasserstein distance (or trust region). This is done in the discrete and continuous multi-armed bandit settings with entropy regularisation. We show that in the small steps limit with respect to the Wasserstein distance $W_2$, policy dynamics are governed by the Fokker-Planck (heat) equation, following the Jordan-Kinderlehrer-Otto result. This means that policies undergo diffusion and advection, concentrating near actions with high reward. This helps elucidate the nature of convergence in the probability matching setup, and provides justification for empirical practices such as Gaussian policy priors and additive gradient noise.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Inexact JKO and proximal-gradient algorithms in the Wasserstein space

    math.OC 2025-05 conditional novelty 7.0 of 10

    Inexact JKO and proximal-gradient schemes in Wasserstein space converge weakly to minimizers with O(1/σ_n) energy rates, provided the error sequences are summable and the stepsize sums diverge.

  2. Wasserstein Adaptive Value Estimation for Actor-Critic Reinforcement Learning

    cs.LG 2025-01 reject novelty 4.0 of 10

    WAVE adds an adaptively weighted Sinkhorn approximation of the Wasserstein distance between successive Q-value distributions to the critic loss in actor-critic reinforcement learning.

Pith tools