Pith. sign in

Paper Citation Record · LEDGER

Wasserstein Policy Optimization

As of 17 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 5 inbound Pith citation observations for arXiv:2505.00663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00663 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:47:05.795761Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:19:53.934535Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:56:44.503712Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy22
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5b7dfde-d8ef-4053-a00f-f73925c912b2 · outbound

This paper cites write newline.

Wasserstein Policy Optimization write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.588081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.588081Z digest=sha256:ea2c9025305ca52d429bd65d8e6b53296bd330fe31900883126cb5a36835fece

Observation 2254d3da-e420-4c79-b6aa-9e78c980eec9 · outbound

This paper cites T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M.

Wasserstein Policy Optimization T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.331926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.593154Z digest=sha256:c8b63267ca25c5fcee1bcabe8a675fc1e6b90d72442291e7e3742b8d46177a39

Observation 5ab9bc6b-cab4-4dbc-b48b-8645b4b1b909 · outbound

This paper cites Wasserstein Robust Reinforcement Learning.

Wasserstein Policy Optimization Wasserstein Robust Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.597189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.597189Z digest=sha256:7e64e3560e30ff136b1e7cc56fcd46a3f128a6a1d5c518edb85ed5f9a88a45f6

Observation 733ba229-d450-4089-88ae-ba561c450cbc · outbound

This paper cites M., Lee, J.

Wasserstein Policy Optimization M., Lee, J

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.601296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.601296Z digest=sha256:b3fb9d97a740d4f7c4482f2ce64060fdaf3983bde45aa44ce2a5c8bf00cdbdf4

Observation d34dd838-1579-4ca8-aba0-b935f4832c21 · outbound

This paper cites Gradient flows: in metric spaces and in the space of probability measures.

Wasserstein Policy Optimization Gradient flows: in metric spaces and in the space of probability measures

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.604877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.604877Z digest=sha256:a0b0575a2f8184db94ff985ae9d9937bc7fc10708a0dfc503e46aeca0b8617dd

Observation 56df3e18-0b0e-4184-b5d8-f649980ff82f · outbound

This paper cites W., Budden, D., Dabney, W., Horgan, D., Tb, D., Muldal, A., Heess, N., and Lillicrap, T.

Wasserstein Policy Optimization W., Budden, D., Dabney, W., Horgan, D., Tb, D., Muldal, A., Heess, N., and Lillicrap, T

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.306496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.608466Z digest=sha256:e232541f3bfc69b314f620185acede328a506a8ed851a6371d7aa2749d3561e1

Observation 41dad690-926f-4135-b459-184cc07b59cb · outbound

This paper cites G., Sutton, R.

Wasserstein Policy Optimization G., Sutton, R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.294957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.612357Z digest=sha256:d353eb798e8c493f2fad9258464285d4c94b68afe766d89d107914783ea2218b

Observation c109af0b-e5a3-41d2-a497-fa0e246179c5 · outbound

This paper cites G., Dabney, W., and Munos, R.

Wasserstein Policy Optimization G., Dabney, W., and Munos, R

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.283686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.616220Z digest=sha256:4437a76420d5be34f830ec255c9b1261d740119ddbb472e7de8626d26eaf22b5

Observation 9662de86-3d7e-48cc-930a-e1e12a352561 · outbound

This paper cites and Brenier, Y.

Wasserstein Policy Optimization and Brenier, Y

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.619989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.619989Z digest=sha256:ebd45c66652c9174f2c0d7d49f5b2b1496cecc89e8eb1e74696da49454a8ce90

Observation d8ac57cd-5b35-4274-82fb-516c7af50691 · outbound

This paper cites Development of free-boundary equilibrium and transport solvers for simulation and real-time interpretation of tokamak experiments.

Wasserstein Policy Optimization Development of free-boundary equilibrium and transport solvers for simulation and real-time interpretation of tokamak experiments

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.266647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.624192Z digest=sha256:24e17203bc389c22210762efde4bc15ce12f04de3f8c17b8e2526b851d21da0d

Observation 59307d2d-0cea-46e5-806f-be758d7ddaf0 · outbound

This paper cites MICo: Improved representations via sampling-based state similarity for Markov decision processes.

Wasserstein Policy Optimization MICo: Improved representations via sampling-based state similarity for Markov decision processes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.628603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.628603Z digest=sha256:3939772f9bffa559de566a41d1efea3c0156a4822955b58bd85cc8b4da561a64

Observation 5b4cddc1-b794-4c27-9d69-bffcd7d184c8 · outbound

This paper cites T., Rubanova, Y., Bettencourt, J., and Duvenaud, D.

Wasserstein Policy Optimization T., Rubanova, Y., Bettencourt, J., and Duvenaud, D

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.634035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.634035Z digest=sha256:8787fa66d9c05df42dda37f4ac5b28a1d20ac934fc97286a554acb97302f837d

Observation c38ecf22-afac-4732-8aba-d3fcf077b551 · outbound

This paper cites Fast and accurate deep network learning by exponential linear units (elus).

Wasserstein Policy Optimization Fast and accurate deep network learning by exponential linear units (elus)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.248832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.638728Z digest=sha256:29a0b633e9a8c9a56010ce7ba4dad91ef9a0961b9c57dac16817d1200bfcf2ff

Observation 6f464a34-6fdf-406c-87b2-0e9a0a2d4222 · outbound

This paper cites Magnetic control of tokamak plasmas through deep reinforcement learning.

Wasserstein Policy Optimization Magnetic control of tokamak plasmas through deep reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.642253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.642253Z digest=sha256:75968e2a874d3b97e8915ae674852dc5a4eabac3dd607f2d0874e0b052f7cf2f

Observation 0495d737-0898-4b96-bce4-c7eec5c1cbae · outbound

This paper cites Experimental research on the TCV tokamak.

Wasserstein Policy Optimization Experimental research on the TCV tokamak

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.645704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.645704Z digest=sha256:865e0d3178ae818f3bec2276e7266bf9f52b02c2fb070e8b74046d1c70666022

Observation 3a3fc4cf-9920-41a2-b044-ee578f2224fb · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning.

Wasserstein Policy Optimization Sigmoid-weighted linear units for neural network function approximation in reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.649683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.649683Z digest=sha256:b18634b44154d482100e3361dfe6f5b9b0366c84d342eee1630ee244790e6324

Observation 2ec59336-bd68-4aaf-a304-ee1b46abba65 · outbound

This paper cites CALE: Continuous Arcade Learning Environment.

Wasserstein Policy Optimization CALE: Continuous Arcade Learning Environment

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:47:05.935740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.653202Z digest=sha256:17f41b04295561a137db2cd519a41179a25be802641bf9656532684574d0f316

Observation b7d4ecf8-0e0e-4cbb-b112-33cb32583843 · outbound

This paper cites Metrics for finite markov decision processes.

Wasserstein Policy Optimization Metrics for finite markov decision processes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.222698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.657534Z digest=sha256:29802b9e5cba3c5b8fd59deefc629a70c23978682dadc0ccc5aace86cca98c43

Observation 6335febe-a02a-473e-9ae8-dedb8a0a42fe · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Wasserstein Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.661528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.661528Z digest=sha256:722ffef755437c86b764aee1a2f7bd9011cad72a61904bb9e354a21d10420e33

Observation 27536610-f19e-42b6-8a56-2f38f5d39127 · outbound

This paper cites H., Tirumala, D., Humplik, J., Wulfmeier, M., Tunyasuvunakool, S., Siegel, N.

Wasserstein Policy Optimization H., Tirumala, D., Humplik, J., Wulfmeier, M., Tunyasuvunakool, S., Siegel, N

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.202530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.666603Z digest=sha256:374331b882b71f3834deaea216c3fa5824989488243c79c58062f71365b6e65d

Observation 857f1670-9d4e-432b-99e7-5be13481fc37 · outbound

This paper cites Wasserstein Unsupervised Reinforcement Learning.

Wasserstein Policy Optimization Wasserstein Unsupervised Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:47:05.922391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.670497Z digest=sha256:e7c06f8580beacb82190d13f22b8ebccff21018a3aada34c013cd429b88bc080

Observation 744ec0f7-4171-4892-9043-dd086adef2b2 · outbound

This paper cites Learning continuous control policies by stochastic value gradients.

Wasserstein Policy Optimization Learning continuous control policies by stochastic value gradients

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.191434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.674549Z digest=sha256:399ee794528a7c3d513a5eb1a090c9c00f35a9872d9d687177382a7732624277

Observation 2fd903b9-9f74-4b6a-b4c6-bf191014cea0 · outbound

This paper cites Acme: A Research Framework for Distributed Reinforcement Learning.

Wasserstein Policy Optimization Acme: A Research Framework for Distributed Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.678159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.678159Z digest=sha256:7617831c7ee35bfc3cd4cd7df8c1b0cab35bb4f7687174d152336d85478d1b37

Observation 0395637c-0a5c-4c90-9a66-d09ac040974d · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.177293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.683628Z digest=sha256:2fa1a4b59e35094f965d4994599627df8b5148e37383b6ccdaa68da46e8544f6

Observation 45aaa500-a603-4a85-a431-74aea021f4e5 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.165323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.689320Z digest=sha256:827a6576dfd5ab6d5c0732938ec5103bb38095779770565da67754c494a7f83d

Observation b48467e3-ffd7-45c6-a333-91bda38df65b · outbound

This paper cites Categorical reparameterization with G umbel-softmax.

Wasserstein Policy Optimization Categorical reparameterization with G umbel-softmax

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.154538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.694063Z digest=sha256:fe4349d4bc3355d56a871a2ebca090e6edcb169118060cd34d1f435e19e1ef1e

Observation acc2e2a9-bed6-428e-bb47-9446929ffd68 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.142634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.698562Z digest=sha256:a303c57887bdd229f7a16d3bcbc70d15efea13511f13c8e357ee00e9ee75ceca

Observation 51a4ad7c-dfc5-4b11-af34-2432dd3181ce · outbound

This paper cites Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions Control.

Wasserstein Policy Optimization Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions Control

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:47:05.898230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.702822Z digest=sha256:582d0550a8a25661e79e68ba530688c6bd29a655c80f3a73fed23951741512c4

Observation 683d7e2a-d004-4fca-a499-b7fcbeebd1f1 · outbound

This paper cites Continuous control with deep reinforcement learning.

Wasserstein Policy Optimization Continuous control with deep reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.706926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.706926Z digest=sha256:f16e6d668d4e859a391331d2cede6a9ce1ea65ccc490fa7a68a54dd0bfab0682

Observation fa09afc5-6421-442b-89e4-cf543840fdda · outbound

This paper cites J., Mnih, A., and Teh, Y.

Wasserstein Policy Optimization J., Mnih, A., and Teh, Y

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.130226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.711300Z digest=sha256:6544754bab72ff67ec1313341a43999301ddd2417cfecbfcc4ca448df3b038cf

Observation 9850b65c-b8ab-4482-8108-b8e4a356824a · outbound

This paper cites and Grosse, R.

Wasserstein Policy Optimization and Grosse, R

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.118288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.715064Z digest=sha256:ffc0197d02fc1494dcff3e90c082fd935c8102098af67fe37c3a818a077da383

Observation ac03dda1-012e-4034-95eb-2c6f0a8b5bb0 · outbound

This paper cites M., Likmeta, A., and Restelli, M.

Wasserstein Policy Optimization M., Likmeta, A., and Restelli, M

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.105754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.718976Z digest=sha256:87c19a279a4506867815c57f4a2a36e8724e9e9059a31b5dc8c6742705007a21

Observation e144be77-e467-494d-945b-4a65c3f7f94a · outbound

This paper cites Efficient Wasserstein Natural Gradients for Reinforcement Learning.

Wasserstein Policy Optimization Efficient Wasserstein Natural Gradients for Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.722747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.722747Z digest=sha256:e97084a28902e61c9934073003216378e59932a95f5294f944af5370401fae42

Observation afd51786-a8c3-48b1-b167-0d308880fc81 · outbound

This paper cites Wasserstein quantum M onte C arlo: a novel approach for solving the quantum many-body schr \"o dinger equation.

Wasserstein Policy Optimization Wasserstein quantum M onte C arlo: a novel approach for solving the quantum many-body schr \"o dinger equation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.094875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.727282Z digest=sha256:2abcb58458840da6a517d70770674e109d3c2417b2a2ea2e51f7a075ef70c1ff

Observation f9fe4337-2e4f-4197-ae50-0dc6f03d9957 · outbound

This paper cites Learning to score behaviors for guided policy optimization.

Wasserstein Policy Optimization Learning to score behaviors for guided policy optimization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.083773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.731386Z digest=sha256:61c2e2155802ad722deea15d61906cba37f3d131c2e542e82e92ae037f26e67b

Observation 75eb8626-fb33-486c-b607-92acf35b5274 · outbound

This paper cites Revisiting Natural Gradient for Deep Networks.

Wasserstein Policy Optimization Revisiting Natural Gradient for Deep Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.734888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.734888Z digest=sha256:239516b753831022bc06a23e3f41c9a360c8ffa11599d547760843a37cfaca1c

Observation d6f13103-9698-4474-9d76-bc0048a81deb · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.073639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.739123Z digest=sha256:d5bc65ab3eb8c5214801d0ef679bbb63c214e0a65e02b33bf67393c79fd1bc3e

Observation 800c788b-db9b-4b7d-a572-59a67930c74c · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.063889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.742610Z digest=sha256:b881a329a69a5b7f74856721659a25d603a54ba828ba495536172ac72c74bbd7

Observation 094c430b-f82f-4804-b88c-8209fbf8b329 · outbound

This paper cites Trust region policy optimization.

Wasserstein Policy Optimization Trust region policy optimization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.054524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.746388Z digest=sha256:9e90f4565ba84d939a5b246e923c4f277edf5e63735adeefb48dbf9e05fe5c4a

Observation e8ff4f7c-2665-4a90-8f0b-8d98c02f2f19 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Wasserstein Policy Optimization High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.750096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.750096Z digest=sha256:86723cf04ebf84027257780af338c50cf84f625bc973e01fd0e71e1951683a2d

Observation 368ea6c0-f820-49db-9008-38cc22728977 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Wasserstein Policy Optimization Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.753537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.753537Z digest=sha256:807996b4e4e3f8a39db06d2200e12bf362aaa2bcc9b64fa1c4bbf86a47d61e2a

Observation 07849974-2fdd-4b13-8a43-d4366ea6a636 · outbound

This paper cites Deterministic policy gradient algorithms.

Wasserstein Policy Optimization Deterministic policy gradient algorithms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.044687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.757112Z digest=sha256:bcbce0dbdb1aa93ecef2e9b769011117e38e963481d156285c2bad5afbf0272b

Observation 49344668-057b-44f7-aa5c-cc03726391f9 · outbound

This paper cites V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control.

Wasserstein Policy Optimization V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.760618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.760618Z digest=sha256:5f191670c07a56c0c0106adaf7064ce7d146b881e6b5db59cba59e61bda8eb6e

Observation 6dee8b22-00dd-4fd2-953b-4b7fa8930f26 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.763904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.763904Z digest=sha256:7306360bb275eb810b6050acbdfb52cb4167a12c56290394da86dbf921ef0771

Observation e000d2e0-7711-4541-8ec6-49ec0d647327 · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

Wasserstein Policy Optimization S., McAllester, D., Singh, S., and Mansour, Y

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.767085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.767085Z digest=sha256:013d58ca9800b63467c45cdd5b77da82644e8ca8ff846203d3ed07e0d8bb9a1b

Observation e5e35a59-79ac-48c1-99df-a0c872ebd9c8 · outbound

This paper cites DeepMind Control Suite.

Wasserstein Policy Optimization DeepMind Control Suite

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.770349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.770349Z digest=sha256:75e1e44c54faebbc833c8e107cc225a3cdf91c4d15c66eaef00e1c6cc969b88d

Observation f4975490-ef45-488d-9b33-09be25f5bf14 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Wasserstein Policy Optimization Mujoco: A physics engine for model-based control

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.773879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.773879Z digest=sha256:3a4e5135cb69f38ed103ee3e21c284a3e1ad2429f5a1b9b9c97d366cc98effbf

Observation 1b4fa37c-7da9-453d-82ff-66061c19dcde · outbound

This paper cites D., Michi, A., Chervonyi, Y., Davies, I., Paduraru, C., Lazic, N., Felici, F., Ewalds, T., Donner, C., Galperti, C., et al.

Wasserstein Policy Optimization D., Michi, A., Chervonyi, Y., Davies, I., Paduraru, C., Lazic, N., Felici, F., Ewalds, T., Donner, C., Galperti, C., et al

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.776825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.776825Z digest=sha256:fb69d43b7c6605ea358031a3d6e55d7fb08e10eeac1c2b1e7e2e43a7d32a199e

Observation 72b0774e-5686-4f6c-81fb-ee1b71650855 · outbound

This paper cites dm\_control: Software and tasks for continuous control.

Wasserstein Policy Optimization dm\_control: Software and tasks for continuous control

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.010833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.779820Z digest=sha256:bdf31e69824bcdee00c7910b38dcbf13f8995ac04cd47cf7ce5d2ff79db59b8d

Observation 76ec6004-2123-427e-af7e-1944d25931b7 · outbound

This paper cites Double q-learning.

Wasserstein Policy Optimization Double q-learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.000229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.783069Z digest=sha256:bf741ed2602c21cd0f005dc60f167f554deaef0d32417b3a32e5d73f769f3b3d

Observation 3aa6ab08-bd3a-4050-a7d9-e48c39a41b19 · outbound

This paper cites and Wiering, M.

Wasserstein Policy Optimization and Wiering, M

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:05.989811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.786087Z digest=sha256:7c17e903c20fcbb75e2b97723ba8200e75e4d9e9198dae6cfaf23c6eeb146ebd

Observation 2014b134-8925-4098-b8d1-a254e7e549b1 · outbound

This paper cites Beyond regression: N ew tools for prediction and analysis in the behavioral sciences.

Wasserstein Policy Optimization Beyond regression: N ew tools for prediction and analysis in the behavioral sciences

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:05.978635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.789226Z digest=sha256:27f547c156d3455deb5ec93112a4fb1928850f53b3c7a6d7eb3c8ce4d9ab93b2

Observation 7ec0bc46-1d29-498e-8564-414468894dd6 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.792655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.792655Z digest=sha256:d37db7299b8707ed0a0346105003d77be49b738c30065b318d28eb0d9a96b96c

Observation 8b443736-3676-4c3d-a32e-2b9135f53f77 · outbound

This paper cites Policy optimization as W asserstein gradient flows.

Wasserstein Policy Optimization Policy optimization as W asserstein gradient flows

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:05.962950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.795761Z digest=sha256:dd87b213316fb9ec58726ba019784701d625d8f0fa56eb1267fc78f65c076eb0

Pith citing papers

Observation ac0cc622-9671-4b93-b2be-8695092a6a38 · inbound

Challenges and opportunities for AI to help deliver fusion energy cites this paper.

Challenges and opportunities for AI to help deliver fusion energy Wasserstein Policy Optimization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:43:25.441685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T00:39:46.244035Z digest=sha256:1bb7a8af24158090f45d9360f61d729895e72fa1fb9a57765f86a678a5191ccc

Observation d42fbe4a-6cef-4066-bedb-7edae218c41c · inbound

A note on convergence of Wasserstein policy optimization cites this paper.

A note on convergence of Wasserstein policy optimization Wasserstein Policy Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.858631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-22T06:36:57.345741Z digest=sha256:e18783716051557e0ffce63c332d47da85606d173d84a327a94eae13bcfeab21

Observation 5238f7d7-76f5-44de-b82b-131b8f3d9832 · inbound

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning cites this paper.

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning Wasserstein Policy Optimization

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:24:00.339861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T22:19:53.934535Z digest=sha256:7f7f1fb537fb9185c9c4b288e80c11a3f80d4aa0e3f61252b61cf56a2609b035

Observation 42fb7b7b-e288-4bc2-a5d4-6af7262093e9 · inbound

Ratio-Variance Regularized Policy Optimization cites this paper.

Ratio-Variance Regularized Policy Optimization Wasserstein Policy Optimization

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.983387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:2eb004a43899db48437c3e14125ec5da99839e7f54111cefe97bc2b5b937cd0b

Observation 2aca0c6c-ab04-45a0-9c52-c2ef9587a81b · inbound

Policy Gradient for Continuous-Time Robust Markov Decision Processes cites this paper.

Policy Gradient for Continuous-Time Robust Markov Decision Processes Wasserstein Policy Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:56:44.505465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T07:18:48.673738Z digest=sha256:5b5b14b6842610e02f6ef26aa971f0c3ffd703c34c90882a9d736b186d844466