Pith. sign in

Paper Citation Record · LEDGER

Wasserstein Policy Optimization

As of 17 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 5 inbound Pith citation observations for arXiv:2505.00663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00663 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:47:05.795761Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:19:53.934535Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:56:44.503712Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy22
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5b7dfde-d8ef-4053-a00f-f73925c912b2 · outbound

This paper cites write newline.

Wasserstein Policy Optimization write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.588081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.588081Z digest=sha256:31a5807d450a294e3ebe39c379c571808823316ae38967fe64ea559e6549d55e

Observation 2254d3da-e420-4c79-b6aa-9e78c980eec9 · outbound

This paper cites T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M.

Wasserstein Policy Optimization T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.331926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.593154Z digest=sha256:7be0dad50c4df7ebe1a9f09e0921ddd1b57faa3572e63cb1c358fe3b3efd71c9

Observation 5ab9bc6b-cab4-4dbc-b48b-8645b4b1b909 · outbound

This paper cites Wasserstein Robust Reinforcement Learning.

Wasserstein Policy Optimization Wasserstein Robust Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.597189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.597189Z digest=sha256:47f6b4e89330b66d1caf2db2c22f8d2b62b821bce0eaecba1a46033cf2804893

Observation 733ba229-d450-4089-88ae-ba561c450cbc · outbound

This paper cites M., Lee, J.

Wasserstein Policy Optimization M., Lee, J

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.601296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.601296Z digest=sha256:e33c75c873c4099bcf4c0928328eb11f32bfa741bc9052571ec4fea277084fdb

Observation d34dd838-1579-4ca8-aba0-b935f4832c21 · outbound

This paper cites Gradient flows: in metric spaces and in the space of probability measures.

Wasserstein Policy Optimization Gradient flows: in metric spaces and in the space of probability measures

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.604877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.604877Z digest=sha256:2592f9e0b78e998c4ec66a45937c6fa441c5b817f1c7532351eab79601a10be2

Observation 56df3e18-0b0e-4184-b5d8-f649980ff82f · outbound

This paper cites W., Budden, D., Dabney, W., Horgan, D., Tb, D., Muldal, A., Heess, N., and Lillicrap, T.

Wasserstein Policy Optimization W., Budden, D., Dabney, W., Horgan, D., Tb, D., Muldal, A., Heess, N., and Lillicrap, T

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.306496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.608466Z digest=sha256:c0a7d6db123455eb2ff84ac1910925df319292db9512b9c237d5e05c47896dbe

Observation 41dad690-926f-4135-b459-184cc07b59cb · outbound

This paper cites G., Sutton, R.

Wasserstein Policy Optimization G., Sutton, R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.294957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.612357Z digest=sha256:245c32b18fc89df4adcda12039537849f666264fd4356ddbf2ac7597589256df

Observation c109af0b-e5a3-41d2-a497-fa0e246179c5 · outbound

This paper cites G., Dabney, W., and Munos, R.

Wasserstein Policy Optimization G., Dabney, W., and Munos, R

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.283686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.616220Z digest=sha256:f243774398e468b0a10f8d89bb01b168397c061c92bee3d4318841ce68d75249

Observation 9662de86-3d7e-48cc-930a-e1e12a352561 · outbound

This paper cites and Brenier, Y.

Wasserstein Policy Optimization and Brenier, Y

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.619989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.619989Z digest=sha256:fbf64cfaa309176d700844a8d0dcecb20bb836484b9ba14dad5acdbfe5de40e9

Observation d8ac57cd-5b35-4274-82fb-516c7af50691 · outbound

This paper cites Development of free-boundary equilibrium and transport solvers for simulation and real-time interpretation of tokamak experiments.

Wasserstein Policy Optimization Development of free-boundary equilibrium and transport solvers for simulation and real-time interpretation of tokamak experiments

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.266647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.624192Z digest=sha256:b63b7711aae143b29907fcfe15cca7f3023adf8cdfbb6227fb8a0b6ed1db5982

Observation 59307d2d-0cea-46e5-806f-be758d7ddaf0 · outbound

This paper cites MICo: Improved representations via sampling-based state similarity for Markov decision processes.

Wasserstein Policy Optimization MICo: Improved representations via sampling-based state similarity for Markov decision processes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.628603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.628603Z digest=sha256:1a8a471bcc8d01a2eeb4c756429b47d650a89834f0a8d56c8856a4635fd2e41a

Observation 5b4cddc1-b794-4c27-9d69-bffcd7d184c8 · outbound

This paper cites T., Rubanova, Y., Bettencourt, J., and Duvenaud, D.

Wasserstein Policy Optimization T., Rubanova, Y., Bettencourt, J., and Duvenaud, D

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.634035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.634035Z digest=sha256:c44f678c499b130932557cb65bbd971bcfe28cec98df05e7c092c90371ad787f

Observation c38ecf22-afac-4732-8aba-d3fcf077b551 · outbound

This paper cites Fast and accurate deep network learning by exponential linear units (elus).

Wasserstein Policy Optimization Fast and accurate deep network learning by exponential linear units (elus)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.248832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.638728Z digest=sha256:fda1d683d0c0ce358406fe5e169c4b854a1bbaee1a074261096afe65c6be5580

Observation 6f464a34-6fdf-406c-87b2-0e9a0a2d4222 · outbound

This paper cites Magnetic control of tokamak plasmas through deep reinforcement learning.

Wasserstein Policy Optimization Magnetic control of tokamak plasmas through deep reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.642253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.642253Z digest=sha256:075cbb264fd4aa4cbd65577309863240a5e4528bb106ba51a16188d2c6b55544

Observation 0495d737-0898-4b96-bce4-c7eec5c1cbae · outbound

This paper cites Experimental research on the TCV tokamak.

Wasserstein Policy Optimization Experimental research on the TCV tokamak

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.645704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.645704Z digest=sha256:86a9edf1fbbb221bf55f49e92131592766907ddc6da2dae679e7d05a5fdddef9

Observation 3a3fc4cf-9920-41a2-b044-ee578f2224fb · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning.

Wasserstein Policy Optimization Sigmoid-weighted linear units for neural network function approximation in reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.649683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.649683Z digest=sha256:0a8a8a51224101013f60914d98d020b445806b2dac0ac3ba630f658ea366c84c

Observation 2ec59336-bd68-4aaf-a304-ee1b46abba65 · outbound

This paper cites CALE: Continuous Arcade Learning Environment.

Wasserstein Policy Optimization CALE: Continuous Arcade Learning Environment

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:47:05.935740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.653202Z digest=sha256:564009eeecddb5ee4d60490eb411cd1f01b39115cabe577e7f2d2518b686c224

Observation b7d4ecf8-0e0e-4cbb-b112-33cb32583843 · outbound

This paper cites Metrics for finite markov decision processes.

Wasserstein Policy Optimization Metrics for finite markov decision processes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.222698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.657534Z digest=sha256:5590918ce3f7449bfd4e3aed718e769c2da0f4f07f6b89b194ceda7fe6a7fa14

Observation 6335febe-a02a-473e-9ae8-dedb8a0a42fe · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Wasserstein Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.661528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.661528Z digest=sha256:cd61898c496cf1963df534e2746e52023acf474a4d576e3a2300f9107143aae0

Observation 27536610-f19e-42b6-8a56-2f38f5d39127 · outbound

This paper cites H., Tirumala, D., Humplik, J., Wulfmeier, M., Tunyasuvunakool, S., Siegel, N.

Wasserstein Policy Optimization H., Tirumala, D., Humplik, J., Wulfmeier, M., Tunyasuvunakool, S., Siegel, N

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.202530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.666603Z digest=sha256:fc4ff6f46c49ad785c99ec818631cc2731408a45557fe0fc8bd961a64e3421de

Observation 857f1670-9d4e-432b-99e7-5be13481fc37 · outbound

This paper cites Wasserstein Unsupervised Reinforcement Learning.

Wasserstein Policy Optimization Wasserstein Unsupervised Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:47:05.922391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.670497Z digest=sha256:813393542c28d0caa7205c8032341a80d3bdbda3c8934f369a87f05c75691ee4

Observation 744ec0f7-4171-4892-9043-dd086adef2b2 · outbound

This paper cites Learning continuous control policies by stochastic value gradients.

Wasserstein Policy Optimization Learning continuous control policies by stochastic value gradients

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.191434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.674549Z digest=sha256:7e80b23610e34c3df0ab498d88698e2aa8690b2e358abf67aedc4db1265e8624

Observation 2fd903b9-9f74-4b6a-b4c6-bf191014cea0 · outbound

This paper cites Acme: A Research Framework for Distributed Reinforcement Learning.

Wasserstein Policy Optimization Acme: A Research Framework for Distributed Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.678159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.678159Z digest=sha256:0aa69bb002d832b4569654f08c0e9fb81b723ef4f4c68ef5a6435c92b74cc93c

Observation 0395637c-0a5c-4c90-9a66-d09ac040974d · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.177293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.683628Z digest=sha256:ffa7578fa5e0eced01aa0552ee5229adedfa7b229adc562edd9667d05efb3f0b

Observation 45aaa500-a603-4a85-a431-74aea021f4e5 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.165323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.689320Z digest=sha256:40861427b03bd3a6e59ee69efe1076290710e4acb4051e6ecd95e643eb1cc1af

Observation b48467e3-ffd7-45c6-a333-91bda38df65b · outbound

This paper cites Categorical reparameterization with G umbel-softmax.

Wasserstein Policy Optimization Categorical reparameterization with G umbel-softmax

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.154538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.694063Z digest=sha256:745efa15f2ddd2417a60ea77914acffc99990db2bd35194957012c6f3ab301fa

Observation acc2e2a9-bed6-428e-bb47-9446929ffd68 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.142634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.698562Z digest=sha256:1c7d61b1688bb536ecb903134e7af6dce4007c9e207ddcb7cecba17733db987e

Observation 51a4ad7c-dfc5-4b11-af34-2432dd3181ce · outbound

This paper cites Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions Control.

Wasserstein Policy Optimization Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions Control

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:47:05.898230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.702822Z digest=sha256:87458475c05cff97636848ae96a3928c20ea32fbc6176f00ceb2bec65c6005db

Observation 683d7e2a-d004-4fca-a499-b7fcbeebd1f1 · outbound

This paper cites Continuous control with deep reinforcement learning.

Wasserstein Policy Optimization Continuous control with deep reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.706926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.706926Z digest=sha256:36c60ae5ec13ccf5d37e10c96c985be3c42a84b00b8b5767f94013488a4f8eac

Observation fa09afc5-6421-442b-89e4-cf543840fdda · outbound

This paper cites J., Mnih, A., and Teh, Y.

Wasserstein Policy Optimization J., Mnih, A., and Teh, Y

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.130226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.711300Z digest=sha256:0a2d35d46199a09d62c2cb1e1a9993408905f59e5af30d99737760470a33fc87

Observation 9850b65c-b8ab-4482-8108-b8e4a356824a · outbound

This paper cites and Grosse, R.

Wasserstein Policy Optimization and Grosse, R

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.118288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.715064Z digest=sha256:fcb0ece7f1ba3fd56209e413c080aae7c2b435074f526884e4d18cb0f462285d

Observation ac03dda1-012e-4034-95eb-2c6f0a8b5bb0 · outbound

This paper cites M., Likmeta, A., and Restelli, M.

Wasserstein Policy Optimization M., Likmeta, A., and Restelli, M

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.105754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.718976Z digest=sha256:881a60af22ed839578ffe21b97285d68c60f05eb785106b0a475d6a9525c1059

Observation e144be77-e467-494d-945b-4a65c3f7f94a · outbound

This paper cites Efficient Wasserstein Natural Gradients for Reinforcement Learning.

Wasserstein Policy Optimization Efficient Wasserstein Natural Gradients for Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.722747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.722747Z digest=sha256:2a57a2b68700fb4a9c535483eee68b6c300fbc3d1307d43e547bfa4ea67ba491

Observation afd51786-a8c3-48b1-b167-0d308880fc81 · outbound

This paper cites Wasserstein quantum M onte C arlo: a novel approach for solving the quantum many-body schr \"o dinger equation.

Wasserstein Policy Optimization Wasserstein quantum M onte C arlo: a novel approach for solving the quantum many-body schr \"o dinger equation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.094875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.727282Z digest=sha256:9cf0fccd8bb789d32e5bd16fc3fd56697692aeb1f2ce10fc3742a186d402d7aa

Observation f9fe4337-2e4f-4197-ae50-0dc6f03d9957 · outbound

This paper cites Learning to score behaviors for guided policy optimization.

Wasserstein Policy Optimization Learning to score behaviors for guided policy optimization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.083773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.731386Z digest=sha256:f8523f68e5c1ef2468fd38a00d587c3d2841291fe072a8675d5f492b15109c77

Observation 75eb8626-fb33-486c-b607-92acf35b5274 · outbound

This paper cites Revisiting Natural Gradient for Deep Networks.

Wasserstein Policy Optimization Revisiting Natural Gradient for Deep Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.734888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.734888Z digest=sha256:7c26519502d6bfd3ff21d8f42ba95f129a3bae0870755d4ecb60fe7c59fe5b03

Observation d6f13103-9698-4474-9d76-bc0048a81deb · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.073639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.739123Z digest=sha256:513b2d024dd1c8bf874f5e78b1087fd1a92bf306894a8a9ffaf5269919073076

Observation 800c788b-db9b-4b7d-a572-59a67930c74c · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.063889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.742610Z digest=sha256:0ecdbc95d2eb4a69475e6e60e024b4479bddadb76fad8b7b5364b13fc520b61a

Observation 094c430b-f82f-4804-b88c-8209fbf8b329 · outbound

This paper cites Trust region policy optimization.

Wasserstein Policy Optimization Trust region policy optimization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.054524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.746388Z digest=sha256:f6db09369e64195ea3ea4e0235dc0281d7d413a39863b8881090b3d6a25c1f15

Observation e8ff4f7c-2665-4a90-8f0b-8d98c02f2f19 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Wasserstein Policy Optimization High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.750096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.750096Z digest=sha256:442060622ac6b532a2725e60babf0557924c35e64d11fea9d3680da276e50dee

Observation 368ea6c0-f820-49db-9008-38cc22728977 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Wasserstein Policy Optimization Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.753537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.753537Z digest=sha256:95122a310b80b7080e0c7e803e71bc55b2bb71375908308a3c475dc271209df9

Observation 07849974-2fdd-4b13-8a43-d4366ea6a636 · outbound

This paper cites Deterministic policy gradient algorithms.

Wasserstein Policy Optimization Deterministic policy gradient algorithms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.044687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.757112Z digest=sha256:6b2b005075f0218b0eef144613dd486b6c3988b99b0bbf3f383dcc9040e7df1b

Observation 49344668-057b-44f7-aa5c-cc03726391f9 · outbound

This paper cites V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control.

Wasserstein Policy Optimization V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.760618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.760618Z digest=sha256:b24229f77b9121312bf94793ddbf30e90e4e306963eb91426daa4d3aa524e4f7

Observation 6dee8b22-00dd-4fd2-953b-4b7fa8930f26 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.763904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.763904Z digest=sha256:7fda28edb4a3ac4217c01aa50cb4dfe85f55c4f6f6eba12d47b648d09cfb1f93

Observation e000d2e0-7711-4541-8ec6-49ec0d647327 · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

Wasserstein Policy Optimization S., McAllester, D., Singh, S., and Mansour, Y

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.767085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.767085Z digest=sha256:726b90db86d51ec652e9be6d867254e60d0ffd95afc8d19117266f64be42f810

Observation e5e35a59-79ac-48c1-99df-a0c872ebd9c8 · outbound

This paper cites DeepMind Control Suite.

Wasserstein Policy Optimization DeepMind Control Suite

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.770349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.770349Z digest=sha256:5c7499e3d439634c3ea3be9be9a196a326500f7c35174b9b9e07d594e7b09afb

Observation f4975490-ef45-488d-9b33-09be25f5bf14 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Wasserstein Policy Optimization Mujoco: A physics engine for model-based control

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.773879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.773879Z digest=sha256:c61925474e5058bbb8a39792ba2d803e6a81eab25ed4d2ea7af83f067144527e

Observation 1b4fa37c-7da9-453d-82ff-66061c19dcde · outbound

This paper cites D., Michi, A., Chervonyi, Y., Davies, I., Paduraru, C., Lazic, N., Felici, F., Ewalds, T., Donner, C., Galperti, C., et al.

Wasserstein Policy Optimization D., Michi, A., Chervonyi, Y., Davies, I., Paduraru, C., Lazic, N., Felici, F., Ewalds, T., Donner, C., Galperti, C., et al

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.776825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.776825Z digest=sha256:e88d88de5f7b2beff6b3b10ef8cf84535722b898df5085fa248451af46f2cebb

Observation 72b0774e-5686-4f6c-81fb-ee1b71650855 · outbound

This paper cites dm\_control: Software and tasks for continuous control.

Wasserstein Policy Optimization dm\_control: Software and tasks for continuous control

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.010833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.779820Z digest=sha256:022ebd6eb8ae83d42f0e3c5f26f9cd82df6cda4958477f47c54caf62830caae6

Observation 76ec6004-2123-427e-af7e-1944d25931b7 · outbound

This paper cites Double q-learning.

Wasserstein Policy Optimization Double q-learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.000229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.783069Z digest=sha256:5465713d69c5575b9296408e68d061a2e0ee4ea89dca3b118312118695de2d58

Observation 3aa6ab08-bd3a-4050-a7d9-e48c39a41b19 · outbound

This paper cites and Wiering, M.

Wasserstein Policy Optimization and Wiering, M

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:05.989811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.786087Z digest=sha256:1157b50960007aed0774511a981ccb45858f21c5a9a049f616bd860d0a8eb22d

Observation 2014b134-8925-4098-b8d1-a254e7e549b1 · outbound

This paper cites Beyond regression: N ew tools for prediction and analysis in the behavioral sciences.

Wasserstein Policy Optimization Beyond regression: N ew tools for prediction and analysis in the behavioral sciences

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:05.978635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.789226Z digest=sha256:23d26130432250d5ce26d1a3fc9132e02fad006a7d62f1e477133b1aa25b60b7

Observation 7ec0bc46-1d29-498e-8564-414468894dd6 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.792655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.792655Z digest=sha256:f37e9aa1ecd180dc11b3327b98059496e2c5f04992e4a11a9d35efa191be1287

Observation 8b443736-3676-4c3d-a32e-2b9135f53f77 · outbound

This paper cites Policy optimization as W asserstein gradient flows.

Wasserstein Policy Optimization Policy optimization as W asserstein gradient flows

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:05.962950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.795761Z digest=sha256:7cab217bdcd885ace82e8315d482a1c8e39d2f9de67a5a0900957f454e4c6cd3

Pith citing papers

Observation ac0cc622-9671-4b93-b2be-8695092a6a38 · inbound

Challenges and opportunities for AI to help deliver fusion energy cites this paper.

Challenges and opportunities for AI to help deliver fusion energy Wasserstein Policy Optimization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:43:25.441685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T00:39:46.244035Z digest=sha256:13ae4f7b030b0fe7af363533ff36a05b5299af365af849a49e240cafa29ed00f

Observation d42fbe4a-6cef-4066-bedb-7edae218c41c · inbound

A note on convergence of Wasserstein policy optimization cites this paper.

A note on convergence of Wasserstein policy optimization Wasserstein Policy Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.858631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T06:36:57.345741Z digest=sha256:560e7a1cea2ca2af35d872e486c17b64352cfd0903f4dc1d664666ad7c67e2de

Observation 5238f7d7-76f5-44de-b82b-131b8f3d9832 · inbound

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning cites this paper.

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning Wasserstein Policy Optimization

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:24:00.339861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T22:19:53.934535Z digest=sha256:65bb852249fe713b296b20b443cb313ef5db5d3bffa90665a4347d9c8a1ac2f0

Observation 42fb7b7b-e288-4bc2-a5d4-6af7262093e9 · inbound

Ratio-Variance Regularized Policy Optimization cites this paper.

Ratio-Variance Regularized Policy Optimization Wasserstein Policy Optimization

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.983387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:63dbc68a623cb21f41dc303ce522cc78ca6241623cf9c7a8cb1b8daf01b5ddae

Observation 2aca0c6c-ab04-45a0-9c52-c2ef9587a81b · inbound

Policy Gradient for Continuous-Time Robust Markov Decision Processes cites this paper.

Policy Gradient for Continuous-Time Robust Markov Decision Processes Wasserstein Policy Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:56:44.505465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T07:18:48.673738Z digest=sha256:830aea477db5aa03d502685136e4098f0bc6bfb841c2d23732ef72094f911357