Pith. sign in

Paper Citation Record · LEDGER

Mirror Descent Policy Optimization

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2005.09814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2005.09814 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:28:42.268615Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

24
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0d75a5a4-6ea1-4e92-b80d-73e6ea25e561 · inbound

Mirror Descent-Ascent for mean-field min-max problems cites this paper.

Mirror Descent-Ascent for mean-field min-max problems Mirror Descent Policy Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:48:53.043539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T03:47:22.184621Z digest=sha256:74876797f5d0bd8098cd58ae89766ab2ec4ee0ef0d892f1b0148933a8ca40ffb

Observation 86f8400b-94fb-4e00-9968-a0de2b76725d · inbound

Efficient Online Reinforcement Learning for Diffusion Policy cites this paper.

Efficient Online Reinforcement Learning for Diffusion Policy Mirror Descent Policy Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T19:28:42.268615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:28:42.268615Z digest=sha256:59f18e22ac3cfa9d4ecee177b41c3bff8efecea6a75215385a1c0363d76ca10c

Observation ce92ea2b-0577-448c-bb91-92038e927280 · inbound

Muon is Scalable for LLM Training cites this paper.

Muon is Scalable for LLM Training Mirror Descent Policy Optimization

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.240735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:0b15585edb10f52165f699c6df2602a0219006b1852ce8fcf1b478810773ecf4

Observation cd57ddcc-8515-48d4-9d6f-ba5a74779574 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Mirror Descent Policy Optimization

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:29:56.916211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:c8266ea861e6f2bf7388946666251e7915ae23caf23c75f78167e9e0ee078d2f

Observation 89fea3c6-8f12-4a8f-99a4-1d118f9ff6c4 · inbound

StaQ it! Growing neural networks for Policy Mirror Descent cites this paper.

StaQ it! Growing neural networks for Policy Mirror Descent Mirror Descent Policy Optimization

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.274823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.274823Z digest=sha256:ba8cfff335117fd1fac0f6d18d9f43a7c89fada871f824bf8c7b443471765b7b

Observation ae7111e3-ab19-47b2-b7ab-59905970265b · inbound

Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning cites this paper.

Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning Mirror Descent Policy Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:56:46.000303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:56:46.000303Z digest=sha256:a929996bf712f533b55cb2353909965d1145a90335d4d68a8810bcd82e687d5c

Observation c6a8c166-7e67-471a-8e93-8ad3cc798fc8 · inbound

On the Effect of Regularization in Policy Mirror Descent cites this paper.

On the Effect of Regularization in Policy Mirror Descent Mirror Descent Policy Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:47.685612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:18:47.685612Z digest=sha256:178714c53a77ad91fe9be2c54fada8d44c2deeb52b623a59a46d8cfb5851f3a3

Observation 4273129e-9f29-4c6b-a718-a5c46792249a · inbound

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives cites this paper.

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives Mirror Descent Policy Optimization

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T17:06:39.772528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T17:05:36.100114Z digest=sha256:b34b06c7977f2c80abba4871c94cdb40e407002c4783befa10b05352db133768

Observation dbc5e0e8-c761-4c7a-8982-1905e36b648b · inbound

Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces cites this paper.

Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces Mirror Descent Policy Optimization

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T21:15:38.709665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T21:14:54.053177Z digest=sha256:947272b50c999f452f009d7e23a06f4b7d95f59480615c765b6b792b9253013d

Observation dfbff5c1-c899-4947-b8b8-8983e3201c42 · inbound

Optimal Posterior Sampling for Policy Identification in Tabular Markov Decision Processes cites this paper.

Optimal Posterior Sampling for Policy Identification in Tabular Markov Decision Processes Mirror Descent Policy Optimization

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:46:14.256716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-07T16:22:51.320973Z digest=sha256:3978ffa27674f4c8273184b4c7184f343906762785631b2c8e106e30fc1fb273

Observation 3e1c3ff1-7248-44ab-b112-1701486e0a21 · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Mirror Descent Policy Optimization

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:46:10.338147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T13:55:22.422923Z digest=sha256:e24ed6beb36cffb88ad74d2148d68c037c26928b444a18a62650c93c97345e84

Observation 4a514623-c37f-4dcd-b26d-f8abf87dcf3d · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Mirror Descent Policy Optimization

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T09:04:04.439050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T09:02:28.407064Z digest=sha256:63b7a8d3b42c62b45922d123f4b662c870e55ac327717a20cb532769dd24bdc2

Observation b3dce0b0-8a55-46f1-ba00-1873240e3db5 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Mirror Descent Policy Optimization

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.271610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:6e37f70736a4cf0e22d70987b521b338da93ee033da858d863a932f2f9e7b3a8

Observation ea2350a4-db8b-4bcd-9470-3335b03c3166 · inbound

Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation cites this paper.

Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation Mirror Descent Policy Optimization

Reference 137

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:48:17.524423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T12:44:29.147095Z digest=sha256:125bcd9fa88fcee4cbb30802a136c9ca3d5239abf6088b4fe7522731fd28fcc7

Observation 3c4c513e-24fe-45ad-a19f-547b0d8f27aa · inbound

Credit Assignment with Resets in Language Model Reasoning cites this paper.

Credit Assignment with Resets in Language Model Reasoning Mirror Descent Policy Optimization

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T21:53:59.238296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T21:50:18.827822Z digest=sha256:51efb6e9269792e0c1626b3e5e9e3e744e5ff88614454fdf6e7479f2a4943bc9

Observation 6e3b7b17-14c5-4708-ae37-b7cc90c6b381 · inbound

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs cites this paper.

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs Mirror Descent Policy Optimization

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:24:32.086752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T09:24:21.047027Z digest=sha256:89aad4c11fa22ff92acc39fd7f0f1b87b57d10d81ee066021e851cb7748f065b

Observation 3ec84ec4-10c1-42ad-b944-123aa27f30a6 · inbound

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs cites this paper.

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs Mirror Descent Policy Optimization

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:24:31.615265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T09:24:21.047027Z digest=sha256:4af00ad6f55ad19fd324e76493cd8835b6cca6c07b02e3f120f76ff6618bc453

Observation e73e4d5e-b46e-4335-bc00-b2ca1e70d098 · inbound

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer cites this paper.

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer Mirror Descent Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T08:30:10.584105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:30:10.584105Z digest=sha256:235ca0658fd97ba44648044409c52c18a251661180761dacec8e7d368dcd70a2

Observation d3c3ed6e-a074-449b-92bc-82328076a10b · inbound

On the Policy Convergence of Policy Mirror Descent Methods cites this paper.

On the Policy Convergence of Policy Mirror Descent Methods Mirror Descent Policy Optimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:578baf31fa4384c4915d1f0f3eb73570af6cd211bc796f797b732c4aa4402105