Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2005.09814.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:28:42.268615Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
24
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 0d75a5a4-6ea1-4e92-b80d-73e6ea25e561 · inbound
Mirror Descent-Ascent for mean-field min-max problems Mirror Descent Policy Optimization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 86f8400b-94fb-4e00-9968-a0de2b76725d · inbound
Efficient Online Reinforcement Learning for Diffusion Policy Mirror Descent Policy Optimization
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce92ea2b-0577-448c-bb91-92038e927280 · inbound
Muon is Scalable for LLM Training Mirror Descent Policy Optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd57ddcc-8515-48d4-9d6f-ba5a74779574 · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Mirror Descent Policy Optimization
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 89fea3c6-8f12-4a8f-99a4-1d118f9ff6c4 · inbound
StaQ it! Growing neural networks for Policy Mirror Descent Mirror Descent Policy Optimization
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae7111e3-ab19-47b2-b7ab-59905970265b · inbound
Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning Mirror Descent Policy Optimization
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a8c166-7e67-471a-8e93-8ad3cc798fc8 · inbound
On the Effect of Regularization in Policy Mirror Descent Mirror Descent Policy Optimization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4273129e-9f29-4c6b-a718-a5c46792249a · inbound
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives Mirror Descent Policy Optimization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dbc5e0e8-c761-4c7a-8982-1905e36b648b · inbound
Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces Mirror Descent Policy Optimization
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dfbff5c1-c899-4947-b8b8-8983e3201c42 · inbound
Optimal Posterior Sampling for Policy Identification in Tabular Markov Decision Processes Mirror Descent Policy Optimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3e1c3ff1-7248-44ab-b112-1701486e0a21 · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Mirror Descent Policy Optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4a514623-c37f-4dcd-b26d-f8abf87dcf3d · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Mirror Descent Policy Optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b3dce0b0-8a55-46f1-ba00-1873240e3db5 · inbound
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Mirror Descent Policy Optimization
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea2350a4-db8b-4bcd-9470-3335b03c3166 · inbound
Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation Mirror Descent Policy Optimization
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3c4c513e-24fe-45ad-a19f-547b0d8f27aa · inbound
Credit Assignment with Resets in Language Model Reasoning Mirror Descent Policy Optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6e3b7b17-14c5-4708-ae37-b7cc90c6b381 · inbound
Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs Mirror Descent Policy Optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3ec84ec4-10c1-42ad-b944-123aa27f30a6 · inbound
Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs Mirror Descent Policy Optimization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e73e4d5e-b46e-4335-bc00-b2ca1e70d098 · inbound
A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer Mirror Descent Policy Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3c3ed6e-a074-449b-92bc-82328076a10b · inbound
On the Policy Convergence of Policy Mirror Descent Methods Mirror Descent Policy Optimization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.