Pith. sign in

Paper Citation Record · LEDGER

StaQ it! Growing neural networks for Policy Mirror Descent

As of 12 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.13862.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13862 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:34:04.168498Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86109e89-1663-4981-a269-60bc7099eba2 · outbound

This paper cites single-task RL.

StaQ it! Growing neural networks for Policy Mirror Descent single-task RL

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.503389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:03.294608Z digest=sha256:464e8eed877cab115b2f1b15e33786f92baaff757330aee5080db4c13e4a4759

Observation f7264f06-876d-4bf3-ac89-6f1a90f7d4ea · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:07.288593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:02.894245Z digest=sha256:e6e763439775653e05a2dcb16af2bedcd69f9091919fdadc3e1d17e4683ecf2f

Observation bf13b061-6f02-4338-947f-3b3082204d4b · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:06.730051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:03.161264Z digest=sha256:8be913c7dc57b8910e7d7657d65bb06b7e6ed9da2246ca1e62c0a6c19f509382

Observation 7014e099-0c24-42b9-a2c3-688f78ce2d9a · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

StaQ it! Growing neural networks for Policy Mirror Descent MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.491366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.491366Z digest=sha256:f7484c893988974c51a49e740952e9d4484635868434be5faa4460df954b21e9

Observation 13004cb4-2553-452c-87f0-0eb8c1412822 · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:07.545576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:02.732844Z digest=sha256:fc4a617815cc9c0ac2ee9bf97572533b7d0ac7dfed53fc858214a7ecec83946a

Observation 8bb50f7b-3cd9-4073-b96f-4cbeae8614ef · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:06.224679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:03.429724Z digest=sha256:ba8a8371f28965ed0cd1a52ef1b7a99b95a4cc59180af9f88955f290781b8d25

Observation c052211d-e640-432d-8c14-0d3c16cdc075 · outbound

This paper cites For PQN, we use the CleanRL implementation (Huang et al., 2022).

StaQ it! Growing neural networks for Policy Mirror Descent For PQN, we use the CleanRL implementation (Huang et al., 2022)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.917222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:03.537326Z digest=sha256:459c371261c103539756ffee8214455af365069c4c3fc2193f811c106648bcc7

Observation 2977c183-1317-43ac-b637-925c2a76c054 · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:05.694862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:03.697384Z digest=sha256:a9b0ce149731bd2e191d68390b6dd791b27b8848143e462f19fef64ee23fef2f

Observation 604c68c7-509d-41e0-a758-99730384f708 · outbound

This paper cites In the MinAtar environ- ments α is linearly annealed from 1 to 0 over the course of learning.∗Humanoid-v4 uses a hidden layer size of256.

StaQ it! Growing neural networks for Policy Mirror Descent In the MinAtar environ- ments α is linearly annealed from 1 to 0 over the course of learning.∗Humanoid-v4 uses a hidden layer size of256

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.424094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:03.814872Z digest=sha256:c29c9ae2eefaa08ba477ecf8322a8c46d4a63ba1943bd214a9a9df23e5140f87

Observation e27c280b-70f5-4b1e-bb4a-d2578ef42c6b · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 18

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:34:05.131796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:04.036926Z digest=sha256:eba900b3ee52feb3a3474f552507ef01597f307f32100764f08bad9f6f598032

Observation 81bef3d0-d1a3-4e91-8a59-3480fe621da4 · outbound

This paper cites Classic and MinAtar hyperparameters are based on the original paper (Gallici et al., 2025), while MuJoCo hyperparameters were found by hyperparameter tuning.

StaQ it! Growing neural networks for Policy Mirror Descent Classic and MinAtar hyperparameters are based on the original paper (Gallici et al., 2025), while MuJoCo hyperparameters were found by hyperparameter tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:04.937775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:04.168498Z digest=sha256:bfed264e509656230f306f37326b0c21b18258db121f7dd0dd6bb861167d1db4

Observation 0b0f0d1f-d8c6-4ac8-bb09-b376ac207a24 · outbound

This paper cites doi: 10.1016/S0167-6377(02)00231-6.

StaQ it! Growing neural networks for Policy Mirror Descent doi: 10.1016/S0167-6377(02)00231-6

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.737396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.737396Z digest=sha256:076eddc08d25ce07adc717bcce25a8180f34098e4dc28417c8610659b183e15b

Observation 89fea3c6-8f12-4a8f-99a4-1d118f9ff6c4 · outbound

This paper cites Mirror Descent Policy Optimization.

StaQ it! Growing neural networks for Policy Mirror Descent Mirror Descent Policy Optimization

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.274823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.274823Z digest=sha256:4cd5ad9baca40614635d1a1f8e8d2d0f059ab6391c3ee061fb832635c282d0c6

Observation a11d5ff7-9cf5-4c62-91ea-fcb73de79e63 · outbound

This paper cites We can see in Fig.

StaQ it! Growing neural networks for Policy Mirror Descent We can see in Fig

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:07.019219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:34:03.027368Z digest=sha256:1fa09f216736708e3f70b191f708fbe069973c8783389a7d2ba2a8deb4aa9a10

Observation 82fdb66c-f27d-4497-abf3-de8584454647 · outbound

This paper cites Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies.

StaQ it! Growing neural networks for Policy Mirror Descent Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.598175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.598175Z digest=sha256:4723f5cec3fd2040ffb6da11a0e68bce8d047de857c272b2e6de81e87eeb22cc

Observation eb783b25-fc5b-4a52-b22e-b3a4fd72315a · outbound

This paper cites Homotopic Policy Mirror Descent: Policy Convergence, Implicit Regularization, and Improved Sample Complexity.

StaQ it! Growing neural networks for Policy Mirror Descent Homotopic Policy Mirror Descent: Policy Convergence, Implicit Regularization, and Improved Sample Complexity

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.124836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.124836Z digest=sha256:fe9029fbd3928a29657423d0c29364d18e85f26b2b40f7e66bab8aeff25d3b78

Observation 726edaa6-805d-4d80-9f87-9537bbeb5c61 · outbound

This paper cites Randomized Ensembled Double Q-Learning: Learning Fast Without a Model.

StaQ it! Growing neural networks for Policy Mirror Descent Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.863748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.863748Z digest=sha256:5ebf5203bf2c31cc0b8fc7daedd0c8eed54d82da1046725a39d090f93d5258a4

Observation 6f1bc58c-a35c-4fa4-a565-7dc90dcc5f57 · outbound

This paper cites Maxmin Q-learning: Controlling the Estimation Bias of Q-learning.

StaQ it! Growing neural networks for Policy Mirror Descent Maxmin Q-learning: Controlling the Estimation Bias of Q-learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.965749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.965749Z digest=sha256:6393e40d6876f485bb6f52062aee8f28fa5e4b0a3dce1e97a46184062f45b3fc

Observation b7c1b16f-eca0-4fbb-9424-7b73cac187dd · outbound

This paper cites van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J.

StaQ it! Growing neural networks for Policy Mirror Descent van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.354750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.354750Z digest=sha256:86657fd250ada4c637794feb9a207954fe46313670b1a8a87e3348c26310af5f

Pith citing papers

No inbound Pith citation observations are available.