Pith. sign in

Paper Citation Record · LEDGER

StaQ it! Growing neural networks for Policy Mirror Descent

As of 10 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.13862.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13862 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:34:04.168498Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86109e89-1663-4981-a269-60bc7099eba2 · outbound

This paper cites single-task RL.

StaQ it! Growing neural networks for Policy Mirror Descent single-task RL

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.503389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:03.294608Z digest=sha256:075704eb854d460ee52167e0b72bdc0ceabb8e5f2bab53dec1abf430c713a418

Observation f7264f06-876d-4bf3-ac89-6f1a90f7d4ea · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:07.288593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:02.894245Z digest=sha256:894979607ca40d9cb8c38c674c48826315137f78108dcb75365af5381e7150e8

Observation bf13b061-6f02-4338-947f-3b3082204d4b · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:06.730051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:03.161264Z digest=sha256:3cd10aeb5525428ffa3c70dca7536e4817ba22ea2c3783cabd4d39e2109108b7

Observation 7014e099-0c24-42b9-a2c3-688f78ce2d9a · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

StaQ it! Growing neural networks for Policy Mirror Descent MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.491366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.491366Z digest=sha256:fa01e9c0e2399038098e2c356c2de37f310d9982797eaafd3b1860667c086637

Observation 13004cb4-2553-452c-87f0-0eb8c1412822 · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:07.545576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:02.732844Z digest=sha256:4de833c6590c301df5166ed46636a9629f9a0b249422c2286be01b2b24375196

Observation 8bb50f7b-3cd9-4073-b96f-4cbeae8614ef · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:06.224679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:03.429724Z digest=sha256:d1f26651eba72c860d442572a875bebf33b6768c45af1d19e92650e5b7587d87

Observation c052211d-e640-432d-8c14-0d3c16cdc075 · outbound

This paper cites For PQN, we use the CleanRL implementation (Huang et al., 2022).

StaQ it! Growing neural networks for Policy Mirror Descent For PQN, we use the CleanRL implementation (Huang et al., 2022)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.917222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:03.537326Z digest=sha256:8a7043d3927df5cd1451f78996bba118ecfbfcea6e25e10cac75014b329cc6c2

Observation 2977c183-1317-43ac-b637-925c2a76c054 · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:34:05.694862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:03.697384Z digest=sha256:5c060076d5b4ee84140a3b8a13bc3527481ac3c7b543e1df42802da17da6afef

Observation 604c68c7-509d-41e0-a758-99730384f708 · outbound

This paper cites In the MinAtar environ- ments α is linearly annealed from 1 to 0 over the course of learning.∗Humanoid-v4 uses a hidden layer size of256.

StaQ it! Growing neural networks for Policy Mirror Descent In the MinAtar environ- ments α is linearly annealed from 1 to 0 over the course of learning.∗Humanoid-v4 uses a hidden layer size of256

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.424094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:03.814872Z digest=sha256:b1442980e06acdba034d584b5048d3fd3cdf96b5cfd5cd351acb90892c9fa537

Observation e27c280b-70f5-4b1e-bb4a-d2578ef42c6b · outbound

This paper cites an unresolved cited work.

StaQ it! Growing neural networks for Policy Mirror Descent Unresolved cited work

Reference 18

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:34:05.131796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:04.036926Z digest=sha256:f5987089e77f3a7238505a758579c695ee80c7005faa04b695f6766e10c0e81f

Observation 81bef3d0-d1a3-4e91-8a59-3480fe621da4 · outbound

This paper cites Classic and MinAtar hyperparameters are based on the original paper (Gallici et al., 2025), while MuJoCo hyperparameters were found by hyperparameter tuning.

StaQ it! Growing neural networks for Policy Mirror Descent Classic and MinAtar hyperparameters are based on the original paper (Gallici et al., 2025), while MuJoCo hyperparameters were found by hyperparameter tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:04.937775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:04.168498Z digest=sha256:7bc280d5bdfbe987849f280a336c4fc9fa0f65d8a327a133d4dcccde214e5f80

Observation 0b0f0d1f-d8c6-4ac8-bb09-b376ac207a24 · outbound

This paper cites doi: 10.1016/S0167-6377(02)00231-6.

StaQ it! Growing neural networks for Policy Mirror Descent doi: 10.1016/S0167-6377(02)00231-6

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.737396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.737396Z digest=sha256:f09b641b40dee89a001a5eed4863e6fcb0ac68d07e5fe0941c4fcd97a6255679

Observation 89fea3c6-8f12-4a8f-99a4-1d118f9ff6c4 · outbound

This paper cites Mirror Descent Policy Optimization.

StaQ it! Growing neural networks for Policy Mirror Descent Mirror Descent Policy Optimization

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.274823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.274823Z digest=sha256:ba8cfff335117fd1fac0f6d18d9f43a7c89fada871f824bf8c7b443471765b7b

Observation a11d5ff7-9cf5-4c62-91ea-fcb73de79e63 · outbound

This paper cites We can see in Fig.

StaQ it! Growing neural networks for Policy Mirror Descent We can see in Fig

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:07.019219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:34:03.027368Z digest=sha256:7afc9607b61ab7aa1c11188a91a69515f05085388ee8b3a300d40f938a07b546

Observation 82fdb66c-f27d-4497-abf3-de8584454647 · outbound

This paper cites Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies.

StaQ it! Growing neural networks for Policy Mirror Descent Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.598175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.598175Z digest=sha256:d992a064c53819855de3e6f5cfb6bbfc7557e13a5ac618aecf1081afa08b0e88

Observation eb783b25-fc5b-4a52-b22e-b3a4fd72315a · outbound

This paper cites Homotopic Policy Mirror Descent: Policy Convergence, Implicit Regularization, and Improved Sample Complexity.

StaQ it! Growing neural networks for Policy Mirror Descent Homotopic Policy Mirror Descent: Policy Convergence, Implicit Regularization, and Improved Sample Complexity

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.124836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.124836Z digest=sha256:acda792d987e80d36a545724a565f9d8aa8f7c7a750ebe0c8693d47388617473

Observation 726edaa6-805d-4d80-9f87-9537bbeb5c61 · outbound

This paper cites Randomized Ensembled Double Q-Learning: Learning Fast Without a Model.

StaQ it! Growing neural networks for Policy Mirror Descent Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.863748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.863748Z digest=sha256:7e0b83ed927904c130aca6c4327c4ba59a3eaa5bde9daf180fb0ca0be2fe45ec

Observation 6f1bc58c-a35c-4fa4-a565-7dc90dcc5f57 · outbound

This paper cites Maxmin Q-learning: Controlling the Estimation Bias of Q-learning.

StaQ it! Growing neural networks for Policy Mirror Descent Maxmin Q-learning: Controlling the Estimation Bias of Q-learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.965749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.965749Z digest=sha256:8e4b95cd0b1245c1bd233bc5d6ee89bcf1f3a25394c1f9d1d6644fa3bea16b44

Observation b7c1b16f-eca0-4fbb-9424-7b73cac187dd · outbound

This paper cites van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J.

StaQ it! Growing neural networks for Policy Mirror Descent van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.354750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:02.354750Z digest=sha256:3c3804b144896f642104ff8561e5adfe4d3b8ac2c510094c69aefaedc4ed2deb

Pith citing papers

No inbound Pith citation observations are available.