Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:04.170613Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2504.16415.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:04.170613Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 653de7a0-e14a-4d30-a844-79114d016b31 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca2878ab-e549-467a-8b06-6bd76f9bb7c8 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL M., Lee, J
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 93917d66-04eb-49ce-8e98-a9e3f8e805b4 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL U., and Aggarwal, V
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c4622e78-580e-4ef6-bc66-5a57b67e4a53 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL First-order methods in optimization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8b4b070-4b30-47ad-8c5e-97c0be17014c · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Stochastic multi-armed-bandit problem with non-stationary rewards
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dcaceccf-a89b-488c-aa53-53760d493ceb · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL S., Ghavamzadeh, M., and Lee, M
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e11a1424-fac8-4ed4-af34-327cebf3369a · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 94b9d0e9-78b0-4e12-881a-09704368f686 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Fast global convergence of natural policy gradient methods with entropy regularization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b8489771-fe9f-4d25-94fe-812dc81e94fa · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Optimizing for the future in non-stationary mdps
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1676c12f-fb0f-4024-b1f2-2becc8f34351 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Stabilizing reinforcement learning in dynamic environment with application to online recommendation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4ac1d03a-6728-4324-918d-4e140bdeb12d · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL and Zhao, L
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3f1ef721-8f11-483f-853b-a48936e43dce · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 85046599-b60b-49c4-ba6e-2e3e49c817e6 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 24689d8c-b7d8-492a-821f-96c1c02e42a6 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ef86e6ac-8325-4271-913e-b48613ffe31c · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL A kernel-based approach to non-stationary reinforcement learning in metric spaces
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 207f88a5-a57e-4d80-8411-061ad28e0c8c · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL M., and Mansour, Y
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c201b2fe-8135-476a-891f-74025763357b · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Dynamic regret of policy optimization in non-stationary environments
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 18ed8316-7435-4fe5-968f-26ddfac355b2 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Non-stationary reinforcement learning under general function approximation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2cd4f8d5-c7b2-4f18-9148-9aa8005410d1 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL A sliding-window algorithm for markov decision processes with arbitrarily changing rewards and transitions, 2018
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a1f95f88-4cb7-4d29-bcea-79402c0638f2 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL On Upper-Confidence Bound Policies for Non-Stationary Bandit Problems
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7444774b-c490-49f7-aaaf-9a767560eb1f · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Transient Non-Stationarity and Generalisation in Deep Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a783cc-88cd-4381-84ea-1d15b75b8b8c · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Near-optimal regret bounds for reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a950b67b-098c-4223-82f4-d8ccde3256a1 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Efficient reinforcement learning for routing jobs in heterogeneous queueing systems
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 94ac8006-84f1-4ffe-89ca-0fdd2be74acf · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6615cda7-1c04-45a2-b2d9-8ba0e1b443be · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL and Qian, P
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eeeb929f-f14a-4ae9-ae11-9f6b52eda3b4 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Towards continual reinforcement learning: A review and perspectives
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec2f33a2-8e55-4d43-8873-7adaaf4d9deb · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL R., Varma, S
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f71f109a-8aa0-4ed3-a96a-09d665fa23b5 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL T., Romberg, J., and Maguluri, S
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2923654e-62a3-433f-a4e0-a7a9239c85be · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c939c320-c385-424b-a86d-8883a306e65f · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Improved regret bound and experience replay in regularized policy iteration
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3e2a53e9-02f4-47ac-a315-6d51d8646982 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL and Rachelson, E
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b373027a-0703-405c-88fe-b887980b8a9c · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Pausing policy learning in non-stationary reinforcement learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1c993825-9e01-4eda-aa3b-696f9a48ed80 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Rl-qn: A reinforcement learning framework for optimal control of queueing systems
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7bc2e48e-317b-48d7-82ac-eef77431d013 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL A Definition of Non-Stationary Bandits
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96129abb-a2ab-488b-9a3a-a86e9ec7b14e · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Nonstationary bandit learning via predictive sampling
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fa0791ca-66de-4294-b7d7-95b4148e8c84 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Average reward reinforcement learning: Foundations, algorithms, and empirical results
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3b38bb50-96f6-4444-9da8-efc50c204f34 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Model-free nonstationary reinforcement learning: Near-optimal regret and applications in multiagent reinforcement learning and inventory control
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7c4e1aa2-4ae1-4cfd-8133-12f97bc87932 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL New insights and perspectives on the natural gradient method
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f9fd8b3d-866e-48f3-aea2-a2e05a6eed6d · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 64f089b1-aab5-4b53-a9fe-86759e30c5f5 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL and Srikant, R
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d8631203-b9e7-4f10-ae79-c455b2175960 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Performance bounds for policy-based average reward reinforcement learning algorithms
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e13e4eda-4d78-4a9e-90b6-7f579e8d7dd8 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Bridging the gap between value and policy based reinforcement learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1c888d79-729b-481e-af21-e2e8af216279 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Variational regret bounds for reinforcement learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0c74fd38-3b86-46e4-b764-21e61e672db4 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL A survey of reinforcement learning algorithms for dynamically varying environments
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7cc4bef4-0572-4bc7-a6af-9ed68e9df8c2 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL and Papadimitriou, C
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 92027896-4a5b-4ed7-91e8-dc88be160dbe · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Reinforcement learning for humanoid robotics
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation de98ca8f-217c-48db-8cd9-5e3ca019b59e · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12bbe108-bada-4cc1-be69-f1a3ae9ba9e0 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Taming Non-stationary Bandits: A Bayesian Approach
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b1eae28-da1a-457a-badc-0404906eb8a7 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9f6fd08f-133a-43e5-8029-9d32cbd1cfcf · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL S., McAllester, D., Singh, S., and Mansour, Y
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 07f8b68d-865c-44b6-a5a9-c4a505f98284 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Efficient Learning in Non-Stationary Linear Markov Decision Processes
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 609a6bef-86e7-4ab2-b128-633b6fc78cab · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Non-asymptotic analysis for single-loop ( N atural) actor-critic with compatible function approximation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6bbed65e-980d-4ddb-80f2-b7a6683b5676 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL and Luo, H
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 38c9fcba-22e8-40c3-8b8f-b6b677a123cf · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL F., Zhang, W., Xu, P., and Gu, Q
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 12087280-777f-4d97-8e43-0fc39d8cad9a · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL M., Golmohammadi, A., Shi, Y., et al
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 855a2975-1c4f-4ce8-97d8-4a99ce175245 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Multi-agent reinforcement learning: A selective overview of theories and algorithms
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fc320535-15a3-44f1-b4ee-59dfacc79af9 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0f857b11-1235-487a-af37-095c7afe56d3 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Nonstationary Reinforcement Learning with Linear Function Approximation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01959a49-5c02-454e-9b45-0d99969c0422 · outbound
Natural Policy Gradient for Average Reward Non-Stationary RL Finite-sample analysis for sarsa with linear function approximation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.