Pith. sign in

Paper Citation Record · LEDGER

Natural Policy Gradient for Average Reward Non-Stationary RL

As of 17 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2504.16415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16415 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:04.170613Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy44
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 653de7a0-e14a-4d30-a844-79114d016b31 · outbound

This paper cites write newline.

Natural Policy Gradient for Average Reward Non-Stationary RL write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:03.948118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:03.948118Z digest=sha256:92faccc615742db0b9da3f89678267440b8f58dcd442cfa52cb28f29102416fe

Observation ca2878ab-e549-467a-8b06-6bd76f9bb7c8 · outbound

This paper cites M., Lee, J.

Natural Policy Gradient for Average Reward Non-Stationary RL M., Lee, J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.907969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.953166Z digest=sha256:0ed06e75f0ad37ec1262996ad1b30cf13ae3fe97e8d05c31596da8c597774f94

Observation 93917d66-04eb-49ce-8e98-a9e3f8e805b4 · outbound

This paper cites U., and Aggarwal, V.

Natural Policy Gradient for Average Reward Non-Stationary RL U., and Aggarwal, V

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.896660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.957028Z digest=sha256:a4b8d57a012709e9c9af436936f466f21f03c0dbd8f82df82353f3b468bc44b5

Observation c4622e78-580e-4ef6-bc66-5a57b67e4a53 · outbound

This paper cites First-order methods in optimization.

Natural Policy Gradient for Average Reward Non-Stationary RL First-order methods in optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:03.961027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:03.961027Z digest=sha256:7fdc513140d64138bf0a9161e91d002b9fed4150738a61260452e41ca571cc49

Observation b8b4b070-4b30-47ad-8c5e-97c0be17014c · outbound

This paper cites Stochastic multi-armed-bandit problem with non-stationary rewards.

Natural Policy Gradient for Average Reward Non-Stationary RL Stochastic multi-armed-bandit problem with non-stationary rewards

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.879255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.964686Z digest=sha256:e739913c2f53b83ab0f0a2d26e9cefb4977525941a3dec5869be17d0a7611fed

Observation dcaceccf-a89b-488c-aa53-53760d493ceb · outbound

This paper cites S., Ghavamzadeh, M., and Lee, M.

Natural Policy Gradient for Average Reward Non-Stationary RL S., Ghavamzadeh, M., and Lee, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.868979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.968267Z digest=sha256:53fa3efdde540e93ec7e6b0fc2f914a6b2938d409a983f00f86609d92b85b2c6

Observation e11a1424-fac8-4ed4-af34-327cebf3369a · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Natural Policy Gradient for Average Reward Non-Stationary RL Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.856558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.972827Z digest=sha256:cdf8bc4d9719bd20566ef933df7e0a67f5475c2da3e436551b18203bcebe057d

Observation 94b9d0e9-78b0-4e12-881a-09704368f686 · outbound

This paper cites Fast global convergence of natural policy gradient methods with entropy regularization.

Natural Policy Gradient for Average Reward Non-Stationary RL Fast global convergence of natural policy gradient methods with entropy regularization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.843310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.977907Z digest=sha256:c8f5b4b7c36a128046dba5e2f214f8a8ef0a2536eb6ed8c8116d79162c9338e6

Observation b8489771-fe9f-4d25-94fe-812dc81e94fa · outbound

This paper cites Optimizing for the future in non-stationary mdps.

Natural Policy Gradient for Average Reward Non-Stationary RL Optimizing for the future in non-stationary mdps

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.830462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.982030Z digest=sha256:2f5dbb79bfbcc2334637d6960191fd88e2cb7976991c1791924e7190b54697ec

Observation 1676c12f-fb0f-4024-b1f2-2becc8f34351 · outbound

This paper cites Stabilizing reinforcement learning in dynamic environment with application to online recommendation.

Natural Policy Gradient for Average Reward Non-Stationary RL Stabilizing reinforcement learning in dynamic environment with application to online recommendation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.817636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.985559Z digest=sha256:37564c15fd6c0537dc8fcbc2dc49db6963099ede701d786b6fe0da1ea97c7773

Observation 4ac1d03a-6728-4324-918d-4e140bdeb12d · outbound

This paper cites and Zhao, L.

Natural Policy Gradient for Average Reward Non-Stationary RL and Zhao, L

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.805055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.989285Z digest=sha256:da329cb4632fc7725a8b70df4cbc97916ed5e54d3d1362b5541ac47a52d6ef90

Observation 3f1ef721-8f11-483f-853b-a48936e43dce · outbound

This paper cites C., Simchi-Levi, D., and Zhu, R.

Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.793923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.993078Z digest=sha256:0f2fed59f30168eeec7257b1e5a058b5ff2e5d79395bcd59e5ca7b36ad25cc28

Observation 85046599-b60b-49c4-ba6e-2e3e49c817e6 · outbound

This paper cites C., Simchi-Levi, D., and Zhu, R.

Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.782754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.996510Z digest=sha256:3090be704f6a1b8f13659f1ea9cc193e5a442fe22080f3f7020caabe4d3d20a5

Observation 24689d8c-b7d8-492a-821f-96c1c02e42a6 · outbound

This paper cites C., Simchi-Levi, D., and Zhu, R.

Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.770734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.000031Z digest=sha256:5b8af14db79cad9cff6304ae2989d35d718da063883b8359eba4d0c92eec54fd

Observation ef86e6ac-8325-4271-913e-b48613ffe31c · outbound

This paper cites A kernel-based approach to non-stationary reinforcement learning in metric spaces.

Natural Policy Gradient for Average Reward Non-Stationary RL A kernel-based approach to non-stationary reinforcement learning in metric spaces

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.757771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.003317Z digest=sha256:52bd6721655e3647ad6e2065ec7f27109515b186c6e4f9ad98759d065cd60d66

Observation 207f88a5-a57e-4d80-8411-061ad28e0c8c · outbound

This paper cites M., and Mansour, Y.

Natural Policy Gradient for Average Reward Non-Stationary RL M., and Mansour, Y

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.744455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.006825Z digest=sha256:e809ec64436dcbbe409df7b493e7fb6756a6dcef64ea4448fead5cb6f93d2eca

Observation c201b2fe-8135-476a-891f-74025763357b · outbound

This paper cites Dynamic regret of policy optimization in non-stationary environments.

Natural Policy Gradient for Average Reward Non-Stationary RL Dynamic regret of policy optimization in non-stationary environments

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.732565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.010119Z digest=sha256:872a8264d3580ce88a75ff97b6ab402234c8db29d45c602e9bfdbc9bee9d45e6

Observation 18ed8316-7435-4fe5-968f-26ddfac355b2 · outbound

This paper cites Non-stationary reinforcement learning under general function approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Non-stationary reinforcement learning under general function approximation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.715426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.013406Z digest=sha256:4f830ab7a7c27eb3107b6a66aaf87f4358a471d1af6bd8b07b25006c41999d56

Observation 2cd4f8d5-c7b2-4f18-9148-9aa8005410d1 · outbound

This paper cites A sliding-window algorithm for markov decision processes with arbitrarily changing rewards and transitions, 2018.

Natural Policy Gradient for Average Reward Non-Stationary RL A sliding-window algorithm for markov decision processes with arbitrarily changing rewards and transitions, 2018

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.702291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.016639Z digest=sha256:4c42cbee8c6af4f39f677a971d93254bdd89a1ee67332a1afc46837b7e68fd7a

Observation a1f95f88-4cb7-4d29-bcea-79402c0638f2 · outbound

This paper cites On Upper-Confidence Bound Policies for Non-Stationary Bandit Problems.

Natural Policy Gradient for Average Reward Non-Stationary RL On Upper-Confidence Bound Policies for Non-Stationary Bandit Problems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.020144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.020144Z digest=sha256:205bfeb80df9d4edc0ad02ebe20e7e4944fc06cd74428d550dd7f5ec249290a9

Observation 7444774b-c490-49f7-aaaf-9a767560eb1f · outbound

This paper cites Transient Non-Stationarity and Generalisation in Deep Reinforcement Learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Transient Non-Stationarity and Generalisation in Deep Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.023886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.023886Z digest=sha256:87e0b4f8db915ee58e5680991ab0d096533980bd730b15f147b9ebd430566839

Observation 80a783cc-88cd-4381-84ea-1d15b75b8b8c · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Near-optimal regret bounds for reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.689556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.027462Z digest=sha256:4c6955bda1527da1de8165b6b37e754a0c71d69d0479ec021324f4ad98e5fd2f

Observation a950b67b-098c-4223-82f4-d8ccde3256a1 · outbound

This paper cites Efficient reinforcement learning for routing jobs in heterogeneous queueing systems.

Natural Policy Gradient for Average Reward Non-Stationary RL Efficient reinforcement learning for routing jobs in heterogeneous queueing systems

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.676827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.031248Z digest=sha256:a68944e52e5588d86286b39bd52ec0f04a45cfde720701dbaa53308c2bdf1916

Observation 94ac8006-84f1-4ffe-89ca-0fdd2be74acf · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.664726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.034636Z digest=sha256:5979d436818d622feca90201bf310e419f11698dd288f20ded149a71fb1aac74

Observation 6615cda7-1c04-45a2-b2d9-8ba0e1b443be · outbound

This paper cites and Qian, P.

Natural Policy Gradient for Average Reward Non-Stationary RL and Qian, P

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.651057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.038113Z digest=sha256:c05a0fbe5a681f53b31b137da29a46916075e0b539821bcd4ee6e230ab671f19

Observation eeeb929f-f14a-4ae9-ae11-9f6b52eda3b4 · outbound

This paper cites Towards continual reinforcement learning: A review and perspectives.

Natural Policy Gradient for Average Reward Non-Stationary RL Towards continual reinforcement learning: A review and perspectives

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.041523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.041523Z digest=sha256:84103506a37984825996e881287f575cdcfb990e11e78ed55fa2319a029932a9

Observation ec2f33a2-8e55-4d43-8873-7adaaf4d9deb · outbound

This paper cites R., Varma, S.

Natural Policy Gradient for Average Reward Non-Stationary RL R., Varma, S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.628726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.045034Z digest=sha256:02f8d2f21a5f590ac400ad809c6d5c6614aaa736e6aab64c07ff9fc141f55c73

Observation f71f109a-8aa0-4ed3-a96a-09d665fa23b5 · outbound

This paper cites T., Romberg, J., and Maguluri, S.

Natural Policy Gradient for Average Reward Non-Stationary RL T., Romberg, J., and Maguluri, S

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.614699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.048824Z digest=sha256:54878521355b45d8da57aba400f4b8c504816dcf6432b8e3137448919807e138

Observation 2923654e-62a3-433f-a4e0-a7a9239c85be · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.601039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.052793Z digest=sha256:af95845069e21da963ab056f43cc3974e2093b706939fef2f1a2d0262a198d1a

Observation c939c320-c385-424b-a86d-8883a306e65f · outbound

This paper cites Improved regret bound and experience replay in regularized policy iteration.

Natural Policy Gradient for Average Reward Non-Stationary RL Improved regret bound and experience replay in regularized policy iteration

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.588694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.056203Z digest=sha256:7a131fad61b1355437413caa633b25830763caa45cd1da7867b7f2dcd0f50faa

Observation 3e2a53e9-02f4-47ac-a315-6d51d8646982 · outbound

This paper cites and Rachelson, E.

Natural Policy Gradient for Average Reward Non-Stationary RL and Rachelson, E

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.575504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.059581Z digest=sha256:dee15b719eca8b157dd5e72f31e00c4c4e68814223eb161fefc75b50e0e4e05b

Observation b373027a-0703-405c-88fe-b887980b8a9c · outbound

This paper cites Pausing policy learning in non-stationary reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Pausing policy learning in non-stationary reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.562565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.062965Z digest=sha256:5c849a53612248a504178b93199ead41f106240ed8bffaf9fc96bde9a04704bb

Observation 1c993825-9e01-4eda-aa3b-696f9a48ed80 · outbound

This paper cites Rl-qn: A reinforcement learning framework for optimal control of queueing systems.

Natural Policy Gradient for Average Reward Non-Stationary RL Rl-qn: A reinforcement learning framework for optimal control of queueing systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.551164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.066666Z digest=sha256:0bc74d862e4c0e7526958944bfcd82ecf5de6005b3270359bd75614587faf7df

Observation 7bc2e48e-317b-48d7-82ac-eef77431d013 · outbound

This paper cites A Definition of Non-Stationary Bandits.

Natural Policy Gradient for Average Reward Non-Stationary RL A Definition of Non-Stationary Bandits

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.069961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.069961Z digest=sha256:22fb66649ec20e3c3e30ff2349b39dfda74865d2b11a4f09c82fcabdf33b1b72

Observation 96129abb-a2ab-488b-9a3a-a86e9ec7b14e · outbound

This paper cites Nonstationary bandit learning via predictive sampling.

Natural Policy Gradient for Average Reward Non-Stationary RL Nonstationary bandit learning via predictive sampling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.538438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.073909Z digest=sha256:8841308c4b8f14d5ed946c77e211c6b1d7f9740c16656356532e781a06a6814d

Observation fa0791ca-66de-4294-b7d7-95b4148e8c84 · outbound

This paper cites Average reward reinforcement learning: Foundations, algorithms, and empirical results.

Natural Policy Gradient for Average Reward Non-Stationary RL Average reward reinforcement learning: Foundations, algorithms, and empirical results

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.527413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.078377Z digest=sha256:820543d972c5c37a0b7952a6f65ffb1128e0e7abd1b2149901fc6d3791829df3

Observation 3b38bb50-96f6-4444-9da8-efc50c204f34 · outbound

This paper cites Model-free nonstationary reinforcement learning: Near-optimal regret and applications in multiagent reinforcement learning and inventory control.

Natural Policy Gradient for Average Reward Non-Stationary RL Model-free nonstationary reinforcement learning: Near-optimal regret and applications in multiagent reinforcement learning and inventory control

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.514105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.082744Z digest=sha256:80e1582d3f03749f8d2b4d9f7f833bd9c3c816ccd9c09e58c1581c5031de30d1

Observation 7c4e1aa2-4ae1-4cfd-8133-12f97bc87932 · outbound

This paper cites New insights and perspectives on the natural gradient method.

Natural Policy Gradient for Average Reward Non-Stationary RL New insights and perspectives on the natural gradient method

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.501063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.087368Z digest=sha256:259886cab8aaba8189f60a8ed37a5fb1d6e4171a0640eab6bd256a16fc88f89e

Observation f9fd8b3d-866e-48f3-aea2-a2e05a6eed6d · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.488647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.091564Z digest=sha256:076082fff6c11a555ed11471380d82396e5575fe82941a1097f7f56a4e23f4a7

Observation 64f089b1-aab5-4b53-a9fe-86759e30c5f5 · outbound

This paper cites and Srikant, R.

Natural Policy Gradient for Average Reward Non-Stationary RL and Srikant, R

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.477406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.095137Z digest=sha256:ce76e0ebc0724d66b63e3b1371d59583acb7ce4204940d93bd36bddec9796119

Observation d8631203-b9e7-4f10-ae79-c455b2175960 · outbound

This paper cites Performance bounds for policy-based average reward reinforcement learning algorithms.

Natural Policy Gradient for Average Reward Non-Stationary RL Performance bounds for policy-based average reward reinforcement learning algorithms

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.465214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.098637Z digest=sha256:c88ed0543fc2252b1a106d5716e65614af66a3fe242e48138e42b2e0df8dfa78

Observation e13e4eda-4d78-4a9e-90b6-7f579e8d7dd8 · outbound

This paper cites Bridging the gap between value and policy based reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Bridging the gap between value and policy based reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.452784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.102009Z digest=sha256:a068632e2f47ebb1215c36a9f655dcac26edef1feb9a40a68a5c766836e4886d

Observation 1c888d79-729b-481e-af21-e2e8af216279 · outbound

This paper cites Variational regret bounds for reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Variational regret bounds for reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.440305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.105527Z digest=sha256:5029cf81a4ebbb1aae1c5b379d3c58f5b91f1d9dde0c8e685d55cb16d2016f01

Observation 0c74fd38-3b86-46e4-b764-21e61e672db4 · outbound

This paper cites A survey of reinforcement learning algorithms for dynamically varying environments.

Natural Policy Gradient for Average Reward Non-Stationary RL A survey of reinforcement learning algorithms for dynamically varying environments

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.428052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.109453Z digest=sha256:0d0f4b719ae9c50cfe3d8f493b3748001a51d3c033979eeab83b2077cdae061a

Observation 7cc4bef4-0572-4bc7-a6af-9ed68e9df8c2 · outbound

This paper cites and Papadimitriou, C.

Natural Policy Gradient for Average Reward Non-Stationary RL and Papadimitriou, C

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.413135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.113199Z digest=sha256:10cc26251ec7c7cf6107f8fd9712e78a03171024bed922ae055e4d45350f2e0c

Observation 92027896-4a5b-4ed7-91e8-dc88be160dbe · outbound

This paper cites Reinforcement learning for humanoid robotics.

Natural Policy Gradient for Average Reward Non-Stationary RL Reinforcement learning for humanoid robotics

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.400798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.117009Z digest=sha256:3df1b045c7d252ad18abe6b6a6b799b5e218db1aa5e4ac922a419cff11f3bcf4

Observation de98ca8f-217c-48db-8cd9-5e3ca019b59e · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.120637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.120637Z digest=sha256:e61867d924ce0c1d9eaa59114b1bbcc7da8fd5fd5bd1d4cc049c3127c67634f2

Observation 12bbe108-bada-4cc1-be69-f1a3ae9ba9e0 · outbound

This paper cites Taming Non-stationary Bandits: A Bayesian Approach.

Natural Policy Gradient for Average Reward Non-Stationary RL Taming Non-stationary Bandits: A Bayesian Approach

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.124835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.124835Z digest=sha256:553b2651fe7667c777cd07590a155fca945422e1d82deb411de2e265924f62a3

Observation 0b1eae28-da1a-457a-badc-0404906eb8a7 · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.381015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.129502Z digest=sha256:b6ec7216f0698d039ef157d5d2c47fceb879359d37ae586928402432284f1e17

Observation 9f6fd08f-133a-43e5-8029-9d32cbd1cfcf · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

Natural Policy Gradient for Average Reward Non-Stationary RL S., McAllester, D., Singh, S., and Mansour, Y

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.369819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.133253Z digest=sha256:c432efd4b6716f0355c58224856afb231e83e13d0478d1fe37ce46915e88d9d9

Observation 07f8b68d-865c-44b6-a5a9-c4a505f98284 · outbound

This paper cites Efficient Learning in Non-Stationary Linear Markov Decision Processes.

Natural Policy Gradient for Average Reward Non-Stationary RL Efficient Learning in Non-Stationary Linear Markov Decision Processes

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:16:04.223437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.136748Z digest=sha256:679764dc3883a119cfdaf692d633c5b688733057149f66a9715ecdf565d868e0

Observation 609a6bef-86e7-4ab2-b128-633b6fc78cab · outbound

This paper cites Non-asymptotic analysis for single-loop ( N atural) actor-critic with compatible function approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Non-asymptotic analysis for single-loop ( N atural) actor-critic with compatible function approximation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.358298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.143651Z digest=sha256:69dca1abf0eef41f4ad619a7605a887992716d3a3cd230985fe0801b51576cd8

Observation 6bbed65e-980d-4ddb-80f2-b7a6683b5676 · outbound

This paper cites and Luo, H.

Natural Policy Gradient for Average Reward Non-Stationary RL and Luo, H

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.345302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.147668Z digest=sha256:0183de8b098f044ee6cc016b185b39c3e422908cc7e490571dc5fb48c78ac4f7

Observation 38c9fcba-22e8-40c3-8b8f-b6b677a123cf · outbound

This paper cites F., Zhang, W., Xu, P., and Gu, Q.

Natural Policy Gradient for Average Reward Non-Stationary RL F., Zhang, W., Xu, P., and Gu, Q

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.332643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.151138Z digest=sha256:bf52f3f6e0497efbf1403cf97fcf3c2fa3a0c36e3972c4ad4d0cf8180aa4a0bd

Observation 12087280-777f-4d97-8e43-0fc39d8cad9a · outbound

This paper cites M., Golmohammadi, A., Shi, Y., et al.

Natural Policy Gradient for Average Reward Non-Stationary RL M., Golmohammadi, A., Shi, Y., et al

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.320024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.155857Z digest=sha256:8f827cf376e1225c15bfaae1423bfed474722f464a5b274d270579e61fe59b42

Observation 855a2975-1c4f-4ce8-97d8-4a99ce175245 · outbound

This paper cites Multi-agent reinforcement learning: A selective overview of theories and algorithms.

Natural Policy Gradient for Average Reward Non-Stationary RL Multi-agent reinforcement learning: A selective overview of theories and algorithms

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.306025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.159451Z digest=sha256:a46c5bfd01dbfccf798212862fce64272878858230e347514be3b1083c0acf14

Observation fc320535-15a3-44f1-b4ee-59dfacc79af9 · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.292073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.162985Z digest=sha256:8deb8eba960587dce9bca5715a4298fd078bba22881328aba18f2bd6944186f0

Observation 0f857b11-1235-487a-af37-095c7afe56d3 · outbound

This paper cites Nonstationary Reinforcement Learning with Linear Function Approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Nonstationary Reinforcement Learning with Linear Function Approximation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.166805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.166805Z digest=sha256:e17ac5b2b2eb86b8788e8aa1b10a5041e67fa75a58fc3a4611e6b940a9575cd3

Observation 01959a49-5c02-454e-9b45-0d99969c0422 · outbound

This paper cites Finite-sample analysis for sarsa with linear function approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Finite-sample analysis for sarsa with linear function approximation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.279920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.170613Z digest=sha256:573be41dc28a2de36d6c898e48d0639b783e745345dbb5efedc29fb84dd00187

Pith citing papers

No inbound Pith citation observations are available.