Pith. sign in

Paper Citation Record · LEDGER

Natural Policy Gradient for Average Reward Non-Stationary RL

As of 17 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2504.16415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16415 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:04.170613Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy44
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 653de7a0-e14a-4d30-a844-79114d016b31 · outbound

This paper cites write newline.

Natural Policy Gradient for Average Reward Non-Stationary RL write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:03.948118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:03.948118Z digest=sha256:92faccc615742db0b9da3f89678267440b8f58dcd442cfa52cb28f29102416fe

Observation ca2878ab-e549-467a-8b06-6bd76f9bb7c8 · outbound

This paper cites M., Lee, J.

Natural Policy Gradient for Average Reward Non-Stationary RL M., Lee, J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.907969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.953166Z digest=sha256:f001f256bf0251e4e31057235e40c7338059ed8a68d961a7aa99241ef61e7fc7

Observation 93917d66-04eb-49ce-8e98-a9e3f8e805b4 · outbound

This paper cites U., and Aggarwal, V.

Natural Policy Gradient for Average Reward Non-Stationary RL U., and Aggarwal, V

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.896660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.957028Z digest=sha256:201344522e6a9f109037bf8ae92cc78619be835996e43abe4d8c22342753a7a9

Observation c4622e78-580e-4ef6-bc66-5a57b67e4a53 · outbound

This paper cites First-order methods in optimization.

Natural Policy Gradient for Average Reward Non-Stationary RL First-order methods in optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:03.961027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:03.961027Z digest=sha256:7fdc513140d64138bf0a9161e91d002b9fed4150738a61260452e41ca571cc49

Observation b8b4b070-4b30-47ad-8c5e-97c0be17014c · outbound

This paper cites Stochastic multi-armed-bandit problem with non-stationary rewards.

Natural Policy Gradient for Average Reward Non-Stationary RL Stochastic multi-armed-bandit problem with non-stationary rewards

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.879255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.964686Z digest=sha256:9c8c9535e33a572346b7acbd1c0daf598270f59efc6a1b5fad3476d7fb0a9d0d

Observation dcaceccf-a89b-488c-aa53-53760d493ceb · outbound

This paper cites S., Ghavamzadeh, M., and Lee, M.

Natural Policy Gradient for Average Reward Non-Stationary RL S., Ghavamzadeh, M., and Lee, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.868979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.968267Z digest=sha256:f32381dc95bef60eda4ddee8c175b09098d43fd876a9e1a6dc036feea64cf4f3

Observation e11a1424-fac8-4ed4-af34-327cebf3369a · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Natural Policy Gradient for Average Reward Non-Stationary RL Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.856558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.972827Z digest=sha256:52d52ccdf4e8e8c5651e1b2e46102f485b527989809ab305c79c73b5078b1560

Observation 94b9d0e9-78b0-4e12-881a-09704368f686 · outbound

This paper cites Fast global convergence of natural policy gradient methods with entropy regularization.

Natural Policy Gradient for Average Reward Non-Stationary RL Fast global convergence of natural policy gradient methods with entropy regularization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.843310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.977907Z digest=sha256:c12c151ab86f8b20de18290bcb44018db0a7cca8a5efc539c2437101d74814e4

Observation b8489771-fe9f-4d25-94fe-812dc81e94fa · outbound

This paper cites Optimizing for the future in non-stationary mdps.

Natural Policy Gradient for Average Reward Non-Stationary RL Optimizing for the future in non-stationary mdps

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.830462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.982030Z digest=sha256:5a4098fc7f8530cd68b7b610953832a4d06ef30f1cb403b5a09cf7a12fa73289

Observation 1676c12f-fb0f-4024-b1f2-2becc8f34351 · outbound

This paper cites Stabilizing reinforcement learning in dynamic environment with application to online recommendation.

Natural Policy Gradient for Average Reward Non-Stationary RL Stabilizing reinforcement learning in dynamic environment with application to online recommendation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.817636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.985559Z digest=sha256:5103159adc30550aacc5ff7c2174fc329057731cd1f644f9c182f3b0a24de87b

Observation 4ac1d03a-6728-4324-918d-4e140bdeb12d · outbound

This paper cites and Zhao, L.

Natural Policy Gradient for Average Reward Non-Stationary RL and Zhao, L

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.805055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.989285Z digest=sha256:0476e5a8845e9863814b0dd227221d530a6d964c11a517190afb2aac82e96a1d

Observation 3f1ef721-8f11-483f-853b-a48936e43dce · outbound

This paper cites C., Simchi-Levi, D., and Zhu, R.

Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.793923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.993078Z digest=sha256:f0e79e909a403b38dee57fa378a535cf0235fef5f3a8271d8cf6760cdad4694e

Observation 85046599-b60b-49c4-ba6e-2e3e49c817e6 · outbound

This paper cites C., Simchi-Levi, D., and Zhu, R.

Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.782754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.996510Z digest=sha256:a6d1fc9796380589d148f51ad17e1eead9f95fed2dd1aa7e4af952d0165c9d69

Observation 24689d8c-b7d8-492a-821f-96c1c02e42a6 · outbound

This paper cites C., Simchi-Levi, D., and Zhu, R.

Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.770734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.000031Z digest=sha256:ec4ba598618eabe0ca2d229867c27c3aa98309e5390400fa1e34513a15b1f06a

Observation ef86e6ac-8325-4271-913e-b48613ffe31c · outbound

This paper cites A kernel-based approach to non-stationary reinforcement learning in metric spaces.

Natural Policy Gradient for Average Reward Non-Stationary RL A kernel-based approach to non-stationary reinforcement learning in metric spaces

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.757771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.003317Z digest=sha256:e5c5f8c32f202a132229371a5804e3abfbb33c9d0ffd34c3a94b97e0894c0abb

Observation 207f88a5-a57e-4d80-8411-061ad28e0c8c · outbound

This paper cites M., and Mansour, Y.

Natural Policy Gradient for Average Reward Non-Stationary RL M., and Mansour, Y

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.744455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.006825Z digest=sha256:26674cdcc559f13014cf0d5ccd28379fa9ded53ed4ff436ab98e8b95fddee8e4

Observation c201b2fe-8135-476a-891f-74025763357b · outbound

This paper cites Dynamic regret of policy optimization in non-stationary environments.

Natural Policy Gradient for Average Reward Non-Stationary RL Dynamic regret of policy optimization in non-stationary environments

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.732565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.010119Z digest=sha256:d38c8fed7ffafe05fee23bb90000551f4b8f6bfc0cd173d1c1128387fa747d8b

Observation 18ed8316-7435-4fe5-968f-26ddfac355b2 · outbound

This paper cites Non-stationary reinforcement learning under general function approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Non-stationary reinforcement learning under general function approximation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.715426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.013406Z digest=sha256:9cf52923ace2e440c882d9a4829e6bd23f64fd8243c938c83a2e7abb005f3a01

Observation 2cd4f8d5-c7b2-4f18-9148-9aa8005410d1 · outbound

This paper cites A sliding-window algorithm for markov decision processes with arbitrarily changing rewards and transitions, 2018.

Natural Policy Gradient for Average Reward Non-Stationary RL A sliding-window algorithm for markov decision processes with arbitrarily changing rewards and transitions, 2018

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.702291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.016639Z digest=sha256:ae59f8ace19e6e5d693656b337fb2076ff3f28ae536f7f4ce1616b446cd1411d

Observation a1f95f88-4cb7-4d29-bcea-79402c0638f2 · outbound

This paper cites On Upper-Confidence Bound Policies for Non-Stationary Bandit Problems.

Natural Policy Gradient for Average Reward Non-Stationary RL On Upper-Confidence Bound Policies for Non-Stationary Bandit Problems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.020144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.020144Z digest=sha256:205bfeb80df9d4edc0ad02ebe20e7e4944fc06cd74428d550dd7f5ec249290a9

Observation 7444774b-c490-49f7-aaaf-9a767560eb1f · outbound

This paper cites Transient Non-Stationarity and Generalisation in Deep Reinforcement Learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Transient Non-Stationarity and Generalisation in Deep Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.023886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.023886Z digest=sha256:87e0b4f8db915ee58e5680991ab0d096533980bd730b15f147b9ebd430566839

Observation 80a783cc-88cd-4381-84ea-1d15b75b8b8c · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Near-optimal regret bounds for reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.689556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.027462Z digest=sha256:1a2628eb291ccc770a7020d61953cdcaa7e1247c1fceb89ef6946919e06dfb34

Observation a950b67b-098c-4223-82f4-d8ccde3256a1 · outbound

This paper cites Efficient reinforcement learning for routing jobs in heterogeneous queueing systems.

Natural Policy Gradient for Average Reward Non-Stationary RL Efficient reinforcement learning for routing jobs in heterogeneous queueing systems

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.676827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.031248Z digest=sha256:bb8758b92830fbfb396913bae1cba49b89234be4e5a292793ea2aa95930447fe

Observation 94ac8006-84f1-4ffe-89ca-0fdd2be74acf · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.664726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.034636Z digest=sha256:d2a98c614629a394d80c00a7cbe811042634aee26e7938c4b65517544e679821

Observation 6615cda7-1c04-45a2-b2d9-8ba0e1b443be · outbound

This paper cites and Qian, P.

Natural Policy Gradient for Average Reward Non-Stationary RL and Qian, P

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.651057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.038113Z digest=sha256:d5a5f8f65381f7eb63188674a7c09f88e8f090aa1920ed2a876d0d04732f9b5b

Observation eeeb929f-f14a-4ae9-ae11-9f6b52eda3b4 · outbound

This paper cites Towards continual reinforcement learning: A review and perspectives.

Natural Policy Gradient for Average Reward Non-Stationary RL Towards continual reinforcement learning: A review and perspectives

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.041523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.041523Z digest=sha256:84103506a37984825996e881287f575cdcfb990e11e78ed55fa2319a029932a9

Observation ec2f33a2-8e55-4d43-8873-7adaaf4d9deb · outbound

This paper cites R., Varma, S.

Natural Policy Gradient for Average Reward Non-Stationary RL R., Varma, S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.628726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.045034Z digest=sha256:b3503f8544d399d7085bba218c4465c8a7dab768f043aced99e04b49876bc8de

Observation f71f109a-8aa0-4ed3-a96a-09d665fa23b5 · outbound

This paper cites T., Romberg, J., and Maguluri, S.

Natural Policy Gradient for Average Reward Non-Stationary RL T., Romberg, J., and Maguluri, S

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.614699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.048824Z digest=sha256:af3b3148d66eed4d54330601c38e9eaafe020ba481d945848baa0a961b9833eb

Observation 2923654e-62a3-433f-a4e0-a7a9239c85be · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.601039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.052793Z digest=sha256:c26b4a290a0a2a6c89cdee978fcc97c23ba2dc91ba229c97d3f65c6879f8856d

Observation c939c320-c385-424b-a86d-8883a306e65f · outbound

This paper cites Improved regret bound and experience replay in regularized policy iteration.

Natural Policy Gradient for Average Reward Non-Stationary RL Improved regret bound and experience replay in regularized policy iteration

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.588694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.056203Z digest=sha256:984ac6b3fe097ed233333d3c7b5aaa73b2b9d1cb9003ebc43bc05c5988dfd446

Observation 3e2a53e9-02f4-47ac-a315-6d51d8646982 · outbound

This paper cites and Rachelson, E.

Natural Policy Gradient for Average Reward Non-Stationary RL and Rachelson, E

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.575504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.059581Z digest=sha256:4b0d8d4815ca11c5c5f156a60b2961668f82a6b613b628da1b580f435d091b47

Observation b373027a-0703-405c-88fe-b887980b8a9c · outbound

This paper cites Pausing policy learning in non-stationary reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Pausing policy learning in non-stationary reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.562565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.062965Z digest=sha256:6eb9a9fa6c6e25c3a24da65fa96cb595877f577492067ae3eabe5f9f71dceed2

Observation 1c993825-9e01-4eda-aa3b-696f9a48ed80 · outbound

This paper cites Rl-qn: A reinforcement learning framework for optimal control of queueing systems.

Natural Policy Gradient for Average Reward Non-Stationary RL Rl-qn: A reinforcement learning framework for optimal control of queueing systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.551164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.066666Z digest=sha256:2b8a7795dece5f9752118fb0931a63487663a71b0926643f7fbe49f5589a3759

Observation 7bc2e48e-317b-48d7-82ac-eef77431d013 · outbound

This paper cites A Definition of Non-Stationary Bandits.

Natural Policy Gradient for Average Reward Non-Stationary RL A Definition of Non-Stationary Bandits

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.069961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.069961Z digest=sha256:0261ead06b224c8797a62a4f14d49f4d44a8c07083f31d5200ede6fd27b30ee3

Observation 96129abb-a2ab-488b-9a3a-a86e9ec7b14e · outbound

This paper cites Nonstationary bandit learning via predictive sampling.

Natural Policy Gradient for Average Reward Non-Stationary RL Nonstationary bandit learning via predictive sampling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.538438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.073909Z digest=sha256:858a9b34726a391f330a2fdec1b735ee12603c9a03a0a7843d02c06b276f17ef

Observation fa0791ca-66de-4294-b7d7-95b4148e8c84 · outbound

This paper cites Average reward reinforcement learning: Foundations, algorithms, and empirical results.

Natural Policy Gradient for Average Reward Non-Stationary RL Average reward reinforcement learning: Foundations, algorithms, and empirical results

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.527413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.078377Z digest=sha256:b43ee2f657313b1d9564a3d2201de9bbf41d914508a087767258b98f9501fc91

Observation 3b38bb50-96f6-4444-9da8-efc50c204f34 · outbound

This paper cites Model-free nonstationary reinforcement learning: Near-optimal regret and applications in multiagent reinforcement learning and inventory control.

Natural Policy Gradient for Average Reward Non-Stationary RL Model-free nonstationary reinforcement learning: Near-optimal regret and applications in multiagent reinforcement learning and inventory control

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.514105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.082744Z digest=sha256:1a255fbbf57d2c74ba04006f2dff5b64acc3ed64cc83f4b7c3a414c1c51545c7

Observation 7c4e1aa2-4ae1-4cfd-8133-12f97bc87932 · outbound

This paper cites New insights and perspectives on the natural gradient method.

Natural Policy Gradient for Average Reward Non-Stationary RL New insights and perspectives on the natural gradient method

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.501063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.087368Z digest=sha256:8040aecd6ad22ecdba4801d1fee09a488f59b6a2aaae44debc79b86a069871b0

Observation f9fd8b3d-866e-48f3-aea2-a2e05a6eed6d · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.488647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.091564Z digest=sha256:c10d035e9c39fdbaa5442f1527860642b39054ae6940610ab345c98234499c8f

Observation 64f089b1-aab5-4b53-a9fe-86759e30c5f5 · outbound

This paper cites and Srikant, R.

Natural Policy Gradient for Average Reward Non-Stationary RL and Srikant, R

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.477406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.095137Z digest=sha256:c3b4cac64fb3728463fcff2dd037dce4879594bf5aad19c77c1266b68a474ba2

Observation d8631203-b9e7-4f10-ae79-c455b2175960 · outbound

This paper cites Performance bounds for policy-based average reward reinforcement learning algorithms.

Natural Policy Gradient for Average Reward Non-Stationary RL Performance bounds for policy-based average reward reinforcement learning algorithms

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.465214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.098637Z digest=sha256:79600058dff0a1bdf000afa6dbe7cab012854e9234d8ddcfbc12a02740edd3d5

Observation e13e4eda-4d78-4a9e-90b6-7f579e8d7dd8 · outbound

This paper cites Bridging the gap between value and policy based reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Bridging the gap between value and policy based reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.452784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.102009Z digest=sha256:19226276e701e88916be7a3740af25db2126bc3ba71ff12653122956fc28ee01

Observation 1c888d79-729b-481e-af21-e2e8af216279 · outbound

This paper cites Variational regret bounds for reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Variational regret bounds for reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.440305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.105527Z digest=sha256:622420686a23d59ae57337f75355d0752a365baea83a523a881bf102e0235f49

Observation 0c74fd38-3b86-46e4-b764-21e61e672db4 · outbound

This paper cites A survey of reinforcement learning algorithms for dynamically varying environments.

Natural Policy Gradient for Average Reward Non-Stationary RL A survey of reinforcement learning algorithms for dynamically varying environments

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.428052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.109453Z digest=sha256:2d422408f87404370e6a5aec7ad0442bf80508a60c02eaa69db6a93b25b99d21

Observation 7cc4bef4-0572-4bc7-a6af-9ed68e9df8c2 · outbound

This paper cites and Papadimitriou, C.

Natural Policy Gradient for Average Reward Non-Stationary RL and Papadimitriou, C

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.413135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.113199Z digest=sha256:15161315b19c9300b86110a0460c193edb78c5d7bccd9aad253349cf8d465dc0

Observation 92027896-4a5b-4ed7-91e8-dc88be160dbe · outbound

This paper cites Reinforcement learning for humanoid robotics.

Natural Policy Gradient for Average Reward Non-Stationary RL Reinforcement learning for humanoid robotics

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.400798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.117009Z digest=sha256:57a722f23d0b06ed361f272a8cc5d85a448023262df8af6ccd207a9cc507ff61

Observation de98ca8f-217c-48db-8cd9-5e3ca019b59e · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.120637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.120637Z digest=sha256:e61867d924ce0c1d9eaa59114b1bbcc7da8fd5fd5bd1d4cc049c3127c67634f2

Observation 12bbe108-bada-4cc1-be69-f1a3ae9ba9e0 · outbound

This paper cites Taming Non-stationary Bandits: A Bayesian Approach.

Natural Policy Gradient for Average Reward Non-Stationary RL Taming Non-stationary Bandits: A Bayesian Approach

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.124835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.124835Z digest=sha256:553b2651fe7667c777cd07590a155fca945422e1d82deb411de2e265924f62a3

Observation 0b1eae28-da1a-457a-badc-0404906eb8a7 · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.381015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.129502Z digest=sha256:739ba19df0b8e197739b0cc99cca96fef6f1479423471c3dae9d2fb161032061

Observation 9f6fd08f-133a-43e5-8029-9d32cbd1cfcf · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

Natural Policy Gradient for Average Reward Non-Stationary RL S., McAllester, D., Singh, S., and Mansour, Y

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.369819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.133253Z digest=sha256:66fd84a178b065e5d94d42af8a54181c26d4426a2ded07b060c20ed9ed017e27

Observation 07f8b68d-865c-44b6-a5a9-c4a505f98284 · outbound

This paper cites Efficient Learning in Non-Stationary Linear Markov Decision Processes.

Natural Policy Gradient for Average Reward Non-Stationary RL Efficient Learning in Non-Stationary Linear Markov Decision Processes

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:16:04.223437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.136748Z digest=sha256:a265aeeeffa362f8f22e07b162879ad47e568bf46918ed452fc05986aeddf9bb

Observation 609a6bef-86e7-4ab2-b128-633b6fc78cab · outbound

This paper cites Non-asymptotic analysis for single-loop ( N atural) actor-critic with compatible function approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Non-asymptotic analysis for single-loop ( N atural) actor-critic with compatible function approximation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.358298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.143651Z digest=sha256:48de5d581b27cd27f4ccb2744e3243711aef858eea6d3da193dbd7fee4504bfe

Observation 6bbed65e-980d-4ddb-80f2-b7a6683b5676 · outbound

This paper cites and Luo, H.

Natural Policy Gradient for Average Reward Non-Stationary RL and Luo, H

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.345302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.147668Z digest=sha256:f365f2936e50394ffa8de5f5522d420c657f9d6fee8eb6f504d3369510df527e

Observation 38c9fcba-22e8-40c3-8b8f-b6b677a123cf · outbound

This paper cites F., Zhang, W., Xu, P., and Gu, Q.

Natural Policy Gradient for Average Reward Non-Stationary RL F., Zhang, W., Xu, P., and Gu, Q

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.332643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.151138Z digest=sha256:2b671c2a1aa33c5bfce75298f728ee006015fba43dbb22857c7891207751e15b

Observation 12087280-777f-4d97-8e43-0fc39d8cad9a · outbound

This paper cites M., Golmohammadi, A., Shi, Y., et al.

Natural Policy Gradient for Average Reward Non-Stationary RL M., Golmohammadi, A., Shi, Y., et al

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.320024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.155857Z digest=sha256:9f914da2352b3f8444aaa06ba5394bceb9f15c9201a8cddc85f839bcc7e2dcb2

Observation 855a2975-1c4f-4ce8-97d8-4a99ce175245 · outbound

This paper cites Multi-agent reinforcement learning: A selective overview of theories and algorithms.

Natural Policy Gradient for Average Reward Non-Stationary RL Multi-agent reinforcement learning: A selective overview of theories and algorithms

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.306025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.159451Z digest=sha256:6a46510487d00505b6b7a0541b5fa2057f996e43786d91374eea236696b74c8b

Observation fc320535-15a3-44f1-b4ee-59dfacc79af9 · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.292073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.162985Z digest=sha256:a73b90766e656e574bb4bb2b961dfd5bea996d3473cbfacf0459235954e2dcc4

Observation 0f857b11-1235-487a-af37-095c7afe56d3 · outbound

This paper cites Nonstationary Reinforcement Learning with Linear Function Approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Nonstationary Reinforcement Learning with Linear Function Approximation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.166805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.166805Z digest=sha256:bbc2e85ae51242df68496d1bfd2bd17649a1763fd862eab4713a98e84b0a7a9a

Observation 01959a49-5c02-454e-9b45-0d99969c0422 · outbound

This paper cites Finite-sample analysis for sarsa with linear function approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Finite-sample analysis for sarsa with linear function approximation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.279920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.170613Z digest=sha256:7d439f939d669db84e5d9cbad416e3e26f823a59c5ae943635dcdcfbaf56821f

Pith citing papers

No inbound Pith citation observations are available.