Pith. sign in

Paper Citation Record · LEDGER

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2506.07040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07040 v4

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:54:12.780877Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef70f0d3-4b29-48c4-9680-c2a66502970f · outbound

This paper cites On the theory of policy gradient methods: Optimality, approximation, and distribution shift.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning On the theory of policy gradient methods: Optimality, approximation, and distribution shift

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:17.474894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:09.616419Z digest=sha256:5d38c63ef2fc33a31ea5cacebef5fbf4b23e5b25b4e93c42777f6031668ef505

Observation 9965448c-ccfe-4d5c-8797-9056a40d7186 · outbound

This paper cites Bounded semigroups of matrices.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Bounded semigroups of matrices

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:17.297043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:09.707205Z digest=sha256:46a751dca55ee3cad4331ce4d615fb84413c9edcd91419e2f7ded8a728a89e22

Observation 9ef97421-6307-40dd-bbb8-6844ed0b88fa · outbound

This paper cites Single sample path-based optimization of markov chains.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Single sample path-based optimization of markov chains

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:17.111742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:09.869727Z digest=sha256:6f43c81f32a65908b7f78ee6be3faf25c7c7744b07e7bf368d403b3c291c7b01

Observation 869a3cd3-5da4-42c8-8d7b-7d5f54645349 · outbound

This paper cites Sample complexity of distributionally robust average-reward reinforcement learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Sample complexity of distributionally robust average-reward reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:09.997942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:09.997942Z digest=sha256:e1dfef55ea344829c98def6c64e399d72ff16afe4649530752a918f25e06a76b

Observation 61e2f1a7-d41a-4e96-be53-613e8558d664 · outbound

This paper cites Distributionally robust stochastic optimization with W asserstein distance.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Distributionally robust stochastic optimization with W asserstein distance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:16.886561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.170591Z digest=sha256:f97c8334a2ef97cd37237e58dc0dcf3ed87030203ee40ef3827a68149713b07c

Observation 67561d46-dd68-4551-9c2d-42cc6c7d39d6 · outbound

This paper cites Distributionally robust stochastic optimization with wasserstein distance.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Distributionally robust stochastic optimization with wasserstein distance

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:16.737358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.326835Z digest=sha256:b3cd0ca64e7b098cb9cc92205e78f1f407475caaf22a0b92335c2892a15d8f1f

Observation 825f67de-894c-4c02-bc88-01753b950a87 · outbound

This paper cites Sim2real in robotics and automation: Applications and challenges.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Sim2real in robotics and automation: Applications and challenges

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:16.572891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.442042Z digest=sha256:ccc6b077034e542063e901f420c872724d32761a47083d2f1e53ce98fc7a776d

Observation 7f309759-8e64-451a-a29d-50e69cfd7db2 · outbound

This paper cites A robust version of the probability ratio test.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning A robust version of the probability ratio test

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:16.386215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.553831Z digest=sha256:f5de1497afa1e7f0f72637162ae194bb147fa032bdc0248d2f2ff2ca66b769b5

Observation ec7c81e9-8789-41b2-8b25-686b96860ea4 · outbound

This paper cites Robust dynamic programming.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robust dynamic programming

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:16.227215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.666749Z digest=sha256:202bb720559441d3ed7712d0133ca7d9e260adf52def6e8f027d2457d130a226

Observation 19f8b178-7a01-43f6-9858-f64fdbe6a7ff · outbound

This paper cites Is q-learning provably efficient? Advances in neural information processing systems , 31, 2018.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Is q-learning provably efficient? Advances in neural information processing systems , 31, 2018

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.999655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.782406Z digest=sha256:ad3c6cc347204dd4c2d2a6a0d7967fc6c3c97ea777962f9fa30e5f4dbcf8af7a

Observation c88f8ffa-82e1-41a6-9ab9-6bcc34a47707 · outbound

This paper cites Learning robust policy against disturbance in transition dynamics via state-conservative policy optimization.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Learning robust policy against disturbance in transition dynamics via state-conservative policy optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.849079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.965109Z digest=sha256:599aacc859ed7f58ea693303fd130db0d1b183897873737e7672a81c76655179

Observation c16151e2-04d4-42a4-a6bf-94912749d28a · outbound

This paper cites Policy gradient for rectangular robust markov decision processes.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Policy gradient for rectangular robust markov decision processes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.670114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.050545Z digest=sha256:c560f956fedb1183b4814196d952e413ee634369eadedb616ee00e659f2b412a

Observation 6b2800b2-64ec-4e46-b8c4-4e056478a7aa · outbound

This paper cites First-order Policy Optimization for Robust Markov Decision Process.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning First-order Policy Optimization for Robust Markov Decision Process

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:11.213481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:11.213481Z digest=sha256:d1cc41dc2cedf4e720dcf1ea1362c871a4adb3747d94a44e491221398e5983de

Observation 1bf65130-af37-4ad7-840a-5349010729bd · outbound

This paper cites Reinforcement learning in robust markov decision processes.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Reinforcement learning in robust markov decision processes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.458906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.344482Z digest=sha256:ebef3277bef346bfec190422daf94b19b44cc1d9c83d947f5061ed4652ec0fd7

Observation 62b951b7-0427-4e8c-8d7e-707f1a52bc06 · outbound

This paper cites Robustness in markov decision problems with uncertain transition matrices.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robustness in markov decision problems with uncertain transition matrices

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.289414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.492619Z digest=sha256:f051ff0c0519ba8f66bdf4f6c33bd26f62396f584e88e985701079c8031380a0

Observation fccc989e-2617-4fa9-b37f-b38b700b2ac1 · outbound

This paper cites Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:54:13.272094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.619397Z digest=sha256:45933cc9b3179894f6e9e9de6d5427c54f59d826ecde19427cb707abbbe04837

Observation 71dc82ed-7b97-47dd-8580-1cefb697f5b9 · outbound

This paper cites Policy optimization for robust average reward mdps.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Policy optimization for robust average reward mdps

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.126484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.789338Z digest=sha256:a5ef456ff7af9340f91e859583f3ff13359e0e040a7776e22549be4d32ad850d

Observation d18695a6-251c-4982-8605-2b4a842c01f6 · outbound

This paper cites u nderhauf, Oliver Brock, Walter Scheirer, Raia Hadsell, Dieter Fox, J \.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning u nderhauf, Oliver Brock, Walter Scheirer, Raia Hadsell, Dieter Fox, J \

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.913558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.912786Z digest=sha256:f0787a0f8cd5393e4e9efeb08f9c3f13f5bcf3f9f0a52565d7ab82f792ef8186

Observation a5de33fd-8f62-47de-b381-294e48cdd827 · outbound

This paper cites Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:11.972128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:11.972128Z digest=sha256:78f297d1d5a785533da657e57d9e7959572959ebe05539b975a3be9c6caaef4b

Observation 93319adb-9ce9-47a1-b0a8-f6ca9d1f2e87 · outbound

This paper cites A finite sample complexity bound for distributionally robust q-learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning A finite sample complexity bound for distributionally robust q-learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.746932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.041666Z digest=sha256:6c5cb6548f16bb2b9542e1aac6e12daea3d147305914d556540f2201c9527e9a

Observation 97c97d62-421b-4f4c-ae57-4b937c851a91 · outbound

This paper cites Sample complexity of variance-reduced distributionally robust q-learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Sample complexity of variance-reduced distributionally robust q-learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.582387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.176391Z digest=sha256:c16601cc08e87c9458e0f1ebfb5bd92594773a536f036766b8d6cfd78c1ff7c4

Observation 734d1291-44f4-4525-afbe-f989b6af7e69 · outbound

This paper cites Robust average-reward markov decision processes.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robust average-reward markov decision processes

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.431435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.288138Z digest=sha256:dd84f16285f7803afa5fe2a06f853212e8690db73ac3c8a18091460e1e9d5a74

Observation b1942abc-4d0c-4301-9d6a-10be62db657d · outbound

This paper cites Robust average-reward reinforcement learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robust average-reward reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.269153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.366541Z digest=sha256:92b5687905ef3b0106895c2686d906618095f0a95e2b42277dd86025bad7e685

Observation 5e0f1919-a5d5-4225-89fe-f4557761741c · outbound

This paper cites Model-free robust average-reward reinforcement learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-free robust average-reward reinforcement learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.072906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.439077Z digest=sha256:9f14f031bfa9fee91f859df893907d5f0298a03c66aa360bbc046bec44b99fb9

Observation 126f3854-5a0f-4100-be64-f9c5f82dfeac · outbound

This paper cites Policy gradient method for robust reinforcement learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Policy gradient method for robust reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:13.920610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.514568Z digest=sha256:a76c4c82dd68a03614be80b943ba0b643030c70dfe93a864a58ac2105e3b5418

Observation 06666b7c-fa97-4d1e-8f68-9f7089ee3c9a · outbound

This paper cites Model-free reinforcement learning in infinite-horizon average-reward markov decision processes.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-free reinforcement learning in infinite-horizon average-reward markov decision processes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:13.733990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.593503Z digest=sha256:4a978d82ca956c7e979e0d50ba3f646469c5d718a3c69940fd992794005973f9

Observation 8431c331-d2a7-44ed-b314-063e4a2c122d · outbound

This paper cites Finite-sample analysis of policy evaluation for robust average reward reinforcement learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Finite-sample analysis of policy evaluation for robust average reward reinforcement learning

Reference 27

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:54:13.058498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.692896Z digest=sha256:4a2b236c806832100007f567ad48b6fa6b7ffb54bfa3bd6a8c4ef8703d9b0731

Observation bf156b68-91e8-421d-837f-e06af6c5b5ae · outbound

This paper cites Natural actor-critic for robust reinforcement learning with function approximation.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Natural actor-critic for robust reinforcement learning with function approximation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:13.578820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.780877Z digest=sha256:54d8efe677f0e2dcc2b7581df4f156ca6f5c3c2f3a650a554a7c6be77bbb6a79

Pith citing papers

No inbound Pith citation observations are available.