Pith. sign in

Paper Citation Record · LEDGER

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

As of 17 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 2 inbound Pith citation observations for arXiv:2505.12462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12462 v3

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:45:55.572614Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:15.232280Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:54:13.187637Z

Reference resolution

98 of 98 outbound references displayed

  • verified exact12
  • verified fuzzy36
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d26834a-d4ce-4fda-9441-52e06de6fbd9 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.012652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.012652Z digest=sha256:21299c88719be1f1af8e43c83018cea6025b2f4570a4afb078e6b556f791a655

Observation 1eb46f3e-0760-4950-887f-39c805aa1223 · outbound

This paper cites Douzero: Mastering doudizhu with self-play deep reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Douzero: Mastering doudizhu with self-play deep reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.021461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.021461Z digest=sha256:3439a8ce4e1fdf66dfbcca5b748285bde509e9a9692b5394f15f6c55b55fdf32

Observation b707fc15-9dc1-4da5-97b7-f4f2f50d5602 · outbound

This paper cites Honor of kings arena: an environment for generalization in competitive reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Honor of kings arena: an environment for generalization in competitive reinforcement learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.026944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.026944Z digest=sha256:63b4e3dfe2e2707d60dc450bbac9af0a524aa0562690e1d891e3a222d4f4a855

Observation c806d400-8e03-4809-81c7-e5ed855a8f02 · outbound

This paper cites On efficient reinforcement learning for full-length game of starcraft ii.Journal of Artificial Intelligence Research, 75:213–260, 2022.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On efficient reinforcement learning for full-length game of starcraft ii.Journal of Artificial Intelligence Research, 75:213–260, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.032504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.032504Z digest=sha256:94b8c16d60f8586cb3d607593e554fb6bd7c95379988fb476d75f033e4bba12b

Observation e3a664c7-e609-4c0f-baaf-8cd19f03444b · outbound

This paper cites Sim-to-real transfer in deep reinforcement learning for robotics: a survey.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sim-to-real transfer in deep reinforcement learning for robotics: a survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.038467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.038467Z digest=sha256:568be788356596537b9d29fb020a7ccf9dbd1b05f8a8e7c0df3b92b9cc734d31

Observation 3fe4120d-b114-4d0f-8c58-fd72d0565cf0 · outbound

This paper cites Sim-to-real transfer of robotic control with dynamics randomization.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sim-to-real transfer of robotic control with dynamics randomization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.043424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.043424Z digest=sha256:d3abb2d28f41fa75ecdf7c70afdbd2798f2d52c17c53215352e923ab4669afd5

Observation f7885c32-810a-4504-8718-63b1ba83a855 · outbound

This paper cites Domain randomiza- tion for transferring deep neural networks from simulation to the real world.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Domain randomiza- tion for transferring deep neural networks from simulation to the real world

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.050564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.050564Z digest=sha256:8fdeebe8e498f6aebcdbda7f874f2bcb19c019a10dcbf10d9ba677af5ec102b4

Observation 4ac9a28b-e264-41ae-81b6-f199731c73b6 · outbound

This paper cites Deep reinforce- ment learning that matters.Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), 2018.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Deep reinforce- ment learning that matters.Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), 2018

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.056318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.056318Z digest=sha256:c1c385d6f05667be55526394d7eb575938d4816a95710228771e1d1aa65c7632

Observation 25bf0507-06e5-4627-80ed-fd0c9b82d074 · outbound

This paper cites EPOpt: Learning Robust Neural Network Policies Using Model Ensembles.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis EPOpt: Learning Robust Neural Network Policies Using Model Ensembles

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.061975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.061975Z digest=sha256:64b24f0770e5fda0c1035539cc73c205a2058865a744f9f383cd23ece0014f52

Observation bd1772ca-b9d8-4a2d-a485-846850f30d09 · outbound

This paper cites A Study on Overfitting in Deep Reinforcement Learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A Study on Overfitting in Deep Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.068985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.068985Z digest=sha256:dc8b508992aeff45f3a337ff25170ed8320514238f7479f27b4c3a72ec3dc083

Observation 08dfd164-3590-4ee2-aa1a-7d2bafadf905 · outbound

This paper cites Solving uncertain Markov decision processes.Carnegie Mellon University, Technical Report, 2001.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Solving uncertain Markov decision processes.Carnegie Mellon University, Technical Report, 2001

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.074365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.074365Z digest=sha256:b020de518b27eeef4339ccf1d50aaa3740e5cffe4f53a7c26c656a0d0af73a37

Observation 88ad3ab9-4db6-42cf-8c00-31354b7910e2 · outbound

This paper cites Robustness in Markov decision problems with uncertain transition matrices.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robustness in Markov decision problems with uncertain transition matrices

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.079730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.079730Z digest=sha256:e45914fc2d9bd9e70fd2d1321b3f6e226e05168f217d9797f0ba7db0fcf7e71f

Observation c2e870ad-258f-4c4e-8d40-b7002998c790 · outbound

This paper cites Robust dynamic programming.Mathematics of Operations Research, 30(2):257–280, 2005.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust dynamic programming.Mathematics of Operations Research, 30(2):257–280, 2005

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.085098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.085098Z digest=sha256:2e0dcfaaf50ed2fc8f034cc75ef17e725c4a44922ff6a1e71245f63f31f494c5

Observation 9a6c6295-57b6-4160-849f-4f7d373254a0 · outbound

This paper cites Robust adversarial reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust adversarial reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.091133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.091133Z digest=sha256:d7862da588e83d28876978ac92c88e86fb8e25bf15d96d6ac858d3b5f09dfc9c

Observation 4182383a-6e32-4580-9a9c-2b84f367c332 · outbound

This paper cites Atia, and Yue Wang.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Atia, and Yue Wang

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.096890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.096890Z digest=sha256:24edc7c7878162042b590453b22d3426cd0a725d1c2034e5faf4ff44ba3b4182

Observation 8e3e992f-0cc6-48ed-bf4e-7d7b9e3a16c6 · outbound

This paper cites A reinforcement learning method for maximizing undiscounted rewards.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A reinforcement learning method for maximizing undiscounted rewards

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.102249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.102249Z digest=sha256:c05300c1cd6e34a25dc8f5aad1c9e8da51fdf5629bedb5053bed955b82c863d5

Observation 48519a1a-69e1-4789-9419-355483dd2feb · outbound

This paper cites True online td (lambda).

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis True online td (lambda)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.108024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.108024Z digest=sha256:fe85b10ff33d579aeb4d5d4f96175dde335b11810b5f4a18f4c562f7702e0666

Observation 06b8b386-9eff-4e76-9fe0-db9c8775cd1c · outbound

This paper cites an unresolved cited work.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.114297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.114297Z digest=sha256:2e60d0c400da0db3f20640405acb3c75eb192c3b1e82c4619f9d1fdba266916f

Observation 10e188f0-2da0-495c-930f-e6ac0e546bfc · outbound

This paper cites Learning algorithms for Markov decision processes with average cost.SIAM Journal on Control and Optimization, 40(3):681–698, 2001.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Learning algorithms for Markov decision processes with average cost.SIAM Journal on Control and Optimization, 40(3):681–698, 2001

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.120109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.120109Z digest=sha256:b30113b74480aca665d5f5f00435ec6c9ab3994b49e8ed50964b5245265ad7a7

Observation 5a9308df-8781-424a-b142-ecf4db2669bb · outbound

This paper cites Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.125399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.125399Z digest=sha256:f678c8e6e7e16eb073d7836c97ebefff0b9c9db7a7319e2ed58d878d2f74e841

Observation 8931bd48-0615-4b72-95af-ce4eed826168 · outbound

This paper cites A dynamic pricing demand response algorithm for smart grid: Reinforcement learning approach.Applied energy, 220:220–230, 2018.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A dynamic pricing demand response algorithm for smart grid: Reinforcement learning approach.Applied energy, 220:220–230, 2018

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.649236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.130372Z digest=sha256:df82adc2b9efb39408c39347863ad2c026d7e364db3edba35c96fd3b1fd3ad29

Observation 47f7954c-59c0-4c84-a761-5655afba3ce8 · outbound

This paper cites an unresolved cited work.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:45:57.625804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.135451Z digest=sha256:6cbf14d89f40b108f2c5193731cd50006382baed4af664b220a9865f6daa999f

Observation 45aec443-a977-4933-b7e0-e2c000978f6b · outbound

This paper cites an unresolved cited work.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:45:57.603775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.140975Z digest=sha256:3ab7ad551bd11a3aedce52df16e26fb50030d1c58e75911c46e22c305413cbc7

Observation 46338669-06e2-477f-b2b1-171aa9535c93 · outbound

This paper cites Learning to trade via direct reinforcement.IEEE transactions on neural Networks, 12(4):875–889, 2001.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Learning to trade via direct reinforcement.IEEE transactions on neural Networks, 12(4):875–889, 2001

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.580920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.147357Z digest=sha256:8dc31409ac89bb49b08e161e509ca43eb6304afa18fe37592f895128b016a8f6

Observation 0d14b53b-01d8-40e8-a516-9178f6e221de · outbound

This paper cites Reinforcement learning in economics and finance.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reinforcement learning in economics and finance

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.560183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.153980Z digest=sha256:55b5626ee7f13aad7419c0ba1f064d6641ef5c395277339dbccea564506fd8a4

Observation 72157e8a-bd7e-4818-ba9f-e85f85f0fa9a · outbound

This paper cites PhD thesis, Université d’Ottawa/University of Ottawa, 2021.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis PhD thesis, Université d’Ottawa/University of Ottawa, 2021

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.538739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.160356Z digest=sha256:e3be813a30c98a78ebb648053994f178f1cccb7a2784992f3ecf57cd768b585f

Observation 3fe792d5-1114-4f2e-ac90-bd4b1d151b1b · outbound

This paper cites Deep reinforcement learning model for stock portfolio management based on data fusion.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Deep reinforcement learning model for stock portfolio management based on data fusion

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.516416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.166554Z digest=sha256:d8b65936930f9369b1015587e8d7fe055d76eb06a8d8bccc71f35682e62d7641

Observation c504f33f-4f56-49b2-9216-1206f387d8e1 · outbound

This paper cites John Wiley & Sons, 2013.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis John Wiley & Sons, 2013

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.171757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.171757Z digest=sha256:64d3e74c1b59b3ab3fda2c86668712eba72f8422f6982820d39fa873a984bf15

Observation 312039e8-b8a7-4799-b573-f31a2ba735b3 · outbound

This paper cites Model-free robust average-reward reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Model-free robust average-reward reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.177706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.177706Z digest=sha256:3f01078caac233b9a45d9b8e5b18a2169045ae8215cfc1440a58076fa24b02d0

Observation ac0ab864-bac9-42f7-8543-14b90c043a3e · outbound

This paper cites Beyond discounted returns: Robust Markov decision processes with average and Blackwell optimality.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Beyond discounted returns: Robust Markov decision processes with average and Blackwell optimality

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.182579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.182579Z digest=sha256:fb5d91ff7a03536f20c6bc1a5e36a9210e9d3d26aa3a0e1bb7262a2c6fb4ceb9

Observation ff63a073-7486-48e5-a640-243d1f8a9835 · outbound

This paper cites Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.187895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.187895Z digest=sha256:6ad1278a329c7364be4729ef12909f6507dd4383657d311b3d015e1b2c53c1eb

Observation 03c3a3cf-0150-4f91-99a2-51666ec12ac3 · outbound

This paper cites Span-Based Optimal Sample Complexity for Average Reward MDPs.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Span-Based Optimal Sample Complexity for Average Reward MDPs

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.598129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.193656Z digest=sha256:9e6c9f6372481688aa38085377c7eb950e20bdfa282436adfe069e21e88ddcce

Observation ebba0abb-edd2-4f8c-bd25-42c2b4be8d86 · outbound

This paper cites Robust average-reward Markov decision processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust average-reward Markov decision processes

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.448638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.199546Z digest=sha256:493823f1e270f60cbf0b35ff2766248ee39ba45aaedb10048546a1cba0983722

Observation f1279985-aebc-48a8-9c57-449479ff1227 · outbound

This paper cites Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.570520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.205307Z digest=sha256:56d5ab08a0dd1295d0bff8b79435e8cfb32ba32a5bdf516a0e86314a00b26e5e

Observation 7df5871f-2f1b-4e21-9a7d-d4dc57982ba2 · outbound

This paper cites A reduction framework for distributionally robust reinforcement learning under average reward.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A reduction framework for distributionally robust reinforcement learning under average reward

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.427393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.210962Z digest=sha256:0b154dcf894083c4d51512ef2034e0b112c7c75b3ee25fff4a84a4909420acc1

Observation 3ef557fc-1712-4ea4-a58e-221803ccd085 · outbound

This paper cites Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.540031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.215882Z digest=sha256:cf269a05c0a6a7376efc705eb618d22bd45b34609859adfe484c55587060d48e

Observation 731997b1-12fc-4481-89e5-315d3905efa7 · outbound

This paper cites Finite-sample analysis of policy evaluation for robust average reward reinforcement learning.arXiv preprint arXiv:2502.16816, 2025.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finite-sample analysis of policy evaluation for robust average reward reinforcement learning.arXiv preprint arXiv:2502.16816, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.221589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.221589Z digest=sha256:42ae335afe759f2e89e4c2c3668638416b0f07239ebbb06c4f752153f832d461

Observation e9fbe852-c90f-46ad-be6b-e15f0270eb98 · outbound

This paper cites Sample complexity of distributionally robust average-reward reinforce- ment learning.arXiv preprint arXiv:2505.10007, 2025.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample complexity of distributionally robust average-reward reinforce- ment learning.arXiv preprint arXiv:2505.10007, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.226946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.226946Z digest=sha256:546829c43910a41665c8ba427a82de620aff7bd3aac0b3f02383581472dfc2da

Observation 689c66a9-3d39-42a9-8118-180e9d37ad66 · outbound

This paper cites Dynamic Programming and Optimal Control 3rd edition, volume II.Belmont, MA: Athena Scientific, 2011.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Dynamic Programming and Optimal Control 3rd edition, volume II.Belmont, MA: Athena Scientific, 2011

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.403509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.233628Z digest=sha256:2afb8064a2330f59221d1fa09b5e7b6f342aef74e94a2879843232b2af78366e

Observation 3bf91cab-7322-4085-89b1-f779c1e5d78b · outbound

This paper cites Fixed points of nonexpanding maps.Bulletin of the American Mathematical Society, 73(6):957–961, 1967.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Fixed points of nonexpanding maps.Bulletin of the American Mathematical Society, 73(6):957–961, 1967

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.374939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.239150Z digest=sha256:4de04fc4be4b9696f9321665a4fb2dec200329386d0d893a2fb3c88abdcb6403

Observation 231e7c11-9dcc-4e25-b2fa-85442057c34b · outbound

This paper cites On the convergence rate of the Halpern-iteration.Optimization Letters, 15(2):405–418, 2021.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On the convergence rate of the Halpern-iteration.Optimization Letters, 15(2):405–418, 2021

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.351750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.244118Z digest=sha256:180893e2a034170247cefc4347d31172a19b8be460f41e2bd1f7dbe824368344

Observation f6508e46-19b5-4062-942a-46c20f72112f · outbound

This paper cites Near-Optimal Sample Complexity for MDPs via Anchoring.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Near-Optimal Sample Complexity for MDPs via Anchoring

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.372929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.249423Z digest=sha256:c55c51d630fec8e4a6b673672da98a9683a9e579349eff186aead249c8908fc0

Observation 9c944db6-5aa7-41bb-a437-86e8f9e5fdf7 · outbound

This paper cites Online robust reinforcement learning with model uncertainty.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Online robust reinforcement learning with model uncertainty

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.327574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.256185Z digest=sha256:43f1ef97ee156edfce737e9737bea3f18e98a169e9bcd4e91fdcda97816b847e

Observation 79954e55-b9b0-4351-b1de-d7a892a9c78a · outbound

This paper cites Policy gradient method for robust reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Policy gradient method for robust reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.305017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.261275Z digest=sha256:9fdbfda6812f9c9159afd33d1c14e54dc5cac81c165ac357d76f3fe58e89fb4b

Observation 4957033f-4b06-4e80-8eec-b8f1f563a3f2 · outbound

This paper cites Minimax-Optimal Multi-Agent Robust Reinforcement Learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Minimax-Optimal Multi-Agent Robust Reinforcement Learning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.347346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.267200Z digest=sha256:cfa50a828d61ed7ad14c8cf5778cd0d3f131e9a8edfffff4ec372a0b680fe5c3

Observation a70223d0-d33b-4d53-ab1a-7619e15e51a2 · outbound

This paper cites An Efficient Solution to s-Rectangular Robust Markov Decision Processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis An Efficient Solution to s-Rectangular Robust Markov Decision Processes

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.273305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.273305Z digest=sha256:c819f0d4dea3615b606bf4af500356479d4347aa5dddf0678284c6eca18137d9

Observation a38a16e3-d437-45c0-b66f-ce924df465ff · outbound

This paper cites Robust Markov decision processes.Mathematics of Operations Research, 38(1):153–183, 2013.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust Markov decision processes.Mathematics of Operations Research, 38(1):153–183, 2013

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.282637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.278368Z digest=sha256:f57a378c33c0e143f2bff32342bf7e6361e90d7117b0306660ae57df79efe9b0

Observation 86f4ac45-2605-42ef-a246-ae1e404bc3ef · outbound

This paper cites Sample complexity of robust reinforcement learning with a generative model.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample complexity of robust reinforcement learning with a generative model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.283480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.283480Z digest=sha256:21f6285ff1051ba3a246182f0ae4f3ef7658d5f3df29cd2961f6fff8dd0fcc6e

Observation b770c689-eddb-4ee5-8fe2-74dabab080e7 · outbound

This paper cites The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.288950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.288950Z digest=sha256:e63f9b7e6262eb524b1470c936ec36124c01baeb01739916f4e3c02876fd29d8

Observation 73e592bd-9389-4697-843d-8bc9fcd08f98 · outbound

This paper cites Improved sample complexity bounds for distributionally robust reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Improved sample complexity bounds for distributionally robust reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.244736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.294763Z digest=sha256:2fcbe69f9bb782ccf5113b54eadf74f509d4c44a7fa5aa8a566b91937eb2d82b

Observation 2aedd444-f311-4d75-82a7-65b9c1fc94ec · outbound

This paper cites Learning and planning in average-reward Markov decision processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Learning and planning in average-reward Markov decision processes

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.223307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.300280Z digest=sha256:178f5cf6d84c6179e89a072470a24f1981bb0036a6f0bfd6a4b3ac8e4a4567f4

Observation f0c00779-a547-44fc-9bcb-341c24aade4a · outbound

This paper cites On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.305675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.305675Z digest=sha256:418c0e624852d8df331edd4739fc19e1f67e997d192f6650b9c4e382d0d2750b

Observation 368010cc-a2a8-46d9-92c0-8c2b32488ac2 · outbound

This paper cites The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.311465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.311465Z digest=sha256:bade3e9429f831ad18d793e8f0e9bd33f4cd659e0b3236625a6c12b2671c6a24

Observation 4fc46e46-8a10-4048-8159-a798b104e511 · outbound

This paper cites Sharper model-free reinforcement learning for average-reward Markov decision processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sharper model-free reinforcement learning for average-reward Markov decision processes

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.200349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.318155Z digest=sha256:964c062703c51245973f17dcbe1167af63e8aa7c6a1e64174a2677ec99174130

Observation 338b6fba-9057-4f73-81b2-e094b64a5a5f · outbound

This paper cites Finite sample analysis of average-reward TD learning andQ-learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finite sample analysis of average-reward TD learning andQ-learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.180856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.323713Z digest=sha256:44ff359212337baf9867dcedc2eb55fab24141219afa9ee093972ecca6b3d779

Observation 16119503-5435-46f9-b367-31935e38d1b4 · outbound

This paper cites A first order method for solving convex bilevel optimization problems.SIAM Journal on Optimization, 27(2):640–660, 2017.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A first order method for solving convex bilevel optimization problems.SIAM Journal on Optimization, 27(2):640–660, 2017

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.330060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.330060Z digest=sha256:0b2741aabde3141f01444eaee2370136a1e85f9bcc6546450e60a79101a86750

Observation 8281b926-a697-47fe-9c30-0197123579d3 · outbound

This paper cites Exact optimal accelerated complexity for fixed-point iterations.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Exact optimal accelerated complexity for fixed-point iterations

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.147868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.335227Z digest=sha256:e410a29b1ffdcc8dce3002ca736ad751b9b7ee67c4608079326364508aa53f28

Observation a810d212-e376-4b4a-bd04-2c33a966d48e · outbound

This paper cites Optimal error bounds for non-expansive fixed-point iterations in normed spaces.Mathematical Programming, 199(1):343–374, 2023.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal error bounds for non-expansive fixed-point iterations in normed spaces.Mathematical Programming, 199(1):343–374, 2023

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.341859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.341859Z digest=sha256:bdfbeb41843fbbe8f3c8edefe56fc5fe5e33218c25325acb831fb03878d12071

Observation 7e7f076f-034e-4a4a-bd36-18af452e264c · outbound

This paper cites Optimal non-asymptotic rates of value iteration for average-reward markov decision processes.arXiv preprint arXiv:2504.09913, 2025.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal non-asymptotic rates of value iteration for average-reward markov decision processes.arXiv preprint arXiv:2504.09913, 2025

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:45:56.251252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.346894Z digest=sha256:b460b709f8c06f972bde7701196282ae17c25124528d23e28a2e01c4ca189f17

Observation 0179c015-5c57-402a-8a54-9d2524c23b05 · outbound

This paper cites Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.352235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.352235Z digest=sha256:42d6cbbafaf93c6adc7facb61c2c492c221d76dec4c235596d3a1a7437058501

Observation 73ea37a3-fa74-4ac4-ad76-62c9d6d337a5 · outbound

This paper cites Minimax optimal and computationally efficient algorithms for distributionally robust offline reinforcement learning.arXiv preprint arXiv:2403.09621, 2024.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Minimax optimal and computationally efficient algorithms for distributionally robust offline reinforcement learning.arXiv preprint arXiv:2403.09621, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.358774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.358774Z digest=sha256:23564d14f4febae7fa1dd6244469bee9c8745870a9d1eea56d9579527519af92

Observation 41b163ad-42a6-4882-bdf9-0dca89716d64 · outbound

This paper cites McGill University (Canada), 2021.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis McGill University (Canada), 2021

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.107802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.365701Z digest=sha256:778ae8ba0c341c86688c9f993e988a160ca608143684a67f1958c39e85ec212d

Observation 5556fe4f-b56f-4df7-8f14-823e3d65a13a · outbound

This paper cites Distribution- ally robustQ-learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Distribution- ally robustQ-learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.085211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.370858Z digest=sha256:83138a4347c8e2752394f8f99e60de70bbd8bcf40fd69332339bfef78f7ac340

Observation a6a48b0e-fe83-437e-a98e-c9454632999b · outbound

This paper cites Truncated Variance Reduced Value Iteration.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Truncated Variance Reduced Value Iteration

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.376217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.376217Z digest=sha256:8432b76160b118d318caa83e0e74d72841d30a0aef9275bcc24b2f16a571a290

Observation afefc198-8f6d-4aa6-af39-c1728aa97aae · outbound

This paper cites Optimal approximation of average reward markov decision processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal approximation of average reward markov decision processes

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.061174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.383742Z digest=sha256:2ebac864d23427b1399696b415da69f8d88827343dbfc797c40e4bbf5bb7f669

Observation 6291b4b6-4fc3-4311-bbb2-2b5e064cb378 · outbound

This paper cites Optimal Sample Complexity for Average Reward Markov Decision Processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal Sample Complexity for Average Reward Markov Decision Processes

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.060257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.389158Z digest=sha256:338123dfb94d7cd820d4b9cce1e22a98733a5325db05cc8b508dbd701f5ff16e

Observation bbcd3ccd-dc07-4cd4-a930-e02b95d2e698 · outbound

This paper cites Optimal Sample Complexity of Reinforcement Learning for Mixing Discounted Markov Decision Processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal Sample Complexity of Reinforcement Learning for Mixing Discounted Markov Decision Processes

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.034715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.394758Z digest=sha256:004e0bbf3f08f6c3bd0d6eda1d98963465737688e93cac40813d7d3528ca8e73

Observation ec06317d-4efe-436b-9f2a-14d16d134575 · outbound

This paper cites Feasible q-learning for average reward reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Feasible q-learning for average reward reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.037853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.400232Z digest=sha256:19f0d7ac4448df16d0aca643b96374ad94c30ca40b796c645258c22b29df5c27

Observation f8309385-1e25-41c3-b22f-454f64601c69 · outbound

This paper cites Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.405318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.405318Z digest=sha256:10f232cee26164435a81db3ed14f3fa606927b0b1be1e8823e35bc1d5cbf8ca1

Observation d8ccf8a8-f416-4469-8e24-c1d9a1b5e8ad · outbound

This paper cites Tightening the dependence on horizon in the sample complexity of q-learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Tightening the dependence on horizon in the sample complexity of q-learning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.017859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.411017Z digest=sha256:e6703b0c85dd6ab375d3a127cfb8cdaea75a1a9c098f33254deeef499d507e8b

Observation 6d3af054-8da4-48b8-95a8-8e61ef415611 · outbound

This paper cites Achieving the asymptotically minimax optimal sample complexity of offline reinforcement learning: A DRO-based approach.Transactions on Machine Learning Research, 2024.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Achieving the asymptotically minimax optimal sample complexity of offline reinforcement learning: A DRO-based approach.Transactions on Machine Learning Research, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.998296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.416316Z digest=sha256:8b475b6e10ad7932d9a77909f3b5dc3ff0d6b8f10ee09a302ebb2289e144a6cb

Observation 3475e849-e554-4f88-a71e-ee3ac9f89f61 · outbound

This paper cites Gambling in a rigged casino: The adversarial multi-armed bandit problem.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Gambling in a rigged casino: The adversarial multi-armed bandit problem

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.421904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.421904Z digest=sha256:a71fd1d12511d27f6b7ec2b727ac0adcf52939ffda7a92fae586821a348bbadf

Observation fbc64807-e11b-4f17-a8a1-5a588ff2601b · outbound

This paper cites What Doubling Tricks Can and Can't Do for Multi-Armed Bandits.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis What Doubling Tricks Can and Can't Do for Multi-Armed Bandits

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.427082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.427082Z digest=sha256:318b2c9cb41a54e433f40cfa622f3c285ff40a01beadf7494aa7db2dbea0e1dc

Observation d620034c-9e67-4f37-a518-62646140225c · outbound

This paper cites Reducing blackwell and average optimality to discounted MDPs via the blackwell discount factor.Advances in Neural Information Processing Systems, 36, 2024.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reducing blackwell and average optimality to discounted MDPs via the blackwell discount factor.Advances in Neural Information Processing Systems, 36, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.965390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.432919Z digest=sha256:63b5f1d83f5b4f65fb32e20e3f9f9cdfee467ed443beb99e1ecacc6c99baa7fe

Observation 3643a70c-a5ad-47f9-a32a-602042336f2d · outbound

This paper cites Finding good policies in average-reward Markov Decision Processes without prior knowledge.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finding good policies in average-reward Markov Decision Processes without prior knowledge

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:55.971177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.437973Z digest=sha256:e93500beaf5dd179b7280fa81139c4cc778658bd9d6c8e47df9351080981c62c

Observation 797ef007-ed44-46d6-93e1-003df84b07f1 · outbound

This paper cites Model-free robust reinforcement learning with sample complexity analysis.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Model-free robust reinforcement learning with sample complexity analysis

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.947229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.443468Z digest=sha256:788ee2847c6b25af8839ee67adbc042eb90a4db02974fdca9b4070b5261b3e6d

Observation bef91495-4b79-44ee-b2d2-151a0645a5ed · outbound

This paper cites Bounded parameter Markov decision processes with average reward criterion.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Bounded parameter Markov decision processes with average reward criterion

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.924798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.449685Z digest=sha256:add466148db22858b0117a30412734b13cc282dae82ea0a95226b3503f3be0ee

Observation fcd4fd70-0b96-4efa-b7ef-a3b957949251 · outbound

This paper cites Bellman optimality of average-reward robust markov decision processes with a constant gain.arXiv preprint arXiv:2509.14203, 2025.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Bellman optimality of average-reward robust markov decision processes with a constant gain.arXiv preprint arXiv:2509.14203, 2025

Reference 78

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:45:55.945028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.456436Z digest=sha256:074f6b75a5c9fa8173c352f84b543e9f5b283bc620c7ed193f6b21bc713b0a4b

Observation cf7f92d9-6f21-470b-8ce1-f1505ebfcdd4 · outbound

This paper cites Solving Long-run Average Reward Robust MDPs via Stochastic Games.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Solving Long-run Average Reward Robust MDPs via Stochastic Games

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.462391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.462391Z digest=sha256:3428ae8d420341a5915f0bdf4b360b5a1ae001613126a77390186bcb2d7d5714

Observation 0323bc47-76a1-48a6-ab37-ffcac4ab488e · outbound

This paper cites Reinforcement learning in robust Markov decision processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reinforcement learning in robust Markov decision processes

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.902355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.470239Z digest=sha256:0a83be9e62f1e289bb09ee59b0d19e75651cc6a8932a5484be8079ad94c35de9

Observation 83e1f644-ffc9-4576-b919-8d896900f98c · outbound

This paper cites Toward theoretical understandings of robust markov decision processes: Sample complexity and asymptotics.The Annals of Statistics, 50(6):3223–3248, 2022.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Toward theoretical understandings of robust markov decision processes: Sample complexity and asymptotics.The Annals of Statistics, 50(6):3223–3248, 2022

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.476517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.476517Z digest=sha256:7567f5810df093e650d675b5e4515209519f9cc191b6ddebae7086114a72e034

Observation 288d8550-4d1c-473e-a74e-e10ee77594f4 · outbound

This paper cites Finite-sample regret bound for distributionally robust offline tabular reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finite-sample regret bound for distributionally robust offline tabular reinforcement learning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.871560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.483191Z digest=sha256:0c8bf565d502fa276046ffbc00bec1e3c8e713e8c76d5e37295272b65d011710

Observation b5d591de-c93f-4a01-835b-3f9633b1cb0e · outbound

This paper cites A finite sample complexity bound for distributionally robustq-learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A finite sample complexity bound for distributionally robustq-learning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.851425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.488217Z digest=sha256:f587bd386e1b033db013575dd808f6d1ae60d7437536c70d3a4de8e76fe55927

Observation ae1f8c9d-162f-4da7-a650-84c95a131da9 · outbound

This paper cites Single-Trajectory Distributionally Robust Reinforcement Learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Single-Trajectory Distributionally Robust Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.493828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.493828Z digest=sha256:3fcc94f30e923ad950806ba5f4dd6ccacd5e24b97c659de870b0c9f2412715e7

Observation 64d8134c-cc7e-45f0-9658-1ea0d8ee2f54 · outbound

This paper cites Sample Complexity of Variance-reduced Distributionally Robust Q-learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample Complexity of Variance-reduced Distributionally Robust Q-learning

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:55.821668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.499344Z digest=sha256:bbefdabf3822a3baa44dfdb365b7be7b33ce87953c6fb8366792b43b09981496

Observation bb58ba8b-5994-49a0-b605-3b8d7e2c0dc6 · outbound

This paper cites Bring your own (non-robust) algorithm to solve robust mdps by estimating the worst kernel.arXiv e-prints, pages arXiv–2306, 2023.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Bring your own (non-robust) algorithm to solve robust mdps by estimating the worst kernel.arXiv e-prints, pages arXiv–2306, 2023

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.831060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.504737Z digest=sha256:24ebb472004ef4b4fbf44e017b263007f00a966f6c40e305b8785f28f0e86265

Observation b1595b8c-b9ef-4564-acbc-dafc5387a8dc · outbound

This paper cites Twice regularized MDPs and the equivalence between robustness and regularization.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Twice regularized MDPs and the equivalence between robustness and regularization

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.812885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.509603Z digest=sha256:a5547d2204f97e879c6d85ea7a5c7f6e1aacf1804ae631c7ff5ced619970571b

Observation 10f49b19-4ad2-4f2e-b373-bfebbe1e3168 · outbound

This paper cites Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.514519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.514519Z digest=sha256:0798d18f2ce37a355903954e8aef8b357d20267f6dfea08fbb2c9ad51069f942

Observation bf25c8fe-24d5-4a71-b7e8-fee6f65a26e0 · outbound

This paper cites Sample Complexity of Offline Distributionally Robust Linear Markov Decision Processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample Complexity of Offline Distributionally Robust Linear Markov Decision Processes

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.519507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.519507Z digest=sha256:6ed77bf42066a103fb5753573c2c36a5ab4eb01ca115655b568a5861431d3a28

Observation 44a08380-d943-469e-81c6-e6543d0afad0 · outbound

This paper cites A unified principle of pessimism for offline reinforcement learning under model mismatch.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A unified principle of pessimism for offline reinforcement learning under model mismatch

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.793150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.525300Z digest=sha256:8fbe0715d287c06858fef853c29ef14be200dfa423fc8623c2e7d068f8dfe6ba

Observation c237809b-e432-4ccc-8c39-684560ff0331 · outbound

This paper cites Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.531036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.531036Z digest=sha256:d485c2c136276851a08da6397881373c6f162072ef289fe800913edfe5d3f5f2

Observation 8004d090-d143-47e0-b57e-770675366687 · outbound

This paper cites Provably near-optimal distributionally robust reinforcement learning in online settings.arXiv preprint arXiv:2508.03768, 2025.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Provably near-optimal distributionally robust reinforcement learning in online settings.arXiv preprint arXiv:2508.03768, 2025

Reference 92

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:45:55.736471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.536974Z digest=sha256:46e9be280ae6ddfda495a0b1b79f863ba0c7213c26546d00e7896cb468e76b95

Observation 1ab2fa1c-d692-4b1d-8c1a-b1ed925655ca · outbound

This paper cites Sample complexity of distributionally robust off-dynamics reinforcement learning with online interaction.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample complexity of distributionally robust off-dynamics reinforcement learning with online interaction

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.542380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.542380Z digest=sha256:75fff768c92a41e173df2cc75c847f6e4f82d16ccfdc76102a61511b522b4fe6

Observation 3913a34b-a13e-439b-af82-cd7f19adc41b · outbound

This paper cites John Wiley & Sons, 2014.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis John Wiley & Sons, 2014

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.548339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.548339Z digest=sha256:850bbbbc7ea7680fb7e7c19d36748cbdaa5027a1931e6e91186372742869bb52

Observation 934d4e26-d97e-477c-85f3-127962baae9c · outbound

This paper cites Average-reward model-free reinforcement learning: a systematic review and literature mapping.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Average-reward model-free reinforcement learning: a systematic review and literature mapping

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.553994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.553994Z digest=sha256:cf3f88ba7484804d3e0bd0851722145bec459fb9023d6b1b9d997636d3a0656b

Observation 929fa1d8-5f50-4949-b9d1-fdd170f0395c · outbound

This paper cites Towards tight bounds on the sample complexity of average-reward mdps.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Towards tight bounds on the sample complexity of average-reward mdps

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.746953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.560171Z digest=sha256:86866677991a032b411453eb88581134f8e917a4eadafc24853e85eb5e428455

Observation 98370f72-38ee-441f-8184-fbd7b96c6cf2 · outbound

This paper cites Stochastic first-order methods for average-reward markov decision processes.Mathematics of Operations Research, 2024.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Stochastic first-order methods for average-reward markov decision processes.Mathematics of Operations Research, 2024

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.726923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.565897Z digest=sha256:51212d91753555f90e9f1330da58560e349347ef15ff43633e9d0a5a11deb5c7

Observation 518ee3bf-f958-46b1-bcba-8b3836481063 · outbound

This paper cites On the generation of Markov decision processes.Journal of the Operational Research Society, 46(3):354–361, 1995.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On the generation of Markov decision processes.Journal of the Operational Research Society, 46(3):354–361, 1995

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.706015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:45:55.572614Z digest=sha256:dbdd10e407ccd3ae262ae7b9c05255d53d8cb496e83cbc301e5e1ca2d06ecad1

Pith citing papers

Observation fccc989e-2617-4fa9-b37f-b38b700b2ac1 · inbound

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning cites this paper.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:54:13.272094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.619397Z digest=sha256:910cf626d3ec02d9ddc0e9b6562c485b34c33446497a6a1bb26e44d5025c2f15

Observation d541f5cd-ef54-46b1-8f06-ffb1343a118f · inbound

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions cites this paper.

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:15.232280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:15.232280Z digest=sha256:950174b59f409a0d5342f725ff1ece6def36186cb95067a3f9664e83fc715362