Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:45:55.572614Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 2 inbound Pith citation observations for arXiv:2505.12462.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:45:55.572614Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:15.232280Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T05:54:13.187637Z
98 of 98 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4d26834a-d4ce-4fda-9441-52e06de6fbd9 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eb46f3e-0760-4950-887f-39c805aa1223 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Douzero: Mastering doudizhu with self-play deep reinforcement learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b707fc15-9dc1-4da5-97b7-f4f2f50d5602 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Honor of kings arena: an environment for generalization in competitive reinforcement learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c806d400-8e03-4809-81c7-e5ed855a8f02 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On efficient reinforcement learning for full-length game of starcraft ii.Journal of Artificial Intelligence Research, 75:213–260, 2022
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a664c7-e609-4c0f-baaf-8cd19f03444b · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sim-to-real transfer in deep reinforcement learning for robotics: a survey
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe4120d-b114-4d0f-8c58-fd72d0565cf0 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sim-to-real transfer of robotic control with dynamics randomization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7885c32-810a-4504-8718-63b1ba83a855 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Domain randomiza- tion for transferring deep neural networks from simulation to the real world
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ac9a28b-e264-41ae-81b6-f199731c73b6 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Deep reinforce- ment learning that matters.Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), 2018
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25bf0507-06e5-4627-80ed-fd0c9b82d074 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis EPOpt: Learning Robust Neural Network Policies Using Model Ensembles
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1772ca-b9d8-4a2d-a485-846850f30d09 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A Study on Overfitting in Deep Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08dfd164-3590-4ee2-aa1a-7d2bafadf905 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Solving uncertain Markov decision processes.Carnegie Mellon University, Technical Report, 2001
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ad3ab9-4db6-42cf-8c00-31354b7910e2 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robustness in Markov decision problems with uncertain transition matrices
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2e870ad-258f-4c4e-8d40-b7002998c790 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust dynamic programming.Mathematics of Operations Research, 30(2):257–280, 2005
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a6c6295-57b6-4160-849f-4f7d373254a0 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust adversarial reinforcement learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4182383a-6e32-4580-9a9c-2b84f367c332 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Atia, and Yue Wang
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e3e992f-0cc6-48ed-bf4e-7d7b9e3a16c6 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A reinforcement learning method for maximizing undiscounted rewards
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48519a1a-69e1-4789-9419-355483dd2feb · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis True online td (lambda)
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06b8b386-9eff-4e76-9fe0-db9c8775cd1c · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e188f0-2da0-495c-930f-e6ac0e546bfc · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Learning algorithms for Markov decision processes with average cost.SIAM Journal on Control and Optimization, 40(3):681–698, 2001
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9308df-8781-424a-b142-ecf4db2669bb · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8931bd48-0615-4b72-95af-ce4eed826168 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A dynamic pricing demand response algorithm for smart grid: Reinforcement learning approach.Applied energy, 220:220–230, 2018
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 47f7954c-59c0-4c84-a761-5655afba3ce8 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 45aec443-a977-4933-b7e0-e2c000978f6b · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 46338669-06e2-477f-b2b1-171aa9535c93 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Learning to trade via direct reinforcement.IEEE transactions on neural Networks, 12(4):875–889, 2001
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0d14b53b-01d8-40e8-a516-9178f6e221de · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reinforcement learning in economics and finance
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72157e8a-bd7e-4818-ba9f-e85f85f0fa9a · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis PhD thesis, Université d’Ottawa/University of Ottawa, 2021
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3fe792d5-1114-4f2e-ac90-bd4b1d151b1b · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Deep reinforcement learning model for stock portfolio management based on data fusion
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c504f33f-4f56-49b2-9216-1206f387d8e1 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis John Wiley & Sons, 2013
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312039e8-b8a7-4799-b573-f31a2ba735b3 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Model-free robust average-reward reinforcement learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac0ab864-bac9-42f7-8543-14b90c043a3e · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Beyond discounted returns: Robust Markov decision processes with average and Blackwell optimality
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff63a073-7486-48e5-a640-243d1f8a9835 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c3a3cf-0150-4f91-99a2-51666ec12ac3 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Span-Based Optimal Sample Complexity for Average Reward MDPs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ebba0abb-edd2-4f8c-bd25-42c2b4be8d86 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust average-reward Markov decision processes
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f1279985-aebc-48a8-9c57-449479ff1227 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7df5871f-2f1b-4e21-9a7d-d4dc57982ba2 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A reduction framework for distributionally robust reinforcement learning under average reward
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3ef557fc-1712-4ea4-a58e-221803ccd085 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 731997b1-12fc-4481-89e5-315d3905efa7 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finite-sample analysis of policy evaluation for robust average reward reinforcement learning.arXiv preprint arXiv:2502.16816, 2025
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9fbe852-c90f-46ad-be6b-e15f0270eb98 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample complexity of distributionally robust average-reward reinforce- ment learning.arXiv preprint arXiv:2505.10007, 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689c66a9-3d39-42a9-8118-180e9d37ad66 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Dynamic Programming and Optimal Control 3rd edition, volume II.Belmont, MA: Athena Scientific, 2011
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3bf91cab-7322-4085-89b1-f779c1e5d78b · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Fixed points of nonexpanding maps.Bulletin of the American Mathematical Society, 73(6):957–961, 1967
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 231e7c11-9dcc-4e25-b2fa-85442057c34b · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On the convergence rate of the Halpern-iteration.Optimization Letters, 15(2):405–418, 2021
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f6508e46-19b5-4062-942a-46c20f72112f · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Near-Optimal Sample Complexity for MDPs via Anchoring
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9c944db6-5aa7-41bb-a437-86e8f9e5fdf7 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Online robust reinforcement learning with model uncertainty
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 79954e55-b9b0-4351-b1de-d7a892a9c78a · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Policy gradient method for robust reinforcement learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4957033f-4b06-4e80-8eec-b8f1f563a3f2 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Minimax-Optimal Multi-Agent Robust Reinforcement Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a70223d0-d33b-4d53-ab1a-7619e15e51a2 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis An Efficient Solution to s-Rectangular Robust Markov Decision Processes
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a38a16e3-d437-45c0-b66f-ce924df465ff · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust Markov decision processes.Mathematics of Operations Research, 38(1):153–183, 2013
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 86f4ac45-2605-42ef-a246-ae1e404bc3ef · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample complexity of robust reinforcement learning with a generative model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b770c689-eddb-4ee5-8fe2-74dabab080e7 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73e592bd-9389-4697-843d-8bc9fcd08f98 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Improved sample complexity bounds for distributionally robust reinforcement learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2aedd444-f311-4d75-82a7-65b9c1fc94ec · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Learning and planning in average-reward Markov decision processes
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f0c00779-a547-44fc-9bcb-341c24aade4a · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 368010cc-a2a8-46d9-92c0-8c2b32488ac2 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc46e46-8a10-4048-8159-a798b104e511 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sharper model-free reinforcement learning for average-reward Markov decision processes
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 338b6fba-9057-4f73-81b2-e094b64a5a5f · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finite sample analysis of average-reward TD learning andQ-learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 16119503-5435-46f9-b367-31935e38d1b4 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A first order method for solving convex bilevel optimization problems.SIAM Journal on Optimization, 27(2):640–660, 2017
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8281b926-a697-47fe-9c30-0197123579d3 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Exact optimal accelerated complexity for fixed-point iterations
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a810d212-e376-4b4a-bd04-2c33a966d48e · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal error bounds for non-expansive fixed-point iterations in normed spaces.Mathematical Programming, 199(1):343–374, 2023
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e7f076f-034e-4a4a-bd36-18af452e264c · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal non-asymptotic rates of value iteration for average-reward markov decision processes.arXiv preprint arXiv:2504.09913, 2025
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0179c015-5c57-402a-8a54-9d2524c23b05 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ea37a3-fa74-4ac4-ad76-62c9d6d337a5 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Minimax optimal and computationally efficient algorithms for distributionally robust offline reinforcement learning.arXiv preprint arXiv:2403.09621, 2024
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41b163ad-42a6-4882-bdf9-0dca89716d64 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis McGill University (Canada), 2021
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5556fe4f-b56f-4df7-8f14-823e3d65a13a · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Distribution- ally robustQ-learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a6a48b0e-fe83-437e-a98e-c9454632999b · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Truncated Variance Reduced Value Iteration
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afefc198-8f6d-4aa6-af39-c1728aa97aae · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal approximation of average reward markov decision processes
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6291b4b6-4fc3-4311-bbb2-2b5e064cb378 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal Sample Complexity for Average Reward Markov Decision Processes
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bbcd3ccd-dc07-4cd4-a930-e02b95d2e698 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal Sample Complexity of Reinforcement Learning for Mixing Discounted Markov Decision Processes
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec06317d-4efe-436b-9f2a-14d16d134575 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Feasible q-learning for average reward reinforcement learning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f8309385-1e25-41c3-b22f-454f64601c69 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8ccf8a8-f416-4469-8e24-c1d9a1b5e8ad · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Tightening the dependence on horizon in the sample complexity of q-learning
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6d3af054-8da4-48b8-95a8-8e61ef415611 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Achieving the asymptotically minimax optimal sample complexity of offline reinforcement learning: A DRO-based approach.Transactions on Machine Learning Research, 2024
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3475e849-e554-4f88-a71e-ee3ac9f89f61 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Gambling in a rigged casino: The adversarial multi-armed bandit problem
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc64807-e11b-4f17-a8a1-5a588ff2601b · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis What Doubling Tricks Can and Can't Do for Multi-Armed Bandits
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d620034c-9e67-4f37-a518-62646140225c · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reducing blackwell and average optimality to discounted MDPs via the blackwell discount factor.Advances in Neural Information Processing Systems, 36, 2024
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3643a70c-a5ad-47f9-a32a-602042336f2d · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finding good policies in average-reward Markov Decision Processes without prior knowledge
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 797ef007-ed44-46d6-93e1-003df84b07f1 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Model-free robust reinforcement learning with sample complexity analysis
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bef91495-4b79-44ee-b2d2-151a0645a5ed · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Bounded parameter Markov decision processes with average reward criterion
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fcd4fd70-0b96-4efa-b7ef-a3b957949251 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Bellman optimality of average-reward robust markov decision processes with a constant gain.arXiv preprint arXiv:2509.14203, 2025
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cf7f92d9-6f21-470b-8ce1-f1505ebfcdd4 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Solving Long-run Average Reward Robust MDPs via Stochastic Games
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0323bc47-76a1-48a6-ab37-ffcac4ab488e · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reinforcement learning in robust Markov decision processes
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 83e1f644-ffc9-4576-b919-8d896900f98c · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Toward theoretical understandings of robust markov decision processes: Sample complexity and asymptotics.The Annals of Statistics, 50(6):3223–3248, 2022
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 288d8550-4d1c-473e-a74e-e10ee77594f4 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b5d591de-c93f-4a01-835b-3f9633b1cb0e · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A finite sample complexity bound for distributionally robustq-learning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ae1f8c9d-162f-4da7-a650-84c95a131da9 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Single-Trajectory Distributionally Robust Reinforcement Learning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d8134c-cc7e-45f0-9658-1ea0d8ee2f54 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample Complexity of Variance-reduced Distributionally Robust Q-learning
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bb58ba8b-5994-49a0-b605-3b8d7e2c0dc6 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Bring your own (non-robust) algorithm to solve robust mdps by estimating the worst kernel.arXiv e-prints, pages arXiv–2306, 2023
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b1595b8c-b9ef-4564-acbc-dafc5387a8dc · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Twice regularized MDPs and the equivalence between robustness and regularization
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 10f49b19-4ad2-4f2e-b373-bfebbe1e3168 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf25c8fe-24d5-4a71-b7e8-fee6f65a26e0 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample Complexity of Offline Distributionally Robust Linear Markov Decision Processes
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a08380-d943-469e-81c6-e6543d0afad0 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A unified principle of pessimism for offline reinforcement learning under model mismatch
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c237809b-e432-4ccc-8c39-684560ff0331 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8004d090-d143-47e0-b57e-770675366687 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Provably near-optimal distributionally robust reinforcement learning in online settings.arXiv preprint arXiv:2508.03768, 2025
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1ab2fa1c-d692-4b1d-8c1a-b1ed925655ca · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample complexity of distributionally robust off-dynamics reinforcement learning with online interaction
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3913a34b-a13e-439b-af82-cd7f19adc41b · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis John Wiley & Sons, 2014
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 934d4e26-d97e-477c-85f3-127962baae9c · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Average-reward model-free reinforcement learning: a systematic review and literature mapping
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 929fa1d8-5f50-4949-b9d1-fdd170f0395c · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Towards tight bounds on the sample complexity of average-reward mdps
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 98370f72-38ee-441f-8184-fbd7b96c6cf2 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Stochastic first-order methods for average-reward markov decision processes.Mathematics of Operations Research, 2024
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 518ee3bf-f958-46b1-bcba-8b3836481063 · outbound
Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On the generation of Markov decision processes.Journal of the Operational Research Society, 46(3):354–361, 1995
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fccc989e-2617-4fa9-b37f-b38b700b2ac1 · inbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d541f5cd-ef54-46b1-8f06-ffb1343a118f · inbound
Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.