Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:42:09.234480Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2504.11997.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:42:09.234480Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bf20c210-1e8f-4150-a5b6-21453d6537d1 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Improved algorithms for linear stochastic bandits
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5f0c9a4d-d416-4a2a-9d54-0666df91980e · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Near-opt imal regret bounds for reinforcement learn- ing
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1ea6fd89-cdc6-48f4-8e41-2abc1102bf37 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Model-based reinforcement learning with value-targeted regression
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 002223ee-3be3-4914-972c-8bb3ae082ce9 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs REGAL: a regularizati on based algorithm for reinforcement learn- ing in weakly communicating MDPs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 31c5fe95-e827-45d8-9aee-9d599ebdc74a · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Learning Infinite- Horizon Average-Reward Linear Mixture MDPs of Bounded Span
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3c9707ba-a4c4-41bd-9b5b-39656b819037 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Efficient bias-span- constrained exploration-exploitation in reinforcement l earning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d93434ac-c213-49e1-85fb-85fb567db066 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Inven tory management in supply chains: a re- inforcement learning approach
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4fda9f0f-7d61-40fc-a096-0838a5d95c09 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Can deep reinforce- ment learning improve inventory management? performance o n lost sales, dual-sourcing, and multi- echelon problems
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 98eaf7ce-6c45-4b91-965b-db82a068ab8d · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Reinforcement learning for long-run a verage cost
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a373189d-04ed-4aea-bb3f-17d2dc51226c · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Sample-effi cient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 24fdec2c-618e-4828-b6ce-4d8d17d54ab2 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Reinforcement Learn- ing for Infinite-Horizon Average-Reward Linear MDPs via App roximation by Discounted-Reward MDPs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e391adbc-7cf0-4878-8159-94883c14b200 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Provably efficient reinforcement learning with linear function approximation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 657a964b-9e31-40c9-b403-68f9bd3a29ce · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Towards tight bounds on th e sample complexity of average-reward MDPs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cedffd33-f0b3-448f-8bf8-2e0fd7300a92 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Reinforcement learning based routin g in networks: Review and classification of approaches
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0c11398c-0252-4f53-9d54-307eef909972 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Sample complexity of reinforcement learning using linearly combined model ensembles
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1dbf06e3-abd3-41d1-b432-3413b6dc6477 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217f2d16-2c1e-46c1-8c2e-141cdeb8a9cc · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Optimal Sample Complexity for Average Reward Markov Decision Processes
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67fd2d37-9401-4a37-83f0-86ede1737eaa · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Learning infinite-horizon average-reward mdps with linear function approximation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 53a1d58c-2ca8-4d3b-b05c-fcaf71db6ffc · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Model-free reinforcement learning in infinite-horizon average-rewar d markov decision processes
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d2731aac-4d0a-458c-b9a5-3e15acf4d278 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Nearly minimax o ptimal regret for learning infinite- horizon average-reward mdps with linear function approxim ation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0345c7d6-3486-466a-958a-fabedd69edf8 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Joint optimiz ation of preventive maintenance and production scheduling for multi-state production systems based on reinforcement learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 990b21e2-c469-466b-a0bc-df1453fd3a51 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Regret minimization for reinforcement learning by evaluating the optimal bias function
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d57704ee-93b7-40f9-8cb2-ac9df73c03b4 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Sharper Model-free Reinf orcement Learning for Average-reward Markov Decision Processes
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 486aa4c3-e495-49a0-8ec6-1b840927354f · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Span-Based Optimal Sample Complexity for Average Reward MDPs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c22267d-d48d-4aa0-9489-d9849a5e0a23 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs If n is odd, we can take φ n = 0 and similar argument holds
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f18c35c9-61d7-4e3d-be48-0c3248b617e2 · outbound
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs By Lemma 6, we have for t≥ 4, Qt u(s, a)≤ r(s, a) + γ[P V t u+1](s, a) + 2β‖ϕ (s, a)‖Λ −1 t + 2(mt−3− mt)
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.