Pith. sign in

Paper Citation Record · LEDGER

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs

As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2504.11997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.11997 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:42:09.234480Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf20c210-1e8f-4150-a5b6-21453d6537d1 · outbound

This paper cites Improved algorithms for linear stochastic bandits.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Improved algorithms for linear stochastic bandits

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.620630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.136777Z digest=sha256:26c24dffad31ba1eb662681bad477fd634b1d3a540e9dec2edb690a58730f3b1

Observation 5f0c9a4d-d416-4a2a-9d54-0666df91980e · outbound

This paper cites Near-opt imal regret bounds for reinforcement learn- ing.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Near-opt imal regret bounds for reinforcement learn- ing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.607023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.141280Z digest=sha256:a236cb4e186d53e47315064ec059124b6b539d62922b0e9b074a4c34bd88f5b1

Observation 1ea6fd89-cdc6-48f4-8e41-2abc1102bf37 · outbound

This paper cites Model-based reinforcement learning with value-targeted regression.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Model-based reinforcement learning with value-targeted regression

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.593152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.145088Z digest=sha256:797154a733c420a696c4b6910951acdc95c419b52b78d1183f0ef1e24d0eb4e4

Observation 002223ee-3be3-4914-972c-8bb3ae082ce9 · outbound

This paper cites REGAL: a regularizati on based algorithm for reinforcement learn- ing in weakly communicating MDPs.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs REGAL: a regularizati on based algorithm for reinforcement learn- ing in weakly communicating MDPs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.579572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.148984Z digest=sha256:c70d099b67831e9313f5f3d88f1620e34191f24db91525a8027ae35a6108fa29

Observation 31c5fe95-e827-45d8-9aee-9d599ebdc74a · outbound

This paper cites Learning Infinite- Horizon Average-Reward Linear Mixture MDPs of Bounded Span.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Learning Infinite- Horizon Average-Reward Linear Mixture MDPs of Bounded Span

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.567360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.153118Z digest=sha256:5272f56d259cac69cbe95bbc64a8b1ad24ac72a631da0fb3825abc3ee7ffa0e2

Observation 3c9707ba-a4c4-41bd-9b5b-39656b819037 · outbound

This paper cites Efficient bias-span- constrained exploration-exploitation in reinforcement l earning.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Efficient bias-span- constrained exploration-exploitation in reinforcement l earning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.554902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.157133Z digest=sha256:2ed41a97ed7a74283080bc68873a109f830d34cde476586ea6866439add3cfd3

Observation d93434ac-c213-49e1-85fb-85fb567db066 · outbound

This paper cites Inven tory management in supply chains: a re- inforcement learning approach.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Inven tory management in supply chains: a re- inforcement learning approach

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.542022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.161662Z digest=sha256:781e5b93091cab1f3f0382a48d0abb1bf8dba7595ddb0a3a03c07ae9b736c767

Observation 4fda9f0f-7d61-40fc-a096-0838a5d95c09 · outbound

This paper cites Can deep reinforce- ment learning improve inventory management? performance o n lost sales, dual-sourcing, and multi- echelon problems.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Can deep reinforce- ment learning improve inventory management? performance o n lost sales, dual-sourcing, and multi- echelon problems

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.529036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.165411Z digest=sha256:1f71c01018a897d4f58052a85a3c0362d519a9bac387c645be96173b619707f3

Observation 98eaf7ce-6c45-4b91-965b-db82a068ab8d · outbound

This paper cites Reinforcement learning for long-run a verage cost.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Reinforcement learning for long-run a verage cost

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.515365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.168956Z digest=sha256:174bf81cdcfa4865d0b1002313b8fb1e47077b23a9c86ee60dacf3799b2d8af9

Observation a373189d-04ed-4aea-bb3f-17d2dc51226c · outbound

This paper cites Sample-effi cient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Sample-effi cient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.502455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.172453Z digest=sha256:ef217d2d05310094c74815b6512190c406dc0e113162ef90f4ea1a83f58d2974

Observation 24fdec2c-618e-4828-b6ce-4d8d17d54ab2 · outbound

This paper cites Reinforcement Learn- ing for Infinite-Horizon Average-Reward Linear MDPs via App roximation by Discounted-Reward MDPs.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Reinforcement Learn- ing for Infinite-Horizon Average-Reward Linear MDPs via App roximation by Discounted-Reward MDPs

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.489896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.175983Z digest=sha256:d50bfb779b279a5ed4d034a07d592e8d2e9874b1068278e82d2388cd49211c11

Observation e391adbc-7cf0-4878-8159-94883c14b200 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Provably efficient reinforcement learning with linear function approximation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.475838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.179550Z digest=sha256:9b09ed611ba125c455288a12f775a8d3bbdc6cc4b0fafe9fee14127337cb78c7

Observation 657a964b-9e31-40c9-b403-68f9bd3a29ce · outbound

This paper cites Towards tight bounds on th e sample complexity of average-reward MDPs.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Towards tight bounds on th e sample complexity of average-reward MDPs

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.461767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.183813Z digest=sha256:c1ce3767b018bcc107b07c22643e23303850b84e60be9ab197af8b784fe08059

Observation cedffd33-f0b3-448f-8bf8-2e0fd7300a92 · outbound

This paper cites Reinforcement learning based routin g in networks: Review and classification of approaches.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Reinforcement learning based routin g in networks: Review and classification of approaches

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.446829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.187438Z digest=sha256:fb77c5c8e66e5e3806724d3c8bfc55c36478ed5b850ace2fe90818deddbdd804

Observation 0c11398c-0252-4f53-9d54-307eef909972 · outbound

This paper cites Sample complexity of reinforcement learning using linearly combined model ensembles.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Sample complexity of reinforcement learning using linearly combined model ensembles

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.431697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.191051Z digest=sha256:657282d859fa11af983024861064ed588011e8c003154044a9a20afd14aa75a1

Observation 1dbf06e3-abd3-41d1-b432-3413b6dc6477 · outbound

This paper cites Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:42:09.194938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:42:09.194938Z digest=sha256:2f27ed47c349bf59165073e9cec49d3a80e37d51a5b93bf42ca7ac13b000fb6b

Observation 217f2d16-2c1e-46c1-8c2e-141cdeb8a9cc · outbound

This paper cites Optimal Sample Complexity for Average Reward Markov Decision Processes.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Optimal Sample Complexity for Average Reward Markov Decision Processes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:42:09.198958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:42:09.198958Z digest=sha256:c33652bdd5b42086ce933c7bb9d63b573f327bb1a7c639bf037bfeebc8020937

Observation 67fd2d37-9401-4a37-83f0-86ede1737eaa · outbound

This paper cites Learning infinite-horizon average-reward mdps with linear function approximation.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Learning infinite-horizon average-reward mdps with linear function approximation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.417503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.203038Z digest=sha256:26c6fb7cb27ff6c920fdb2a901189dce55c2b0867bbbb60479bc070e08bd3e60

Observation 53a1d58c-2ca8-4d3b-b05c-fcaf71db6ffc · outbound

This paper cites Model-free reinforcement learning in infinite-horizon average-rewar d markov decision processes.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Model-free reinforcement learning in infinite-horizon average-rewar d markov decision processes

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.402404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.206780Z digest=sha256:651608a148f980ce9fcd7700410cd44c9713a02b17e6212bd2ebf20fc18eccbd

Observation d2731aac-4d0a-458c-b9a5-3e15acf4d278 · outbound

This paper cites Nearly minimax o ptimal regret for learning infinite- horizon average-reward mdps with linear function approxim ation.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Nearly minimax o ptimal regret for learning infinite- horizon average-reward mdps with linear function approxim ation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.388691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.210522Z digest=sha256:98aaa03d1105537f58c8062ed9fb340dddd1e83936c387dd88fbe2c680517650

Observation 0345c7d6-3486-466a-958a-fabedd69edf8 · outbound

This paper cites Joint optimiz ation of preventive maintenance and production scheduling for multi-state production systems based on reinforcement learning.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Joint optimiz ation of preventive maintenance and production scheduling for multi-state production systems based on reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.373977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.214369Z digest=sha256:eaa9db812fc0979e9eefb73702c0a73b2488374eee5d4eab328c435317afbd32

Observation 990b21e2-c469-466b-a0bc-df1453fd3a51 · outbound

This paper cites Regret minimization for reinforcement learning by evaluating the optimal bias function.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Regret minimization for reinforcement learning by evaluating the optimal bias function

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.360091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.217863Z digest=sha256:07505d5f32cbf8146cb820292e23fd658c3c37e0293c03db7c538018b60ff47c

Observation d57704ee-93b7-40f9-8cb2-ac9df73c03b4 · outbound

This paper cites Sharper Model-free Reinf orcement Learning for Average-reward Markov Decision Processes.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Sharper Model-free Reinf orcement Learning for Average-reward Markov Decision Processes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.346138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.221490Z digest=sha256:a70b69f360ec86d843750ed444bbc96542f516913f96671a3cc08c84ab6e9ec7

Observation 486aa4c3-e495-49a0-8ec6-1b840927354f · outbound

This paper cites Span-Based Optimal Sample Complexity for Average Reward MDPs.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Span-Based Optimal Sample Complexity for Average Reward MDPs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:42:09.225245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:42:09.225245Z digest=sha256:09180bd82b21f0d875b18c0c711613e613875be187b4fd46adb69371360e9134

Observation 5c22267d-d48d-4aa0-9489-d9849a5e0a23 · outbound

This paper cites If n is odd, we can take φ n = 0 and similar argument holds.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs If n is odd, we can take φ n = 0 and similar argument holds

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.331024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.229998Z digest=sha256:3d86483297b03d483e8a159d29025515618f819e04dc70a60ddaa2a27ec48b47

Observation f18c35c9-61d7-4e3d-be48-0c3248b617e2 · outbound

This paper cites By Lemma 6, we have for t≥ 4, Qt u(s, a)≤ r(s, a) + γ[P V t u+1](s, a) + 2β‖ϕ (s, a)‖Λ −1 t + 2(mt−3− mt).

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs By Lemma 6, we have for t≥ 4, Qt u(s, a)≤ r(s, a) + γ[P V t u+1](s, a) + 2β‖ϕ (s, a)‖Λ −1 t + 2(mt−3− mt)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.314883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.234480Z digest=sha256:d68b0050a82814b09d1050a02496599154fe54c1c0d266e5488f4aafe383b4b3

Pith citing papers

No inbound Pith citation observations are available.