Pith. sign in

Paper Citation Record · LEDGER

Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2212.00603.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.00603 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:48:32.824182Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T18:55:59.140544Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32ab1a26-20b5-443f-9542-807e6181f9e3 · inbound

Near-Optimal Sample Complexity for MDPs via Anchoring cites this paper.

Near-Optimal Sample Complexity for MDPs via Anchoring Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T22:48:32.824182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:48:32.824182Z digest=sha256:256114674d587d1d442da3ef8b273811b7d9f2212ba29cc1fd980eab99dd25a8

Observation a5de33fd-8f62-47de-b381-294e48cdd827 · inbound

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning cites this paper.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:11.972128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:11.972128Z digest=sha256:78f297d1d5a785533da657e57d9e7959572959ebe05539b975a3be9c6caaef4b

Observation f54689f6-5d42-400b-a9bf-f0e1dfc99477 · inbound

A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model cites this paper.

A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:57.865812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:57.865812Z digest=sha256:8411fdcc0ab8f0a01055662e58b361abdd527436c55a13f30915661e5c114bac

Observation 287162d5-da44-4a40-a323-5ea354c2632d · inbound

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies cites this paper.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:50:57.135525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:22fc67cb667f14b06941fe60dd12c9f4ec3256d907054043dab2e7f8498cd864

Observation 74e3f11d-0540-4342-bec9-7f8b67a8098c · inbound

Learning in Markovian bandits with non-observable states and constrained decision epochs cites this paper.

Learning in Markovian bandits with non-observable states and constrained decision epochs Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:59.142261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:ff85c43b4538795bfba36ffe263f6d1ecd7a47185ab5064860bc325a658cad44