Pith. sign in

Paper Citation Record · LEDGER

Asymptotically optimal regret in communicating Markov decision processes

As of 15 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2505.18064.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18064 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:53.055380Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b97480b8-281b-492d-a57e-531ff7779aa3 · outbound

This paper cites The regret lower bound for communicating Markov Decision Processes.

Asymptotically optimal regret in communicating Markov decision processes The regret lower bound for communicating Markov Decision Processes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.317140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.317140Z digest=sha256:1984d1fe263508b84a6585f57f0a5e953b8a3c7be147ef5103eca3ca91242fa8

Observation c1177cdb-2798-4a64-9447-616a5d833687 · outbound

This paper cites Thompson Sampling: An Asymptotically Optimal Finite Time Analysis.

Asymptotically optimal regret in communicating Markov decision processes Thompson Sampling: An Asymptotically Optimal Finite Time Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.699755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.699755Z digest=sha256:697bc0119c99fe19749e7b4edd17168e284dcdfbc14bab3ee0f89e455d0c1837

Observation 59780be3-1494-4337-9bc4-08bd81a43143 · outbound

This paper cites Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities.

Asymptotically optimal regret in communicating Markov decision processes Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.867553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.867553Z digest=sha256:4073c0b3f09fe3fd862c17d403af8995953072d1b595c33ae8735f9ad77d4509

Observation c8f64276-0b8f-4855-882f-859cd0eab264 · outbound

This paper cites 2 2.1.1 Randomized policies, their gain, bias & gap functions.

Asymptotically optimal regret in communicating Markov decision processes 2 2.1.1 Randomized policies, their gain, bias & gap functions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:54.230041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:41:53.055380Z digest=sha256:d85e122c690f1ad0937b3cc9ba424f36d4703ba5156b2e7f2b8314d12462a518

Observation 4f545917-448a-4995-90a0-eae6717909f2 · outbound

This paper cites Shipra Agrawal and Navin Goyal.

Asymptotically optimal regret in communicating Markov decision processes Shipra Agrawal and Navin Goyal

Reference 1988

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:41:54.076804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:41:52.015182Z digest=sha256:a51cc01b449562e95ab1a3d7215c5bece977a5709afa4e9e82684dfbb982c9b1

Observation e787acf3-7951-4e31-b64d-2be6d3e1ae14 · outbound

This paper cites OptimisminReinforcementLearningand Kullback-Leibler Divergence.2010 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 115–122, September.

Asymptotically optimal regret in communicating Markov decision processes OptimisminReinforcementLearningand Kullback-Leibler Divergence.2010 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 115–122, September

Reference 2006

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:54.432856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:41:52.388109Z digest=sha256:79f242fc38fcd966d4f8ffc5acad24034e33f4b1918f3e8666bbb8b937b27ca2

Observation aa5e2e6b-1d62-480e-ada3-dfe14b0cbdb4 · outbound

This paper cites Optimism in Reinforcement Learning and Kullback-Leibler Divergence.

Asymptotically optimal regret in communicating Markov decision processes Optimism in Reinforcement Learning and Kullback-Leibler Divergence

Reference 2010

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:41:53.465716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:41:52.449262Z digest=sha256:37f40164523601c77a5bd7e23d184609e76a72117946102ca4b7eba00d0123aa

Observation 4dcb854f-c078-4950-baf7-133786970846 · outbound

This paper cites Analysis of Thompson Sampling for the multi-armed bandit problem.

Asymptotically optimal regret in communicating Markov decision processes Analysis of Thompson Sampling for the multi-armed bandit problem

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.078339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.078339Z digest=sha256:2953618164ea08d9bb101f7ef32c2e7dbebce27b1de40e6bc6b6a54030678abf

Observation f77ad02a-93d7-4232-9cd3-b69d95b5a342 · outbound

This paper cites Improved Analysis of UCRL2 with Empirical Bernstein Inequality.

Asymptotically optimal regret in communicating Markov decision processes Improved Analysis of UCRL2 with Empirical Bernstein Inequality

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.529881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.529881Z digest=sha256:0bd42ae9462cf775369da01860f990477d8d558f0f13d4a0982f20130fb620ac

Observation 25ad5aad-bcea-4ec1-a0a9-b4747667d607 · outbound

This paper cites Regret Analysis in Deterministic Reinforcement Learning.

Asymptotically optimal regret in communicating Markov decision processes Regret Analysis in Deterministic Reinforcement Learning

Reference 2021

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:41:53.279499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:41:52.936313Z digest=sha256:ff319ef01106de868a107b9b55a0191ad4a660d0ccf071499a7cbadb59303728

Observation 7f096932-4b03-4d01-afc6-f1398f33e713 · outbound

This paper cites _eprint: 2502.06480.

Asymptotically optimal regret in communicating Markov decision processes _eprint: 2502.06480

Reference 2025

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:41:53.838157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:41:52.215124Z digest=sha256:c37007b948386f34b821f9d75b93da5e3f02bddd7ef0f61c8468d366a3ee7d64

Pith citing papers

No inbound Pith citation observations are available.