Pith. sign in

Paper Citation Record · LEDGER

Asymptotically optimal regret in communicating Markov decision processes

As of 10 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2505.18064.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18064 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:53.055380Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b97480b8-281b-492d-a57e-531ff7779aa3 · outbound

This paper cites The regret lower bound for communicating Markov Decision Processes.

Asymptotically optimal regret in communicating Markov decision processes The regret lower bound for communicating Markov Decision Processes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.317140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.317140Z digest=sha256:0961246376205965c3a8a5497f3933d5d5659a80b1eae72e1b81fea7f8bc6c2b

Observation c1177cdb-2798-4a64-9447-616a5d833687 · outbound

This paper cites Thompson Sampling: An Asymptotically Optimal Finite Time Analysis.

Asymptotically optimal regret in communicating Markov decision processes Thompson Sampling: An Asymptotically Optimal Finite Time Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.699755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.699755Z digest=sha256:25374a457024d0b2250ae81fcb14a73cdcff96fdfcc7a710abbbf4ae3f74725a

Observation 59780be3-1494-4337-9bc4-08bd81a43143 · outbound

This paper cites Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities.

Asymptotically optimal regret in communicating Markov decision processes Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.867553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.867553Z digest=sha256:d66ac587d58d4aafbb23b97cbb4da42770edbd62347d41c3474334a74d2fe978

Observation c8f64276-0b8f-4855-882f-859cd0eab264 · outbound

This paper cites 2 2.1.1 Randomized policies, their gain, bias & gap functions.

Asymptotically optimal regret in communicating Markov decision processes 2 2.1.1 Randomized policies, their gain, bias & gap functions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:54.230041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:41:53.055380Z digest=sha256:35541cb2a85f97803e6b8287520a9642920d48b7e1d2db47565c74c22a5dae1f

Observation 4f545917-448a-4995-90a0-eae6717909f2 · outbound

This paper cites Shipra Agrawal and Navin Goyal.

Asymptotically optimal regret in communicating Markov decision processes Shipra Agrawal and Navin Goyal

Reference 1988

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:41:54.076804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:41:52.015182Z digest=sha256:722d6c4b23bbb5af72d683db7fdf021d837f414e4b99ab7b341600785012f1d7

Observation e787acf3-7951-4e31-b64d-2be6d3e1ae14 · outbound

This paper cites OptimisminReinforcementLearningand Kullback-Leibler Divergence.2010 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 115–122, September.

Asymptotically optimal regret in communicating Markov decision processes OptimisminReinforcementLearningand Kullback-Leibler Divergence.2010 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 115–122, September

Reference 2006

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:54.432856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:41:52.388109Z digest=sha256:b4eefce8cc616f00e930b271cd6d2bc28bba04d26b03a886bd37d0a546a77bac

Observation aa5e2e6b-1d62-480e-ada3-dfe14b0cbdb4 · outbound

This paper cites Optimism in Reinforcement Learning and Kullback-Leibler Divergence.

Asymptotically optimal regret in communicating Markov decision processes Optimism in Reinforcement Learning and Kullback-Leibler Divergence

Reference 2010

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:41:53.465716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:41:52.449262Z digest=sha256:43361d447b268ffcc34b77df2e004e6da7258758c8c02a7503c81cf0450171cd

Observation 4dcb854f-c078-4950-baf7-133786970846 · outbound

This paper cites Analysis of Thompson Sampling for the multi-armed bandit problem.

Asymptotically optimal regret in communicating Markov decision processes Analysis of Thompson Sampling for the multi-armed bandit problem

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.078339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.078339Z digest=sha256:5262b53046052670b0bcf4f3114243c591e34efebe551a849af19110bb7f66bd

Observation f77ad02a-93d7-4232-9cd3-b69d95b5a342 · outbound

This paper cites Improved Analysis of UCRL2 with Empirical Bernstein Inequality.

Asymptotically optimal regret in communicating Markov decision processes Improved Analysis of UCRL2 with Empirical Bernstein Inequality

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.529881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.529881Z digest=sha256:fc2d6b1f329fc474f2194d8599b6914989941fd3ee0fa9e82b15694d520e0458

Observation 25ad5aad-bcea-4ec1-a0a9-b4747667d607 · outbound

This paper cites Regret Analysis in Deterministic Reinforcement Learning.

Asymptotically optimal regret in communicating Markov decision processes Regret Analysis in Deterministic Reinforcement Learning

Reference 2021

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:41:53.279499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:41:52.936313Z digest=sha256:1432a58807000d16117c030598d43b5be581e93fe6a33b70eb31ca90fcc3f89e

Observation 7f096932-4b03-4d01-afc6-f1398f33e713 · outbound

This paper cites _eprint: 2502.06480.

Asymptotically optimal regret in communicating Markov decision processes _eprint: 2502.06480

Reference 2025

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:41:53.838157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:41:52.215124Z digest=sha256:7fe1657de55253912f965271b7f4bc17dd7f890f46531e22e338d6363cb1b28b

Pith citing papers

No inbound Pith citation observations are available.