Pith. sign in

Paper Citation Record · LEDGER

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2607.17201.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17201 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:54:01.147324Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7151ce36-21b7-4331-9ea6-7279ad572995 · outbound

This paper cites Al Marjani and A.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Al Marjani and A

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.493874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.493874Z digest=sha256:dbc9b77fa1eb3eef29263a2a7455757a449664e4fcf9d12ba37b8d6288c4081f

Observation 06e02925-65d5-44b7-a941-788000b5515e · outbound

This paper cites Al Marjani, A.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Al Marjani, A

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.529153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.529153Z digest=sha256:cda0e721e0999c7538c8564b331ba03d10e0e4d4defa6d321dde6def3eec0e9a

Observation eb8f7d5f-4cd3-4159-bab6-483cb77255c7 · outbound

This paper cites Policy Testing in Markov Decision Processes.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Policy Testing in Markov Decision Processes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.581627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.581627Z digest=sha256:62a0a7c3b76b6eced8efb5a180c7167c3aedddaf52c4005e59a29431b94173e3

Observation 5db72b6e-b558-44fb-84bb-66656e05f635 · outbound

This paper cites Berge.Topological Spaces: Including a Treatment of Multi-Valued Functions, Vector Spaces, and Convexity.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Berge.Topological Spaces: Including a Treatment of Multi-Valued Functions, Vector Spaces, and Convexity

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.645968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.645968Z digest=sha256:53bdb8793050bf2d1162b0fe118aebc3bfd94c450576b2a884b14101e60b788d

Observation d909eea6-b31c-4031-a7f4-2b717132859b · outbound

This paper cites The regret lower bound for communicating Markov Decision Processes.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning The regret lower bound for communicating Markov Decision Processes

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.706383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.706383Z digest=sha256:a38e98e442aaf8a583b366920dfb64459a5653fce9ca5f9d4f5a2edaab6873ca

Observation 57780a37-c1e1-4079-bbe6-c2a72904419b · outbound

This paper cites Boucheron, G.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Boucheron, G

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.770229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.770229Z digest=sha256:850325c4af3dfd3348c999e1f3751a1daeb60753201df9cd7e3d27b0a1cd4b20

Observation 35b958e0-4312-4ee4-8b38-d8bd4a94e922 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.826541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.826541Z digest=sha256:7e609f86a39e3e6846076141a208d826b69b0dfa6ad8f3ce3a20fc1ebedabf1a

Observation cc0fdbbe-2b6f-468b-9fab-af3a2ff57e66 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 8

Resolution
verified exact
doi, observed 2026-08-01T18:59:12.738112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-01T18:53:58.871081Z digest=sha256:3d92b19f6279f91695e4cb170e13bb5a164df7eb218d72414fecdd15b4ba5d7a

Observation c7876929-fb53-45a4-9151-039a347cdc81 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.945774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.945774Z digest=sha256:a64cdd93179ed85505f48154a9f9541ec69314ab74f05bc31f2bde24c5f35dc5

Observation 89863d45-3678-4211-85ce-8f2e6f2a8cd6 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.009892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.009892Z digest=sha256:15e6b7c0303f4effcfa748d2520a8ff0f6e57744e88c6d3de53200ee6ee4c368

Observation 8f0bcc96-f663-41c7-b677-5d9384688533 · outbound

This paper cites Degenne and W.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Degenne and W

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.116553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.116553Z digest=sha256:5d5937e85f65231ae8f5b086424b8e8f79efcdb8f845342feeef2ccbbd14f2cd

Observation b9d2103a-99e4-478c-a38f-eebcc5d336bb · outbound

This paper cites Degenne, W.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Degenne, W

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.167718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.167718Z digest=sha256:374ab26304810a7f2509595d8ccf2f768e072e1416678264e5988ebc062ca5e6

Observation 3d740a2c-859d-47ca-859f-242de7a9d63c · outbound

This paper cites Garivier and E.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Garivier and E

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.262719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.262719Z digest=sha256:0ef65ff318986602bf55c66490badf0b0f4ad9d095842d5719c1ee422436c863

Observation 03a5005e-b5c8-4b51-affd-a8ca7f56596f · outbound

This paper cites Thresholding Bandit for Dose-ranging: The Impact of Monotonicity.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Thresholding Bandit for Dose-ranging: The Impact of Monotonicity

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.352381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.352381Z digest=sha256:73d556407230019706fc4a87fdf8f53fe3ed5b1261d63795cf8055f60dc7c028

Observation ea5a2da7-1be2-42c6-90c7-e54f0ed970c6 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.442412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.442412Z digest=sha256:78edafb380ba2a8dcb07ab1bda07cc9e90aba02dafd85a8141d6d4dcadd960cb

Observation 0932b424-0528-462e-8a90-dcd389e59d9f · outbound

This paper cites Jonsson, E.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Jonsson, E

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.452766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.452766Z digest=sha256:42e8da465e5eef4f4fd5b59c859df06eb287be66419804dfde5b03874ce08c93

Observation 1d2e93fe-fe6a-480f-ab5a-d8737b0ea7cb · outbound

This paper cites Jourdan and A.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Jourdan and A

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.511504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.511504Z digest=sha256:6741401ebcdf8d423abd381f6b3a02dc3ac20e353166586e77e5f90e29c8430a

Observation 1f0be86c-19be-4cb5-8176-6cacd2e5c25f · outbound

This paper cites Jourdan, R.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Jourdan, R

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.578201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.578201Z digest=sha256:07a0bdfe6f4a55403884a7ebc126b705578f0a2af6610cadeedcf66a74dcc9b0

Observation 57eccf5b-b2b7-4de4-970c-f440032526d9 · outbound

This paper cites Kaufmann, P.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Kaufmann, P

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.692268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.692268Z digest=sha256:81e310fac6c6e781e5651c292541f195a606475807f076c065ea632fee4aa4eb

Observation d906c45c-3b27-448f-be8d-5343fa621c18 · outbound

This paper cites Lazzaro and C.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Lazzaro and C

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.790934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.790934Z digest=sha256:d18be40cb43e11d13170639e3f0f24efd59911b82d2259893359b11b88d382d3

Observation 842bcfb4-453a-442e-882f-d3e4cf19f9a7 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.854229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.854229Z digest=sha256:e8d5918041fc53598a22eca7c96434460d64ffe0d89cffde6796da070f509731

Observation dc50e33e-8360-45a6-9365-1d9cb6cbf796 · outbound

This paper cites Poiani, M.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Poiani, M

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.941681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.941681Z digest=sha256:c71db99962c46c38aabbc36fb2e300321f7a48319728da474b3e246011131804

Observation 10d17296-4c7f-4280-b064-9b74ca0f8251 · outbound

This paper cites Poiani, M.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Poiani, M

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.031409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.031409Z digest=sha256:d569d7b20db16140c20ea2b807e9e86b32638b820981d15448f91fe71213fa7f

Observation a3557330-acb1-4377-9e61-ef785a6d2b92 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.132912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.132912Z digest=sha256:c917ee1efc6e3ff966c0e073a85059b94f761f65576101cff62b22bd3da007fd

Observation 661319d6-b19e-4270-8ce8-866f141e0195 · outbound

This paper cites Adaptive Exploration for Multi-Reward Multi-Policy Evaluation.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Adaptive Exploration for Multi-Reward Multi-Policy Evaluation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.193317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.193317Z digest=sha256:9a119574768ac36a9c795991a468010606a64aa608b91809cad021c337b51dfa

Observation 1b4326d2-0013-4adf-a545-e70733671d2e · outbound

This paper cites Russo and A.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Russo and A

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.274126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.274126Z digest=sha256:10457409c4876ae7503d5d107e14187871c0803e16c84615a5160a20b7108c5f

Observation 224278e4-ab14-4781-9513-404ffd466581 · outbound

This paper cites Russo and F.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Russo and F

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.432674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.432674Z digest=sha256:fe4bb6543e4c1b2cffa8b92f2b02eb916370a036a6f716cbcdb2abb76913af01

Observation f0043074-c72f-4304-8351-3d4601eaf280 · outbound

This paper cites Pure Exploration with Feedback Graphs.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Pure Exploration with Feedback Graphs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.626153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.626153Z digest=sha256:c52eae2d6ccce20afa64493d1f281d76162246cb045c08ea55ae5816fcd954d5

Observation 2db139a3-37fb-4422-b44f-c4ca9c94165a · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.721970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.721970Z digest=sha256:891ebb2f5735eb8525d49af6b64e734df241623e055deb7122a81fbe0f00da6e

Observation 50d7c8a2-7619-4abe-9d3d-2db699ddff46 · outbound

This paper cites Asymptotically Optimal Sequential Testing with Markovian Data.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Asymptotically Optimal Sequential Testing with Markovian Data

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.796549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.796549Z digest=sha256:c5e03c4be6f2598b3026c2c5a3b9cf8c806a855fcb931d2845ec15a05e1e9e4e

Observation c8a9d37c-bec9-4521-9c69-4ffcc2cfc9e0 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.837113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.837113Z digest=sha256:8730071be6298a4d4ce9e507520869a94f768893dbb039641fcdbfd12697e843

Observation a29c040a-498a-4498-9d0d-dd22925ba511 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.883532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.883532Z digest=sha256:596fd95aa1d89c094d6c275d3f61da69c3612d40e7d2d147e071dde47a44773a

Observation 73eadf8b-3c6d-4b0a-bfa7-1eb5171673f6 · outbound

This paper cites Taupin, Y.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Taupin, Y

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.937068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.937068Z digest=sha256:6adc7ce914d6cb693ae3f3abdc89cdc64965116a8fa04bcb90eeeb66c38e8f75

Observation 78ae6ba3-d871-475b-9c6b-9e31467066b0 · outbound

This paper cites Tuynman and R.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Tuynman and R

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.975367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.975367Z digest=sha256:c41e65cefe94e7a65d5bd7c1f35c47d266fe534595af53d8597290cc70e2da06

Observation dc1d2a00-e8be-4281-88aa-fd7bc1184742 · outbound

This paper cites Zalinescu.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Zalinescu

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:01.031230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:01.031230Z digest=sha256:08a2093ee39dd649904e2d6192de51bc399b5e358334c6a7e5f2e0010ad11990

Observation e79e6aa5-05e1-4189-afa2-4290b4bd6dd3 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:01.101340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:01.101340Z digest=sha256:65c8a6072b7f511c7717fd49d7c3731967d653a58061501b66d91dc69ba0f2dd

Observation c115e5c2-e857-4ce1-92bb-b79578246156 · outbound

This paper cites Lemma 42.Let Ω⋆(M) := ( ω∈Ω(M) inf M′∈Alt(M) X s,a ω(s, a)KL(P(s, a), P′(s, a)) = (T⋆(M))−1 ).

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Lemma 42.Let Ω⋆(M) := ( ω∈Ω(M) inf M′∈Alt(M) X s,a ω(s, a)KL(P(s, a), P′(s, a)) = (T⋆(M))−1 )

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:01.147324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:01.147324Z digest=sha256:d92f7da9b852187227e1f87d6fe59c2d7995404395c7ea23f4a67da017b483b8

Pith citing papers

No inbound Pith citation observations are available.