Pith. sign in

Paper Citation Record · LEDGER

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs

As of 11 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2506.06521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06521 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:09:21.939768Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T18:10:52.644951Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-09T06:45:40.553981Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact3
  • verified fuzzy38
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 752a21c9-eded-434d-be7b-0f168f5d613f · outbound

This paper cites Navigating to the best policy in markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Navigating to the best policy in markov decision processes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:23.062037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.645313Z digest=sha256:f9b51e7e60b2fbed05e6e0fa875a37e43b33ff7c4ccbe21a6b557eab85a55f77

Observation 3e9af9e9-27e2-47d8-8e8b-e7c0ff6e7545 · outbound

This paper cites Logarithmic online regret bounds for undiscounted reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Logarithmic online regret bounds for undiscounted reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.650637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.650637Z digest=sha256:e44297b1b971f0ccd6270c78e8a08dcd2e0fcc0dea25914b9a39a528c9a3a2ad

Observation 2f7ec633-71fd-414e-bb52-31eb330fa6ce · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Finite-time analysis of the multiarmed bandit problem

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:23.022547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.655458Z digest=sha256:7a4d6e77d64cf759cf40d17f2b04e1c9786222f77730815b2356ac1ebd57d93e

Observation 1560998c-4618-4634-a281-39ba3a04691c · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near-optimal regret bounds for reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.660747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.660747Z digest=sha256:7a921be98e27b41a401646573d43b19080263a7d924583c9f2c936a1d0825bd3

Observation e5655449-346d-46ee-9f2c-2fe69a778fac · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Minimax regret bounds for reinforcement learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.665444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.665444Z digest=sha256:b52c6967aee9b7d11028faa0236437279dcfe7b7ed120e8d70dd0b7e4474677c

Observation e54fec24-9f68-4cf5-aa49-5f8fe58a873c · outbound

This paper cites REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.670031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.670031Z digest=sha256:e520838d2c084dc69516df96dc88a6d81eee0841cd7e06de9f9818651d216b5f

Observation c29630a4-0f78-40e9-887c-b334b8b794a4 · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.675666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.675666Z digest=sha256:8fe5122b97d9768752d47b47844dd95779220c988f00d6d69a8f21a937191bae

Observation 8f3663fe-4d78-4094-b5ae-d7cf90cf1a0d · outbound

This paper cites Top-k off-policy correction for a reinforce recommender system.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Top-k off-policy correction for a reinforce recommender system

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.961449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.680902Z digest=sha256:0fe99a7072b77dd84547a47d3ccf56cfe2147438a6e6af5e20f9ad1641a26580

Observation d1ac95b1-9f22-4f96-ae91-10a8d6604ec1 · outbound

This paper cites Variance-Aware Sparse Linear Bandits.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Variance-Aware Sparse Linear Bandits

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:22.085200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.685486Z digest=sha256:37435f654a92ab73bdd4e6b415285c8e8e7e50821b3fd3c56cd6301e1556dca0

Observation cec2c58e-e241-40a0-9c94-fdd2d08dfd2b · outbound

This paper cites Policy certificates: Towards accountable reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Policy certificates: Towards accountable reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.912679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.690575Z digest=sha256:a701322e66ff9d30c0bfc214a7331fdfbc46f147d1167f45d8da4471944c6122

Observation 6fa8dfe2-4779-4705-9abb-c97da5ade988 · outbound

This paper cites Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.695323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.695323Z digest=sha256:d4a725d319d33f068ebc7fa6b375f64a668da0626b90afdcb22f5c2922af9591

Observation 60b87a59-dc4d-49ed-a24e-ebb846250806 · outbound

This paper cites Gap-dependent bounds for two-player markov games.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Gap-dependent bounds for two-player markov games

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.865549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.700715Z digest=sha256:7284ce7e25358b3a3fe09d41433a8bd3915803a5dfdbc703e6acec081051f2e8

Observation a168d2be-fb72-4c2d-a587-6dbfc37b7b0a · outbound

This paper cites Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.842175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.705363Z digest=sha256:82fc9e3fe280ae8dafe56df2af34360b57c9b2af54b8976335eb2b667234bda3

Observation d22e8a7b-3d5e-480c-8a43-fe6b483cb5ad · outbound

This paper cites Efficient bias-span-constrained exploration-exploitation in reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Efficient bias-span-constrained exploration-exploitation in reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.820123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.709935Z digest=sha256:1bb6dc2621f881fbf538c1a68396f67df73cfc641c4f36428de6da1cdb1715cf

Observation ec4381ac-2eb4-48cf-9940-8968b0f5725d · outbound

This paper cites Logarithmic regret for reinforcement learning with linear function approximation.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Logarithmic regret for reinforcement learning with linear function approximation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.800170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.714557Z digest=sha256:54d25583eb0dfb5695597c3fad54000079999a8510b434e9c8ec97c8ea048d44

Observation 0076f097-b362-4663-b39d-45d547528fa1 · outbound

This paper cites Tackling heavy-tailed rewards in reinforcement learning with function approximation: Minimax optimal and instance-dependent regret bounds.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Tackling heavy-tailed rewards in reinforcement learning with function approximation: Minimax optimal and instance-dependent regret bounds

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.780075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.719607Z digest=sha256:d6cf1156cf168e91ddce6c53a58e96d969fa05da71db9efd75624fc1d7b1d7bd

Observation 4c146e2b-e269-42fb-9c88-5f62ed713de8 · outbound

This paper cites Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.723948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.723948Z digest=sha256:739e5ca4653ec5ae6764a10d820ba091e9e715643035232ad87a2f3321bcb9eb

Observation aaf290a2-8bd0-4d6a-927c-49934c46ac58 · outbound

This paper cites Reward-free exploration for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reward-free exploration for reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.751095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.728273Z digest=sha256:93d24215484c74ae652a31f7cfc674899425ceba09f507b7f065cc7ebadba553

Observation 4001edbd-21f0-4c19-b55c-fc39d77c45e5 · outbound

This paper cites Planning in markov decision processes with gap-dependent sample complexity.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Planning in markov decision processes with gap-dependent sample complexity

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.729062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.732692Z digest=sha256:1143653bb945720383f3f50093e620ad0288a46c4fe7224afdbabf6ae68370bf

Observation e9893598-862c-4221-92c3-14e79de722bd · outbound

This paper cites Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.708304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.737357Z digest=sha256:2233d12300778dd462fa517ed632ecec1a7cd37c7c5317eb511f49011fe1b2c5

Observation e3b9b6b2-9766-4c14-9ef0-847e3dce72d3 · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Asymptotically efficient adaptive allocation rules

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.742030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.742030Z digest=sha256:b95bbd51da869b2bb362129020dae1fe7a57fd24a0de368bcc34dbee8a0ca716

Observation e475c0b6-34e8-4d90-b0ca-53a9292c24d2 · outbound

This paper cites Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.675579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.746775Z digest=sha256:25a41459d20ea94cf0270af130237813759f6e6830b219691944fdbba1523af9

Observation 0374572d-fe69-4408-9c01-0759e382bb4f · outbound

This paper cites Continuous control with deep reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Continuous control with deep reinforcement learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.751752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.751752Z digest=sha256:7e6bd1eb6a5bc3cd9fa18d75c0aea19222251704e755b508a76d1ed8bb20cdb1

Observation d0a8208e-7d51-4449-b65d-cd43cbc729b9 · outbound

This paper cites Deep reinforcement learning for dynamic treatment regimes on medical registry data.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Deep reinforcement learning for dynamic treatment regimes on medical registry data

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.655567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.756482Z digest=sha256:39970f9032c6ea3e00008b087d20cde1db54a74ddcf4b290c7426b29f053660e

Observation dfa9340b-fa18-408b-bcbb-1d0b634b2f37 · outbound

This paper cites Adaptive Sampling for Best Policy Identification in Markov Decision Processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Adaptive Sampling for Best Policy Identification in Markov Decision Processes

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:22.039979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.761176Z digest=sha256:4cf9d2d163569ab502823a6ab10b613eb795ef57763c663e77a754371a88160b

Observation 82315765-0338-43ad-bf3a-22ed3b6f355b · outbound

This paper cites Empirical Bernstein Bounds and Sample Variance Penalization.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Empirical Bernstein Bounds and Sample Variance Penalization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.766080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.766080Z digest=sha256:ff4e3c7448c45c358b571375527cd0ee76e0a0fc6fe5bf726e7eac2f6e8e22e8

Observation da871a9a-82bc-4c25-9641-0600ed1a9997 · outbound

This paper cites Ucb momentum q-learning: Correcting the bias without forgetting.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Ucb momentum q-learning: Correcting the bias without forgetting

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.634394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.771025Z digest=sha256:47323e42944d6a4bcf0bfbbad7038b97365f97808dae3910aacf1891bbbea81f

Observation 7e4070e2-adca-46e5-9b69-c550d4538112 · outbound

This paper cites Reinforcement learning for optimized trade execution.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning for optimized trade execution

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.612325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.775540Z digest=sha256:6cf88469602f02577e1bceea5f46d7b052fcc6fb2bad600efbe19d04262b69d3

Observation 840d6741-43ff-467c-9513-0fb223349c7e · outbound

This paper cites On instance-dependent bounds for offline reinforcement learning with linear function approximation.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs On instance-dependent bounds for offline reinforcement learning with linear function approximation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.592062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.779948Z digest=sha256:72e4d9dd85f379c9b9fa4a40ae24cad39c103c947b554095b355a10bdd356cc7

Observation 83035fc7-3068-41b5-9a09-6380cec19577 · outbound

This paper cites Exploration in structured reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Exploration in structured reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.572522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.784415Z digest=sha256:077a910ad0906efef090e056c31ee207afd642ceb4597bb25d28704571cc061e

Observation 143457cc-037a-4666-9f5a-426af5f8d651 · outbound

This paper cites Why is posterior sampling better than optimism for reinforcement learning? In International conference on machine learning, pages 2701--2710.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Why is posterior sampling better than optimism for reinforcement learning? In International conference on machine learning, pages 2701--2710

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.552330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.788946Z digest=sha256:c6066dbea9bdce6795ecd834d11e9132b9189dd9a8536d86cd1393579641b00a

Observation dc82e5fe-f4e1-49a8-9b08-93ef8f346b0c · outbound

This paper cites Reinforcement learning in linear mdps: Constant regret and representation selection.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning in linear mdps: Constant regret and representation selection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.530992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.793430Z digest=sha256:c36922073b2d8841436ffd198825dbb4f470741f3cb20a059e579438108d543a

Observation fdcbba15-75cf-4a2c-8f4e-d664249465b1 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Mastering the game of go with deep neural networks and tree search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.797839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.797839Z digest=sha256:8b9707a5a5ae85bd7e1836aad184f3929066e2a557511b1f2c2e1ec818bd52ec

Observation 1c0afb3a-986d-4ccd-8bb4-e9c44f172cfb · outbound

This paper cites Non-asymptotic gap-dependent regret bounds for tabular mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Non-asymptotic gap-dependent regret bounds for tabular mdps

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.802739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.802739Z digest=sha256:46d53cb0d3532067558cee4ec34660c5b5323c38695f143e725ddd6334b24377

Observation 53f03b79-cce4-4b1a-a750-b79a831414fd · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning: An introduction, volume 1

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.806964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.806964Z digest=sha256:7d8534f0ac516dcfc38b7021ea58579677dfdbe1962815a270aabd5d92a5115f

Observation ce9bbdc6-dc6b-45b5-b181-dd2192b8093f · outbound

This paper cites Variance-aware regret bounds for undiscounted reinforcement learning in mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Variance-aware regret bounds for undiscounted reinforcement learning in mdps

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.481576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.811736Z digest=sha256:47e0df25731685ce29ae29c3d03f4920072cf9ccf3cf234db830a0bc301a6aa7

Observation a7757793-65e0-4df2-a5bc-6aca707c4cd1 · outbound

This paper cites Optimistic linear programming gives logarithmic regret for irreducible mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Optimistic linear programming gives logarithmic regret for irreducible mdps

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.460484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.816074Z digest=sha256:e8ecb4e151cf2bfdd9d5262fd048554f6859ce6d1c9bcd7b264ea31861ea1842

Observation 87b74cce-14e2-47e2-a05b-8814cb258c4c · outbound

This paper cites Near instance-optimal pac reinforcement learning for deterministic mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near instance-optimal pac reinforcement learning for deterministic mdps

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.441791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.820781Z digest=sha256:704aab9e5dc569faf3be1db443729afdb03ac5c3af4028a60596be404010a1ef

Observation 646c21e6-60f6-4282-ba5a-dbe02e9614a8 · outbound

This paper cites Optimistic pac reinforcement learning: the instance-dependent view.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Optimistic pac reinforcement learning: the instance-dependent view

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.422810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.825041Z digest=sha256:1c7ef75672709982beec31b93d6ae23c30d4b4d74adbb4878f4caa78beb181e7

Observation 665fbd0f-d338-43cc-89fe-2f62a84c2986 · outbound

This paper cites Reinforcement learning with logarithmic regret and policy switches.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning with logarithmic regret and policy switches

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.403477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.830888Z digest=sha256:495d93e1dacbed41d859f2bba1ef4cddb73b4108133063b407c37517719ed035

Observation a4fba1ab-3745-4ef1-9a60-f3f745eabfc8 · outbound

This paper cites Instance-dependent near-optimal policy identification in linear mdps via online experiment design.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Instance-dependent near-optimal policy identification in linear mdps via online experiment design

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.387794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.835885Z digest=sha256:86cf9500b87c1cf9b3e2d39ca8405ba7152bf7fbcfd26ad017708f0f0af66ea4

Observation c96a5531-24de-4071-ac6b-bb1c7376ee50 · outbound

This paper cites First-order regret in reinforcement learning with linear function approximation: A robust estimation approach.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs First-order regret in reinforcement learning with linear function approximation: A robust estimation approach

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.371017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.841054Z digest=sha256:47f625dbe786abd82d16335a46ae4573fad81706b05e5b2566471c88c54d8dff

Observation 343a71d3-ab56-4d4c-a37e-13f8cab7b505 · outbound

This paper cites Beyond no regret: Instance-dependent pac reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Beyond no regret: Instance-dependent pac reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.352868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.846110Z digest=sha256:67f2950e6ad79490666d14b7c95970f65ee84185f9748fe54333c1fb1ce0dc8c

Observation 587e52c9-d96d-4887-80a9-2f0250ffdba0 · outbound

This paper cites On gap-dependent bounds for offline reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs On gap-dependent bounds for offline reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.335760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.851766Z digest=sha256:75aff0ee42185970796fc2fee485e602c41f2f92a89f6a78949506408ed54963

Observation 0c2295a7-13c1-4934-a9f2-7106bdeb9912 · outbound

This paper cites Near-optimal randomized exploration for tabular markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near-optimal randomized exploration for tabular markov decision processes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.318255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.857671Z digest=sha256:4d2d07a47b99837e1423d9f1ea3900a6bc3076d0616ae85fcf473adc1eddadf2

Observation ee3fccd1-e11e-4db4-9784-a506b0ca06b9 · outbound

This paper cites Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.300451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.863405Z digest=sha256:451b788d1722dc2f266e0f5a6a0c247a88a58012fa80fa8b7b382d32c632ada7

Observation c561d791-c61e-41ee-922e-0c6d2e04dd84 · outbound

This paper cites Q-learning with logarithmic regret.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Q-learning with logarithmic regret

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.279268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.868774Z digest=sha256:fb1687a91d57186f748e5f91eaec1b9dc028716275d8316fe36434b26219b258

Observation e09e3872-2afb-4001-84da-90d806ce96b6 · outbound

This paper cites Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.875678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.875678Z digest=sha256:02bdb0f20d045ab82c38c8c87034d028df4c11a4dbc5b2967bfb853e193235df

Observation 1fdf739b-4ba0-4d61-acb9-a387c1a54368 · outbound

This paper cites Regret minimization for reinforcement learning by evaluating the optimal bias function.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Regret minimization for reinforcement learning by evaluating the optimal bias function

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.246652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.882653Z digest=sha256:33fbbc443371dea69febff7e3a5991a0be6ec9eeb31fa4f22a869d327b3225dd

Observation 03367650-80df-4da6-97d3-4005db1289cb · outbound

This paper cites Almost optimal model-free reinforcement learningvia reference-advantage decomposition.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Almost optimal model-free reinforcement learningvia reference-advantage decomposition

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.224318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.889570Z digest=sha256:cd4b0962b90d707ff30b23bd7b3575aa694787a901688931df3dfe1ccbc2fdb1

Observation a62ee762-7be0-4634-a9f7-806d703c4041 · outbound

This paper cites Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.203614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.897012Z digest=sha256:0c7ab3dc1857a552bd4aa3620ccbdd66b0378fa6f8fee1f6b9478219333c1b98

Observation eb10c7a5-ddd7-49d9-9810-6017f98d517a · outbound

This paper cites Improved variance-aware confidence sets for linear bandits and linear mixture mdp.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Improved variance-aware confidence sets for linear bandits and linear mixture mdp

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.184382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.903919Z digest=sha256:1b8b6af8e59f6747ae3b0a28849718e024aa723695d4c978f7eac44b7f6f6d6e

Observation 437ff0ee-9e1e-4079-b4c7-1dd3ac6a4efa · outbound

This paper cites Horizon-free reinforcement learning in polynomial time: the power of stationary policies.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Horizon-free reinforcement learning in polynomial time: the power of stationary policies

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.165150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.911067Z digest=sha256:f51b5efc86552c401b666b3f870ec59ca077bb4853ced9a86d5ab2417306aca9

Observation 744f8b4a-6ecc-44b7-bd4a-b4cd2893f646 · outbound

This paper cites Settling the sample complexity of online reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Settling the sample complexity of online reinforcement learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.917081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.917081Z digest=sha256:38debc09aecda82a790def4ee1bc635872998361a3a2e22a7d36f6ea3af2abe8

Observation b1f3bd36-1d48-4a40-8458-ac15aa3701e9 · outbound

This paper cites Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:21.992642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.923994Z digest=sha256:0bd5bc6a54938e5b70c9885b562747fb3c3e7984c5eabe0c97a669d179e3b8e9

Observation a74689ed-21ec-46a7-a988-4e5560f21227 · outbound

This paper cites Nearly minimax optimal reinforcement learning for linear mixture markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Nearly minimax optimal reinforcement learning for linear mixture markov decision processes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.133022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.931913Z digest=sha256:49739489792a9bd0279cb831274454daead1861bdfdc13446d7a4a741e0041a7

Observation 46abb972-bf09-441b-a206-dd59d2d18436 · outbound

This paper cites Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.939768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.939768Z digest=sha256:15a2f743a1a85a45ce493fc65df5304d55bbaee90db9aef186ceb396e5b18574

Pith citing papers

Observation 229aedfe-0a81-41e4-9f90-efbe3fc4d041 · inbound

On-line Learning in Tree MDPs by Treating Policies as Bandit Arms cites this paper.

On-line Learning in Tree MDPs by Treating Policies as Bandit Arms Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:45:40.556056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T18:10:52.644951Z digest=sha256:46762b1b8b510e270347f3e2b0963ee21b2c79d5287f2da207a8762a09109f69