Pith. sign in

Paper Citation Record · LEDGER

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.19854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19854 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:40:27.181643Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4801e300-7b59-41b0-8bfe-f7a0fcf0f73d · outbound

This paper cites Optimistic posterior sampling for reinforcement learning: worst-case regret bounds.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Optimistic posterior sampling for reinforcement learning: worst-case regret bounds

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:22.728512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:22.728512Z digest=sha256:ae67ba6afb3fe977e9eed1e59e93f6a99e5f7ce2463364da7529acb4af0d7cce

Observation 6c269259-52a8-433e-bf90-68ac84d15294 · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Minimax regret bounds for reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:22.814825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:22.814825Z digest=sha256:61aa888af8c4d10dc7f36535fbe023bdd13b0ef9dafbda9deed0733eb3dcf529

Observation fce1e969-1212-4c88-9cc1-30a58ceec6a6 · outbound

This paper cites Regal: a regularization based algorithm for reinforcement learning in weakly communicating mdps.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Regal: a regularization based algorithm for reinforcement learning in weakly communicating mdps

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:22.931295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:22.931295Z digest=sha256:91318193fb077e6d3bb5d6435234747d248567da86755d033e65d472fd5157a7

Observation 7ab63c10-7f08-477c-abc8-699b27b6aeab · outbound

This paper cites Brafman and Moshe Tennenholtz.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Brafman and Moshe Tennenholtz

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.010077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.010077Z digest=sha256:e627fd6ef40e4868df7f07e8b0ec7153fbd97c652a4c37ee6c40053586308501

Observation 9174a010-34e2-4312-9c41-61633472f8a2 · outbound

This paper cites Provably Efficient Exploration in Policy Optimization.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Provably Efficient Exploration in Policy Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.151097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.151097Z digest=sha256:6e277672d82a8966c7d02402c1760b212857ddaa334f5b438fc2fd6f4eeb30b1

Observation e614e620-efcd-49f5-8143-c4dec2944f3a · outbound

This paper cites Implicit finite-horizon approximation and efficient optimal algorithms for stochastic shortest path.Advances in Neural Information Processing Systems, 34, 2021.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Implicit finite-horizon approximation and efficient optimal algorithms for stochastic shortest path.Advances in Neural Information Processing Systems, 34, 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.271360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.271360Z digest=sha256:3f74a220440b0c563c0503c3dabcb75e639975f1614876a7c40894a485b58cc8

Observation 15128eb3-93ae-4a2b-b3f8-ae7da0dcd47e · outbound

This paper cites Sample complexity of episodic fixed-horizon reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Sample complexity of episodic fixed-horizon reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.400362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.400362Z digest=sha256:7bd116abd2907f82558c98e386a14a4f49da553452f16491d0851785a12a8fc4

Observation 00b6e90a-50bb-482a-ba23-3db55c9c7e5c · outbound

This paper cites Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.552816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.552816Z digest=sha256:3ec6aea95efc3ecec185791627a7f19fb7a04509ca444796d530f7a80ae2e1ad

Observation 31b74b9f-edf2-4c00-9f0f-fa38aafa09cd · outbound

This paper cites Policy certificates: Towards accountable reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Policy certificates: Towards accountable reinforcement learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.740350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.740350Z digest=sha256:4fb891da0dcf34f8922fdf2c5577532cc483960daaed1bbc209b99429ef259c0

Observation 4c7bf4ac-ee8b-40aa-8ac7-de0c812c5dc2 · outbound

This paper cites Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.836348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.836348Z digest=sha256:096c12e22e86a95a2847e7babced4a2fcc8e2dd2be35d87c5c6c0f1678eebc1e

Observation 5faac2ce-e8c0-48ab-846a-259c55210303 · outbound

This paper cites Near optimal exploration-exploitation in non- communicating markov decision processes.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near optimal exploration-exploitation in non- communicating markov decision processes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.893443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.893443Z digest=sha256:2f6d91a9fc1906bd43f907a3126dd27ea68080c629025283ad1473643ecf99b1

Observation e095135e-2139-42b5-83c1-2f0f2157a5a9 · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-optimal regret bounds for reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.944648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.944648Z digest=sha256:433d884a080372b3d7560ce16d846d9a0f4bbb448be4983ba35a99d840209195

Observation 7c6d83dd-0910-48af-a758-a4812b4c8f5b · outbound

This paper cites Open problem: The dependence of sample complexity lower bounds on planning horizon.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Open problem: The dependence of sample complexity lower bounds on planning horizon

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.994297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.994297Z digest=sha256:59280b7c230ffd96737847384ee3b58f90a6289a7fe6e128a3e78fbf1f4ece8c

Observation 7c7db996-bc4e-4ab8-8353-0bc018a3330a · outbound

This paper cites Is Q-learning provably efficient? InAdvances in Neural Information Processing Systems, pages 4863–4873, 2018.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Is Q-learning provably efficient? InAdvances in Neural Information Processing Systems, pages 4863–4873, 2018

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.079508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.079508Z digest=sha256:eac8f646b761f88b458c05603b08f6eac006985803bff69d33a61eb38ef2c076

Observation 98353e53-0d49-439a-8d18-a6ecee66b4d0 · outbound

This paper cites PhD thesis, University of London London, England, 2003.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence PhD thesis, University of London London, England, 2003

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.157975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.157975Z digest=sha256:38761e85b98627337e4f2d644742143d3eab682a31ca54e95d9e1bd28631fe52

Observation 489bb316-9a4c-4425-a6ed-eabe06adb878 · outbound

This paper cites Near-optimal reinforcement learning in polynominal time.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-optimal reinforcement learning in polynominal time

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.268544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.268544Z digest=sha256:5a9c5f96909ee515d74258256d274e019e9c24a9e2f28a158afcf5f1cf47decf

Observation fa7e42a6-9880-4103-857a-5b5834278656 · outbound

This paper cites Near-bayesian exploration in polynomial time.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-bayesian exploration in polynomial time

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.384506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.384506Z digest=sha256:c9649a9a5dbf0ecb1f8b2b811601b6742e96b0ccf21be4f6ab26d46350fefc71

Observation d2da9399-ed90-4a4c-970b-8993b1be5470 · outbound

This paper cites Pac bounds for discounted mdps.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Pac bounds for discounted mdps

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.436716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.436716Z digest=sha256:e8cbe070546fb4ae6c2ebd6e56e5ab4a10f4b56b2ff0a053949af1667cb763e3

Observation 59bf1383-2265-46ad-8a9c-9bb2823128b5 · outbound

This paper cites Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning.Advances in Neural Information Processing Systems, 34, 2021.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning.Advances in Neural Information Processing Systems, 34, 2021

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.528200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.528200Z digest=sha256:f2ab6f8fb51436646addf7c9a92e7df1c301517bd6c9a6d843854740d69727ad

Observation 28ca7571-374a-4447-bd9f-0c3aaceb0e04 · outbound

This paper cites Horizon-free learning for Markov decision processes and games: Stochasti- cally bounded rewards and improved bounds.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Horizon-free learning for Markov decision processes and games: Stochasti- cally bounded rewards and improved bounds

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.622410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.622410Z digest=sha256:74016f1d310af7363a1f53505dd1a7a7d1a60be808dd7b4ea9a1fd6d4ee13dbf

Observation ff56f70c-8e43-440a-ada9-b1c6b75b78d2 · outbound

This paper cites Settling the horizon-dependence of sample complexity in reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Settling the horizon-dependence of sample complexity in reinforcement learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.723668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.723668Z digest=sha256:eac4363cb4e8fe7f0917a460e6f510030e3df3407013e85295a502ebd2e43bdd

Observation bd057107-90c4-4b3a-a149-94e9ae10c6db · outbound

This paper cites Empirical Bernstein bounds and sample variance penalization.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Empirical Bernstein bounds and sample variance penalization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.839310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.839310Z digest=sha256:01ad6df16a4e05df4d0a1d68424089f933ab83d2fa4762ac0599f34bc45ff4c0

Observation 03d8f37b-4764-4631-8a16-006a14384333 · outbound

This paper cites Ucb momentum q-learning: Correcting the bias without forgetting.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Ucb momentum q-learning: Correcting the bias without forgetting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.923691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.923691Z digest=sha256:e7b64af227409b3edfe64ef25f441bbd4562215c82efbec930e0d48f7c6e2b84

Observation 6f199159-8fd5-4a68-8c4a-4f7a24c25b5c · outbound

This paper cites A Unifying View of Optimism in Episodic Reinforcement Learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence A Unifying View of Optimism in Episodic Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.011502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.011502Z digest=sha256:89718f12abc6737d41c1efc705535049a04c06a82598add98add3fc53e7ec6ea

Observation 25829e35-b5ec-41be-90bf-7e226d8fc955 · outbound

This paper cites (more) efficient reinforcement learning via posterior sampling.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence (more) efficient reinforcement learning via posterior sampling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.130168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.130168Z digest=sha256:32e81870ec9af891ab7dbfd2a1e1e9c733572a61b93c43946c65d539d1f7f21a

Observation 245af5ec-f3da-421a-ba75-7cf266ec160d · outbound

This paper cites Why is posterior sampling better than optimism for reinforcement learning? InProceedings of the 34th International Conference on Machine Learning-Volume 70.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Why is posterior sampling better than optimism for reinforcement learning? InProceedings of the 34th International Conference on Machine Learning-Volume 70

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.306243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.306243Z digest=sha256:1d7e7e48262b69a535471ef182f9770596e21b3b6a3fcbc40793dac53f348741

Observation b6a3e0ad-2587-48ca-ac34-ade72daa83ed · outbound

This paper cites Towards Tractable Optimism in Model-Based Reinforcement Learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Towards Tractable Optimism in Model-Based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.432600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.432600Z digest=sha256:2c968afa7457d1ad80eefcfdb56bc2d2031ccf022b4f4d6be7f74d40cd3525d0

Observation 245c824c-5e34-434d-898a-a73e470f401f · outbound

This paper cites Nearly horizon-free offline reinforcement learning.Advances in neural information processing systems, 34, 2021.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Nearly horizon-free offline reinforcement learning.Advances in neural information processing systems, 34, 2021

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.606430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.606430Z digest=sha256:24404158c5b955ab736977d0c6c35a533b681f2c2e22859eac7d26a1d9578621

Observation d2f5be01-be32-4e1e-b9a4-f79376a246d1 · outbound

This paper cites Worst-case regret bounds for exploration via randomized value functions.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Worst-case regret bounds for exploration via randomized value functions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.723620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.723620Z digest=sha256:6c8c17dc6b411e7818b4fb88ee60b381bec9df7cc33af28b5ef882063e91cba5

Observation 4e8d873f-f2c7-4b34-90fe-a89e81992529 · outbound

This paper cites Non-asymptotic gap-dependent regret bounds for tabular MDPs.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Non-asymptotic gap-dependent regret bounds for tabular MDPs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.800628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.800628Z digest=sha256:761812d31115499870c7deb8ba5e02e830ab3b7711d2cac54dec20ce1602c8a2

Observation 75b6acb7-2315-4367-8bdb-d6e9152453a7 · outbound

This paper cites PAC model-free reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence PAC model-free reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.889128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.889128Z digest=sha256:d3436f40068ea99127d8b072b952a573b4b2145b9a94797203ff58e30ebf7e57

Observation adc7cf24-c48e-4411-acde-1175eb5923a7 · outbound

This paper cites An analysis of model-based interval estimation for markov decision processes.Journal of Computer and System Sciences, 74(8):1309–1331, 2008.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence An analysis of model-based interval estimation for markov decision processes.Journal of Computer and System Sciences, 74(8):1309–1331, 2008

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.980493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.980493Z digest=sha256:9a7a4824472cb30053a903143fc4633a26b845eb14cb253dae7f9e3626bbc91f

Observation 59520a41-9a07-4041-af11-e2acddf13094 · outbound

This paper cites Model-based reinforcement learning with nearly tight exploration complexity bounds.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Model-based reinforcement learning with nearly tight exploration complexity bounds

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.075780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.075780Z digest=sha256:0b3caa978a456c5ae16918613ed3da64fff7caa096a7647788d36a166efecbd8

Observation d18ebd5e-b905-4ede-a069-f6526fb93e1a · outbound

This paper cites Variance-Aware Regret Bounds for Undiscounted Reinforcement Learning in MDPs.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Variance-Aware Regret Bounds for Undiscounted Reinforcement Learning in MDPs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.138252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.138252Z digest=sha256:a2a41a95c7d854c8666ef7d86ed73cdac8e3f60caaf87503f883fadaf8a5f76f

Observation 8f2af213-6957-4163-b8e0-5c1528a40e63 · outbound

This paper cites Stochastic shortest path: Minimax, parameter-free and towards horizon-free regret.Advances in Neural Information Processing Systems, 34, 2021.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Stochastic shortest path: Minimax, parameter-free and towards horizon-free regret.Advances in Neural Information Processing Systems, 34, 2021

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.236474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.236474Z digest=sha256:f08ce5f85e30d7e3a32a430acef58f2ec06546e73fea1deccb1b99f069bab1b4

Observation d572841b-2c0e-452e-bc99-aa6a402735a8 · outbound

This paper cites Is long horizon reinforcement learning more difficult than short horizon reinforcement learning? InAdvances in Neural Information Processing Systems, 2020.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Is long horizon reinforcement learning more difficult than short horizon reinforcement learning? InAdvances in Neural Information Processing Systems, 2020

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.311969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.311969Z digest=sha256:a76fe0731b53264e409b4700aeee377955fcc1e02cf0b7d25d45e7272e757c93

Observation 041767f7-ac81-47bd-a7a1-123a5d521c32 · outbound

This paper cites Near-Optimal Randomized Exploration for Tabular Markov Decision Processes.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-Optimal Randomized Exploration for Tabular Markov Decision Processes

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.383570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.383570Z digest=sha256:637171d0213505db22b64068b9e3af418dbcb5954fc896959a2f657edc6b3bed

Observation de55941f-e943-4f1b-875e-15a01520adff · outbound

This paper cites $Q$-learning with Logarithmic Regret.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence $Q$-learning with Logarithmic Regret

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.463961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.463961Z digest=sha256:89089adc510a5fe2afe8506fcf08cea353b2a98bdad1832b574693a0ed2918e6

Observation e0baaf33-cc64-4e9b-afe5-77e15b5ab775 · outbound

This paper cites Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.560944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.560944Z digest=sha256:89b9f1a3102243f35d6c05fd0fde4711db58dcbc6c4346870d92534a5c1ea341

Observation c354c13d-6c92-473b-8046-e54aba83750e · outbound

This paper cites Horizon-free reinforcement learning in polynomial time: the power of stationary policies.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Horizon-free reinforcement learning in polynomial time: the power of stationary policies

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.624890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.624890Z digest=sha256:2ca28bfcc0df30a3e00be13e8f357cc16e2148166998b7075f85ea4ed9705877

Observation 06a19116-357c-426b-8a81-1be9e439aa0f · outbound

This paper cites Variance-aware confidence set: Variance- dependent bound for linear bandits and horizon-free bound for linear mixture mdp.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Variance-aware confidence set: Variance- dependent bound for linear bandits and horizon-free bound for linear mixture mdp

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.701935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.701935Z digest=sha256:ed717c45d604e0bae08c92b2a3f726ac6b65f3b7f897d3afd2fc0f88e1891c52

Observation e022117d-be36-46b2-9812-c2b709538011 · outbound

This paper cites Almost optimal model-free reinforcement learning via reference-advantage decomposition.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Almost optimal model-free reinforcement learning via reference-advantage decomposition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.781590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.781590Z digest=sha256:a2d67bc4efa06eb3de2579d5f6eaa5fa8382d8444fd657d8bb9a952b658840d6

Observation 4bf24377-46c4-429f-a828-bbbcb4807ab2 · outbound

This paper cites The horizon-truncation lemma implies that the optimal H-step value is withinOpυqof the optimalH 1-step value.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence The horizon-truncation lemma implies that the optimal H-step value is withinOpυqof the optimalH 1-step value

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.849097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.849097Z digest=sha256:ec81f89f1a6ef8b271b998f9d94f4c8b97fc8b16b5ff9c75651e95b12c92102d

Observation 4f34ac99-42d8-4c7e-927e-d1aa34d12138 · outbound

This paper cites The proof is a backward induction on h.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence The proof is a backward induction on h

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.932895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.932895Z digest=sha256:f6b1c8e895fb339f4161d9eb9aa4f98fab5360a684c464f4e907796b8da32612

Observation 65ed91f3-fcdf-41d2-8e0d-e25b251464ab · outbound

This paper cites The process is stopped when the trajectory reaches the unlearned set Ok or when the frozen count of a learned pair doubles.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence The process is stopped when the trajectory reaches the unlearned set Ok or when the frozen count of a learned pair doubles

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.996404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.996404Z digest=sha256:d6be1cb4232d65a29afef185604c3737aa14309237725d318b1aa5d85c2e552d

Observation 7d4e9a66-26ff-4174-b235-29c98665c644 · outbound

This paper cites This requires a variance closure argument for the optimistic values and the optimistic gaps.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence This requires a variance closure argument for the optimistic values and the optimistic gaps

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.111719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.111719Z digest=sha256:2d4113faa855ac8e53f24b577881f76388d29ebe998730268747d0218f78d890

Observation 56d6f6be-6c0f-41bb-ac7e-04de21c57a2d · outbound

This paper cites V ˚ H1`1.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence V ˚ H1`1

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.181643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.181643Z digest=sha256:29e9246130b6d811dd9ee74103bef85e5e6fd4c90056f6774d6f8551d3c0077f

Pith citing papers

No inbound Pith citation observations are available.