Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T22:47:51.676289Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 100 of 125 outbound references and 0 inbound Pith citation observations for arXiv:2607.06935.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T22:47:51.676289Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 125 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8c96fc50-d110-4b28-9264-fd9372aeb26c · outbound
Mathematical methods of reinforcement learning Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Pro- gramming
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3840bc46-ce83-427a-8711-a6671b20ec41 · outbound
Mathematical methods of reinforcement learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2bfeddfa-607d-448f-9da5-04af7a69f915 · outbound
Mathematical methods of reinforcement learning MIT press, 2022
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 607eb37a-f2f6-4e05-9e06-7f2f2c32d65d · outbound
Mathematical methods of reinforcement learning Deep belief markov models for pomdp inference.Neural networks, page 108386, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b7e8116b-faf7-47b4-86fc-94da8d8ad568 · outbound
Mathematical methods of reinforcement learning A markovian decision process.Journal of mathematics and me- chanics, 6(5):679–684, 1957
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dc7347e8-d4ba-4def-b8db-b5e8c793d64d · outbound
Mathematical methods of reinforcement learning An upper bound on the loss from approximate optimal-value functions.Machine Learning, 16(3):227–233, 1994
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8a84f11f-0456-4c3c-bdbf-91ffa2febd5a · outbound
Mathematical methods of reinforcement learning Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0a247d84-7460-4b2c-a4ed-fa142a938d6b · outbound
Mathematical methods of reinforcement learning Improved and generalized upper bounds on the complexity of policy iteration.Advances in Neural Information Processing Systems, 26, 2013
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 594cc6c6-d866-44d8-ad15-56587c71c4c0 · outbound
Mathematical methods of reinforcement learning Springer, 2018
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d187ea6-6994-44fc-a1b1-2d32ad828c69 · outbound
Mathematical methods of reinforcement learning From Convex Optimization to MDPs: A Review of First-Order, Second-Order and Quasi-Newton Methods for MDPs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4e9e4919-ca7a-480d-899a-9bf521a6209d · outbound
Mathematical methods of reinforcement learning A method for solving the convex programming problem with con- vergence rate o (1/k2)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 50718e16-f93f-4a33-a320-008081266679 · outbound
Mathematical methods of reinforcement learning Springer Science & Business Media, 2013
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f6c4c399-d58c-4373-a8cc-19127ff8c3eb · outbound
Mathematical methods of reinforcement learning Some methods of speeding up the convergence of iteration methods
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 46b874da-b796-45de-b692-1b04b4cb6653 · outbound
Mathematical methods of reinforcement learning A first-order approach to accelerated value iteration.Operations Research, 71(2):517–535, 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b407913f-067d-4bb8-8d91-1324e0ef9f71 · outbound
Mathematical methods of reinforcement learning Pid accelerated value iter- ation algorithm
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2c8bdf71-bd6f-4350-9a2a-e9b4219b939b · outbound
Mathematical methods of reinforcement learning A unified view of entropy-regularized Markov decision processes
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 09e0cabb-2143-49bd-98f3-a877dfec2952 · outbound
Mathematical methods of reinforcement learning Generative adversarial imitation learning.Advances in neural information processing systems, 29, 2016
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f7070b96-34e2-494c-b0b0-d1c1feeab671 · outbound
Mathematical methods of reinforcement learning Variational policy gradient method for reinforcement learning with general utilities
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5622de66-6cb8-415f-97d0-f0ce8f0e848a · outbound
Mathematical methods of reinforcement learning Stochastic Optimization under Hidden Convexity
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dbe4491c-7c82-4134-b8e0-b2170fed9987 · outbound
Mathematical methods of reinforcement learning Model-based reinforcement learning with a generative model is minimax optimal
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b3016a84-d9d8-4c85-a4cb-4843e363023c · outbound
Mathematical methods of reinforcement learning Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1a4aa522-d10d-4def-85f3-f438933b7681 · outbound
Mathematical methods of reinforcement learning Breaking the sample size barrier in model-based reinforcement learning with a generative model.Advances in neural information processing systems, 33:12861–12872, 2020
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 54d16de7-e98c-48af-a929-5d175ed29e32 · outbound
Mathematical methods of reinforcement learning Near-optimal time and sample complexities for solving markov decision processes with a generative model.Advances in Neural Information Processing Systems, 31, 2018
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6be2a1cc-dc5f-4921-b98d-dd3ec061397b · outbound
Mathematical methods of reinforcement learning Reinforcement learning: Theory and algorithms.CS Dept., UW Seattle, Seattle, WA, USA, Tech
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 96cf21f2-ea97-4172-a02d-64d3893fa985 · outbound
Mathematical methods of reinforcement learning Optimal sample complexity for average reward markov decision processes
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0487e0b8-1526-48f8-9c37-f4a48152808e · outbound
Mathematical methods of reinforcement learning Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction.Advances in neural information processing systems, 33:7031–7043, 2020
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4493884f-f773-4eba-832e-1c6fd0507e4b · outbound
Mathematical methods of reinforcement learning From dirichlet to rubin: Optimistic exploration in rl without bonuses
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d468340e-52af-41dd-b1dd-b2e2251eb0b1 · outbound
Mathematical methods of reinforcement learning Pac bounds for discounted mdps
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 03ef91a2-a2c4-4a05-be39-4e6be5e7af6d · outbound
Mathematical methods of reinforcement learning Assouad, fano, and le cam
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation df529439-d5cf-4154-b3bb-2420c9b74978 · outbound
Mathematical methods of reinforcement learning Towards tight bounds on the sample complexity of average-reward mdps
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 43a93845-8091-4c44-b98a-ccbbf5f8f77a · outbound
Mathematical methods of reinforcement learning pages 17, 18
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e3272491-4db6-4eea-8e57-369a0a88a7b1 · outbound
Mathematical methods of reinforcement learning Freedman’s inequality for matrix martingales
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 14c9d5ae-ad9d-41aa-a718-6799b0080122 · outbound
Mathematical methods of reinforcement learning Is q-learning minimax optimal? a tight sample complexity analysis.Operations Research, 72(1): 222–236, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ce3c1e5a-e422-4beb-b2a8-c467b44733cf · outbound
Mathematical methods of reinforcement learning Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4c92c3c3-4836-42e6-9c2c-7c45aecd9f16 · outbound
Mathematical methods of reinforcement learning Finite-sample convergence rates for q-learning and indirect algorithms.Advances in neural information processing systems, 11, 1998
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e0b3c3f8-23c4-41e9-8d6d-655bff861b99 · outbound
Mathematical methods of reinforcement learning A statistical analysis of polyak-ruppert averaged q-learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f6c1b55-1e44-4b51-b1d4-082933adde0a · outbound
Mathematical methods of reinforcement learning Variance-reduced $Q$-learning is minimax optimal
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation feb90f15-4dac-4b57-bfa7-c55160ea031b · outbound
Mathematical methods of reinforcement learning Randomized linear programming solves the markov decision problem in nearly linear (sometimes sublinear) time.Mathematics of Operations Research, 45 (2):517–546, 2020
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2e0b7dab-3863-4abd-9700-f2acaced553d · outbound
Mathematical methods of reinforcement learning Efficiently solving mdps with stochastic mirror descent
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bd3b8edd-bdc9-461a-9a3d-e6a4e7b5c7e3 · outbound
Mathematical methods of reinforcement learning Solving matrix games with near-optimal matvec complexity.arXiv e-prints, pages arXiv–2601, 2026
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fec0dc65-3bd8-4650-b840-7325450d65b3 · outbound
Mathematical methods of reinforcement learning Minimumcostflows, mdps, andℓ1-regressioninnearlylinear time for dense instances
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6ad6faa2-2f65-49c6-8c8c-e0f01027688e · outbound
Mathematical methods of reinforcement learning Efficient global planning in large mdps via stochas- tic primal-dual optimization
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d98744b4-1240-4078-a40c-33ffb1653502 · outbound
Mathematical methods of reinforcement learning Tight high probability bounds for linear stochastic approximation with fixed stepsize.Advances in Neural Information Processing Systems, 34:30063–30074,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9a65bc3-ddaf-4f36-9980-b3134b9e1a5e · outbound
Mathematical methods of reinforcement learning Statistical inferenceforlinearstochasticapproximationwithmarkoviannoise.Advances in Neural Information Processing Systems, 38:174565–174626, 2026
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 863847bc-a59a-4fd4-8e9a-07ad930a98f9 · outbound
Mathematical methods of reinforcement learning Finite-time high-probability bounds for polyak–ruppert averaged iterates of linear stochastic ap- proximation.Mathematics of Operations Research, 50(2):935–964, 2025
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 02e7238c-bee3-44b2-9004-eb6d39afb31a · outbound
Mathematical methods of reinforcement learning Learning to predict by the methods of temporal differences
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bfb97434-0e1c-4dd3-a150-90249a079648 · outbound
Mathematical methods of reinforcement learning Improved high-probability bounds for the temporal difference learning algorithm via exponential stability
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0a4bfb33-aabc-4e9c-9539-b06517c8d8f6 · outbound
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ef7e5892-8397-48f7-8764-622c5dcfcec8 · outbound
Mathematical methods of reinforcement learning Residual algorithms: Reinforcement learning with function approx- imation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a5451273-2d1f-47dd-9455-1dbcb9cc0c3a · outbound
Mathematical methods of reinforcement learning A convergento(n)temporal- difference algorithm for off-policy learning with linear function approximation.Ad- vances in neural information processing systems, 21, 2008
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 58568cbb-93ce-4db7-81e0-575fa443bc02 · outbound
Mathematical methods of reinforcement learning S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and E
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c87c284a-cceb-44e7-8b02-e9540456c859 · outbound
Mathematical methods of reinforcement learning Gaussian approximation for two-timescale linear stochastic approximation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e41f464a-319b-49fc-a105-45a192fac53e · outbound
Mathematical methods of reinforcement learning Some aspects of the sequential design of experiments.Bulletin of the American Mathematical Society, 58(5):527–535, 1952
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7b959147-979c-4d20-9f6a-c3687757c0bf · outbound
Mathematical methods of reinforcement learning Introduction to multi-armed bandits.Foundations and Trends in Machine Learning, 12(1-2):1–286, 2019
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 23df21be-6c13-4283-ac22-b146bde277c6 · outbound
Mathematical methods of reinforcement learning Regret analysis of stochastic and non- stochastic multi-armed bandit problems.Foundations and Trends in Machine Learn- ing, 5(1):1–122, 2012
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 47bc7bf8-28c9-4c32-9409-32c93948d2a8 · outbound
Mathematical methods of reinforcement learning Asymptoticallyefficientadaptiveallocationrules
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2a30c6b4-4d98-45f2-a1d0-21d84d106c8e · outbound
Mathematical methods of reinforcement learning Finite-time analysis of the multi- armed bandit problem.Machine Learning, 47(2-3):235–256, 2002
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e6d3077f-f3c9-49ff-a166-00df0f9538d1 · outbound
Mathematical methods of reinforcement learning Onthelikelihoodthatoneunknownprobabilityexceedsanother in view of the evidence of two samples.Biometrika, 25(3/4):285–294, 1933
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 29f09b7d-f6e8-4f16-b3d6-9a49501a13c5 · outbound
Mathematical methods of reinforcement learning A tutorial on thompson sampling.Foundations and Trends®in Machine Learning, 11(1):1–96,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e2f8921d-b4fd-4696-be06-b55b748311c8 · outbound
Mathematical methods of reinforcement learning Cambridge University Press,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation efca1d98-2947-4f77-9b0e-b9d756863152 · outbound
Mathematical methods of reinforcement learning An empirical evaluation of thompson sampling.Ad- vances in neural information processing systems, 24, 2011
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d423b61c-bdb3-4f7e-94ee-efa0dbf881b9 · outbound
Mathematical methods of reinforcement learning Further optimal regret bounds for thompson sam- pling
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3c3b4746-6292-457a-92ed-0c14e1feec65 · outbound
Mathematical methods of reinforcement learning Analysis of thompson sampling for the multi-armed bandit problem
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 38ea0d67-7489-4b00-a8eb-0af066c4cb07 · outbound
Mathematical methods of reinforcement learning Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a2a1c426-57ee-474a-b6ba-4300d2ed6a4e · outbound
Mathematical methods of reinforcement learning Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 979c3d86-060e-44e1-af69-b9d76768720a · outbound
Mathematical methods of reinforcement learning Near-optimal regret bounds for reinforcement learning.Journal of Machine Learning Research, 11:1563–1600, 2010
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 658b9dd7-2d3d-4471-b1b6-bfacffd5ecc0 · outbound
Mathematical methods of reinforcement learning A unifying view of optimism in episodic reinforce- ment learning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f807cfed-f852-4d1f-b30b-0ed459468af8 · outbound
Mathematical methods of reinforcement learning Minimax regret bounds for reinforcement learning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 52656bc2-427d-4740-8e89-23bada35e115 · outbound
Mathematical methods of reinforcement learning Rusu, Joel Veness, Marc G
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f5001757-cc5a-44e4-83b6-102cd1caed61 · outbound
Mathematical methods of reinforcement learning Asynchronous methods for deep reinforcement learning
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2bc3cfb2-7dcf-4522-92ec-8d4245b58e78 · outbound
Mathematical methods of reinforcement learning Trust region policy optimization
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fb2e76d8-792e-4984-bf33-350c04862815 · outbound
Mathematical methods of reinforcement learning Generalization and exploration via randomized value functions
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5f7bc0a9-979f-42ac-b971-81f27e2b1922 · outbound
Mathematical methods of reinforcement learning Is q-learning provably efficient?Advances in neural information processing systems, 31, 2018
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 62b6891e-6bcc-4c72-84e4-ea8c406cacd2 · outbound
Mathematical methods of reinforcement learning Almost optimal model-free reinforce- ment learningvia reference-advantage decomposition.Advances in Neural Information Processing Systems, 33:15198–15207, 2020
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4c2dc31f-1b9f-47a1-9fd4-518414f1f8ca · outbound
Mathematical methods of reinforcement learning (more) efficient reinforcement learning via posterior sampling.Advances in Neural Information Processing Systems, 26, 2013
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 43600ac8-cd84-486c-b1a9-c8c9a9d5e4ae · outbound
Mathematical methods of reinforcement learning Optimistic posterior sampling for reinforcement learn- ing: Worst-caseregretbounds
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b1a56240-a299-4146-b4a0-0f903b189cb8 · outbound
Mathematical methods of reinforcement learning Optimistic pos- terior sampling for reinforcement learning with few samples and tight guarantees
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a52d8616-c9df-447f-85d0-ccdeea1502c1 · outbound
Mathematical methods of reinforcement learning Deep exploration via randomized value functions.Journal of machine learning research, 20(124):1–62,
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c5f613ff-65f0-4b19-abbf-5344212064b4 · outbound
Mathematical methods of reinforcement learning Worst-case regret bounds for exploration via randomized value func- tions.Advances in neural information processing systems, 32, 2019
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee221d4e-9358-46e7-8e9b-ff9a6167f679 · outbound
Mathematical methods of reinforcement learning Improved worst-case regret bounds for randomized least-squares value iteration
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8b7ce566-a3cb-43f5-b7ab-a007538120c3 · outbound
Mathematical methods of reinforcement learning Near-optimal randomized exploration for tabular markov decision processes.Advances in neural information processing systems, 35:6358–6371, 2022
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd5a70ac-0eba-40e5-a724-ef79cb96a95a · outbound
Mathematical methods of reinforcement learning Finite-time bounds for fitted value iteration
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4d41807c-a370-4f7a-8279-977d1e35412e · outbound
Mathematical methods of reinforcement learning Error bounds for approximate policy iteration
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cecc9e32-113b-4da4-9e6b-c1bc17acc3e6 · outbound
Mathematical methods of reinforcement learning Kernel-based reinforcement learning.Machine learn- ing, 49(2):161–178, 2002
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8cda6af4-4335-461f-a18d-36c3d25cfcdd · outbound
Mathematical methods of reinforcement learning Kernel-based reinforcement learning: A finite-time analysis
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 99c9189c-d339-4b0c-8a72-91fb4b45f762 · outbound
Mathematical methods of reinforcement learning Adaptive discretization for episodic reinforcement learning in metric spaces.Proceedings of the ACM on Measurement and Analysis of Computing Systems, 3(3):1–44, 2019
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d0d4ad1a-f395-4342-a425-e3f90a0f01a5 · outbound
Mathematical methods of reinforcement learning Provably efficient reinforcement learning with linear function approximation
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c462ef11-47a3-4f1e-bd2c-9c342e4f963a · outbound
Mathematical methods of reinforcement learning Linear least-squares algorithms for temporal difference learning.Machine learning, 22(1):33–57, 1996
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d142eb7-5900-4cab-b928-4404be97753c · outbound
Mathematical methods of reinforcement learning Reinforcement learning of motor skills with policy gradients.Neural Networks, 21:682–697, 2008
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 221dcbd7-c929-41e7-828b-9211b9e2bb69 · outbound
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 326af9d8-7ac9-4992-bebb-6aece3e97238 · outbound
Mathematical methods of reinforcement learning Konda and John N
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 808f8a33-eef2-446c-9f53-1c5fda42994a · outbound
Mathematical methods of reinforcement learning Neurocomputing , volume =
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5672a2f2-c10d-4a5d-bd8f-30cc7aa3bdeb · outbound
Mathematical methods of reinforcement learning Arulkumaran, M
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 24823e60-4c89-4f5a-b96b-bdefb12188c6 · outbound
Mathematical methods of reinforcement learning Trust region policy opti- mization via entropy regularization for Kullback–Leibler divergence constraint.Neu- rocomputing, 589:127716, 2024
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c501ec6a-0c56-43ba-b1f9-526212f30e55 · outbound
Mathematical methods of reinforcement learning Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes.Mathematical Program- ming, 198(1):1059–1106, 2023
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d760a2a1-0f15-48a9-ad3c-69d5e7d59b67 · outbound
Mathematical methods of reinforcement learning Proximal Policy Optimization Algorithms
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 979c69d2-4ca7-4ef3-b852-00ca17d33eeb · outbound
Mathematical methods of reinforcement learning Improving proximal policy optimization with alpha divergence.Neurocomputing, 534:94–105,
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4e4ac8e6-8628-4fe0-ad4a-87600a2dfaae · outbound
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 89263e8a-a2e3-4f87-ae49-77870af71df6 · outbound
Mathematical methods of reinforcement learning Reinforcement Learning from Human Feedback: A Statistical Perspective
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 110d358f-45d6-4fc6-a555-e67f35b831ad · outbound
Mathematical methods of reinforcement learning When Do Off-Policy and On-Policy Policy Gradient Methods Align?
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.