Pith. sign in

Paper Citation Record · LEDGER

Mathematical methods of reinforcement learning

As of 18 August 2026, this Paper Citation Record lists 100 of 125 outbound references and 0 inbound Pith citation observations for arXiv:2607.06935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06935 v1

Coverage vector

measured 100 of 125 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T22:47:51.676289Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 125 outbound references displayed

  • verified exact15
  • verified fuzzy79
  • unresolved1
  • parse uncertain2
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c96fc50-d110-4b28-9264-fd9372aeb26c · outbound

This paper cites Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Pro- gramming.

Mathematical methods of reinforcement learning Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Pro- gramming

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.502493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:9428f5ddcf8c24fe88f480314eb1ede1d8c6f0db5a56f8230c7d2d3afabf322c

Observation 3840bc46-ce83-427a-8711-a6671b20ec41 · outbound

This paper cites The Landscape of Agentic Reinforcement Learning for LLMs: A Survey.

Mathematical methods of reinforcement learning The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.774476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:85a0c5ec7c3b288dfc119cc30966a2046649da4e7be3015e4fb2850d54e87dd6

Observation 2bfeddfa-607d-448f-9da5-04af7a69f915 · outbound

This paper cites MIT press, 2022.

Mathematical methods of reinforcement learning MIT press, 2022

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.534166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:fa5641a47fc9b3b1fe765c14cb019592eb339036ec9e17c5cd81375c95a973f8

Observation 607eb37a-f2f6-4e05-9e06-7f2f2c32d65d · outbound

This paper cites Deep belief markov models for pomdp inference.Neural networks, page 108386, 2025.

Mathematical methods of reinforcement learning Deep belief markov models for pomdp inference.Neural networks, page 108386, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.530914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:4f306b0f6bf2715f9a156edbf291dec6af9578b79689c802dc81a6605e656552

Observation b7e8116b-faf7-47b4-86fc-94da8d8ad568 · outbound

This paper cites A markovian decision process.Journal of mathematics and me- chanics, 6(5):679–684, 1957.

Mathematical methods of reinforcement learning A markovian decision process.Journal of mathematics and me- chanics, 6(5):679–684, 1957

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.615465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:a22023bb379ed2ab549eebe8c90b58e6280b9603a8a87482b1de76e565895269

Observation dc7347e8-d4ba-4def-b8db-b5e8c793d64d · outbound

This paper cites An upper bound on the loss from approximate optimal-value functions.Machine Learning, 16(3):227–233, 1994.

Mathematical methods of reinforcement learning An upper bound on the loss from approximate optimal-value functions.Machine Learning, 16(3):227–233, 1994

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.621061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:f28bdf3c1a4c9127816a6185150356273b2f62b0cb15c673b2b2c1c2d9f99621

Observation 8a84f11f-0456-4c3c-bdbf-91ffa2febd5a · outbound

This paper cites an unresolved cited work.

Mathematical methods of reinforcement learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-07-09T23:06:37.542427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:e9df3ecce224e4427ea0f56a260016494a3ab134689c49f34589a829a42fe4d2

Observation 0a247d84-7460-4b2c-a4ed-fa142a938d6b · outbound

This paper cites Improved and generalized upper bounds on the complexity of policy iteration.Advances in Neural Information Processing Systems, 26, 2013.

Mathematical methods of reinforcement learning Improved and generalized upper bounds on the complexity of policy iteration.Advances in Neural Information Processing Systems, 26, 2013

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.547431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:100adcc15dd7ebb9bfcbb1439063f065de890d9a2a633c93e8eaccb624cf72b3

Observation 594cc6c6-d866-44d8-ad15-56587c71c4c0 · outbound

This paper cites Springer, 2018.

Mathematical methods of reinforcement learning Springer, 2018

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.569487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:9d9ca7aa0ecddf9bf47aef3d2eae43c19c846a157ab522c945fed8b840dda919

Observation 2d187ea6-6994-44fc-a1b1-2d32ad828c69 · outbound

This paper cites From Convex Optimization to MDPs: A Review of First-Order, Second-Order and Quasi-Newton Methods for MDPs.

Mathematical methods of reinforcement learning From Convex Optimization to MDPs: A Review of First-Order, Second-Order and Quasi-Newton Methods for MDPs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.745852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:459d810d588bf530ab482c52a3af9a57da4daab472e9450d1e9be280a1bc6f5e

Observation 4e9e4919-ca7a-480d-899a-9bf521a6209d · outbound

This paper cites A method for solving the convex programming problem with con- vergence rate o (1/k2).

Mathematical methods of reinforcement learning A method for solving the convex programming problem with con- vergence rate o (1/k2)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.600268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:f4c0d4eadf42dc002ba16cae5b2cd56524f266bdc2922df99347f1b22daa899b

Observation 50718e16-f93f-4a33-a320-008081266679 · outbound

This paper cites Springer Science & Business Media, 2013.

Mathematical methods of reinforcement learning Springer Science & Business Media, 2013

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:56:37.504613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:33d634208ac36856b0c55d4d043eade61aa65c2938b6305d018bec43c7813976

Observation f6c4c399-d58c-4373-a8cc-19127ff8c3eb · outbound

This paper cites Some methods of speeding up the convergence of iteration methods.

Mathematical methods of reinforcement learning Some methods of speeding up the convergence of iteration methods

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.601878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:42efeee8d36337c4f034bd89e82984947e80c57e651f1707b3218b51d7e8fea3

Observation 46b874da-b796-45de-b692-1b04b4cb6653 · outbound

This paper cites A first-order approach to accelerated value iteration.Operations Research, 71(2):517–535, 2023.

Mathematical methods of reinforcement learning A first-order approach to accelerated value iteration.Operations Research, 71(2):517–535, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.556361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:6acaf155f314d3858a5830d85e8d75ee90a32644689fdc8e80b3a3d85daa72f2

Observation b407913f-067d-4bb8-8d91-1324e0ef9f71 · outbound

This paper cites Pid accelerated value iter- ation algorithm.

Mathematical methods of reinforcement learning Pid accelerated value iter- ation algorithm

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.508074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:9d0e1c2e69575a6d5290cf244e505b4b7f846649e905c2917695e51202bc3dc8

Observation 2c8bdf71-bd6f-4350-9a2a-e9b4219b939b · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Mathematical methods of reinforcement learning A unified view of entropy-regularized Markov decision processes

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.741188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:7fc8971a0814fe61ceace033fd195a4326da30d5854304a442e03823636e2bfb

Observation 09e0cabb-2143-49bd-98f3-a877dfec2952 · outbound

This paper cites Generative adversarial imitation learning.Advances in neural information processing systems, 29, 2016.

Mathematical methods of reinforcement learning Generative adversarial imitation learning.Advances in neural information processing systems, 29, 2016

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.603531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:a24da9b8dd6f7a26828df9a520a8cfd5c4e74f795ed7aacdb0a6ff1956087787

Observation f7070b96-34e2-494c-b0b0-d1c1feeab671 · outbound

This paper cites Variational policy gradient method for reinforcement learning with general utilities.

Mathematical methods of reinforcement learning Variational policy gradient method for reinforcement learning with general utilities

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.624435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:1c4c3a371acdf189f77efc3a088d3f45a802a5419677b4144904e132a3e9e58d

Observation 5622de66-6cb8-415f-97d0-f0ce8f0e848a · outbound

This paper cites Stochastic Optimization under Hidden Convexity.

Mathematical methods of reinforcement learning Stochastic Optimization under Hidden Convexity

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.748081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:044f6cb0799dc01ae8ee7d8b7d6f653431699b30ce9c7dd0d8d9e12ae9e1bae4

Observation dbe4491c-7c82-4134-b8e0-b2170fed9987 · outbound

This paper cites Model-based reinforcement learning with a generative model is minimax optimal.

Mathematical methods of reinforcement learning Model-based reinforcement learning with a generative model is minimax optimal

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.566394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:0c633b607f2dd9343f7b4556d502e014216bac55b511cfd2a9ec8b2f8360c488

Observation b3016a84-d9d8-4c85-a4cb-4843e363023c · outbound

This paper cites Minimax pac bounds on the sample complexity of reinforcement learning with a generative model.

Mathematical methods of reinforcement learning Minimax pac bounds on the sample complexity of reinforcement learning with a generative model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.584620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:1b277c1afb27a1e2944c7bad70cd1e894a82f12e94cca141392a24a57d97e838

Observation 1a4aa522-d10d-4def-85f3-f438933b7681 · outbound

This paper cites Breaking the sample size barrier in model-based reinforcement learning with a generative model.Advances in neural information processing systems, 33:12861–12872, 2020.

Mathematical methods of reinforcement learning Breaking the sample size barrier in model-based reinforcement learning with a generative model.Advances in neural information processing systems, 33:12861–12872, 2020

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.626147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:fe79b5639f4e468a9b7bcc6e8f760f6e1de6ef000be8ec2c2c768154cd202c39

Observation 54d16de7-e98c-48af-a929-5d175ed29e32 · outbound

This paper cites Near-optimal time and sample complexities for solving markov decision processes with a generative model.Advances in Neural Information Processing Systems, 31, 2018.

Mathematical methods of reinforcement learning Near-optimal time and sample complexities for solving markov decision processes with a generative model.Advances in Neural Information Processing Systems, 31, 2018

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.634321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:98d502376514e4dcb8b836d740dfae6381e71640097711f1682c84481b8209c0

Observation 6be2a1cc-dc5f-4921-b98d-dd3ec061397b · outbound

This paper cites Reinforcement learning: Theory and algorithms.CS Dept., UW Seattle, Seattle, WA, USA, Tech.

Mathematical methods of reinforcement learning Reinforcement learning: Theory and algorithms.CS Dept., UW Seattle, Seattle, WA, USA, Tech

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.513427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:7b18d5d79209bc8b7b0439d4ea838953996c28dd8eecee7cc69f38fa9755b696

Observation 96cf21f2-ea97-4172-a02d-64d3893fa985 · outbound

This paper cites Optimal sample complexity for average reward markov decision processes.

Mathematical methods of reinforcement learning Optimal sample complexity for average reward markov decision processes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.544080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:ce47fdbe7a2578515d82731859065629e69c48e64f438cd85e19cdb0520042f7

Observation 0487e0b8-1526-48f8-9c37-f4a48152808e · outbound

This paper cites Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction.Advances in neural information processing systems, 33:7031–7043, 2020.

Mathematical methods of reinforcement learning Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction.Advances in neural information processing systems, 33:7031–7043, 2020

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.522052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:19ebdfc427e23df1c2a28779e3f9549eaa8d8e0ae2eabe9421deb48fc34ca54d

Observation 4493884f-f773-4eba-832e-1c6fd0507e4b · outbound

This paper cites From dirichlet to rubin: Optimistic exploration in rl without bonuses.

Mathematical methods of reinforcement learning From dirichlet to rubin: Optimistic exploration in rl without bonuses

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.532574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:21c643c4066d771d48824f721cae54dbdbea180a492826388f1f5fd101084c75

Observation d468340e-52af-41dd-b1dd-b2e2251eb0b1 · outbound

This paper cites Pac bounds for discounted mdps.

Mathematical methods of reinforcement learning Pac bounds for discounted mdps

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.540846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:011b8d482e8e6126753a997e835684bfc8d124ec63d24e68e99cb31056b90970

Observation 03ef91a2-a2c4-4a05-be39-4e6be5e7af6d · outbound

This paper cites Assouad, fano, and le cam.

Mathematical methods of reinforcement learning Assouad, fano, and le cam

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.629352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:0655ad013e4a7368f6ad95e466d15780a73f8289c25ebb5acdf7042150a48b8f

Observation df529439-d5cf-4154-b3bb-2420c9b74978 · outbound

This paper cites Towards tight bounds on the sample complexity of average-reward mdps.

Mathematical methods of reinforcement learning Towards tight bounds on the sample complexity of average-reward mdps

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.558089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:9229fe52d52592410571f1aa07805889f9b4e9f34d9aa23fb98d7f895ff6b068

Observation 43a93845-8091-4c44-b98a-ccbbf5f8f77a · outbound

This paper cites pages 17, 18.

Mathematical methods of reinforcement learning pages 17, 18

Reference 31

Resolution
parse uncertain
raw_fallback, observed 2026-07-09T23:06:37.535734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:77e26ebe12a2c6b37babd1cd1b037b3766b922f08ca14b92276de61600a47191

Observation e3272491-4db6-4eea-8e57-369a0a88a7b1 · outbound

This paper cites Freedman’s inequality for matrix martingales.

Mathematical methods of reinforcement learning Freedman’s inequality for matrix martingales

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.499104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:a5d888369f52f26025fe06454cda2268b27415b074e5adddea95336f587ba3c5

Observation 14c9d5ae-ad9d-41aa-a718-6799b0080122 · outbound

This paper cites Is q-learning minimax optimal? a tight sample complexity analysis.Operations Research, 72(1): 222–236, 2024.

Mathematical methods of reinforcement learning Is q-learning minimax optimal? a tight sample complexity analysis.Operations Research, 72(1): 222–236, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.517104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:ba69df24fb81591ba215a095a3c8a61150cf3c8b8eccbc1c164b38e9b274f2cf

Observation ce3c1e5a-e422-4beb-b2a8-c467b44733cf · outbound

This paper cites Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning.

Mathematical methods of reinforcement learning Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.755639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:f1a427b97168a652a25138f54c79dec01510605846ed553b365c53488ada4d6c

Observation 4c92c3c3-4836-42e6-9c2c-7c45aecd9f16 · outbound

This paper cites Finite-sample convergence rates for q-learning and indirect algorithms.Advances in neural information processing systems, 11, 1998.

Mathematical methods of reinforcement learning Finite-sample convergence rates for q-learning and indirect algorithms.Advances in neural information processing systems, 11, 1998

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.529072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:e06e6cf67c565174ec5b80dbd44ac749441838f74c276cb39bbc6f3e513cc236

Observation e0b3c3f8-23c4-41e9-8d6d-655bff861b99 · outbound

This paper cites A statistical analysis of polyak-ruppert averaged q-learning.

Mathematical methods of reinforcement learning A statistical analysis of polyak-ruppert averaged q-learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.500740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:b141fb210283c330f37353dab460d5a09e0a1a81fa88979543d1fcd2c294ac6d

Observation 9f6c1b55-1e44-4b51-b1d4-082933adde0a · outbound

This paper cites Variance-reduced $Q$-learning is minimax optimal.

Mathematical methods of reinforcement learning Variance-reduced $Q$-learning is minimax optimal

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.736273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:08b7974bc1dd6cde08f9352d93b235838271620526179d290e58e1183a96a017

Observation feb90f15-4dac-4b57-bfa7-c55160ea031b · outbound

This paper cites Randomized linear programming solves the markov decision problem in nearly linear (sometimes sublinear) time.Mathematics of Operations Research, 45 (2):517–546, 2020.

Mathematical methods of reinforcement learning Randomized linear programming solves the markov decision problem in nearly linear (sometimes sublinear) time.Mathematics of Operations Research, 45 (2):517–546, 2020

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.504323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:d5fdc3c7953e2b6fafe6d078760ace1f6ef6a8abe3aeadc666a840f54cff7014

Observation 2e0b7dab-3863-4abd-9700-f2acaced553d · outbound

This paper cites Efficiently solving mdps with stochastic mirror descent.

Mathematical methods of reinforcement learning Efficiently solving mdps with stochastic mirror descent

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:56:37.502992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:53478d92baea68dddbcef081920e4aec332d27330fd760136b95135ca186f6a7

Observation bd3b8edd-bdc9-461a-9a3d-e6a4e7b5c7e3 · outbound

This paper cites Solving matrix games with near-optimal matvec complexity.arXiv e-prints, pages arXiv–2601, 2026.

Mathematical methods of reinforcement learning Solving matrix games with near-optimal matvec complexity.arXiv e-prints, pages arXiv–2601, 2026

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.605225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:6e2c5e3c1eacb2531dece16ea2fccc78aeda4bc3daaff8931b5708de0d2d0b08

Observation fec0dc65-3bd8-4650-b840-7325450d65b3 · outbound

This paper cites Minimumcostflows, mdps, andℓ1-regressioninnearlylinear time for dense instances.

Mathematical methods of reinforcement learning Minimumcostflows, mdps, andℓ1-regressioninnearlylinear time for dense instances

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.577819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:679c964123a404466f329ee8eaad017846bf0c3c6c9c7c49ff12d4843d191825

Observation 6ad6faa2-2f65-49c6-8c8c-e0f01027688e · outbound

This paper cites Efficient global planning in large mdps via stochas- tic primal-dual optimization.

Mathematical methods of reinforcement learning Efficient global planning in large mdps via stochas- tic primal-dual optimization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.579458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:947c033a3d56be72234a67a80e6fd4b4ad23b82e1ff07899b8eaee27fe91c018

Observation d98744b4-1240-4078-a40c-33ffb1653502 · outbound

This paper cites Tight high probability bounds for linear stochastic approximation with fixed stepsize.Advances in Neural Information Processing Systems, 34:30063–30074,.

Mathematical methods of reinforcement learning Tight high probability bounds for linear stochastic approximation with fixed stepsize.Advances in Neural Information Processing Systems, 34:30063–30074,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.617443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:c214d639671d5162d96e8a33c876e34789817f02abee3a7e667c133192737dc3

Observation d9a65bc3-ddaf-4f36-9980-b3134b9e1a5e · outbound

This paper cites Statistical inferenceforlinearstochasticapproximationwithmarkoviannoise.Advances in Neural Information Processing Systems, 38:174565–174626, 2026.

Mathematical methods of reinforcement learning Statistical inferenceforlinearstochasticapproximationwithmarkoviannoise.Advances in Neural Information Processing Systems, 38:174565–174626, 2026

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.572970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:fe0807d701e4aef73e76cef13d485b6afe6ca453cce49d7dc3224744a5d6b317

Observation 863847bc-a59a-4fd4-8e9a-07ad930a98f9 · outbound

This paper cites Finite-time high-probability bounds for polyak–ruppert averaged iterates of linear stochastic ap- proximation.Mathematics of Operations Research, 50(2):935–964, 2025.

Mathematical methods of reinforcement learning Finite-time high-probability bounds for polyak–ruppert averaged iterates of linear stochastic ap- proximation.Mathematics of Operations Research, 50(2):935–964, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.527414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:18e0def3bcc2e9e08bbaff18684fec81c03e5f4922fd8dd54f9c261a71087459

Observation 02e7238c-bee3-44b2-9004-eb6d39afb31a · outbound

This paper cites Learning to predict by the methods of temporal differences.

Mathematical methods of reinforcement learning Learning to predict by the methods of temporal differences

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.636237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:d80f2e3d0f10ce49710543b66b3c6753af791171177334c6d7f8bfb2e09557d6

Observation bfb97434-0e1c-4dd3-a150-90249a079648 · outbound

This paper cites Improved high-probability bounds for the temporal difference learning algorithm via exponential stability.

Mathematical methods of reinforcement learning Improved high-probability bounds for the temporal difference learning algorithm via exponential stability

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.627821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:efeb4ed5ba64a24f30bf79a7228bc8ec62c6c77f3c3ffb05c1bb2816fd882d73

Observation 0a4bfb33-aabc-4e9c-9539-b06517c8d8f6 · outbound

This paper cites pages 28.

Mathematical methods of reinforcement learning pages 28

Reference 48

Resolution
parse uncertain
raw_fallback, observed 2026-07-09T23:06:37.632533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:30a13b9346c2aa4c84f3b2ccbe5d77b9fd14152496a1a645e82374bbc8006600

Observation ef7e5892-8397-48f7-8764-622c5dcfcec8 · outbound

This paper cites Residual algorithms: Reinforcement learning with function approx- imation.

Mathematical methods of reinforcement learning Residual algorithms: Reinforcement learning with function approx- imation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.613814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:f3f78b3599eebe152e6ab600360d531a982db981797f1593a8a064bf9666be5f

Observation a5451273-2d1f-47dd-9455-1dbcb9cc0c3a · outbound

This paper cites A convergento(n)temporal- difference algorithm for off-policy learning with linear function approximation.Ad- vances in neural information processing systems, 21, 2008.

Mathematical methods of reinforcement learning A convergento(n)temporal- difference algorithm for off-policy learning with linear function approximation.Ad- vances in neural information processing systems, 21, 2008

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.559853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:4d8a75fc9df0efaef633d92e68667d8504079292ecc86c15fce9e3689b57814a

Observation 58568cbb-93ce-4db7-81e0-575fa443bc02 · outbound

This paper cites S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and E.

Mathematical methods of reinforcement learning S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and E

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.588139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:eba2c5d8406a790f8a13fba67f4f028c825376524c0eca7df98800eaa74ee351

Observation c87c284a-cceb-44e7-8b02-e9540456c859 · outbound

This paper cites Gaussian approximation for two-timescale linear stochastic approximation.

Mathematical methods of reinforcement learning Gaussian approximation for two-timescale linear stochastic approximation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.563024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:2c8de637878f1c3f5fe19d70c04d809c2d5d8b312bf40e1e6d35174ca084ad8f

Observation e41f464a-319b-49fc-a105-45a192fac53e · outbound

This paper cites Some aspects of the sequential design of experiments.Bulletin of the American Mathematical Society, 58(5):527–535, 1952.

Mathematical methods of reinforcement learning Some aspects of the sequential design of experiments.Bulletin of the American Mathematical Society, 58(5):527–535, 1952

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.586393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:7ea4e9f5d11cb58854160b3a34ba4373fc914fe1886e51d1fba9175c1df80530

Observation 7b959147-979c-4d20-9f6a-c3687757c0bf · outbound

This paper cites Introduction to multi-armed bandits.Foundations and Trends in Machine Learning, 12(1-2):1–286, 2019.

Mathematical methods of reinforcement learning Introduction to multi-armed bandits.Foundations and Trends in Machine Learning, 12(1-2):1–286, 2019

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.590019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:6af9fa9fb1b6ff6dcd4142235390f11f547ba8ce921ca6e32a4f7311fe541d5f

Observation 23df21be-6c13-4283-ac22-b146bde277c6 · outbound

This paper cites Regret analysis of stochastic and non- stochastic multi-armed bandit problems.Foundations and Trends in Machine Learn- ing, 5(1):1–122, 2012.

Mathematical methods of reinforcement learning Regret analysis of stochastic and non- stochastic multi-armed bandit problems.Foundations and Trends in Machine Learn- ing, 5(1):1–122, 2012

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.598532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:c59ed96d661896abc9daaeac87650ae732840fa3f58c515e6ce60d5055db3de0

Observation 47bc7bf8-28c9-4c32-9409-32c93948d2a8 · outbound

This paper cites Asymptoticallyefficientadaptiveallocationrules.

Mathematical methods of reinforcement learning Asymptoticallyefficientadaptiveallocationrules

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.610463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:3703d6c41bc74f4a6f78cd6b4b829616edf58ccc0c567a1634a8229af80cb692

Observation 2a30c6b4-4d98-45f2-a1d0-21d84d106c8e · outbound

This paper cites Finite-time analysis of the multi- armed bandit problem.Machine Learning, 47(2-3):235–256, 2002.

Mathematical methods of reinforcement learning Finite-time analysis of the multi- armed bandit problem.Machine Learning, 47(2-3):235–256, 2002

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.630993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:82fd9c11aaad98ad82749ac419e5ceaabbb9eccb99f6a1952208b14f504b965c

Observation e6d3077f-f3c9-49ff-a166-00df0f9538d1 · outbound

This paper cites Onthelikelihoodthatoneunknownprobabilityexceedsanother in view of the evidence of two samples.Biometrika, 25(3/4):285–294, 1933.

Mathematical methods of reinforcement learning Onthelikelihoodthatoneunknownprobabilityexceedsanother in view of the evidence of two samples.Biometrika, 25(3/4):285–294, 1933

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.571273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:d669da83efa0980fd8f234488eb14db3c2ee27e986a0d0396cd115c2b16d02d5

Observation 29f09b7d-f6e8-4f16-b3d6-9a49501a13c5 · outbound

This paper cites A tutorial on thompson sampling.Foundations and Trends®in Machine Learning, 11(1):1–96,.

Mathematical methods of reinforcement learning A tutorial on thompson sampling.Foundations and Trends®in Machine Learning, 11(1):1–96,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.582894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:4d27655650899d7e014b1fff996c6c91e66055b39c1c6ad224bf181c0562d4ca

Observation e2f8921d-b4fd-4696-be06-b55b748311c8 · outbound

This paper cites Cambridge University Press,.

Mathematical methods of reinforcement learning Cambridge University Press,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.568008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:c25a6aa8bacde7be3c3b41692cf7aa89553d0fdf80ec22e0cd6c5fdc5b7ec1cc

Observation efca1d98-2947-4f77-9b0e-b9d756863152 · outbound

This paper cites An empirical evaluation of thompson sampling.Ad- vances in neural information processing systems, 24, 2011.

Mathematical methods of reinforcement learning An empirical evaluation of thompson sampling.Ad- vances in neural information processing systems, 24, 2011

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.596758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:852beea8234e41c68b8a395bdfdd8b1193abdb11fbd9feb0d38209ebcde49343

Observation d423b61c-bdb3-4f7e-94ee-efa0dbf881b9 · outbound

This paper cites Further optimal regret bounds for thompson sam- pling.

Mathematical methods of reinforcement learning Further optimal regret bounds for thompson sam- pling

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.574497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:45cc158dafd68e74553aba8e1b305977d627ffc9de5bf975e1181a5941eb8e42

Observation 3c3b4746-6292-457a-92ed-0c14e1feec65 · outbound

This paper cites Analysis of thompson sampling for the multi-armed bandit problem.

Mathematical methods of reinforcement learning Analysis of thompson sampling for the multi-armed bandit problem

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.552759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:2bdf82fd1cb4c2eca3f72c4e4ae27dbeac4930b509ea3d038c96158c3fba5584

Observation 38ea0d67-7489-4b00-a8eb-0af066c4cb07 · outbound

This paper cites Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems.

Mathematical methods of reinforcement learning Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T22:56:37.733876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:a33eb6935c1a0f777214ac31418e4ef23006b53fffc258a243133ff1b4576c7e

Observation a2a1c426-57ee-474a-b6ba-4300d2ed6a4e · outbound

This paper cites Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited.

Mathematical methods of reinforcement learning Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.554424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:e4350adca479f547ee88de72c8f694e00b8815b4c0d07c59e4bb131401954976

Observation 979c3d86-060e-44e1-af69-b9d76768720a · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.Journal of Machine Learning Research, 11:1563–1600, 2010.

Mathematical methods of reinforcement learning Near-optimal regret bounds for reinforcement learning.Journal of Machine Learning Research, 11:1563–1600, 2010

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.591710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:f5cd22f5366e61f1ab781bd84a3f3bbd23f747064259f222bf06b249741c2b24

Observation 658b9dd7-2d3d-4471-b1b6-bfacffd5ecc0 · outbound

This paper cites A unifying view of optimism in episodic reinforce- ment learning.

Mathematical methods of reinforcement learning A unifying view of optimism in episodic reinforce- ment learning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.612156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:c7cedf573230b5ae7262b63c3f395790b847243e97da31d4f9bc9fe9fb7bf48e

Observation f807cfed-f852-4d1f-b30b-0ed459468af8 · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Mathematical methods of reinforcement learning Minimax regret bounds for reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.606967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:ea5ca71c008c2b78f6096dcea81f153e7d9bf9cd13c4cd3a12f9b44716d46032

Observation 52656bc2-427d-4740-8e89-23bada35e115 · outbound

This paper cites Rusu, Joel Veness, Marc G.

Mathematical methods of reinforcement learning Rusu, Joel Veness, Marc G

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.561434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:09fc1936aa469f9b53b02d2bdafd6cefc013dd13d9fd2b1e94e295681495c2d8

Observation f5001757-cc5a-44e4-83b6-102cd1caed61 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Mathematical methods of reinforcement learning Asynchronous methods for deep reinforcement learning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.593511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:d02dac0c96d95b3a1acfebc7b59dd633323347cf8fc0b0ba4a7659d5a8b81678

Observation 2bc3cfb2-7dcf-4522-92ec-8d4245b58e78 · outbound

This paper cites Trust region policy optimization.

Mathematical methods of reinforcement learning Trust region policy optimization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.595136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:e5462bbf6d9482c87526d804d3e0dfc8e2cb8296883b24f4294eaf0cd5359ee3

Observation fb2e76d8-792e-4984-bf33-350c04862815 · outbound

This paper cites Generalization and exploration via randomized value functions.

Mathematical methods of reinforcement learning Generalization and exploration via randomized value functions

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:56:37.508406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:1e1b81ba4ab3a8775c77095c5b659549b495c9421afa9cef6feede0787ca7ba2

Observation 5f7bc0a9-979f-42ac-b971-81f27e2b1922 · outbound

This paper cites Is q-learning provably efficient?Advances in neural information processing systems, 31, 2018.

Mathematical methods of reinforcement learning Is q-learning provably efficient?Advances in neural information processing systems, 31, 2018

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.576159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:94311521572296073a7505b2e17328a8aca2ebb042e4afa52a474835773c68c7

Observation 62b6891e-6bcc-4c72-84e4-ea8c406cacd2 · outbound

This paper cites Almost optimal model-free reinforce- ment learningvia reference-advantage decomposition.Advances in Neural Information Processing Systems, 33:15198–15207, 2020.

Mathematical methods of reinforcement learning Almost optimal model-free reinforce- ment learningvia reference-advantage decomposition.Advances in Neural Information Processing Systems, 33:15198–15207, 2020

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.581204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:32a5d7cbf130caa337aa2b9a609da8074acde86cb28942bcfd46227d8198c887

Observation 4c2dc31f-1b9f-47a1-9fd4-518414f1f8ca · outbound

This paper cites (more) efficient reinforcement learning via posterior sampling.Advances in Neural Information Processing Systems, 26, 2013.

Mathematical methods of reinforcement learning (more) efficient reinforcement learning via posterior sampling.Advances in Neural Information Processing Systems, 26, 2013

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.564706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:0e0e7cd257869099867815dfabbb1a0d4f194c94c366206f1b513d53195f8d6d

Observation 43600ac8-cd84-486c-b1a9-c8c9a9d5e4ae · outbound

This paper cites Optimistic posterior sampling for reinforcement learn- ing: Worst-caseregretbounds.

Mathematical methods of reinforcement learning Optimistic posterior sampling for reinforcement learn- ing: Worst-caseregretbounds

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.622702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:a883edbd6c05507c68f31f5133ddb0c5b052cf03b6dad3969159ee854149408f

Observation b1a56240-a299-4146-b4a0-0f903b189cb8 · outbound

This paper cites Optimistic pos- terior sampling for reinforcement learning with few samples and tight guarantees.

Mathematical methods of reinforcement learning Optimistic pos- terior sampling for reinforcement learning with few samples and tight guarantees

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.518797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:88dd0023542873fb940e00629f4a75ba802be6ac276ee215c064cbda326b299a

Observation a52d8616-c9df-447f-85d0-ccdeea1502c1 · outbound

This paper cites Deep exploration via randomized value functions.Journal of machine learning research, 20(124):1–62,.

Mathematical methods of reinforcement learning Deep exploration via randomized value functions.Journal of machine learning research, 20(124):1–62,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.537575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:6908b1155dc14d6fff388f5e031e8be1524d8b2a61ad03737430a9ff58e2639b

Observation c5f613ff-65f0-4b19-abbf-5344212064b4 · outbound

This paper cites Worst-case regret bounds for exploration via randomized value func- tions.Advances in neural information processing systems, 32, 2019.

Mathematical methods of reinforcement learning Worst-case regret bounds for exploration via randomized value func- tions.Advances in neural information processing systems, 32, 2019

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.525461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:243e7e9f59ea01c0a2c3ef26825e29380c8e99be1f606e3dd0ac3c2f007a0160

Observation ee221d4e-9358-46e7-8e9b-ff9a6167f679 · outbound

This paper cites Improved worst-case regret bounds for randomized least-squares value iteration.

Mathematical methods of reinforcement learning Improved worst-case regret bounds for randomized least-squares value iteration

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.523757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:0c450b3027868e3c86f05d2538e2104778e59b802d304f89b701ceba6143f375

Observation 8b7ce566-a3cb-43f5-b7ab-a007538120c3 · outbound

This paper cites Near-optimal randomized exploration for tabular markov decision processes.Advances in neural information processing systems, 35:6358–6371, 2022.

Mathematical methods of reinforcement learning Near-optimal randomized exploration for tabular markov decision processes.Advances in neural information processing systems, 35:6358–6371, 2022

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.619182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:ee179897378b28f7e35464ec459fdc45a12b50a5e412ccc9eb25b9d41090a0de

Observation fd5a70ac-0eba-40e5-a724-ef79cb96a95a · outbound

This paper cites Finite-time bounds for fitted value iteration.

Mathematical methods of reinforcement learning Finite-time bounds for fitted value iteration

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.545725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:b3cc90c99f18fdc8b2086e4baea70f8d1fdd6ea7b3bebfb91d1f1ae46a62e3bf

Observation 4d41807c-a370-4f7a-8279-977d1e35412e · outbound

This paper cites Error bounds for approximate policy iteration.

Mathematical methods of reinforcement learning Error bounds for approximate policy iteration

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.539240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:3eb5255fb7c9c958a4d96c2e54319faea575b9ae96da3abced34ad3672a7b738

Observation cecc9e32-113b-4da4-9e6b-c1bc17acc3e6 · outbound

This paper cites Kernel-based reinforcement learning.Machine learn- ing, 49(2):161–178, 2002.

Mathematical methods of reinforcement learning Kernel-based reinforcement learning.Machine learn- ing, 49(2):161–178, 2002

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.608890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:f78161480e09f5aa170d378f8a89e041bc1316cbdbfd855440ed5f6872a3ff97

Observation 8cda6af4-4335-461f-a18d-36c3d25cfcdd · outbound

This paper cites Kernel-based reinforcement learning: A finite-time analysis.

Mathematical methods of reinforcement learning Kernel-based reinforcement learning: A finite-time analysis

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.506146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:20a36fba8997f6f303695cce0ba748916c0763199f6b9aa5e9af5913ddcfdaf8

Observation 99c9189c-d339-4b0c-8a72-91fb4b45f762 · outbound

This paper cites Adaptive discretization for episodic reinforcement learning in metric spaces.Proceedings of the ACM on Measurement and Analysis of Computing Systems, 3(3):1–44, 2019.

Mathematical methods of reinforcement learning Adaptive discretization for episodic reinforcement learning in metric spaces.Proceedings of the ACM on Measurement and Analysis of Computing Systems, 3(3):1–44, 2019

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:56:38.249924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:6b412a4eb867770dae7ab7c5dbf965b8c7eb65852d71aaf318d1b497c2b0c799

Observation d0d4ad1a-f395-4342-a425-e3f90a0f01a5 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Mathematical methods of reinforcement learning Provably efficient reinforcement learning with linear function approximation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:56:38.247742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:c4d0359a2f08eb272a482f3b9bcffb217be72a4a16f6bcdbd7cc5075c20b0d3b

Observation c462ef11-47a3-4f1e-bd2c-9c342e4f963a · outbound

This paper cites Linear least-squares algorithms for temporal difference learning.Machine learning, 22(1):33–57, 1996.

Mathematical methods of reinforcement learning Linear least-squares algorithms for temporal difference learning.Machine learning, 22(1):33–57, 1996

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T22:56:38.251963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:738877686b6b0cc26b2b34eed18dbb0187ec8e9c854a010005094983de1622fb

Observation 2d142eb7-5900-4cab-b928-4404be97753c · outbound

This paper cites Reinforcement learning of motor skills with policy gradients.Neural Networks, 21:682–697, 2008.

Mathematical methods of reinforcement learning Reinforcement learning of motor skills with policy gradients.Neural Networks, 21:682–697, 2008

Reference 89

Resolution
verified exact
doi, observed 2026-07-09T22:56:37.583878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:8f7c906edafd46193dc064651f2f3b6db3716c7f6e2c6556f02d18ba22840cd3

Observation 221dcbd7-c929-41e7-828b-9211b9e2bb69 · outbound

This paper cites Williams.

Mathematical methods of reinforcement learning Williams

Reference 90

Resolution
malformed identifier
raw_fallback, observed 2026-07-09T22:56:38.244791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:8e2d23035784f72fafc9753d72239339fbc993689e4bd9a584ec4896367d0063

Observation 326af9d8-7ac9-4992-bebb-6aece3e97238 · outbound

This paper cites Konda and John N.

Mathematical methods of reinforcement learning Konda and John N

Reference 91

Resolution
verified exact
doi, observed 2026-07-09T22:56:37.595607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:b62c04fe4656b2f2c2d9a2f63f9f215f72e3926ee7e4f51411c2fd077d07389c

Observation 808f8a33-eef2-446c-9f53-1c5fda42994a · outbound

This paper cites Neurocomputing , volume =.

Mathematical methods of reinforcement learning Neurocomputing , volume =

Reference 92

Resolution
verified exact
doi, observed 2026-07-09T22:56:37.579375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:be540de837b19369caaca931f31832525ed9fa2a7ca1a580463c9e4492eb9e7e

Observation 5672a2f2-c10d-4a5d-bd8f-30cc7aa3bdeb · outbound

This paper cites Arulkumaran, M.

Mathematical methods of reinforcement learning Arulkumaran, M

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T22:56:37.588235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:209c3164be33d894be1756e76ec4c97f6948b9ffb8d6293efe881e2ad535aae0

Observation 24823e60-4c89-4f5a-b96b-bdefb12188c6 · outbound

This paper cites Trust region policy opti- mization via entropy regularization for Kullback–Leibler divergence constraint.Neu- rocomputing, 589:127716, 2024.

Mathematical methods of reinforcement learning Trust region policy opti- mization via entropy regularization for Kullback–Leibler divergence constraint.Neu- rocomputing, 589:127716, 2024

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-09T22:56:37.582084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:76b422d3740c2fe6bfae8abfe460440a0bce8dbeeb1c029191100a2d30fbce42

Observation c501ec6a-0c56-43ba-b1f9-526212f30e55 · outbound

This paper cites Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes.Mathematical Program- ming, 198(1):1059–1106, 2023.

Mathematical methods of reinforcement learning Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes.Mathematical Program- ming, 198(1):1059–1106, 2023

Reference 95

Resolution
verified exact
doi, observed 2026-07-09T22:56:37.585701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:d8bb3403505fb32b731cb58db04d991cff6690829a14c131e4c4931f92bb8ed9

Observation d760a2a1-0f15-48a9-ad3c-69d5e7d59b67 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mathematical methods of reinforcement learning Proximal Policy Optimization Algorithms

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.772074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:a02ae3f7fc04332c38b8df577c3a71db09a931d76ac66bb09935cf5106d59f50

Observation 979c69d2-4ca7-4ef3-b852-00ca17d33eeb · outbound

This paper cites Improving proximal policy optimization with alpha divergence.Neurocomputing, 534:94–105,.

Mathematical methods of reinforcement learning Improving proximal policy optimization with alpha divergence.Neurocomputing, 534:94–105,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T23:06:37.551094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:0458be4788d5c359116d0cee614de6332340a107233e5f9553f22a03aa0a5261

Observation 4e4ac8e6-8628-4fe0-ad4a-87600a2dfaae · outbound

This paper cites pages 54.

Mathematical methods of reinforcement learning pages 54

Reference 98

Resolution
verified exact
doi, observed 2026-07-09T22:56:37.597170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:d3ab78ff00468e1d784cecfe7ae40f765d5e9c7c3581883ef3bda6463a5711a5

Observation 89263e8a-a2e3-4f87-ae49-77870af71df6 · outbound

This paper cites Reinforcement Learning from Human Feedback: A Statistical Perspective.

Mathematical methods of reinforcement learning Reinforcement Learning from Human Feedback: A Statistical Perspective

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.791390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:acdf9cf1330574b943acbaf145f9c9dcfa764dab3908eca8bbb324b1d4e6c443

Observation 110d358f-45d6-4fc6-a555-e67f35b831ad · outbound

This paper cites When Do Off-Policy and On-Policy Policy Gradient Methods Align?.

Mathematical methods of reinforcement learning When Do Off-Policy and On-Policy Policy Gradient Methods Align?

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.784291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:8bf2d721e864af6ea80fe693046fb853bb9201c90f1cdbd6d1c52f1c5ddf4767

Pith citing papers

No inbound Pith citation observations are available.