Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T05:47:47.782246Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 4 inbound Pith citation observations for arXiv:2510.16132.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T05:47:47.782246Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:51:04.104314Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-28T23:12:46.704827Z
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b9ca54e1-5211-4ff7-b235-5dcb51fba50b · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies MIT press
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d4f5148d-91eb-4074-a1f2-38c03889ed58 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Mastering the game of Go without human knowledge.Nature, 550(7676):354
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8a448e8a-4c02-4048-9657-4a91c76ea079 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies End-to-endtrainingofdeepvisuomotor policies.Journal of Machine Learning Research, 17(39):1–40
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9d108c9a-200e-497d-8d61-963183a0fbb4 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Reinforcementlearningbasedrecommendersystems: A survey.ACM Comput
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d7366329-0358-4172-8f7d-34903b80e443 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c35feb33-52df-43db-987d-a563ef8fbe98 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Q-learning.Machine learning, 8(3-4):279–292
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e5b9b4e2-5e66-4551-a898-5304b62d629b · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A stochastic approximation method.The Annals of Mathematical Statistics, pages 400–407
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 157f3967-b039-4b1b-9099-9916e2018c17 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Rusu, Joel Veness, Marc G
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1181646d-c12d-4e10-b14e-3ead5ebda138 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Asynchronous stochastic approximation and Q-learning.Machine learning, 16(3): 185–202
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 21406b28-d5f0-454f-8f93-947fa0b0fb0f · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies The ODE method for convergence of stochastic approximation and reinforcement learning.SIAM Journal on Control and Optimization, 38(2):447–469
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4f33f359-c396-4228-ac20-4634d1e519c0 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Springer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 69317bae-e396-4799-bbcb-16b1bdd2ae3a · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies The asymptotic convergence-rate of q-learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 24089893-345c-45a7-854c-6cae3e62a881 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Beck and R
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 570ab767-c2e4-4b4e-a7c6-0e05eb209442 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Beck and R
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1d1d9469-7edb-4682-a7fc-1e8a2efc5e8f · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A Lyapunov theory for finite-sample guarantees of Markovian stochastic approximation.Operations Research, 72(4): 1352–1367
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8ce43545-3d7f-4f83-92c8-18ef34751507 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Final iteration convergence bound of Q-learning: Switching system approach.IEEE Transactions on Automatic Control, 69(7):4765–4772
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e8da3c8f-4da7-4b4a-a18c-38ae475180bb · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e442f492-2650-42c9-ba2e-7ec42922dc85 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Variance-reduced $Q$-learning is minimax optimal
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 61420611-9398-41f8-aeb5-d6e9bc60dedb · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Learning rates for Q-learning.Journal of Machine Learning Research, 5(Dec):1–25
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b1ca6e2e-e954-4653-8d8d-eb5a147e7428 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Finite-time analysis of asynchronous stochastic approximation and Q-learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 09a9234b-3834-4dd3-80cd-9e9f5f5358f2 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Sample complexity of asynchronous Q-learning: sharper analysis and variance reduction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9c53fb4e-5279-4c18-80c0-f61414134787 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A statistical analysis of polyak-ruppert averaged q-learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6dcf852c-6970-439c-aa58-1840e8e02047 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5cfcae1b-57c4-414e-ac26-5bfec2989721 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Is q-learning provably efficient? Advances in neural information processing systems, 31
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b42121ae-070b-4859-85e6-37be898e1cdc · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ed1504d7-e858-4fff-a53e-2466ad2dd7e9 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Deep reinforcement learning with double q-learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9e833408-06e9-418f-96ec-5a7613cf462c · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Value-differencebasedexploration: Adaptivecontrolbetween 𝜖-greedy and softmax
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8af9cfa4-0eaf-474c-9f98-5f54ea7fce24 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Dueling network architectures for deep reinforcement learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 504677f2-d2b5-4080-b107-5f49defaca19 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Finite-time error bounds for linear stochastic approximation and TD-learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 45ec2924-76f6-4489-92ff-f4bb5e0e0c9c · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A finite-time analysis of temporal difference learning with linear function approximation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 591423c9-fbd6-41c6-80e2-eb427e05eac3 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Finite-sample analysis for SARSA with linear function approximation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 115ce0fd-40e2-49ad-b329-ac3e2c9f221d · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Stochastic approximation with unbounded Markovian noise: A general-purpose theorem
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2dc4bd93-7a98-4e6a-95e1-214a03ed1f59 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Concentration of contractive stochastic approximation and reinforcement learning.Stochastic Systems, 12(4):411–430
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c22e91bc-90c0-4dd5-84e3-bcbde789ed42 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Convergence of stochastic iterative dynamic programming algorithms
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 549421a8-182b-4158-a414-b8e9113d13e8 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A unified switching system perspective and convergence analysis of Q-learning algorithms.Advances in Neural Information Processing Systems, 33:15556–15567
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2291be45-5c55-4511-92aa-63c321465bd3 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Zapq-learning.AdvancesinNeuralInformationProcessingSystems, 30
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 748b2d33-e46f-4732-b6ff-a5c2bfed9d16 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Instance-optimality in optimal valueestimation: Adaptivityviavariance-reducedq-learning.IEEETransactionsonInformationTheory
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 97155cf2-5cfc-4627-92be-fd8d4a93244e · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c6cf3182-ab9f-48a8-9962-df2c645d9fdd · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies An analysis of reinforcement learning with function approximation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a66069be-9dd5-4796-b973-04b0ab453e23 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Target network and truncation overcome the deadly triad in q-learning.SIAM Journal on Mathematics of Data Science, 5(4):1078–1101
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9d91f3b9-ec2b-4d8f-8d14-07cffe222769 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies TheprojectedBellmanequationinreinforcementlearning.IEEETransactionsonAutomatic Control
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4ab33dd2-86ac-4cf5-a0e3-9665646bf8d8 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies The blessing of heterogeneity in federated q-learning: Linear speedup and beyond.Journal of Machine Learning Research, 26(26):1–85
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a22bbfe9-ea11-4a87-81d7-eb28731b3e5a · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Federated reinforcement learning: Linearspeedupundermarkoviansampling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 402a0e6a-93fe-4ae9-b620-c6c70f702e1b · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Q-learning with logarithmic regret
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 21cb9915-568d-449f-9b24-36abc0302b99 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Online Q-learning using connectionist systems.University of Cambridge, Department of Engineering, Cambridge, UK, 37
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3743101c-2284-4072-9633-6c278603e637 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Convergence results for single-step on-policy reinforcement-learning algorithms.Machine learning, 38:287–308
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f1313c6c-bed2-4cc5-92e7-f24c505ec7ff · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies On the convergence of sarsa with linear function approximation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aaa09663-c42c-4b61-8e0b-319dedf49445 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A finite-time analysis of two time-scale actor-critic methods.Advances in Neural Information Processing Systems, 33:17617–17628
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 93e23c6f-f8fb-4c1a-8ac7-6f2aa4ea2a2e · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Finite sample analysis of two-time-scale natural actor-critic algorithm.IEEE Transactions on Automatic Control
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation acb7a39c-9226-4f1e-b96a-81adc984d3bd · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A finite-sample analysis of payoff-based independent learning in zero-sum stochastic games.Advances in Neural Information Processing Systems, 36:75826–75883
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2ad4678d-86a4-4efa-8b2e-a74eb0860be8 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Two-timescale Q-learning with function approximation in zero-sum stochastic games
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 43e28953-6f72-4311-83c6-2652b3797038 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies John Wiley & Sons
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 17531437-0612-4f18-8223-8197579f5c6e · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Athena Scientific
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f045ed24-6843-4e61-a346-690ec6b8da2a · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Surlesopérationsdanslesensemblesabstraitsetleurapplicationauxéquationsintégrales
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1b4409d4-49d8-4fd8-b8ba-36b530e55e52 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d409e13f-b061-4f0c-a581-be87d7af282d · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies AmericanMathematical Soc
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8728ed8e-5e97-4e17-a6a6-ae862df3e1fb · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Sample and communication-efficient decentralized actor-critic algorithms with finite-time analysis
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5c50e7e5-bc53-44dc-a4e1-0ec69633b0d4 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Sample efficient stochastic policy extra-gradient algorithm for zero-sum markov game
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cb5597b9-418c-47d6-9b89-e05227925f81 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Sample complexity bounds for two timescale value-based reinforcement learningalgorithms
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e2347e59-b47f-4bc9-a945-94f7cbc3e8b9 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies On finite-time convergence of actor-critic algorithm.IEEE Journal on Selected Areas in Information Theory, 2(2):652–664
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 46293fee-e626-4f3f-9100-5ebd516e70e4 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Cambridge University Press
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f07419b0-f028-41e5-be68-0f962d2e291a · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Sharper model-free reinforcement learning for average-reward markov decision processes
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 287162d5-da44-4a40-a323-5ea354c2632d · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d622e720-2f6d-4c5e-b93e-a71ca2230213 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e369d986-826f-4399-9bc9-07aa4076dc84 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Solution representations for poisson’s equation, martingale structure, and the markov chain central limit theorem.Stochastic Systems, 14(1):47–68
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b5fb187-8840-4a1e-97c6-c5166a59f520 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Generalized inverses and their application to applied probability problems.Linear Algebra and its Applications, 45:157–198
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f4ab5305-30f9-4d4b-8ede-20ecd76d15f9 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A liapounov bound for solutions of the poisson equation.The Annals of Probability, pages 916–931
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2c1f914b-abe2-4ad2-84ac-2b3148387d25 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies An approximate policy iteration viewpoint of actor–critic algorithms.Automatica, 179:112395
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 08bc700b-47cd-4237-a278-1707a8269e22 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Boundedness of iterates in Q-learning.Systems & control letters, 55(4):347–349
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f571dc90-2911-4582-a792-e271205f56f8 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Minimax pac bounds on the sample complexity of reinforcement learning with a generative model.Machine learning, 91(3):325–349
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f9b421a6-7f4d-4dcd-94a5-6b8e8acaace5 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Online learning and online convex optimization.Foundations and Trends®in Machine Learning, 4(2):107–194
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e12f59b6-387d-4bc2-a5a8-e52a4656242c · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Beck.First-Order Methods in Optimization
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f747d203-dc94-4722-9b94-a8d9a373b005 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 74c0dc09-8578-4e6a-b435-cacb90495b4d · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies To this end, define𝑧𝑘 := Í𝑘 𝑛=0 P 𝑛𝑦 for any𝑘≥0
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b12d8ff-faa6-40ba-ac89-6a0cb3114080 · outbound
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies 𝑟𝑏∑︁ 𝑖=0 𝑟𝑏 +1 𝑖+1 𝑃𝑖 𝜋𝑏 (𝑠 ′′, 𝑠′) # 𝜋𝑏 (𝑎 ′|𝑠 ′) (Change of variable:𝑖=𝑗−1) = 1 2𝑟𝑏+1 ∑︁ 𝑠′′ ∈ S 𝑝(𝑠 ′′ |𝑠, 𝑎)
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3b08014c-0621-417b-acd6-e070e5b31058 · inbound
Auto-exploration for online reinforcement learning A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb299ed5-5f90-49d5-84a4-6ed886b883ac · inbound
Auto-exploration for online reinforcement learning A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1628a8a3-f589-4358-93ca-05282a1c6e37 · inbound
Achieving $\epsilon^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e15c10b6-921f-4eb5-aa07-40dbad77416e · inbound
Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.