Pith. sign in

Paper Citation Record · LEDGER

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies

As of 15 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 4 inbound Pith citation observations for arXiv:2510.16132.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.16132 v2

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T05:47:47.782246Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:51:04.104314Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-28T23:12:46.704827Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact9
  • verified fuzzy59
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b9ca54e1-5211-4ff7-b235-5dcb51fba50b · outbound

This paper cites MIT press.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies MIT press

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.898472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:875a9258b2662d77f434755dfd9025a0960fcf115d86fec3dc8d29170556049c

Observation d4f5148d-91eb-4074-a1f2-38c03889ed58 · outbound

This paper cites Mastering the game of Go without human knowledge.Nature, 550(7676):354.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Mastering the game of Go without human knowledge.Nature, 550(7676):354

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.883528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:85b7dcf1d23bb5669f0646d8ce7b728ddaf2ccea34ce39b3b7b4736839c4b3d7

Observation 8a448e8a-4c02-4048-9657-4a91c76ea079 · outbound

This paper cites End-to-endtrainingofdeepvisuomotor policies.Journal of Machine Learning Research, 17(39):1–40.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies End-to-endtrainingofdeepvisuomotor policies.Journal of Machine Learning Research, 17(39):1–40

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.831147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:4c745e74130cc00cc2763f16820a7e734ab06049a63662e5c6a0011654f894ed

Observation 9d108c9a-200e-497d-8d61-963183a0fbb4 · outbound

This paper cites Reinforcementlearningbasedrecommendersystems: A survey.ACM Comput.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Reinforcementlearningbasedrecommendersystems: A survey.ACM Comput

Reference 4

Resolution
verified exact
doi, observed 2026-05-18T05:50:56.401270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:8a07939b7e94e7090c33b32d38a7999c4d8c4e3f1a3c2eaf69f5167eb686a9ab

Observation d7366329-0358-4172-8f7d-34903b80e443 · outbound

This paper cites an unresolved cited work.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-18T05:50:57.920556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:06441773efc71be6cf068a2f02a1b1448ae1d3de20fc018565e8d87b30016e54

Observation c35feb33-52df-43db-987d-a563ef8fbe98 · outbound

This paper cites Q-learning.Machine learning, 8(3-4):279–292.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Q-learning.Machine learning, 8(3-4):279–292

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.860221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:ec1fb855e70f397407bcd7b55eefcad99cd9fb9d8d11d75ae24d9b2b5764d53b

Observation e5b9b4e2-5e66-4551-a898-5304b62d629b · outbound

This paper cites A stochastic approximation method.The Annals of Mathematical Statistics, pages 400–407.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A stochastic approximation method.The Annals of Mathematical Statistics, pages 400–407

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.857739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:a9f3e33caf5237db37f0bde5a38043e58098a056444415eeb3807b4f91a340cf

Observation 157f3967-b039-4b1b-9099-9916e2018c17 · outbound

This paper cites Rusu, Joel Veness, Marc G.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Rusu, Joel Veness, Marc G

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.915996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:8dec4a4f2fb5b7f3292994faa389172817be25a02eb137b2e523ce42c87934f2

Observation 1181646d-c12d-4e10-b14e-3ead5ebda138 · outbound

This paper cites Asynchronous stochastic approximation and Q-learning.Machine learning, 16(3): 185–202.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Asynchronous stochastic approximation and Q-learning.Machine learning, 16(3): 185–202

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.781622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:36f8341d95ad9e5e2d706ae496abd74a6d4b57325deb10230593ff3517a306dc

Observation 21406b28-d5f0-454f-8f93-947fa0b0fb0f · outbound

This paper cites The ODE method for convergence of stochastic approximation and reinforcement learning.SIAM Journal on Control and Optimization, 38(2):447–469.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies The ODE method for convergence of stochastic approximation and reinforcement learning.SIAM Journal on Control and Optimization, 38(2):447–469

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.871074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:35b2c6ec61911bf0493e7a832ddf42c61e8672e5823432b3083d570f49223418

Observation 4f33f359-c396-4228-ac20-4634d1e519c0 · outbound

This paper cites Springer.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Springer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.821547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:01d711b91ffc0b364485ec3795dd6fca059a1622165c44cf7a0a59995aec1f78

Observation 69317bae-e396-4799-bbcb-16b1bdd2ae3a · outbound

This paper cites The asymptotic convergence-rate of q-learning.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies The asymptotic convergence-rate of q-learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.779073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:ace117cf169e57fbe7a6b3f9d21c73a67b58d09c9ef6fb149b3c48c01aabb9df

Observation 24089893-345c-45a7-854c-6cae3e62a881 · outbound

This paper cites Beck and R.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Beck and R

Reference 13

Resolution
verified exact
doi, observed 2026-05-18T05:50:56.404442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:b91f80e0d423379c296c1fe991e6d5d1e667e55df68f5edb437e08755968d8f0

Observation 570ab767-c2e4-4b4e-a7c6-0e05eb209442 · outbound

This paper cites Beck and R.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Beck and R

Reference 14

Resolution
verified exact
doi, observed 2026-05-18T05:50:56.408356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:4ed90e6fd7e04c63adcfa99c580d0160f6fc00ec9294dd16698612fb115cca7d

Observation 1d1d9469-7edb-4682-a7fc-1e8a2efc5e8f · outbound

This paper cites A Lyapunov theory for finite-sample guarantees of Markovian stochastic approximation.Operations Research, 72(4): 1352–1367.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A Lyapunov theory for finite-sample guarantees of Markovian stochastic approximation.Operations Research, 72(4): 1352–1367

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.880442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:2c3f7183ea2cadba08d7e4c39feecd3cd54ae823077b656ef047343724b001bc

Observation 8ce43545-3d7f-4f83-92c8-18ef34751507 · outbound

This paper cites Final iteration convergence bound of Q-learning: Switching system approach.IEEE Transactions on Automatic Control, 69(7):4765–4772.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Final iteration convergence bound of Q-learning: Switching system approach.IEEE Transactions on Automatic Control, 69(7):4765–4772

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.783901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:d7dc78511657139dd4a985b53a47459c01d1e5fa21d80700dc197244533a91e6

Observation e8da3c8f-4da7-4b4a-a18c-38ae475180bb · outbound

This paper cites Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:50:57.144619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:a3eb684d1590876b457104cdcb16927d01a8e8cf20347d1c5a6067a084d74fde

Observation e442f492-2650-42c9-ba2e-7ec42922dc85 · outbound

This paper cites Variance-reduced $Q$-learning is minimax optimal.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Variance-reduced $Q$-learning is minimax optimal

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:50:57.149874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:24724ef2e5aad99f4dccc168bf0f80113f8f75945613795cf516e23ea1f02571

Observation 61420611-9398-41f8-aeb5-d6e9bc60dedb · outbound

This paper cites Learning rates for Q-learning.Journal of Machine Learning Research, 5(Dec):1–25.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Learning rates for Q-learning.Journal of Machine Learning Research, 5(Dec):1–25

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.812070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:acaef4c2bf3746552bee7f5ceaf1ae21882e29223c227da4dc8efb0b54e2ca01

Observation b1ca6e2e-e954-4653-8d8d-eb5a147e7428 · outbound

This paper cites Finite-time analysis of asynchronous stochastic approximation and Q-learning.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Finite-time analysis of asynchronous stochastic approximation and Q-learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.923273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:3e3f05598feb1a8ea90628db5daa52cd8f6063b4b48b381afaa61747e1646ff7

Observation 09a9234b-3834-4dd3-80cd-9e9f5f5358f2 · outbound

This paper cites Sample complexity of asynchronous Q-learning: sharper analysis and variance reduction.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Sample complexity of asynchronous Q-learning: sharper analysis and variance reduction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.854241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:24ad8a5cd8b73e55eee6083aac1c4a83f2171b42df150aa2c07c860c7304fede

Observation 9c53fb4e-5279-4c18-80c0-f61414134787 · outbound

This paper cites A statistical analysis of polyak-ruppert averaged q-learning.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A statistical analysis of polyak-ruppert averaged q-learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.809770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:b8f04ac1f43ede891ddabd328ccc35152636ca699fff2ff534649669b2f92303

Observation 6dcf852c-6970-439c-aa58-1840e8e02047 · outbound

This paper cites an unresolved cited work.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-18T05:50:57.910984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:29a9527257ebe4dad62b15b635f30324bf3105505bfba5d0617f1e9dce531b9c

Observation 5cfcae1b-57c4-414e-ac26-5bfec2989721 · outbound

This paper cites Is q-learning provably efficient? Advances in neural information processing systems, 31.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Is q-learning provably efficient? Advances in neural information processing systems, 31

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.846985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:83f1fb471570f1ab190c189cf29afcfff84297da1a21f33cf8884719cf2f438e

Observation b42121ae-070b-4859-85e6-37be898e1cdc · outbound

This paper cites Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:50:57.163269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:dcdc707ec3ea399d7cf1efdc613e39c271ca654cfa6ea3e3a0d8ddf04af8996b

Observation ed1504d7-e858-4fff-a53e-2466ad2dd7e9 · outbound

This paper cites Deep reinforcement learning with double q-learning.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Deep reinforcement learning with double q-learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.844681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:3258dd3e7b5ca883637921b26747675a59a011967d5b0afd4fc3a81ae8713f5c

Observation 9e833408-06e9-418f-96ec-5a7613cf462c · outbound

This paper cites Value-differencebasedexploration: Adaptivecontrolbetween 𝜖-greedy and softmax.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Value-differencebasedexploration: Adaptivecontrolbetween 𝜖-greedy and softmax

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.892340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:3f86a2c539319e8238218c5e84e18f0c031401a903eb8dd7850b887124877f9e

Observation 8af9cfa4-0eaf-474c-9f98-5f54ea7fce24 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Dueling network architectures for deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.850964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:57456637d95f8dca3b29ad8f140ade96fa4b13b93c1e52211a85089b92f464a5

Observation 504677f2-d2b5-4080-b107-5f49defaca19 · outbound

This paper cites Finite-time error bounds for linear stochastic approximation and TD-learning.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Finite-time error bounds for linear stochastic approximation and TD-learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.772706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:4692de6b9ad04a026bb7235bce3bf46743183b5d6a538d32d69a3ee098b231dd

Observation 45ec2924-76f6-4489-92ff-f4bb5e0e0c9c · outbound

This paper cites A finite-time analysis of temporal difference learning with linear function approximation.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A finite-time analysis of temporal difference learning with linear function approximation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.800197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:94cf07094e5976db87a245d8665dfa858c644ae4297bddc821f49dc9f33b6fdf

Observation 591423c9-fbd6-41c6-80e2-eb427e05eac3 · outbound

This paper cites Finite-sample analysis for SARSA with linear function approximation.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Finite-sample analysis for SARSA with linear function approximation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.889331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:933352b113e315e3b3681c5abcef65ec9d0e68d86c4dc0c3941e8c00aa9ade2a

Observation 115ce0fd-40e2-49ad-b329-ac3e2c9f221d · outbound

This paper cites Stochastic approximation with unbounded Markovian noise: A general-purpose theorem.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Stochastic approximation with unbounded Markovian noise: A general-purpose theorem

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.874646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:e902ab79e0e6259d6434cb4f411416cac16899951d85c0d6ac8510f9eb1a0f6b

Observation 2dc4bd93-7a98-4e6a-95e1-214a03ed1f59 · outbound

This paper cites Concentration of contractive stochastic approximation and reinforcement learning.Stochastic Systems, 12(4):411–430.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Concentration of contractive stochastic approximation and reinforcement learning.Stochastic Systems, 12(4):411–430

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.839260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:d171d448c6650f5955e3eda8c539622ee1b5abdb6df97791339a9bd6a44fb903

Observation c22e91bc-90c0-4dd5-84e3-bcbde789ed42 · outbound

This paper cites Convergence of stochastic iterative dynamic programming algorithms.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Convergence of stochastic iterative dynamic programming algorithms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.903822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:9d94e5ddba8f9385b51ee4161d6453102050ab1779a82dfa3d8e2663ea9a272f

Observation 549421a8-182b-4158-a414-b8e9113d13e8 · outbound

This paper cites A unified switching system perspective and convergence analysis of Q-learning algorithms.Advances in Neural Information Processing Systems, 33:15556–15567.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A unified switching system perspective and convergence analysis of Q-learning algorithms.Advances in Neural Information Processing Systems, 33:15556–15567

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.877311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:e9d56319b674fdc4ddf219fb6f9a8d8286e753ee8761458fb0528f0177cc3ef6

Observation 2291be45-5c55-4511-92aa-63c321465bd3 · outbound

This paper cites Zapq-learning.AdvancesinNeuralInformationProcessingSystems, 30.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Zapq-learning.AdvancesinNeuralInformationProcessingSystems, 30

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.828854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:9a395167114c7d68fa307146f91fa290506cae0c3af674a78e5f97c9f28f2cf0

Observation 748b2d33-e46f-4732-b6ff-a5c2bfed9d16 · outbound

This paper cites Instance-optimality in optimal valueestimation: Adaptivityviavariance-reducedq-learning.IEEETransactionsonInformationTheory.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Instance-optimality in optimal valueestimation: Adaptivityviavariance-reducedq-learning.IEEETransactionsonInformationTheory

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.786485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:d74bdae334b14eddf1695188b114cef3051ef512e81eb4a3a5591e31435e2a79

Observation 97155cf2-5cfc-4627-92be-fd8d4a93244e · outbound

This paper cites Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:50:57.154573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:748d62c0b037154ad03a8df4d7b0f91cf9b2c463beaecbca89051bec24a9578f

Observation c6cf3182-ab9f-48a8-9962-df2c645d9fdd · outbound

This paper cites An analysis of reinforcement learning with function approximation.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies An analysis of reinforcement learning with function approximation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.793365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:0f48d94299bf26ac58f74af78c04bf938851a4137b3e07a531f6c6b1e6d60cde

Observation a66069be-9dd5-4796-b973-04b0ab453e23 · outbound

This paper cites Target network and truncation overcome the deadly triad in q-learning.SIAM Journal on Mathematics of Data Science, 5(4):1078–1101.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Target network and truncation overcome the deadly triad in q-learning.SIAM Journal on Mathematics of Data Science, 5(4):1078–1101

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.817177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:466306a7268f32fa274e94c2bed520526df69594f72ea67a9153bcfbdf800026

Observation 9d91f3b9-ec2b-4d8f-8d14-07cffe222769 · outbound

This paper cites TheprojectedBellmanequationinreinforcementlearning.IEEETransactionsonAutomatic Control.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies TheprojectedBellmanequationinreinforcementlearning.IEEETransactionsonAutomatic Control

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.788459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:e96dc98f3ab757d4f48b8fbd95085c0af6ee80c2bf04409ac8aeb023ac0a6abd

Observation 4ab33dd2-86ac-4cf5-a0e3-9665646bf8d8 · outbound

This paper cites The blessing of heterogeneity in federated q-learning: Linear speedup and beyond.Journal of Machine Learning Research, 26(26):1–85.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies The blessing of heterogeneity in federated q-learning: Linear speedup and beyond.Journal of Machine Learning Research, 26(26):1–85

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.833760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:e67d09557e3729faa9de6206b85e82b7c7a37bdbf5e5a3b7231992c9e167b51a

Observation a22bbfe9-ea11-4a87-81d7-eb28731b3e5a · outbound

This paper cites Federated reinforcement learning: Linearspeedupundermarkoviansampling.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Federated reinforcement learning: Linearspeedupundermarkoviansampling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.775212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:7d8759566c60222f643c1de42a5bfdd0f2fa4ca1aaed352859aed249cea2743b

Observation 402a0e6a-93fe-4ae9-b620-c6c70f702e1b · outbound

This paper cites Q-learning with logarithmic regret.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Q-learning with logarithmic regret

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.765694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:bcdc44909c6a3dfa96f0fb4156d7ac978a8cb1bba8dc363ab2f07aed74aa18bb

Observation 21cb9915-568d-449f-9b24-36abc0302b99 · outbound

This paper cites Online Q-learning using connectionist systems.University of Cambridge, Department of Engineering, Cambridge, UK, 37.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Online Q-learning using connectionist systems.University of Cambridge, Department of Engineering, Cambridge, UK, 37

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.791161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:f7f283aed73910e746696f0ead26344d7bfcd72e4cb7521f3a88cb60e7f4ddf8

Observation 3743101c-2284-4072-9633-6c278603e637 · outbound

This paper cites Convergence results for single-step on-policy reinforcement-learning algorithms.Machine learning, 38:287–308.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Convergence results for single-step on-policy reinforcement-learning algorithms.Machine learning, 38:287–308

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.757679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:1449bff2ac05ebeb3fad61b40bbbbd5d325d51ad5b082641dd2e44b589438e16

Observation f1313c6c-bed2-4cc5-92e7-f24c505ec7ff · outbound

This paper cites On the convergence of sarsa with linear function approximation.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies On the convergence of sarsa with linear function approximation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.826434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:cfcb59019e1469f7cdb0817799d018bbaffa1eaad84e300f97ee1998f952c177

Observation aaa09663-c42c-4b61-8e0b-319dedf49445 · outbound

This paper cites A finite-time analysis of two time-scale actor-critic methods.Advances in Neural Information Processing Systems, 33:17617–17628.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A finite-time analysis of two time-scale actor-critic methods.Advances in Neural Information Processing Systems, 33:17617–17628

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.901466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:bc11e03ac23fe81031bf033cc1f15fd3c6df66e73e88755195b21cb966d1b5b5

Observation 93e23c6f-f8fb-4c1a-8ac7-6f2aa4ea2a2e · outbound

This paper cites Finite sample analysis of two-time-scale natural actor-critic algorithm.IEEE Transactions on Automatic Control.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Finite sample analysis of two-time-scale natural actor-critic algorithm.IEEE Transactions on Automatic Control

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.824056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:d4148038fe080e3d99cb59dca3e4b83c28ed2a38c4c190b05f39980ba0961892

Observation acb7a39c-9226-4f1e-b96a-81adc984d3bd · outbound

This paper cites A finite-sample analysis of payoff-based independent learning in zero-sum stochastic games.Advances in Neural Information Processing Systems, 36:75826–75883.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A finite-sample analysis of payoff-based independent learning in zero-sum stochastic games.Advances in Neural Information Processing Systems, 36:75826–75883

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.760039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:f155ab4a4851c3684971cd7a55cf87799d8b13057a2fbf2b6c5481670e396a97

Observation 2ad4678d-86a4-4efa-8b2e-a74eb0860be8 · outbound

This paper cites Two-timescale Q-learning with function approximation in zero-sum stochastic games.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Two-timescale Q-learning with function approximation in zero-sum stochastic games

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:50:56.393420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:1b14a1057957f1e93eeed92683a6f45cb13dbd0bba1b73436ab4fc0c9d099012

Observation 43e28953-6f72-4311-83c6-2652b3797038 · outbound

This paper cites John Wiley & Sons.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies John Wiley & Sons

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.913503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:4e02a4e6804fc9f8ed4901a27fc60c28cfe2e30bb1268e678ca05d28f35eae96

Observation 17531437-0612-4f18-8223-8197579f5c6e · outbound

This paper cites Athena Scientific.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Athena Scientific

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.819386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:955571d090e3031e83a2aef102a4463e0704fbdf90acddc789112d738caa69d4

Observation f045ed24-6843-4e61-a346-690ec6b8da2a · outbound

This paper cites Surlesopérationsdanslesensemblesabstraitsetleurapplicationauxéquationsintégrales.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Surlesopérationsdanslesensemblesabstraitsetleurapplicationauxéquationsintégrales

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.895407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:1e31cf4d0f2bdd568754203622e7a9d393b3c806ff1bb75b4e5e5c658a486e67

Observation 1b4409d4-49d8-4fd8-b8ba-36b530e55e52 · outbound

This paper cites On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:50:57.158526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:a6ce84696247dd3f245a167f8614390c9eb2230c9332bb2a48cc8ad4b4a4c3c0

Observation d409e13f-b061-4f0c-a581-be87d7af282d · outbound

This paper cites AmericanMathematical Soc.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies AmericanMathematical Soc

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.762568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:66bc7990d37c6e0a078c1599e70797ba84dc3a3b779235b8ab04bb3befe57847

Observation 8728ed8e-5e97-4e17-a6a6-ae862df3e1fb · outbound

This paper cites Sample and communication-efficient decentralized actor-critic algorithms with finite-time analysis.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Sample and communication-efficient decentralized actor-critic algorithms with finite-time analysis

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.886546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:48d0c81ffcc876cde16fb5e2b3662342a1afc7dcca8257207dc9ddbac0d7553c

Observation 5c50e7e5-bc53-44dc-a4e1-0ec69633b0d4 · outbound

This paper cites Sample efficient stochastic policy extra-gradient algorithm for zero-sum markov game.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Sample efficient stochastic policy extra-gradient algorithm for zero-sum markov game

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.836205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:c1a2e98636047bedf18b54c590dc059227fc3448e608fa37fd721cff4751730c

Observation cb5597b9-418c-47d6-9b89-e05227925f81 · outbound

This paper cites Sample complexity bounds for two timescale value-based reinforcement learningalgorithms.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Sample complexity bounds for two timescale value-based reinforcement learningalgorithms

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.797716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:493dfec5a495a56ae6b594aae427059daf7ab20608ffaf36661247e126d53564

Observation e2347e59-b47f-4bc9-a945-94f7cbc3e8b9 · outbound

This paper cites On finite-time convergence of actor-critic algorithm.IEEE Journal on Selected Areas in Information Theory, 2(2):652–664.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies On finite-time convergence of actor-critic algorithm.IEEE Journal on Selected Areas in Information Theory, 2(2):652–664

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.814713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:a3e81df3c73f9a2ed83db26b2859426b681cdb278a74acec078b7d052aa91b17

Observation 46293fee-e626-4f3f-9100-5ebd516e70e4 · outbound

This paper cites Cambridge University Press.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Cambridge University Press

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.842054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:211da85bacd668ef9290d7c802f8b7c984e4a51c5f5bdd4f04e5df17d06635ee

Observation f07419b0-f028-41e5-be68-0f962d2e291a · outbound

This paper cites Sharper model-free reinforcement learning for average-reward markov decision processes.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Sharper model-free reinforcement learning for average-reward markov decision processes

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.768576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:44b421f0037ac64c6e06fd1369850c6bb565785b2c9790fd2b053c3d925b1413

Observation 287162d5-da44-4a40-a323-5ea354c2632d · outbound

This paper cites Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:50:57.135525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:c6758c72fd733927bc8de4d3631cf418f469e2fe4d93edecc5244a57968bfb37

Observation d622e720-2f6d-4c5e-b93e-a71ca2230213 · outbound

This paper cites an unresolved cited work.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-18T05:50:57.795648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:016368c5203e643988254a07047bc80bf9595b366d922ce7879c632675ba0c51

Observation e369d986-826f-4399-9bc9-07aa4076dc84 · outbound

This paper cites Solution representations for poisson’s equation, martingale structure, and the markov chain central limit theorem.Stochastic Systems, 14(1):47–68.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Solution representations for poisson’s equation, martingale structure, and the markov chain central limit theorem.Stochastic Systems, 14(1):47–68

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.908573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:f1140313fd820f7facf14cb498877ba54e677b56cbc80d56a89232cf6bf1688a

Observation 8b5fb187-8840-4a1e-97c6-c5166a59f520 · outbound

This paper cites Generalized inverses and their application to applied probability problems.Linear Algebra and its Applications, 45:157–198.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Generalized inverses and their application to applied probability problems.Linear Algebra and its Applications, 45:157–198

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.863120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:069410330d603dd2837dd2a175e7718087d23bccc7c7066f2047ae0961e0126b

Observation f4ab5305-30f9-4d4b-8ede-20ecd76d15f9 · outbound

This paper cites A liapounov bound for solutions of the poisson equation.The Annals of Probability, pages 916–931.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies A liapounov bound for solutions of the poisson equation.The Annals of Probability, pages 916–931

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.865774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:67af772b8ee387c8a9d54b2905df02aaa6620a977207ead580fccc2e3e1b6d32

Observation 2c1f914b-abe2-4ad2-84ac-2b3148387d25 · outbound

This paper cites An approximate policy iteration viewpoint of actor–critic algorithms.Automatica, 179:112395.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies An approximate policy iteration viewpoint of actor–critic algorithms.Automatica, 179:112395

Reference 68

Resolution
malformed identifier
arxiv_id, observed 2026-05-18T05:50:57.139934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:e598e1d53cb303b1b398fe8131aef36acc71e3d57f31803ba99ebfe439b51901

Observation 08bc700b-47cd-4237-a278-1707a8269e22 · outbound

This paper cites Boundedness of iterates in Q-learning.Systems & control letters, 55(4):347–349.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Boundedness of iterates in Q-learning.Systems & control letters, 55(4):347–349

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.807325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:c8018d4571e6a5161d3c197d7daf90544a950fcb1e3e036434ac0c90de23b84f

Observation f571dc90-2911-4582-a792-e271205f56f8 · outbound

This paper cites Minimax pac bounds on the sample complexity of reinforcement learning with a generative model.Machine learning, 91(3):325–349.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Minimax pac bounds on the sample complexity of reinforcement learning with a generative model.Machine learning, 91(3):325–349

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.802443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:9b6d3290aba8b71555c43b0f845c7ec58bac2bdf9c9cce73379069387f060d79

Observation f9b421a6-7f4d-4dcd-94a5-6b8e8acaace5 · outbound

This paper cites Online learning and online convex optimization.Foundations and Trends®in Machine Learning, 4(2):107–194.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Online learning and online convex optimization.Foundations and Trends®in Machine Learning, 4(2):107–194

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.804965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:30ede6525e844015ce3041b59241d23e4df9658d711b40073bc9551abace837e

Observation e12f59b6-387d-4bc2-a5a8-e52a4656242c · outbound

This paper cites Beck.First-Order Methods in Optimization.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Beck.First-Order Methods in Optimization

Reference 72

Resolution
metadata mismatch
doi, observed 2026-05-18T05:50:56.397550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:895fd3c553106786a142e576092bfcec7554078ed168c186084ef4814b71b9dd

Observation f747d203-dc94-4722-9b94-a8d9a373b005 · outbound

This paper cites an unresolved cited work.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-18T05:50:57.905947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:e3253403cda585928f3cdefca9864395b511fe4f556215425e8301d81aaef6a0

Observation 74c0dc09-8578-4e6a-b435-cacb90495b4d · outbound

This paper cites To this end, define𝑧𝑘 := Í𝑘 𝑛=0 P 𝑛𝑦 for any𝑘≥0.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies To this end, define𝑧𝑘 := Í𝑘 𝑛=0 P 𝑛𝑦 for any𝑘≥0

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.918307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:88b0313fdec17fff48a32998253243e4a7ce91d5887d6b7c6f069ad539854cd3

Observation 8b12d8ff-faa6-40ba-ac89-6a0cb3114080 · outbound

This paper cites 𝑟𝑏∑︁ 𝑖=0 𝑟𝑏 +1 𝑖+1 𝑃𝑖 𝜋𝑏 (𝑠 ′′, 𝑠′) # 𝜋𝑏 (𝑎 ′|𝑠 ′) (Change of variable:𝑖=𝑗−1) = 1 2𝑟𝑏+1 ∑︁ 𝑠′′ ∈ S 𝑝(𝑠 ′′ |𝑠, 𝑎).

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies 𝑟𝑏∑︁ 𝑖=0 𝑟𝑏 +1 𝑖+1 𝑃𝑖 𝜋𝑏 (𝑠 ′′, 𝑠′) # 𝜋𝑏 (𝑎 ′|𝑠 ′) (Change of variable:𝑖=𝑗−1) = 1 2𝑟𝑏+1 ∑︁ 𝑠′′ ∈ S 𝑝(𝑠 ′′ |𝑠, 𝑎)

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:50:57.868662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:93c6a1e64761a993d026582d15da21a6e6f3b97e32d8215a46423d26c07e043d

Pith citing papers

Observation 3b08014c-0621-417b-acd6-e070e5b31058 · inbound

Auto-exploration for online reinforcement learning cites this paper.

Auto-exploration for online reinforcement learning A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T18:18:59.262679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:18:59.262679Z digest=sha256:f6669526b52cab7ee31146ca1113354fbafbbf7d3aa5b97081b0fda31b4e75aa

Observation cb299ed5-5f90-49d5-84a4-6ed886b883ac · inbound

Auto-exploration for online reinforcement learning cites this paper.

Auto-exploration for online reinforcement learning A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T06:51:04.104314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:51:04.104314Z digest=sha256:f0f9aee1b479655daa822dd8a50c66693ef8f4cc8125db1fe3aa6828bfbfa799

Observation 1628a8a3-f589-4358-93ca-05282a1c6e37 · inbound

Achieving $\epsilon^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions cites this paper.

Achieving $\epsilon^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:29:23.755661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T19:28:32.795407Z digest=sha256:a678ce477040241b7c3492636d6389a4e7794d971bfd0644bcdb427ef6dac84c

Observation e15c10b6-921f-4eb5-aa07-40dbad77416e · inbound

Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework cites this paper.

Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:12:46.706182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T23:11:02.699220Z digest=sha256:bca2186fbe867a0a6093851f2ecd3292c72875be7fc6fa316fe9e1f2cfc2365a