Pith. sign in

Paper Citation Record · LEDGER

Variance-reduced $Q$-learning is minimax optimal

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:1906.04697.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1906.04697 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:48:32.814848Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:56:37.735104Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a7e16fd6-1b54-4c91-a3da-ba2c5e5aec17 · inbound

Near-Optimal Sample Complexity for MDPs via Anchoring cites this paper.

Near-Optimal Sample Complexity for MDPs via Anchoring Variance-reduced $Q$-learning is minimax optimal

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T22:48:32.814848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:48:32.814848Z digest=sha256:99471f5b84bb2bd1a7922b7115aedc1596bdd3fdabf14641ee62d336aee51431

Observation db9b0175-46a3-4991-819e-a16e72750e78 · inbound

From Set Convergence to Pointwise Convergence: Finite-Time Guarantees for Average-Reward Q-Learning with Adaptive Stepsizes cites this paper.

From Set Convergence to Pointwise Convergence: Finite-Time Guarantees for Average-Reward Q-Learning with Adaptive Stepsizes Variance-reduced $Q$-learning is minimax optimal

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:35:00.870887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T17:34:49.191496Z digest=sha256:ea101ef2ea270c6198b3d31c0c94c6c7736f31c73a3c421c8f60b9cb76f2929d

Observation 4a183801-060a-4eee-905a-2a31da41e93c · inbound

Statistical and Algorithmic Foundations of Reinforcement Learning cites this paper.

Statistical and Algorithmic Foundations of Reinforcement Learning Variance-reduced $Q$-learning is minimax optimal

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-06T16:15:15.902771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:15:15.902771Z digest=sha256:25baa2587871f40b0840684df3322ef88a85e9fab3c717fa3f58f2cdd85325b1

Observation e442f492-2650-42c9-ba2e-7ec42922dc85 · inbound

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies cites this paper.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Variance-reduced $Q$-learning is minimax optimal

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:50:57.149874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:781131603f889c2c70993251cbbe66a98b016a99858072f44d1a05c98ba3ec86

Observation d976f237-ad61-48dd-99a0-9f29ce3c1dce · inbound

Gaussian Approximation for Asynchronous Q-learning cites this paper.

Gaussian Approximation for Asynchronous Q-learning Variance-reduced $Q$-learning is minimax optimal

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:25:59.876792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:11:48.816106Z digest=sha256:5681010fd9b4760befbc04382fca9c211fa116efd9c918f5ec0c8ad4fea9753b

Observation 55c8a0bd-c112-4d6b-882e-435756089d32 · inbound

Achieving $\epsilon^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions cites this paper.

Achieving $\epsilon^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions Variance-reduced $Q$-learning is minimax optimal

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:29:23.748333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T19:28:32.795407Z digest=sha256:069e9447554932586cecc5301283ab8784483ec992ab2ed5880f89737332fb89

Observation 0ccccb12-8fc1-4a95-8d3b-35cbb733e050 · inbound

On Gaussian approximation for entropy-regularized Q-learning with function approximation cites this paper.

On Gaussian approximation for entropy-regularized Q-learning with function approximation Variance-reduced $Q$-learning is minimax optimal

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:50.614532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T22:11:31.021067Z digest=sha256:c50b53578530d1c88f4ce8124500825413fde2b2980a1c2a6fded82209ae0b10

Observation 10b65a33-d25c-4af0-8a5f-a075caa5b6f8 · inbound

Randomization for Faster Exact Optimization of Discounted Markov Decision Processes cites this paper.

Randomization for Faster Exact Optimization of Discounted Markov Decision Processes Variance-reduced $Q$-learning is minimax optimal

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:26:55.007464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T03:30:29.891537Z digest=sha256:e9b1a7443ffb3b13da1b70aae1ad44ffc465a216133ba2a650b1bb9167932491

Observation aa90a590-bddc-475e-9c15-7f8b4377e926 · inbound

Stationary Robust Mean-Field Games under Model Mismatches cites this paper.

Stationary Robust Mean-Field Games under Model Mismatches Variance-reduced $Q$-learning is minimax optimal

Reference 280

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.736437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T10:50:40.841967Z digest=sha256:0adfbe75d518e31ef48385fdc35e35a4a155d5e08e91db0844781335c11cbb30

Observation ed66b643-142d-45c6-8141-7f23ca46de9a · inbound

Minimax PAC Bounds for Learning in Exogenous Contextual MDPs cites this paper.

Minimax PAC Bounds for Learning in Exogenous Contextual MDPs Variance-reduced $Q$-learning is minimax optimal

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:10:05.254216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T21:39:39.968655Z digest=sha256:83d84c8ccb83e4a47e9a12f7a2a3c25aa8a2fd14c8b21ad8cc1d1665c07bad9d

Observation 9f6c1b55-1e44-4b51-b1d4-082933adde0a · inbound

Mathematical methods of reinforcement learning cites this paper.

Mathematical methods of reinforcement learning Variance-reduced $Q$-learning is minimax optimal

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.736273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:c571df2fea1dee88dafd3081105b6b070fb2ac0406e55b27a78d6838686ce0e3