Pith. sign in

Paper Citation Record · LEDGER

Variance-reduced $Q$-learning is minimax optimal

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:1906.04697.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1906.04697 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:48:32.814848Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:56:37.735104Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a7e16fd6-1b54-4c91-a3da-ba2c5e5aec17 · inbound

Near-Optimal Sample Complexity for MDPs via Anchoring cites this paper.

Near-Optimal Sample Complexity for MDPs via Anchoring Variance-reduced $Q$-learning is minimax optimal

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T22:48:32.814848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:48:32.814848Z digest=sha256:99471f5b84bb2bd1a7922b7115aedc1596bdd3fdabf14641ee62d336aee51431

Observation db9b0175-46a3-4991-819e-a16e72750e78 · inbound

From Set Convergence to Pointwise Convergence: Finite-Time Guarantees for Average-Reward Q-Learning with Adaptive Stepsizes cites this paper.

From Set Convergence to Pointwise Convergence: Finite-Time Guarantees for Average-Reward Q-Learning with Adaptive Stepsizes Variance-reduced $Q$-learning is minimax optimal

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:35:00.870887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T17:34:49.191496Z digest=sha256:798bf51a6f2b70fd5d68308bd94845a4be4d06aee69e85adf4eb5bf85bb24a62

Observation 4a183801-060a-4eee-905a-2a31da41e93c · inbound

Statistical and Algorithmic Foundations of Reinforcement Learning cites this paper.

Statistical and Algorithmic Foundations of Reinforcement Learning Variance-reduced $Q$-learning is minimax optimal

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-06T16:15:15.902771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:15:15.902771Z digest=sha256:25baa2587871f40b0840684df3322ef88a85e9fab3c717fa3f58f2cdd85325b1

Observation e442f492-2650-42c9-ba2e-7ec42922dc85 · inbound

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies cites this paper.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Variance-reduced $Q$-learning is minimax optimal

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:50:57.149874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:a43aaaa93205aef1ed3b96baf3a77b7f811af6ce6ebbfa20ce94b3f44983b01a

Observation d976f237-ad61-48dd-99a0-9f29ce3c1dce · inbound

Gaussian Approximation for Asynchronous Q-learning cites this paper.

Gaussian Approximation for Asynchronous Q-learning Variance-reduced $Q$-learning is minimax optimal

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:25:59.876792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:11:48.816106Z digest=sha256:25d434f86bdbc62fae2a154d74fccbd0f61c67923e21b87364de73feb965bcd1

Observation 55c8a0bd-c112-4d6b-882e-435756089d32 · inbound

Achieving $\epsilon^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions cites this paper.

Achieving $\epsilon^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions Variance-reduced $Q$-learning is minimax optimal

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:29:23.748333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T19:28:32.795407Z digest=sha256:1da5e55197f8dc810229d751ce6c3c11965db1f26524bc4a9ead6be66e22e2eb

Observation 0ccccb12-8fc1-4a95-8d3b-35cbb733e050 · inbound

On Gaussian approximation for entropy-regularized Q-learning with function approximation cites this paper.

On Gaussian approximation for entropy-regularized Q-learning with function approximation Variance-reduced $Q$-learning is minimax optimal

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:50.614532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T22:11:31.021067Z digest=sha256:17c8675ba742b63ea1c6676d42967b78fdab25b334076d5e4dbded439614f985

Observation 10b65a33-d25c-4af0-8a5f-a075caa5b6f8 · inbound

Randomization for Faster Exact Optimization of Discounted Markov Decision Processes cites this paper.

Randomization for Faster Exact Optimization of Discounted Markov Decision Processes Variance-reduced $Q$-learning is minimax optimal

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:26:55.007464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T03:30:29.891537Z digest=sha256:0f626a0709b1af470d1ea00acec7327c31e03aa6c6ae394a4891dd382f2e4517

Observation aa90a590-bddc-475e-9c15-7f8b4377e926 · inbound

Stationary Robust Mean-Field Games under Model Mismatches cites this paper.

Stationary Robust Mean-Field Games under Model Mismatches Variance-reduced $Q$-learning is minimax optimal

Reference 280

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.736437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T10:50:40.841967Z digest=sha256:3ba8df2c1b4e5849fe8f35d559e16f37c999b21f61e35e73f5f562d1fe4aca8f

Observation ed66b643-142d-45c6-8141-7f23ca46de9a · inbound

Minimax PAC Bounds for Learning in Exogenous Contextual MDPs cites this paper.

Minimax PAC Bounds for Learning in Exogenous Contextual MDPs Variance-reduced $Q$-learning is minimax optimal

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:10:05.254216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-25T21:39:39.968655Z digest=sha256:6b3b5e7b2bc989896a2b2dc34822c045617aa1d34ad71343f1bdce28ece03d99

Observation 9f6c1b55-1e44-4b51-b1d4-082933adde0a · inbound

Mathematical methods of reinforcement learning cites this paper.

Mathematical methods of reinforcement learning Variance-reduced $Q$-learning is minimax optimal

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:56:37.736273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-09T22:47:51.676289Z digest=sha256:2f17b09af5c7c3bd9199c5f691dbe74080db6b9d01399f2995acf301fd33bf5e