Pith. sign in

Paper Citation Record · LEDGER

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values

As of 12 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.09523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09523 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:01:47.191696Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82aaab77-0338-46e0-bb06-92e787f30d06 · outbound

This paper cites write newline.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:44.735602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:44.735602Z digest=sha256:737ac7e5811b8deebb4a414466ff69cefca63daf5c11fed59057c24ab9159495

Observation a2eb5b45-b5dc-4f01-b32a-fa21b259404f · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:51.426569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:44.855222Z digest=sha256:0ccc10d9a19b95b01eaaed227ea20ef9cc251fb1c0f8d568b3355e37452fb433

Observation 49aac7d9-f0e4-4fd7-be52-3ff2dcc6abd5 · outbound

This paper cites Sur les op \'e rations dans les ensembles abstraits et leur application aux \'e quations int \'e grales.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Sur les op \'e rations dans les ensembles abstraits et leur application aux \'e quations int \'e grales

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:51.113175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:44.976043Z digest=sha256:b74897e6e8a44b8f8b47a53298e6fd513a264ad3d97449911c46d422a2442d11

Observation 88c31345-0586-4a52-b1f5-a7038a19a76d · outbound

This paper cites Dynamic Programming.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Dynamic Programming

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:45.055174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:45.055174Z digest=sha256:969a0c9dd7cafbd0ae7ccf50595a650a6193e6fe8cf8f1b8f2f695ac8bc3d99f

Observation 7450fe15-d8bf-428c-9bda-665639b9468a · outbound

This paper cites Bertsekas and John N.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Bertsekas and John N

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.904931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.192835Z digest=sha256:def4eb150dac02baf353a89d18764365260eac650e78a733cf091651dabe0d8d

Observation fbe8b06b-4a0d-4413-a3e5-73037c103b4f · outbound

This paper cites Multistep Credit Assignment in Deep Reinforcement Learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Multistep Credit Assignment in Deep Reinforcement Learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.775974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.304714Z digest=sha256:34ecce21801a8f9ce50c9f3a4dde497dc78ef1a774591d1f9ef4f6a1cc9a7003

Observation 49119f6f-aff3-4fbb-8560-2f69162cbcff · outbound

This paper cites ChainerRL : A deep reinforcement learning library.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values ChainerRL : A deep reinforcement learning library

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.632827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.368932Z digest=sha256:cbdcdcdb75f4621f00c8560ad9d6524a8d3bed78e00c61e99f9cab52b182f98d

Observation 7dfa3418-b4fe-4d31-83b9-87ac0e446569 · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:50.491388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.486551Z digest=sha256:56998bd9406ff874acb45cc618c209a96b88f49f340c3b1939ca199de4b7b2fa

Observation d742ebfd-07fc-4337-90c9-25fd85e9f059 · outbound

This paper cites Kingma and Jimmy Ba.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Kingma and Jimmy Ba

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.360877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.604642Z digest=sha256:99b630716a74b4207bd941295c83e9b879a90e399edab5bf00d279c8e966ce3f

Observation 58a7a71b-afa8-4ac7-82b8-9c0fbb5dfb7a · outbound

This paper cites Orr, and Klaus-Robert M \"u ller.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Orr, and Klaus-Robert M \"u ller

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.213833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.697423Z digest=sha256:04169d3c8cae147fb1181fa83fd7a5380456e09a5d01656d864e862de982606f

Observation 9a3e49df-698b-4a37-8b30-74688ca8b8b5 · outbound

This paper cites Rusu, Joel Veness, Marc G.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Rusu, Joel Veness, Marc G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.064504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.818776Z digest=sha256:ba68d95f856f8f9adeca698c09bc51315165666dcfd970d70923c04ca2dba278

Observation c9259be3-76b0-4b10-a6a0-c393a0c51e4d · outbound

This paper cites Towards model-free RL algorithms that scale well with unstructured data.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Towards model-free RL algorithms that scale well with unstructured data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:01:47.519541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.917099Z digest=sha256:9f84e34afae8ab9cc60767a2cf866e1a246ced82150589fc1a1f9fd5fe275b84

Observation 03ae75ac-1be6-466a-bae3-d767da002530 · outbound

This paper cites Revisiting Rainbow : Promoting more insightful and inclusive deep reinforcement learning research.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Revisiting Rainbow : Promoting more insightful and inclusive deep reinforcement learning research

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.937782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.020479Z digest=sha256:c69f0744a073dccb470e4e116be21a08270556a883b8aeee046202c06fff2938

Observation b6d1f872-1467-4ef7-83ea-07d58b4a3290 · outbound

This paper cites Rummery and Mahesan Niranjan.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Rummery and Mahesan Niranjan

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.834689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.115279Z digest=sha256:a6d30f3f4843df540e8eab631c61c94fd3d2a2267d02ddd8250e478ea283caa8

Observation 49ad9446-5ee9-4e3d-b989-db0a0616c173 · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:49.703269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.207382Z digest=sha256:4dc35ff54654a1d0ff48ba771be4b8bce75552788e98f7e0e64cb3ba76005ff0

Observation 8cff4409-b4e7-4c88-bd2e-fb44c043a8b3 · outbound

This paper cites Littman, and Csaba Szepesv \'a ri.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Littman, and Csaba Szepesv \'a ri

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.575152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.290297Z digest=sha256:b6de19d779bf3e2a6932c5d2aa4adafe12aaa6f91aea262cbc92795f12a65f4b

Observation 58931b2f-9137-4d8e-a2c7-a1c07d8fb3ce · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:49.427621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.377144Z digest=sha256:02abe4595ef5a7b514af85249dad086ce0c19bebb5b87d5daef42faadd74c94c

Observation 90ab8c3d-19a9-4e2f-8f9d-8f470a925902 · outbound

This paper cites Sutton and Andrew G.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Sutton and Andrew G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.289687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.468196Z digest=sha256:439f6149bd643e07b6bab21111e77f6b768d51bd8302a4209850d2a72c52d08f

Observation 380e6ba9-f74a-4281-accb-a82f3045e7c2 · outbound

This paper cites VA-learning as a more efficient alternative to Q-learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values VA-learning as a more efficient alternative to Q-learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.135279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.563594Z digest=sha256:5073b91f517711626ecfe4bca151d9118c60e8c540dfcea88382af3df174d070

Observation 41bad30e-a29f-44fd-993e-b5942ea11e2f · outbound

This paper cites Double Q-learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Double Q-learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.994867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.638334Z digest=sha256:6dd5446929095eab9d864d560c595f968d91396faf62f22565484246c7ba67cd

Observation 2413f22d-5b75-47ec-b6f8-1940c20fd601 · outbound

This paper cites Insights in Reinforcement Learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Insights in Reinforcement Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.840394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.723290Z digest=sha256:633b5fd5734b2015cee2176cb9ba9a3d11618086ddeeed631ca76d7f15ef5279

Observation 344721d1-d006-4cd5-be36-51e75d1f6521 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Dueling network architectures for deep reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.710073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.799484Z digest=sha256:82761d3f88f3e7ef1079fc91dcba74a3ba7eb42a7a464b51c40f28e3052777ae

Observation 3bc160fe-f2bf-477d-8e61-9713e3fe5b9d · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:48.572660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.815493Z digest=sha256:84d9129079a11054ee2b7313c2f981212233303f96d6c52087a17602e685a829

Observation 15cdb5f2-f4e9-43b5-91fa-fc55735cdd2f · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:48.406295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.827140Z digest=sha256:f1518fd2f128b25648f8002b703cd2cda6f68393f444cc0abed8f4752ba5b1e7

Observation a390f672-f32c-42d0-aa0d-86a298d88ebc · outbound

This paper cites Wiering and Hado van Hasselt.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Wiering and Hado van Hasselt

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.101275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.894642Z digest=sha256:8b054aea33d36b5beea32920bed795ad38d78adf846f90b3bfa68fcd22cac000

Observation d490c7c4-2992-4a36-8d87-4706443e9f22 · outbound

This paper cites Wiering and Hado van Hasselt.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Wiering and Hado van Hasselt

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:47.826751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T18:01:47.067674Z digest=sha256:0ac8a0301babcc5ed1cd15146ffc50b6a7f703b99572727db0065d075a0a44c2

Observation 85b62dba-ea0f-4746-a7cd-3f2f496889b6 · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:47.191696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:47.191696Z digest=sha256:927597b607ec61bb8578767709bab427037b18d1be3d2256184b1dbee1e2bed3

Pith citing papers

No inbound Pith citation observations are available.