Pith. sign in

Paper Citation Record · LEDGER

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values

As of 13 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.09523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09523 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:01:47.191696Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82aaab77-0338-46e0-bb06-92e787f30d06 · outbound

This paper cites write newline.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:44.735602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:44.735602Z digest=sha256:737ac7e5811b8deebb4a414466ff69cefca63daf5c11fed59057c24ab9159495

Observation a2eb5b45-b5dc-4f01-b32a-fa21b259404f · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:51.426569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:44.855222Z digest=sha256:627f39ecf184e8a2a8b641d2842c9d6a4668a5b1d79104e55fe5d360d96595ce

Observation 49aac7d9-f0e4-4fd7-be52-3ff2dcc6abd5 · outbound

This paper cites Sur les op \'e rations dans les ensembles abstraits et leur application aux \'e quations int \'e grales.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Sur les op \'e rations dans les ensembles abstraits et leur application aux \'e quations int \'e grales

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:51.113175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:44.976043Z digest=sha256:603574448e3a838236a1c5843f70b6a08f61e2d19b0b256d05864c17c541e15b

Observation 88c31345-0586-4a52-b1f5-a7038a19a76d · outbound

This paper cites Dynamic Programming.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Dynamic Programming

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:45.055174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:45.055174Z digest=sha256:969a0c9dd7cafbd0ae7ccf50595a650a6193e6fe8cf8f1b8f2f695ac8bc3d99f

Observation 7450fe15-d8bf-428c-9bda-665639b9468a · outbound

This paper cites Bertsekas and John N.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Bertsekas and John N

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.904931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.192835Z digest=sha256:56ceef726a192f4c16d937e02a9d931ad072e96b6d380d9170587b21a36efad2

Observation fbe8b06b-4a0d-4413-a3e5-73037c103b4f · outbound

This paper cites Multistep Credit Assignment in Deep Reinforcement Learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Multistep Credit Assignment in Deep Reinforcement Learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.775974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.304714Z digest=sha256:66a925c11438e17bd1f80bb70a9e80b92e53af355f576223ced3ea300ec3a00d

Observation 49119f6f-aff3-4fbb-8560-2f69162cbcff · outbound

This paper cites ChainerRL : A deep reinforcement learning library.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values ChainerRL : A deep reinforcement learning library

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.632827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.368932Z digest=sha256:5438186a9cadc90ef6c1c3d4287bceadfb2528a1b678cba0aab2f4cc9502df15

Observation 7dfa3418-b4fe-4d31-83b9-87ac0e446569 · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:50.491388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.486551Z digest=sha256:e45b7d10cd96fb181e43199d78723476a278c50bd113a810b46f480488426c9d

Observation d742ebfd-07fc-4337-90c9-25fd85e9f059 · outbound

This paper cites Kingma and Jimmy Ba.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Kingma and Jimmy Ba

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.360877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.604642Z digest=sha256:d756c2403f155f73ad5c1f3bf95742ae7faff082eda130fbd32dee08f87150fc

Observation 58a7a71b-afa8-4ac7-82b8-9c0fbb5dfb7a · outbound

This paper cites Orr, and Klaus-Robert M \"u ller.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Orr, and Klaus-Robert M \"u ller

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.213833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.697423Z digest=sha256:10bc30d309c9837595c29d38d9bc9adfbdcec9ec13860538385d943f34b49d63

Observation 9a3e49df-698b-4a37-8b30-74688ca8b8b5 · outbound

This paper cites Rusu, Joel Veness, Marc G.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Rusu, Joel Veness, Marc G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:50.064504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.818776Z digest=sha256:556f513e683e3ed0ceddd4732a14898330ca84d16f72cefa6caac68f28f2ed38

Observation c9259be3-76b0-4b10-a6a0-c393a0c51e4d · outbound

This paper cites Towards model-free RL algorithms that scale well with unstructured data.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Towards model-free RL algorithms that scale well with unstructured data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:01:47.519541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:45.917099Z digest=sha256:6d0da96621b8b9dcc971c1552663e4f9343199584958da7b06033c6f7e6ab380

Observation 03ae75ac-1be6-466a-bae3-d767da002530 · outbound

This paper cites Revisiting Rainbow : Promoting more insightful and inclusive deep reinforcement learning research.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Revisiting Rainbow : Promoting more insightful and inclusive deep reinforcement learning research

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.937782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.020479Z digest=sha256:b7fa265ab86e35f8bd5ef97df92e9bb0c43a3025e15b412e042d425dd389039d

Observation b6d1f872-1467-4ef7-83ea-07d58b4a3290 · outbound

This paper cites Rummery and Mahesan Niranjan.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Rummery and Mahesan Niranjan

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.834689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.115279Z digest=sha256:f33fe9d6b086602c969d4b01a94932747583d0e792eb482d465efb8d589c2b79

Observation 49ad9446-5ee9-4e3d-b989-db0a0616c173 · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:49.703269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.207382Z digest=sha256:0be8136a70958d86f4d44f794dd9143f41781f7f4f29ef55d466905bbfd801ab

Observation 8cff4409-b4e7-4c88-bd2e-fb44c043a8b3 · outbound

This paper cites Littman, and Csaba Szepesv \'a ri.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Littman, and Csaba Szepesv \'a ri

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.575152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.290297Z digest=sha256:c4714369d7ddd618a1ccc2db1d786ff4903bae4be1058ba817df1b4d6c783ea1

Observation 58931b2f-9137-4d8e-a2c7-a1c07d8fb3ce · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:49.427621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.377144Z digest=sha256:f993ceea46fb68631e052cd7d2edf82f0d05c08aefcbdb38b1b5664203ae2306

Observation 90ab8c3d-19a9-4e2f-8f9d-8f470a925902 · outbound

This paper cites Sutton and Andrew G.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Sutton and Andrew G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.289687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.468196Z digest=sha256:0542f40a948e97ee9bdbcb0a0dd5eb4e88ee35d79706934f79c0d2321b7b8b4d

Observation 380e6ba9-f74a-4281-accb-a82f3045e7c2 · outbound

This paper cites VA-learning as a more efficient alternative to Q-learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values VA-learning as a more efficient alternative to Q-learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:49.135279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.563594Z digest=sha256:1e72b154d36b81f818f4dd7a846c278e5aa8038568edb91717b2eadb80288d47

Observation 41bad30e-a29f-44fd-993e-b5942ea11e2f · outbound

This paper cites Double Q-learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Double Q-learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.994867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.638334Z digest=sha256:d6b7885d6e580dd6b9e42e35b0ccf2c5a7668541241711d25f79f59c83ae720b

Observation 2413f22d-5b75-47ec-b6f8-1940c20fd601 · outbound

This paper cites Insights in Reinforcement Learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Insights in Reinforcement Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.840394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.723290Z digest=sha256:e0d5354616614e57f3919772e12781a2f32849d7dd54905fc499ac396f668052

Observation 344721d1-d006-4cd5-be36-51e75d1f6521 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Dueling network architectures for deep reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.710073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.799484Z digest=sha256:52de9554d7ff095de0616a348a668c90083fb8b9aaab3703df04208a01f0b0e8

Observation 3bc160fe-f2bf-477d-8e61-9713e3fe5b9d · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:48.572660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.815493Z digest=sha256:8e44d74e1578db4ebefab44bde457d03e7e820863b501d74acab933b301d8536

Observation 15cdb5f2-f4e9-43b5-91fa-fc55735cdd2f · outbound

This paper cites an unresolved cited work.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:01:48.406295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.827140Z digest=sha256:387f73d0522ab1897a9980f28230480f6b2376ab40f240177732c2a26a5d5a72

Observation a390f672-f32c-42d0-aa0d-86a298d88ebc · outbound

This paper cites Wiering and Hado van Hasselt.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Wiering and Hado van Hasselt

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:48.101275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:46.894642Z digest=sha256:33e0c859f212fd3874bdb4e7a3db0cc1428ffb9b5c08b179f9d2cad184510e93

Observation d490c7c4-2992-4a36-8d87-4706443e9f22 · outbound

This paper cites Wiering and Hado van Hasselt.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Wiering and Hado van Hasselt

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:01:47.826751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T18:01:47.067674Z digest=sha256:c39ebdce4e5a7a2960f2ac93eb68b9bbd480a4a6dfc2615218e7b4a6197fd0c4

Observation 85b62dba-ea0f-4746-a7cd-3f2f496889b6 · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:01:47.191696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:01:47.191696Z digest=sha256:927597b607ec61bb8578767709bab427037b18d1be3d2256184b1dbee1e2bed3

Pith citing papers

No inbound Pith citation observations are available.