Pith. sign in

Paper Citation Record · LEDGER

Simplifying Deep Temporal Difference Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2407.04811.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.04811 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:23:16.184323Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1cbf3a39-31ac-4f64-b5ea-4588a997c971 · inbound

Plasticity Loss in Deep Reinforcement Learning: A Survey cites this paper.

Plasticity Loss in Deep Reinforcement Learning: A Survey Simplifying Deep Temporal Difference Learning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T18:03:18.157374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T18:02:30.199552Z digest=sha256:e654e5bbd79f7dd239281a3d59f9443fa9b6c72d710b26ba7e210cbd3ece3efb

Observation f23f055e-cb93-49fb-bd93-46373db78d0a · inbound

Hadamax Encoding: Elevating Performance in Model-Free Atari cites this paper.

Hadamax Encoding: Elevating Performance in Model-Free Atari Simplifying Deep Temporal Difference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:16.184323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:16.184323Z digest=sha256:9f4b40ee3e391ed3c5e77232774ebaad65f66498a12410957094f48034dde3cd

Observation f9805582-53ad-478a-984d-dfa78092c0f0 · inbound

Universal Value-Function Uncertainties cites this paper.

Universal Value-Function Uncertainties Simplifying Deep Temporal Difference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:55.363030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:49:55.363030Z digest=sha256:8dc671ade13b5aff05e503610d227d3351ffb4f9fff0bbbd6a12e52536b7cc2a

Observation 3349cc25-8617-4799-b1a8-445fb5b43491 · inbound

FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control cites this paper.

FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control Simplifying Deep Temporal Difference Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:33.176135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:33.176135Z digest=sha256:15aa761d57ad139234dd79218722407be689b63ea26a63792cb5d8682f8043e7

Observation f5e31804-8085-4696-9cd2-b998016c275e · inbound

The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks cites this paper.

The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks Simplifying Deep Temporal Difference Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:10:52.373117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:10:52.373117Z digest=sha256:d916c4edef1c2888ab90148b5fdfcf5ca287b0d7d859a0257f292499d52cbe76

Observation 625ffa4f-d64e-42a5-9af4-316d21545853 · inbound

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models cites this paper.

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models Simplifying Deep Temporal Difference Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:32.888434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:32.888434Z digest=sha256:56ebad87ab02fc0b4c8b444a94b8bce3281f2c72094e09a82b51652a8998ff50

Observation ab8b6e32-c3dd-4343-aa03-25fb4bda2d62 · inbound

Benchmarking Massively Parallelized Multi-Task Reinforcement Learning for Robotics Tasks cites this paper.

Benchmarking Massively Parallelized Multi-Task Reinforcement Learning for Robotics Tasks Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:03:23.236918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:03:23.236918Z digest=sha256:74ee91bf65a6e8477551b31ab32d70243888f1209b19bec33886ede462cd3ab7

Observation 0a5445b3-83e8-45f8-bdab-b98429e1a043 · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Simplifying Deep Temporal Difference Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:02.239780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:02.239780Z digest=sha256:60696ecff3949935e7724b3aa10b947f90be82e1348544fb3c66b8bd5b8cbd88

Observation 4c61f33e-1cc3-470c-b869-fd968e9c35dc · inbound

Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning cites this paper.

Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:57.233093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:57.233093Z digest=sha256:9a04235f60ee2b105137164ae743ee8a9b58c5ae45495c8b132f27cf14e1ec16

Observation 1877c9b2-7022-41d7-9ccd-9958b2bb4981 · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.774894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T07:41:48.634351Z digest=sha256:9ebfc720df8c75df82a1472b96c2bbe5260dc0c4d6f28e5e0646131b224d18e8

Observation 18f36599-0f5f-4134-bf72-988583957402 · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T05:37:38.429736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:37:38.429736Z digest=sha256:64929c3e43430872440810244a52eb94e7f4c1bf84a7270f0bba6ae826d5ecd1

Observation 700bd6c9-d87a-4d83-b166-9073f9a2d975 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Simplifying Deep Temporal Difference Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:15:49.847929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:04:56.512544Z digest=sha256:f869f0c31071956d1cad09af6475949fc59b732088e059885482a569cb4222a5

Observation 66d0fc71-81f6-4aee-b9ff-a19f156460cf · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Simplifying Deep Temporal Difference Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:12:41.290942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T17:08:31.770889Z digest=sha256:315e28d05d08c5cc4d8129bf4f6d1266777c82ca2998de9bca57bd45aa31c972

Observation 99db9eed-557d-4a49-9d21-11518d8db892 · inbound

A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations cites this paper.

A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations Simplifying Deep Temporal Difference Learning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:36:26.179257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T10:27:55.337253Z digest=sha256:29d50dd31e45b84089dd63315369f2c325a8ac78e957f5cd2a9d1092bab224f9

Observation a98f02b8-939f-483e-a53d-b0622f339fc3 · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:24:27.808540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:21:15.933174Z digest=sha256:30f3435b64c4c9cd4d4e8c93625f80e9ca22edcd43ca08abbfd21cdd284947cb

Observation 7fb16431-0e58-43f7-b718-674279663c0f · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling Simplifying Deep Temporal Difference Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.638679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:17:59.698127Z digest=sha256:8fa78930caf73a8be4a9a85f077f974706ec8eba594c5cc65d964fe56343cc7a

Observation e7d2bc9c-9f63-4cf8-a1f9-e1d79c416930 · inbound

Goal-Conditioned Agents that Learn Everything All at Once cites this paper.

Goal-Conditioned Agents that Learn Everything All at Once Simplifying Deep Temporal Difference Learning

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T05:00:21.532243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T04:59:48.867927Z digest=sha256:69e16a7c0f3a9ed6adda1a87111b230bd335dfa86d70947cca1cc4a0a6d96914

Observation 610deba9-2a2d-4ea0-baf2-60ece410dbf5 · inbound

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion cites this paper.

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion Simplifying Deep Temporal Difference Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:05:49.709954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T00:54:12.099045Z digest=sha256:e832a35ac9c3b1d295509b9713ec8bfab60581eb638085d2409be1f02e9113fa

Observation 55d89c66-9695-4d48-b86b-6f1e692f7511 · inbound

Task diversity produces systematic transfer but inhibits continual reinforcement learning cites this paper.

Task diversity produces systematic transfer but inhibits continual reinforcement learning Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.240548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T18:58:56.444461Z digest=sha256:f82940f58b0a04b21de87d8972bf76d6518b032d1c576234a93bbc5166ec1816

Observation 8fbaf87d-acd4-4039-b25b-b2f5d992d2c6 · inbound

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning cites this paper.

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning Simplifying Deep Temporal Difference Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T07:41:45.477737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T07:37:08.423086Z digest=sha256:1aa8406628d4f0b5eab6a549421321e3b832170ad20f0c35c3131f74b875444f

Observation 03362496-b669-41a4-b8d4-ec3f5ad205a7 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Simplifying Deep Temporal Difference Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.706984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:96ec7c3c79e483307ca939e553cd872138274d87df51586941f3f411ca511cb3

Observation 3aef960c-934a-4938-8528-e304a4dd1541 · inbound

Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning cites this paper.

Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning Simplifying Deep Temporal Difference Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:51:17.007544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:51:17.007544Z digest=sha256:1014b5d706bf843d446fdb1951eed55d560ff4eb1289f1c821afa7fd32504d87