Pith. sign in

Paper Citation Record · LEDGER

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

As of 16 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2411.14019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14019 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T15:42:59.404858Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5c84053-f80e-4b8f-84c1-453c19f3cc2e · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.123943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.123943Z digest=sha256:7ea20702d70e85f2d0329896ddd7251a6a9beb32b0d57515bb725052f5fcd96b

Observation ecdca55a-7bb2-4742-aaef-1a34505bbcd0 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.744959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.128254Z digest=sha256:0de5bd43e2e0b399da7159e3dadcb1e47b449b0424bcd81535d9da5dadcc9fec

Observation 381d1499-5fe6-4378-a685-f75d8aa4fea9 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Playing Atari with Deep Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.132112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.132112Z digest=sha256:2e9edccfa7d0bace36432125fab61e68fa9ec7fb53eecdb913ae34b050df34c5

Observation 24cd254b-6099-4857-9682-86463961ca40 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Dota 2 with Large Scale Deep Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.136421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.136421Z digest=sha256:1a1d198d7abe9766e5976cff62781b0300a3996c30cdf56a23a10610bf815d83

Observation 24576388-8192-482d-9d48-c44095a59020 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.734943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.140358Z digest=sha256:44a7853940a2376f1d7f255a78dfc60d305bf2422efc1fa95fb36608c8cceb15

Observation fb9eaa60-77f6-4d12-9569-17dc64c5a582 · outbound

This paper cites & V an Roy, B.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & V an Roy, B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.724610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.144006Z digest=sha256:e0b2b1e3952f362f12a538961035b15372b789d788d897cf3cb8a9bccefffe49

Observation 99d7a4e1-9531-47c0-9a6c-36dc7f0e74f7 · outbound

This paper cites & Williams, R.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Williams, R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.714635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.147694Z digest=sha256:1ed8391ac9e16af6ca989c52641e7b8bf175e47151f152596050d0c97fe442c7

Observation 83f9c814-d6f3-4b7a-a427-2955f06f1e62 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.151129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.151129Z digest=sha256:04855ee6efbeb4edb3ccc9a417f1a74f34581d71b9a453d1d48054fa5f5c82d9

Observation 6aca79d4-1d02-4420-aef9-2238ca940b1d · outbound

This paper cites & Sutton, R.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Sutton, R

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.697294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.154534Z digest=sha256:6fc7ced5660c7c9ac18a3bd3b67fac4457e04b306d639cbfe1c3b3a7c74c231d

Observation f87e1e2e-1a36-4a68-acc0-825f438294a9 · outbound

This paper cites Algorithms for reinforcement learning (Springer nature, 2022).

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Algorithms for reinforcement learning (Springer nature, 2022)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.687679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.157947Z digest=sha256:fcde0327039195faa80dc40f161eab953129b62ce4fb9559d12ad84bba082d84

Observation 7a9d7340-920e-4246-9be3-e5a2dc69fb58 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.676781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.161517Z digest=sha256:b80b6109f346c89fc11ce7f6796689f558a315e6cada055a95ee610ea50ed2fd

Observation 32a26af1-7437-4fd2-8094-5ba61c633612 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.666649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.164937Z digest=sha256:09a27796d022d534aa7e507a42df4ac4f1e64a25a251f250a13d3ddca7522d7c

Observation 9e936308-dedb-45db-bfe0-127856906239 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.657054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.168279Z digest=sha256:7e3e992aea60309aee8b38f1438e39698d40f9237e38a1fa173766142c41e28c

Observation cec608e5-bf70-4f60-aa52-e78108fc2ecd · outbound

This paper cites & Silver, D.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Silver, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.646626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.175545Z digest=sha256:66b1c189253d9d225027b5848a08a3728faf6095a0336c7aadab40cfbda47a58

Observation b2894724-a0e9-4ba2-9528-9381dd00129d · outbound

This paper cites Error bounds for approximate policy iteration.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Error bounds for approximate policy iteration

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.636542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.178859Z digest=sha256:cdfb7cde34839d0e3f058a2e2207d50a3c7ee0d044848a05bf39e67895e644f5

Observation 5391d55f-7134-4852-851d-8a29e11141e1 · outbound

This paper cites Prioritized Experience Replay.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Prioritized Experience Replay

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.182279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.182279Z digest=sha256:72a8676bf6d791ac1931c37e28a267daae3321ed6ec016f51c4bed401e7acf58

Observation 76a94055-7c46-4536-8752-8daad7f98cb2 · outbound

This paper cites & Pilarski, P.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Pilarski, P

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.625890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.185862Z digest=sha256:bc30d1f922df6cd559a2be5f40806191e50f56302a7bbac85ed3276bb257657f

Observation 9b2e9293-a66c-4e0a-a9e8-bf86c7793da1 · outbound

This paper cites & Wen, Z.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Wen, Z

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.615362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.189256Z digest=sha256:3534b887fd53020d69112466134f2206b96c02e1c9eb9ca2b9a7e617c14b105e

Observation 4e491de7-3065-4406-9cef-a33d81355653 · outbound

This paper cites How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.192879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.192879Z digest=sha256:1a8ab0d2beab51e26f45a656f1e327145ee5983ab6758f8c9880a91caea7a8eb

Observation 3e042fc0-aab3-41b1-8c54-9faf9ea65bcd · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.604663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.196488Z digest=sha256:12aa407f02b5b161e270670565bee1cea8d41e0c3afcaa3b75f6d2b6561da465

Observation 15c7bbf3-4bd7-4f00-991f-56d3cfba2e47 · outbound

This paper cites Optiongan: Learning joint reward-policy options using gen erative adversarial inverse reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Optiongan: Learning joint reward-policy options using gen erative adversarial inverse reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.592919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.199768Z digest=sha256:dc58d6b067447775ce0247d1f4968bae5f1650e65bcfb949d9697879e550e772

Observation 25bf3420-43d4-42aa-a0ea-de2cbe694ad5 · outbound

This paper cites Discovering hierarchy in reinforcement learnin g with hexq.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Discovering hierarchy in reinforcement learnin g with hexq

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.580832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.203160Z digest=sha256:8af1edea1c20be3135e18b46530e3a90ec5c77679204070034d13c5c96636e4a

Observation 0de67471-4fdf-4303-af8b-bdaa098d74ad · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.570128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.206660Z digest=sha256:199a1d3cb748bd56ac240ed6e8263b9cc4ac2922b5a3dfc23af855cae6a4abf6

Observation 2ee35c95-28e3-4b5c-9f77-d23cdeac66fd · outbound

This paper cites & Shimkin, N.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Shimkin, N

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.558865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.209835Z digest=sha256:fcd2d1b56ed490fb1a6f330e7dde0e19a6fbd0362538c6e041235aad8138587f

Observation f9f7ba4d-e162-4c3c-94cd-6b9179d2d483 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.547755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.213035Z digest=sha256:632bd6cedc6f32f160bd7aee348de93aea310fbabf8c72a6d737896562d50fa0

Observation 08db37ee-5cc7-4356-ac02-a81ee85b08a1 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.216225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.216225Z digest=sha256:e9f591ba2206d68cdde6322d13c4a3c20e9b5eb38b483d864d5f6f6ae2073050

Observation 107ff731-acc9-4c2d-a8ee-4c29de087b79 · outbound

This paper cites Human-level control through deep reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Human-level control through deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.529899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.219702Z digest=sha256:92027153d7e3e97c85f79b7281a09402ab37ab5c7b6dd4759cfad77ddbd24c97

Observation 4e8acf22-2e3c-486b-a08d-25bec08e83d0 · outbound

This paper cites A markovian decision process.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition A markovian decision process

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.223137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.223137Z digest=sha256:c185557861a6f437b5e995a8a38d2de5edcae575ed6debd024ae5a6847dbacc0

Observation a7040e14-bba6-4992-a7ad-7b1cee181c07 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.226498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.226498Z digest=sha256:8499fe370328967d55b35e940af4191a76fba1b5e1b3c088bbb21edb10f4b683

Observation e20b45e4-f732-4a9c-b065-32c7bdef6a0b · outbound

This paper cites S., McAllester, D.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition S., McAllester, D

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.513143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.229784Z digest=sha256:b2127a38fd0f179f155794079f2eb29dab253f384a4dc6de0f51e3cf52dd8534

Observation 5ebf1381-c522-491a-b65c-3f76c5b8b3e9 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.503178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.232992Z digest=sha256:e221854a3914c9c72d05aafabe6dd34ad9dc59444fd1cc60c2c4151d8a6bd697

Observation ff5edea6-8a09-4782-aeef-b1888b8d85ab · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Asynchronous methods for deep reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.236178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.236178Z digest=sha256:dfe3292572428fa09741e76cb512cb8dafd8bffc364dc508071a3f9818a83a97

Observation f2de305b-222f-454b-8927-97ea947fc5d7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Proximal Policy Optimization Algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.239471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.239471Z digest=sha256:c4f94050ab962a874b3900d7deed6ae96eb4d65b47233cb7d7d4dc3ea23de3ec

Observation 870b958d-5b33-4066-97d1-fdef6dbd1879 · outbound

This paper cites S., Barto, A.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition S., Barto, A

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.243122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.243122Z digest=sha256:f1af3b3e7f9552d896b18171e0b15ce9ae5ced0d9fc0f5d5fdd16fe6676befce

Observation e860f8dc-57df-468f-a7eb-c6b993051fe4 · outbound

This paper cites Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.246831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.246831Z digest=sha256:505388a6319949e6e5bcc298962a90886fef2e6ef358743e9d4858cafb7cc666

Observation 96c5fc6b-ea2c-407a-b573-4d9f66b394f2 · outbound

This paper cites Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:42:59.410682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.250552Z digest=sha256:1062357cf27e4f5da657432b0ee5d3131fdb73054651089153b979665b65a271

Pith citing papers

Observation 96c5fc6b-ea2c-407a-b573-4d9f66b394f2 · inbound

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition cites this paper.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:42:59.410682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.250552Z digest=sha256:1062357cf27e4f5da657432b0ee5d3131fdb73054651089153b979665b65a271