Pith. sign in

Paper Citation Record · LEDGER

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

As of 16 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2411.14019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14019 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T15:42:59.404858Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5c84053-f80e-4b8f-84c1-453c19f3cc2e · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.123943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.123943Z digest=sha256:9c56a1c294e7a85aeff1085670bb32a81267a62ec2ce359ad3331ebb7b424328

Observation ecdca55a-7bb2-4742-aaef-1a34505bbcd0 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.744959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.128254Z digest=sha256:6ec871339852e097132436a701d30acacf11cf292e9ab1576003cfc2542c258a

Observation 381d1499-5fe6-4378-a685-f75d8aa4fea9 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Playing Atari with Deep Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.132112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.132112Z digest=sha256:518929a55460fa09431dd1c397bd6cb629286f04dcbf7e967be84196595a3988

Observation 24cd254b-6099-4857-9682-86463961ca40 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Dota 2 with Large Scale Deep Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.136421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.136421Z digest=sha256:4223bf7cad00140ece3c29e744fe33516e58bb58b6c531e949e70086eeae5e4a

Observation 24576388-8192-482d-9d48-c44095a59020 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.734943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.140358Z digest=sha256:e967a7683a9395f07bcd8cd5d31ad32733e6ea552712b447216ae2c94cf28861

Observation fb9eaa60-77f6-4d12-9569-17dc64c5a582 · outbound

This paper cites & V an Roy, B.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & V an Roy, B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.724610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.144006Z digest=sha256:1eaf01da42c1345f6857b48bf23f8d3826e780b9bb5bb4779e979bc04d921b88

Observation 99d7a4e1-9531-47c0-9a6c-36dc7f0e74f7 · outbound

This paper cites & Williams, R.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Williams, R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.714635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.147694Z digest=sha256:657d25f9926a6113575c8ca0084fe10ec288a90920dabe0d2beca593fc71b6c6

Observation 83f9c814-d6f3-4b7a-a427-2955f06f1e62 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.151129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.151129Z digest=sha256:13faba56256f91a394078148e262c5fdb6a1e2695f7b3843be23ca1bfd1710dd

Observation 6aca79d4-1d02-4420-aef9-2238ca940b1d · outbound

This paper cites & Sutton, R.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Sutton, R

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.697294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.154534Z digest=sha256:7453bb267cfcd11785b1d0b774853be9f5d5f3372981d871751fe111a8e93942

Observation f87e1e2e-1a36-4a68-acc0-825f438294a9 · outbound

This paper cites Algorithms for reinforcement learning (Springer nature, 2022).

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Algorithms for reinforcement learning (Springer nature, 2022)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.687679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.157947Z digest=sha256:ed20cbaed48b4dd0e4f0e09684bd799cea6d311ab3df160fb5d61d2a7b94fdaa

Observation 7a9d7340-920e-4246-9be3-e5a2dc69fb58 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.676781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.161517Z digest=sha256:b099ad90bffe86cf75282513570f4c0584820424e572a1000b2a59cff9268cce

Observation 32a26af1-7437-4fd2-8094-5ba61c633612 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.666649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.164937Z digest=sha256:fa27ff9ffa4ba8af347cda8d7f87c69bb891f553a07f838b416d1fd54a362685

Observation 9e936308-dedb-45db-bfe0-127856906239 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.657054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.168279Z digest=sha256:9ab7782ef50c47c7a7a10948849f21a15bc11d4c69ce06021bf8a1a842256b0a

Observation cec608e5-bf70-4f60-aa52-e78108fc2ecd · outbound

This paper cites & Silver, D.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Silver, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.646626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.175545Z digest=sha256:b63d4ac4bf79f1893873142084978361f0ae70e8654a65fe5a42e7bdc4b14299

Observation b2894724-a0e9-4ba2-9528-9381dd00129d · outbound

This paper cites Error bounds for approximate policy iteration.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Error bounds for approximate policy iteration

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.636542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.178859Z digest=sha256:308642391e16084e21a5f02e3d4318258777fd65c709a3632b769c3e6d7840fe

Observation 5391d55f-7134-4852-851d-8a29e11141e1 · outbound

This paper cites Prioritized Experience Replay.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Prioritized Experience Replay

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.182279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.182279Z digest=sha256:ff14cd365628dcad9a6bec55797780c30a2a7a161a873b8ffda8b4c068b16550

Observation 76a94055-7c46-4536-8752-8daad7f98cb2 · outbound

This paper cites & Pilarski, P.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Pilarski, P

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.625890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.185862Z digest=sha256:4e6526c9ae641651d97402dcee53f64ed61fc2332e33367a4776be32ab012403

Observation 9b2e9293-a66c-4e0a-a9e8-bf86c7793da1 · outbound

This paper cites & Wen, Z.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Wen, Z

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.615362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.189256Z digest=sha256:f7a03e2487c3ea86eb69cb779080dafb72d05ece3912a7fd40580bbabd4e5946

Observation 4e491de7-3065-4406-9cef-a33d81355653 · outbound

This paper cites How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.192879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.192879Z digest=sha256:89cc8db7ab310cf4d427fb6a5bd1eb9f1e0e37500c7d22a578b23867fc5bd5b3

Observation 3e042fc0-aab3-41b1-8c54-9faf9ea65bcd · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.604663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.196488Z digest=sha256:12a2d2bfd7f53cc7717423aaa9be5b94cc57b8cd2f6ac6439999badfd5635747

Observation 15c7bbf3-4bd7-4f00-991f-56d3cfba2e47 · outbound

This paper cites Optiongan: Learning joint reward-policy options using gen erative adversarial inverse reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Optiongan: Learning joint reward-policy options using gen erative adversarial inverse reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.592919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.199768Z digest=sha256:5df4af5612fc22d3d06fab2d0a3c0584182420336126dc0922c0884d6e7e9116

Observation 25bf3420-43d4-42aa-a0ea-de2cbe694ad5 · outbound

This paper cites Discovering hierarchy in reinforcement learnin g with hexq.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Discovering hierarchy in reinforcement learnin g with hexq

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.580832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.203160Z digest=sha256:ff43d29381443d2c534519d899188c46ee0f5545682b7e6d25334fcf33669eac

Observation 0de67471-4fdf-4303-af8b-bdaa098d74ad · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.570128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.206660Z digest=sha256:4581b5b4a1d1df53aac26361aed130935e8b63a6a33a8f12a803c63651860a81

Observation 2ee35c95-28e3-4b5c-9f77-d23cdeac66fd · outbound

This paper cites & Shimkin, N.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Shimkin, N

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.558865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.209835Z digest=sha256:02e560b116b1cf3337bd65c4d62c37183cd837fb22cb371515bfd55b153cad17

Observation f9f7ba4d-e162-4c3c-94cd-6b9179d2d483 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.547755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.213035Z digest=sha256:e2a0fa4be6d3c7239de152073de86cca5b6fbe80f40cf97a010febad7d0af381

Observation 08db37ee-5cc7-4356-ac02-a81ee85b08a1 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.216225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.216225Z digest=sha256:8d8949dee233cf4e2de8ee93b1406948994d471cb3143953706a42c2200c9add

Observation 107ff731-acc9-4c2d-a8ee-4c29de087b79 · outbound

This paper cites Human-level control through deep reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Human-level control through deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.529899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.219702Z digest=sha256:7f54204439c21e07751c3f27a190583b8850ef7f7eb79e2e07c0cab1df824852

Observation 4e8acf22-2e3c-486b-a08d-25bec08e83d0 · outbound

This paper cites A markovian decision process.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition A markovian decision process

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.223137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.223137Z digest=sha256:92830ee38d9ce0be763912121a090d423e429f5dd5733190cd38bb90d8ba2024

Observation a7040e14-bba6-4992-a7ad-7b1cee181c07 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.226498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.226498Z digest=sha256:d0a53cab3db9074934b6c9a01f4f0bd7b13930cc377c89812d0f0ee6b8568d37

Observation e20b45e4-f732-4a9c-b065-32c7bdef6a0b · outbound

This paper cites S., McAllester, D.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition S., McAllester, D

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:42:59.513143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.229784Z digest=sha256:1deb85afefd6f7130e2f4147cc21074aa70164b581df292ed61a4d153152975f

Observation 5ebf1381-c522-491a-b65c-3f76c5b8b3e9 · outbound

This paper cites an unresolved cited work.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:42:59.503178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.232992Z digest=sha256:94132dd29724e653b7ced58f12a754223d4b443e38b528193080dc4fecf8f159

Observation ff5edea6-8a09-4782-aeef-b1888b8d85ab · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Asynchronous methods for deep reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.236178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.236178Z digest=sha256:914f8a3662452c824ee83319e76e82f85f75ff8edc9ed9ffabddaaa64d48b426

Observation f2de305b-222f-454b-8927-97ea947fc5d7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Proximal Policy Optimization Algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.239471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.239471Z digest=sha256:7ee730c005e041d8bb55556cbf3c33918624dc63292c23681ce0e16ab62e1ffc

Observation 870b958d-5b33-4066-97d1-fdef6dbd1879 · outbound

This paper cites S., Barto, A.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition S., Barto, A

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.243122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.243122Z digest=sha256:f66dbb36c4baad20c43c055e8726c3e564fb60cd06fefe16422782f7daf0a962

Observation e860f8dc-57df-468f-a7eb-c6b993051fe4 · outbound

This paper cites Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.246831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.246831Z digest=sha256:4a4e2e79d3ede591b547994676c50c142d9b62b672c3f7b97c118c1dabe23c32

Observation 96c5fc6b-ea2c-407a-b573-4d9f66b394f2 · outbound

This paper cites Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:42:59.410682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.250552Z digest=sha256:fe08e9ea2c0fcf50ec00644d01f28d6dff0472ce99901369fbcc148c6afb8abe

Pith citing papers

Observation 96c5fc6b-ea2c-407a-b573-4d9f66b394f2 · inbound

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition cites this paper.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:42:59.410682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:42:59.250552Z digest=sha256:fe08e9ea2c0fcf50ec00644d01f28d6dff0472ce99901369fbcc148c6afb8abe