Pith. sign in

Paper Citation Record · LEDGER

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments

As of 13 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2507.00030.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00030 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:12:49.665075Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e1a3798-0d25-4f56-b343-8670aff436f4 · outbound

This paper cites Bellemare and others.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Bellemare and others

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.530869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.475803Z digest=sha256:99f5e7c3e1ea557d618952879bc183ec7b9a3d40cbff6681368f4c4be5c5490a

Observation 3567d400-933f-469b-b9bb-2ca1b7f304a1 · outbound

This paper cites Frame skip is a powerful parameter for learning to play Atari.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Frame skip is a powerful parameter for learning to play Atari

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.398444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.543454Z digest=sha256:39c92c6bb7305e63c9d7d0418fd1e1e0d06acd1c0b553c2752d1f79fd02999b3

Observation 335db7ef-2e9f-4867-8c0d-da5f348a8d89 · outbound

This paper cites Gilbert and Timothy D.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Gilbert and Timothy D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.272890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.647439Z digest=sha256:0bdada30014b7c89af051e0781bfbfba05fa9d48910590f791d1f7a15caed6a6

Observation a555fbe8-65b9-41fd-8310-f78551845d19 · outbound

This paper cites Lakshminarayanan and others.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Lakshminarayanan and others

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.099339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.777519Z digest=sha256:5a15f236f7e149e3f3c0abcb5d6c9d834d75bcd81edfe73f1e9ddcc152528b30

Observation 491d555f-1113-4195-b209-78488a0a87e3 · outbound

This paper cites A contextual-bandit approach to personalized news article recommendation.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments A contextual-bandit approach to personalized news article recommendation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.919715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.881659Z digest=sha256:4e7b6cb0b33e16136d73df0d7c5ffdbbc84263466b70f4d87f7e9018d48d0749

Observation 6bf59c52-fba4-4fd7-a67d-3803fe9168ce · outbound

This paper cites Human-level control through deep reinforcement learning.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Human-level control through deep reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.740473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.954347Z digest=sha256:c37b81213faa5e77463f71659b434dd6f071d718c70316a3fb864ec69a354982

Observation b9ff5d85-7e5b-45f0-82f3-613ccedff340 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Asynchronous methods for deep reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.590134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.014473Z digest=sha256:ee6a978919b7f958d0ed42fb50764be47c909b244ffc391eb021a46a266ada16

Observation 0e8db6ed-34d1-402c-9815-b9aabcb59659 · outbound

This paper cites Mastering the game of Go with deep neural networks and tree search.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Mastering the game of Go with deep neural networks and tree search

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.419941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.154324Z digest=sha256:d6f02b11f71988bd3300ff02e0f9b91b9420f9e27f8c64485fe4b80ea517c738

Observation 26e464ee-730c-4a27-95b4-e2089cef3724 · outbound

This paper cites Sutton and others.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Sutton and others

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.250126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.316626Z digest=sha256:43eb32c3cff13989ac914b22b226aadf9679163a4dde64881f969e373107f626

Observation 4aabe205-466b-4c2b-a5c5-00452fb8a38b · outbound

This paper cites Effect of scalar leptoquarks on the rare decays of B_s meson.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Effect of scalar leptoquarks on the rare decays of B_s meson

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:12:49.836013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.524676Z digest=sha256:99addbe7bce6ecef17b5887e8cc9c92849edbc5b2b5cef81e72a364c98ffe4f7

Observation 9daf9023-c8d9-4a86-93e8-5808381c8ac9 · outbound

This paper cites A gradient estimate for nonlocal minimal graphs.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments A gradient estimate for nonlocal minimal graphs

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:12:50.085655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.665075Z digest=sha256:8329974a0bcd57335819ed5e620e791d108795f3670e4d027b392bcbb318652d

Pith citing papers

No inbound Pith citation observations are available.