Pith. sign in

Paper Citation Record · LEDGER

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments

As of 17 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2507.00030.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00030 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:12:49.665075Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e1a3798-0d25-4f56-b343-8670aff436f4 · outbound

This paper cites Bellemare and others.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Bellemare and others

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.530869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.475803Z digest=sha256:249e6c4f16b02b9ceaa8d5be26ee67f490c50b101edfa16d6cb1c8f37a5e1ff5

Observation 3567d400-933f-469b-b9bb-2ca1b7f304a1 · outbound

This paper cites Frame skip is a powerful parameter for learning to play Atari.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Frame skip is a powerful parameter for learning to play Atari

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.398444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.543454Z digest=sha256:1337acf60bd2b1d193d08fde3104d830006e636be1327b1b25097b194de9e0bd

Observation 335db7ef-2e9f-4867-8c0d-da5f348a8d89 · outbound

This paper cites Gilbert and Timothy D.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Gilbert and Timothy D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.272890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.647439Z digest=sha256:11efea7493709c05bc22983144cdeedbb7c9d84388096c12eeaed06313469d2a

Observation a555fbe8-65b9-41fd-8310-f78551845d19 · outbound

This paper cites Lakshminarayanan and others.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Lakshminarayanan and others

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.099339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.777519Z digest=sha256:e82397c0f8601227649904f5f4cbd8e9e49bead286d82a981bf4eeac14513a5d

Observation 491d555f-1113-4195-b209-78488a0a87e3 · outbound

This paper cites A contextual-bandit approach to personalized news article recommendation.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments A contextual-bandit approach to personalized news article recommendation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.919715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.881659Z digest=sha256:0450cfaabdcf55a5adcb15478782f627105feafd6268f657cf6f489a6b22cb5e

Observation 6bf59c52-fba4-4fd7-a67d-3803fe9168ce · outbound

This paper cites Human-level control through deep reinforcement learning.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Human-level control through deep reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.740473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.954347Z digest=sha256:000bd19cf3c8893cd2ec9e0d12640d7b43e2277c33650a5759c4b7c48c3946d9

Observation b9ff5d85-7e5b-45f0-82f3-613ccedff340 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Asynchronous methods for deep reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.590134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.014473Z digest=sha256:a9528ed448b49f694726c5e34836a71b4f55dea951cc974308b5d19b63e057a2

Observation 0e8db6ed-34d1-402c-9815-b9aabcb59659 · outbound

This paper cites Mastering the game of Go with deep neural networks and tree search.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Mastering the game of Go with deep neural networks and tree search

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.419941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.154324Z digest=sha256:b71661acb1b921c82628641c59a245484a4edcc7e38be06dde1f94ed259dd3a8

Observation 26e464ee-730c-4a27-95b4-e2089cef3724 · outbound

This paper cites Sutton and others.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Sutton and others

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.250126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.316626Z digest=sha256:fcd14e529860d1060e575c74ad876634eae0ddeb7e55fbb56f4557b28ed11520

Observation 4aabe205-466b-4c2b-a5c5-00452fb8a38b · outbound

This paper cites Effect of scalar leptoquarks on the rare decays of B_s meson.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Effect of scalar leptoquarks on the rare decays of B_s meson

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:12:49.836013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.524676Z digest=sha256:177136b726e7c739ca4f0cd8ef23d9591fd510ae6d4d9a5e14a7ea298e2fcd92

Observation 9daf9023-c8d9-4a86-93e8-5808381c8ac9 · outbound

This paper cites A gradient estimate for nonlocal minimal graphs.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments A gradient estimate for nonlocal minimal graphs

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:12:50.085655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.665075Z digest=sha256:18652d1ba2d7ad683669b762c90daa368a7625573796cb3d2903af562caa0f3c

Pith citing papers

No inbound Pith citation observations are available.