Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T15:52:16.912981Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2606.04029.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T15:52:16.912981Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 92ccab8a-f15e-4d28-9476-90046d9d69f8 · outbound
Position: Deployed Reinforcement Learning should be Continual Solving Rubik's Cube with a Robot Hand
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ccad928-0bce-4314-9663-761a3bbb83a8 · outbound
Position: Deployed Reinforcement Learning should be Continual Concrete Problems in AI Safety
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7e0f079a-08c9-4156-aba2-3f4718d3a61a · outbound
Position: Deployed Reinforcement Learning should be Continual Dota 2 with Large Scale Deep Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 68d1d2e2-d849-4f29-9a5b-f21463cf11a4 · outbound
Position: Deployed Reinforcement Learning should be Continual Step-size Optimization for Continual Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b58becd8-7027-4610-86aa-c9d62f85c28e · outbound
Position: Deployed Reinforcement Learning should be Continual Can Context Bridge the Reality Gap? Sim-to-Real Transfer of Context-Aware Policies
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 935d409f-7b00-4f5d-b544-c4f65d3bc493 · outbound
Position: Deployed Reinforcement Learning should be Continual Accessed: 2026-01-11
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae664dc-5123-422c-ba52-fc1bebe5c3c7 · outbound
Position: Deployed Reinforcement Learning should be Continual Accessed: 2026-01-28
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec9c1a80-2451-4081-a511-6be0cac50fcf · outbound
Position: Deployed Reinforcement Learning should be Continual In-Context Learning can Perform Continual Learning Like Humans.arXiv preprint 2509.22764,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7963899a-59d5-47d1-8b4b-416028041cfb · outbound
Position: Deployed Reinforcement Learning should be Continual Accessed: 2026-01-27
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7230b9cf-49f5-4521-825d-c050ed5b3d97 · outbound
Position: Deployed Reinforcement Learning should be Continual Risk-sensitive Actor-Critic with Static Spectral Risk Measures for Online and Offline Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a1215ce-7207-4bd3-b2cd-d9872d7d139d · outbound
Position: Deployed Reinforcement Learning should be Continual ualberta.ca/RLAI/rewardhypothesis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f31dce2-94e3-45c3-91de-18106c1baa28 · outbound
Position: Deployed Reinforcement Learning should be Continual The Alberta Plan for AI Research
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ad5cb1a3-aa96-40c4-bc86-1be31eb72f77 · outbound
Position: Deployed Reinforcement Learning should be Continual Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 57da014b-a1ab-4e2f-860b-dd73df5d1e51 · outbound
Position: Deployed Reinforcement Learning should be Continual On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 62716def-181e-40de-9008-4a5dd09ed3e7 · outbound
Position: Deployed Reinforcement Learning should be Continual The damping coefficient bt grows in a noisily quadratic manner over time, as shown in Figure 3 (Gaussian noise σ= 0.02 )
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.