Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:50.611860Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.12611.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:50.611860Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:10:34.080907Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T14:10:34.228624Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ddf5ae27-cae6-459e-841f-70fc763c6c82 · outbound
Action-Dependent Optimality-Preserving Reward Shaping write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b55c4b9-a2d0-44c2-a7c4-013f73a3350a · outbound
Action-Dependent Optimality-Preserving Reward Shaping M., and Sun, W
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6b15fafb-5b05-4b37-81d1-8e482873f1a3 · outbound
Action-Dependent Optimality-Preserving Reward Shaping Never Give Up: Learning Directed Exploration Strategies
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b331bd9-9092-4ae6-8d2c-c755b77c1b72 · outbound
Action-Dependent Optimality-Preserving Reward Shaping E., Harutyunyan, A., and Bowling, M
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7c905b99-d911-426d-bfd3-65d10de7b0da · outbound
Action-Dependent Optimality-Preserving Reward Shaping Unifying count-based exploration and intrinsic motivation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3e0ac42e-1ac1-47ef-97bd-ff0fa586086b · outbound
Action-Dependent Optimality-Preserving Reward Shaping G., Naddaf, Y., Veness, J., and Bowling, M
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf407efd-e69f-410b-bb47-91a37e68a718 · outbound
Action-Dependent Optimality-Preserving Reward Shaping Large-Scale Study of Curiosity-Driven Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a6a0034-6e01-457c-9678-0f0114402c43 · outbound
Action-Dependent Optimality-Preserving Reward Shaping Exploration by Random Network Distillation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 614c7343-e290-4abd-aa96-3f6e1c3c24ce · outbound
Action-Dependent Optimality-Preserving Reward Shaping Exploration by random network distillation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fb88c7e8-eb47-48c8-995f-7874f20e4dcf · outbound
Action-Dependent Optimality-Preserving Reward Shaping Redeeming Intrinsic Rewards via Constrained Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d3c1b780-762a-4e8f-9bc1-6e9e59620960 · outbound
Action-Dependent Optimality-Preserving Reward Shaping Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 97586bd5-9a50-497c-8f83-6a27b2eee651 · outbound
Action-Dependent Optimality-Preserving Reward Shaping Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6966b1ed-462c-4234-b13b-1d0a63efdee5 · outbound
Action-Dependent Optimality-Preserving Reward Shaping C., Gupta, N., Villalobos-Arias, L., Potts, C
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a3b3e1e7-56ae-49f8-900b-39f27c32fb3c · outbound
Action-Dependent Optimality-Preserving Reward Shaping C., Villalobos-Arias, L., Wang, J., Jhala, A., and Roberts, D
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f1094e15-1266-42f5-bd67-f61342225bbd · outbound
Action-Dependent Optimality-Preserving Reward Shaping Reward shaping in episodic reinforcement learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d25cfc5b-aaa1-476c-a831-6a966b3ed76d · outbound
Action-Dependent Optimality-Preserving Reward Shaping Expressing arbitrary reward functions as potential-based advice
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e7f1cac0-c55f-4638-91ff-6f89bcc20d57 · outbound
Action-Dependent Optimality-Preserving Reward Shaping On stationary point convergence of ppo-clip
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aaddb901-bcf8-4fd1-b9dd-56faae739cb2 · outbound
Action-Dependent Optimality-Preserving Reward Shaping Ppo-rnd, 2022
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6fa5b1fa-2e1a-4fef-9a16-73fc9deac4c4 · outbound
Action-Dependent Optimality-Preserving Reward Shaping Beyond surprise: Improving exploration through surprise novelty
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fc12e4b1-c349-4748-8bf5-e203bb443b40 · outbound
Action-Dependent Optimality-Preserving Reward Shaping Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation afedd25e-13a0-46e6-9c19-5908e9586e39 · outbound
Action-Dependent Optimality-Preserving Reward Shaping A., Veness, J., Bellemare, M
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b3c308f-984b-4341-8d9e-207d743da03f · outbound
Action-Dependent Optimality-Preserving Reward Shaping Y., Harada, D., and Russell, S
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 19c6aab4-1a57-4b81-a9fc-6959b5a2a1e5 · outbound
Action-Dependent Optimality-Preserving Reward Shaping A., and Darrell, T
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8069108d-907d-413b-9d89-3854b37ec643 · outbound
Action-Dependent Optimality-Preserving Reward Shaping and Williams, R
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9a69146f-af2a-4e22-a416-66ac72f24ef4 · outbound
Action-Dependent Optimality-Preserving Reward Shaping and Alstr m, P
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 09fc3cab-872e-4a18-bee7-bbdef8c8a9da · outbound
Action-Dependent Optimality-Preserving Reward Shaping Proximal Policy Optimization Algorithms
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deb75815-c322-44b3-85aa-7a84b315e637 · outbound
Action-Dependent Optimality-Preserving Reward Shaping Potential-based shaping and q-value initialization are equivalent
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9906cebb-dc28-46c7-b4c8-231b4a42f07b · outbound
Action-Dependent Optimality-Preserving Reward Shaping Automatic intrinsic reward shaping for exploration in deep reinforcement learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 02fe87b0-9663-4e20-8075-9adda41ed7ad · inbound
Minding Motivation: The Effect of Intrinsic Motivation on Agent Behaviors Action-Dependent Optimality-Preserving Reward Shaping
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.