Pith. sign in

Paper Citation Record · LEDGER

Action-Dependent Optimality-Preserving Reward Shaping

As of 16 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.12611.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12611 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:50.611860Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:10:34.080907Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T14:10:34.228624Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ddf5ae27-cae6-459e-841f-70fc763c6c82 · outbound

This paper cites write newline.

Action-Dependent Optimality-Preserving Reward Shaping write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.523066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.523066Z digest=sha256:eb6b7f46e28e0c024a951304cd8ceebb38cce985d58e92aff75a7cc9a4333959

Observation 8b55c4b9-a2d0-44c2-a7c4-013f73a3350a · outbound

This paper cites M., and Sun, W.

Action-Dependent Optimality-Preserving Reward Shaping M., and Sun, W

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.878481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.527688Z digest=sha256:423813f55a1993b0f42a0c6fbe6768b50ab8e01ee9b7f676dec80700f9b5526e

Observation 6b15fafb-5b05-4b37-81d1-8e482873f1a3 · outbound

This paper cites Never Give Up: Learning Directed Exploration Strategies.

Action-Dependent Optimality-Preserving Reward Shaping Never Give Up: Learning Directed Exploration Strategies

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.531331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.531331Z digest=sha256:6f4986a5f8903d80156132b79c7e36bde1fbe84c7e84d1bb5ebf6f21ee45024a

Observation 6b331bd9-9092-4ae6-8d2c-c755b77c1b72 · outbound

This paper cites E., Harutyunyan, A., and Bowling, M.

Action-Dependent Optimality-Preserving Reward Shaping E., Harutyunyan, A., and Bowling, M

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.869496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.535027Z digest=sha256:154fbe7b7ef013730951cbdc5511b947381e3628a52119c08fb22dcc10d02c5b

Observation 7c905b99-d911-426d-bfd3-65d10de7b0da · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Action-Dependent Optimality-Preserving Reward Shaping Unifying count-based exploration and intrinsic motivation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.860414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.538576Z digest=sha256:c912143cda186c9cb6a538cee8e35f687140d790688614b911651c86c9d598e7

Observation 3e0ac42e-1ac1-47ef-97bd-ff0fa586086b · outbound

This paper cites G., Naddaf, Y., Veness, J., and Bowling, M.

Action-Dependent Optimality-Preserving Reward Shaping G., Naddaf, Y., Veness, J., and Bowling, M

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.541853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.541853Z digest=sha256:408f9e0e1b8e315f710ff9069f955df71e1c2436c91a5a51c9655d68de2cbec1

Observation cf407efd-e69f-410b-bb47-91a37e68a718 · outbound

This paper cites Large-Scale Study of Curiosity-Driven Learning.

Action-Dependent Optimality-Preserving Reward Shaping Large-Scale Study of Curiosity-Driven Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.545127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.545127Z digest=sha256:1e280cf09feaf7ff30e97f21e252e5826bbe64c4c1b70e0bab3262227a30403d

Observation 1a6a0034-6e01-457c-9678-0f0114402c43 · outbound

This paper cites Exploration by Random Network Distillation.

Action-Dependent Optimality-Preserving Reward Shaping Exploration by Random Network Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.548845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.548845Z digest=sha256:793c5b440844d43cf4a4f6a3d7ebbcdd4973c9267acbba4e42430eb0089a5cc8

Observation 614c7343-e290-4abd-aa96-3f6e1c3c24ce · outbound

This paper cites Exploration by random network distillation.

Action-Dependent Optimality-Preserving Reward Shaping Exploration by random network distillation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.844480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.552559Z digest=sha256:7ee556d18fe40258ca7aa83744819a4bb3e27d0075b97bc5b5ea0b639c6970e9

Observation fb88c7e8-eb47-48c8-995f-7874f20e4dcf · outbound

This paper cites Redeeming Intrinsic Rewards via Constrained Optimization.

Action-Dependent Optimality-Preserving Reward Shaping Redeeming Intrinsic Rewards via Constrained Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:40:50.653534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.555494Z digest=sha256:bb79ef3694e12d29e0bb00496fe0cab1433c64e851ac91c87edc5d77f4a94d76

Observation d3c1b780-762a-4e8f-9bc1-6e9e59620960 · outbound

This paper cites an unresolved cited work.

Action-Dependent Optimality-Preserving Reward Shaping Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:40:50.835012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.559072Z digest=sha256:35a9ab0429d2f4b82ab24f49219911ab3b6cb7bf81fa833865f44a94e347a680

Observation 97586bd5-9a50-497c-8f83-6a27b2eee651 · outbound

This paper cites an unresolved cited work.

Action-Dependent Optimality-Preserving Reward Shaping Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:40:50.825588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.562316Z digest=sha256:d2356d009f5951423492d5a6bc8fd8558bdefee236abb08faa318aba5d13cc4d

Observation 6966b1ed-462c-4234-b13b-1d0a63efdee5 · outbound

This paper cites C., Gupta, N., Villalobos-Arias, L., Potts, C.

Action-Dependent Optimality-Preserving Reward Shaping C., Gupta, N., Villalobos-Arias, L., Potts, C

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.815185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.565273Z digest=sha256:1e6192159dea38dc5da453f648b2387e46ba82ec2ce02d894eacc668aaa9946c

Observation a3b3e1e7-56ae-49f8-900b-39f27c32fb3c · outbound

This paper cites C., Villalobos-Arias, L., Wang, J., Jhala, A., and Roberts, D.

Action-Dependent Optimality-Preserving Reward Shaping C., Villalobos-Arias, L., Wang, J., Jhala, A., and Roberts, D

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.805896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.568247Z digest=sha256:0cbae14e5b91be531b28742bcda0abca9c6fdb427737453906b28955f3aade95

Observation f1094e15-1266-42f5-bd67-f61342225bbd · outbound

This paper cites Reward shaping in episodic reinforcement learning.

Action-Dependent Optimality-Preserving Reward Shaping Reward shaping in episodic reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.795831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.571195Z digest=sha256:6ce9fa96f2d1247bcd21d174b889adeb77860e2eaa024649be88147d81bd7e9f

Observation d25cfc5b-aaa1-476c-a831-6a966b3ed76d · outbound

This paper cites Expressing arbitrary reward functions as potential-based advice.

Action-Dependent Optimality-Preserving Reward Shaping Expressing arbitrary reward functions as potential-based advice

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.785619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.573997Z digest=sha256:09aae40f49724c56c1ea5898c74328867bf1878be29c9673403947f95230df54

Observation e7f1cac0-c55f-4638-91ff-6f89bcc20d57 · outbound

This paper cites On stationary point convergence of ppo-clip.

Action-Dependent Optimality-Preserving Reward Shaping On stationary point convergence of ppo-clip

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.776059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.577203Z digest=sha256:46280cfe3f19c421fd1288cf9c3b50225dd5e9f536f9c8c5dea6e7bb3da50e57

Observation aaddb901-bcf8-4fd1-b9dd-56faae739cb2 · outbound

This paper cites Ppo-rnd, 2022.

Action-Dependent Optimality-Preserving Reward Shaping Ppo-rnd, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.766887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.580076Z digest=sha256:69734d234d09d2c5e860cf73ae93ad60f0cf919dbc2619307bbfed75556fa8a8

Observation 6fa5b1fa-2e1a-4fef-9a16-73fc9deac4c4 · outbound

This paper cites Beyond surprise: Improving exploration through surprise novelty.

Action-Dependent Optimality-Preserving Reward Shaping Beyond surprise: Improving exploration through surprise novelty

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.757619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.583277Z digest=sha256:0c92307c44e9b9df2fae1884a5e900a4c391a6f45c2d904acaec30be74b59a31

Observation fc12e4b1-c349-4748-8bf5-e203bb443b40 · outbound

This paper cites an unresolved cited work.

Action-Dependent Optimality-Preserving Reward Shaping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:40:50.748317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.587330Z digest=sha256:c7f383e077863ad15c897319617559606eab6902bda4a13f218c91e3465ddcb9

Observation afedd25e-13a0-46e6-9c19-5908e9586e39 · outbound

This paper cites A., Veness, J., Bellemare, M.

Action-Dependent Optimality-Preserving Reward Shaping A., Veness, J., Bellemare, M

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.590500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.590500Z digest=sha256:7a8bebe1e50f47afae7f1ab027bb8216e876bd12a1047274069aaf896737aa54

Observation 3b3c308f-984b-4341-8d9e-207d743da03f · outbound

This paper cites Y., Harada, D., and Russell, S.

Action-Dependent Optimality-Preserving Reward Shaping Y., Harada, D., and Russell, S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.733116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.593381Z digest=sha256:3c521c27cfdcc7038a99a7e45ddaaa0a9fda472fc783e3d6aa9551c12f12bdf4

Observation 19c6aab4-1a57-4b81-a9fc-6959b5a2a1e5 · outbound

This paper cites A., and Darrell, T.

Action-Dependent Optimality-Preserving Reward Shaping A., and Darrell, T

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.596486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.596486Z digest=sha256:8366894d56c223d3f0d928ec904586f5a7ae76a7a37f07b061ca78c151f26aa1

Observation 8069108d-907d-413b-9d89-3854b37ec643 · outbound

This paper cites and Williams, R.

Action-Dependent Optimality-Preserving Reward Shaping and Williams, R

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.718099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.599690Z digest=sha256:8a5b009ebc04a847ded49bf751466e109436d8f6b438ee740882a67304913ccf

Observation 9a69146f-af2a-4e22-a416-66ac72f24ef4 · outbound

This paper cites and Alstr m, P.

Action-Dependent Optimality-Preserving Reward Shaping and Alstr m, P

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.708827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.602648Z digest=sha256:49314275cccd7991f4f47c7e8aa4cf169f68d651a0db9039e8ade0d16c3dc30d

Observation 09fc3cab-872e-4a18-bee7-bbdef8c8a9da · outbound

This paper cites Proximal Policy Optimization Algorithms.

Action-Dependent Optimality-Preserving Reward Shaping Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.605584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.605584Z digest=sha256:b6f49ca0bcdba670f8a17e089ffc3a2c68b6825a4ef630fec165b50b141dba07

Observation deb75815-c322-44b3-85aa-7a84b315e637 · outbound

This paper cites Potential-based shaping and q-value initialization are equivalent.

Action-Dependent Optimality-Preserving Reward Shaping Potential-based shaping and q-value initialization are equivalent

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.699687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.608682Z digest=sha256:f23f7b94634cebe9c22acd7b14a4cb20d29c7c4ac65ea7a6b0a2cd6843d96f29

Observation 9906cebb-dc28-46c7-b4c8-231b4a42f07b · outbound

This paper cites Automatic intrinsic reward shaping for exploration in deep reinforcement learning.

Action-Dependent Optimality-Preserving Reward Shaping Automatic intrinsic reward shaping for exploration in deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.689800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.611860Z digest=sha256:d39834f80ed58fcb56f055701bc5a191de6fa8cb565c583a41e726710979a99f

Pith citing papers

Observation 02fe87b0-9663-4e20-8075-9adda41ed7ad · inbound

Minding Motivation: The Effect of Intrinsic Motivation on Agent Behaviors cites this paper.

Minding Motivation: The Effect of Intrinsic Motivation on Agent Behaviors Action-Dependent Optimality-Preserving Reward Shaping

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:10:34.234283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T14:10:34.080907Z digest=sha256:c0f3f95459ed939aa14591273612edb7d77cf702a6f0f188c849504143f3573d