Pith. sign in

Paper Citation Record · LEDGER

Action-Dependent Optimality-Preserving Reward Shaping

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.12611.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12611 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:50.611860Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:10:34.080907Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T14:10:34.228624Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ddf5ae27-cae6-459e-841f-70fc763c6c82 · outbound

This paper cites write newline.

Action-Dependent Optimality-Preserving Reward Shaping write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.523066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.523066Z digest=sha256:eb6b7f46e28e0c024a951304cd8ceebb38cce985d58e92aff75a7cc9a4333959

Observation 8b55c4b9-a2d0-44c2-a7c4-013f73a3350a · outbound

This paper cites M., and Sun, W.

Action-Dependent Optimality-Preserving Reward Shaping M., and Sun, W

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.878481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.527688Z digest=sha256:e1fe71fbeeded4ebfaa6ad45915b244ffe8a4162e214d6c49771867c999bead0

Observation 6b15fafb-5b05-4b37-81d1-8e482873f1a3 · outbound

This paper cites Never Give Up: Learning Directed Exploration Strategies.

Action-Dependent Optimality-Preserving Reward Shaping Never Give Up: Learning Directed Exploration Strategies

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.531331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.531331Z digest=sha256:6f4986a5f8903d80156132b79c7e36bde1fbe84c7e84d1bb5ebf6f21ee45024a

Observation 6b331bd9-9092-4ae6-8d2c-c755b77c1b72 · outbound

This paper cites E., Harutyunyan, A., and Bowling, M.

Action-Dependent Optimality-Preserving Reward Shaping E., Harutyunyan, A., and Bowling, M

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.869496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.535027Z digest=sha256:5e64a97fb4af87a6a3380f101662fbf7cb43d1ac8f243d89ea9c0eda60800f2d

Observation 7c905b99-d911-426d-bfd3-65d10de7b0da · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Action-Dependent Optimality-Preserving Reward Shaping Unifying count-based exploration and intrinsic motivation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.860414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.538576Z digest=sha256:2f0da05360f352f2d72783d6c388ac845612aff05c0c1dd0483784ffec7067e8

Observation 3e0ac42e-1ac1-47ef-97bd-ff0fa586086b · outbound

This paper cites G., Naddaf, Y., Veness, J., and Bowling, M.

Action-Dependent Optimality-Preserving Reward Shaping G., Naddaf, Y., Veness, J., and Bowling, M

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.541853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.541853Z digest=sha256:408f9e0e1b8e315f710ff9069f955df71e1c2436c91a5a51c9655d68de2cbec1

Observation cf407efd-e69f-410b-bb47-91a37e68a718 · outbound

This paper cites Large-Scale Study of Curiosity-Driven Learning.

Action-Dependent Optimality-Preserving Reward Shaping Large-Scale Study of Curiosity-Driven Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.545127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.545127Z digest=sha256:1e280cf09feaf7ff30e97f21e252e5826bbe64c4c1b70e0bab3262227a30403d

Observation 1a6a0034-6e01-457c-9678-0f0114402c43 · outbound

This paper cites Exploration by Random Network Distillation.

Action-Dependent Optimality-Preserving Reward Shaping Exploration by Random Network Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.548845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.548845Z digest=sha256:793c5b440844d43cf4a4f6a3d7ebbcdd4973c9267acbba4e42430eb0089a5cc8

Observation 614c7343-e290-4abd-aa96-3f6e1c3c24ce · outbound

This paper cites Exploration by random network distillation.

Action-Dependent Optimality-Preserving Reward Shaping Exploration by random network distillation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.844480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.552559Z digest=sha256:6424435c76dd716c0b10a8fcc2b7bfaf13e1ff34a1fbeeeb4bbd577dbaac3bfa

Observation fb88c7e8-eb47-48c8-995f-7874f20e4dcf · outbound

This paper cites Redeeming Intrinsic Rewards via Constrained Optimization.

Action-Dependent Optimality-Preserving Reward Shaping Redeeming Intrinsic Rewards via Constrained Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:40:50.653534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.555494Z digest=sha256:ebe50a7b372617b90babe40328d33ee0e763c2f7755cdd891b0158cc0b8fe9f3

Observation d3c1b780-762a-4e8f-9bc1-6e9e59620960 · outbound

This paper cites an unresolved cited work.

Action-Dependent Optimality-Preserving Reward Shaping Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:40:50.835012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.559072Z digest=sha256:12b5f94258f45771885f6e5364fd99fa91cd1476eac3444799cb0211fb3925cc

Observation 97586bd5-9a50-497c-8f83-6a27b2eee651 · outbound

This paper cites an unresolved cited work.

Action-Dependent Optimality-Preserving Reward Shaping Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:40:50.825588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.562316Z digest=sha256:c6fb59e089a42e0674932cdba5159a41fb176b5fad8483eda563ad6ae4b693ba

Observation 6966b1ed-462c-4234-b13b-1d0a63efdee5 · outbound

This paper cites C., Gupta, N., Villalobos-Arias, L., Potts, C.

Action-Dependent Optimality-Preserving Reward Shaping C., Gupta, N., Villalobos-Arias, L., Potts, C

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.815185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.565273Z digest=sha256:ec3f03f3fca0b45639b4db5350e89350050cd61699d99c16f6cf8673b3fb3ced

Observation a3b3e1e7-56ae-49f8-900b-39f27c32fb3c · outbound

This paper cites C., Villalobos-Arias, L., Wang, J., Jhala, A., and Roberts, D.

Action-Dependent Optimality-Preserving Reward Shaping C., Villalobos-Arias, L., Wang, J., Jhala, A., and Roberts, D

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.805896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.568247Z digest=sha256:180199c1049868bcff44f445146ba67f6a7b05e1f94e8d51dcc024b7edb8cc30

Observation f1094e15-1266-42f5-bd67-f61342225bbd · outbound

This paper cites Reward shaping in episodic reinforcement learning.

Action-Dependent Optimality-Preserving Reward Shaping Reward shaping in episodic reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.795831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.571195Z digest=sha256:3cf68b42e7bb6c0376526401cb472b7e8a25e5191631f1095ddc9304b65d641f

Observation d25cfc5b-aaa1-476c-a831-6a966b3ed76d · outbound

This paper cites Expressing arbitrary reward functions as potential-based advice.

Action-Dependent Optimality-Preserving Reward Shaping Expressing arbitrary reward functions as potential-based advice

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.785619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.573997Z digest=sha256:8034d54c08274952e798cb73f6043f78b4ff01f167710ccfe7ccf2fc532e3112

Observation e7f1cac0-c55f-4638-91ff-6f89bcc20d57 · outbound

This paper cites On stationary point convergence of ppo-clip.

Action-Dependent Optimality-Preserving Reward Shaping On stationary point convergence of ppo-clip

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.776059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.577203Z digest=sha256:9cc486a765bc0c663720c74d9dd2a1ef5965d65393037d612e09aa34de6f4317

Observation aaddb901-bcf8-4fd1-b9dd-56faae739cb2 · outbound

This paper cites Ppo-rnd, 2022.

Action-Dependent Optimality-Preserving Reward Shaping Ppo-rnd, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.766887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.580076Z digest=sha256:1c5db73bfa390728b531ea952d6be9915909d37b64d1340828149eb5958a5365

Observation 6fa5b1fa-2e1a-4fef-9a16-73fc9deac4c4 · outbound

This paper cites Beyond surprise: Improving exploration through surprise novelty.

Action-Dependent Optimality-Preserving Reward Shaping Beyond surprise: Improving exploration through surprise novelty

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.757619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.583277Z digest=sha256:c63b8b7e18c78736f3cb27aa43141c726a11e462c1f883651d364044d7c9faa8

Observation fc12e4b1-c349-4748-8bf5-e203bb443b40 · outbound

This paper cites an unresolved cited work.

Action-Dependent Optimality-Preserving Reward Shaping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:40:50.748317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.587330Z digest=sha256:2acba71760346a9fc1396cff4d925bfa97ac0137a9f3192899a8f195507bcba5

Observation afedd25e-13a0-46e6-9c19-5908e9586e39 · outbound

This paper cites A., Veness, J., Bellemare, M.

Action-Dependent Optimality-Preserving Reward Shaping A., Veness, J., Bellemare, M

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.590500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.590500Z digest=sha256:7a8bebe1e50f47afae7f1ab027bb8216e876bd12a1047274069aaf896737aa54

Observation 3b3c308f-984b-4341-8d9e-207d743da03f · outbound

This paper cites Y., Harada, D., and Russell, S.

Action-Dependent Optimality-Preserving Reward Shaping Y., Harada, D., and Russell, S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.733116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.593381Z digest=sha256:6839f31f46ea6ccbedec3b6018af3c9e78ac1de31c0a67cdf2c1899127847b9a

Observation 19c6aab4-1a57-4b81-a9fc-6959b5a2a1e5 · outbound

This paper cites A., and Darrell, T.

Action-Dependent Optimality-Preserving Reward Shaping A., and Darrell, T

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.596486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.596486Z digest=sha256:8366894d56c223d3f0d928ec904586f5a7ae76a7a37f07b061ca78c151f26aa1

Observation 8069108d-907d-413b-9d89-3854b37ec643 · outbound

This paper cites and Williams, R.

Action-Dependent Optimality-Preserving Reward Shaping and Williams, R

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.718099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.599690Z digest=sha256:328144e39dcedeeb6cf9344f9aebb2de87f99724c5dce56fc9c865761e31272f

Observation 9a69146f-af2a-4e22-a416-66ac72f24ef4 · outbound

This paper cites and Alstr m, P.

Action-Dependent Optimality-Preserving Reward Shaping and Alstr m, P

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.708827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.602648Z digest=sha256:399ab95736b37cecc36db2dc5e1360f979175635dc5aa643a14532b2e9f85efc

Observation 09fc3cab-872e-4a18-bee7-bbdef8c8a9da · outbound

This paper cites Proximal Policy Optimization Algorithms.

Action-Dependent Optimality-Preserving Reward Shaping Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.605584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.605584Z digest=sha256:b6f49ca0bcdba670f8a17e089ffc3a2c68b6825a4ef630fec165b50b141dba07

Observation deb75815-c322-44b3-85aa-7a84b315e637 · outbound

This paper cites Potential-based shaping and q-value initialization are equivalent.

Action-Dependent Optimality-Preserving Reward Shaping Potential-based shaping and q-value initialization are equivalent

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.699687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.608682Z digest=sha256:ee706b003ba626eb72fb9f6be481ab3614dd558f47c7df95fd4cf841c09bd2ac

Observation 9906cebb-dc28-46c7-b4c8-231b4a42f07b · outbound

This paper cites Automatic intrinsic reward shaping for exploration in deep reinforcement learning.

Action-Dependent Optimality-Preserving Reward Shaping Automatic intrinsic reward shaping for exploration in deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:50.689800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:40:50.611860Z digest=sha256:4e6a8e8687d31630786524a4c83f05dd6193a769acdc22e4ee9f70825b14de0b

Pith citing papers

Observation 02fe87b0-9663-4e20-8075-9adda41ed7ad · inbound

Minding Motivation: The Effect of Intrinsic Motivation on Agent Behaviors cites this paper.

Minding Motivation: The Effect of Intrinsic Motivation on Agent Behaviors Action-Dependent Optimality-Preserving Reward Shaping

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:10:34.234283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T14:10:34.080907Z digest=sha256:6d21fe758b794db8b26a7d65770cfffbca0c821fefd266880bba16fb4f77256e