Pith. sign in

Paper Citation Record · LEDGER

Adaptive Reward Design for Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2412.10917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10917 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:37:31.607767Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:30:33.390291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:30:33.977487Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8b4a402-f74d-4c99-b096-e5f8f7c13658 · outbound

This paper cites Control synthesis from linear temporal logic specifications using model-free reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Control synthesis from linear temporal logic specifications using model-free reinforcement learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.056244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.416772Z digest=sha256:9a44ff72a8fc6fc3ea7c6d49889ada417aa900a1d4b866561f0fd62212471ddd

Observation 5005d4f1-7527-460b-a175-e0e51bcb431c · outbound

This paper cites OpenAI Gym.

Adaptive Reward Design for Reinforcement Learning OpenAI Gym

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.422543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.422543Z digest=sha256:50180dc62598f73a207cf8e0955ded53b8d7b9b5bc7213ffbfeb8c26de949de1

Observation 50838048-cec8-4079-b4fe-6631b96abc7d · outbound

This paper cites Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications.

Adaptive Reward Design for Reinforcement Learning Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.042967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.427324Z digest=sha256:9c91837f0d4379b971bee7d645267da6713ad1878ae5155687fc30d4129c84e1

Observation 6028a52c-44e4-4ea1-9500-054d80fe0e5a · outbound

This paper cites Learning minimally-violating continuous control for infeasible linear temporal logic specifications.

Adaptive Reward Design for Reinforcement Learning Learning minimally-violating continuous control for infeasible linear temporal logic specifications

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.029062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.433079Z digest=sha256:d459e46cd006337d77258be185dfad27ecf69f12c8e69c7e91e7baeac3812efa

Observation a6937f62-cfca-4933-a5f8-f260ab2aaa7d · outbound

This paper cites Ltl and beyond: Formal languages for reward function specification in reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Ltl and beyond: Formal languages for reward function specification in reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.016510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.440473Z digest=sha256:a1c557a2cdb2982b344b06f2cd9b7161bb18d27bd36baec688999a1b890240e6

Observation ccd9e90b-0801-42ff-9ef7-04b9a0d11362 · outbound

This paper cites Foundations for restraining bolts: Reinforcement learning with ltlf/ldlf restraining specifications.

Adaptive Reward Design for Reinforcement Learning Foundations for restraining bolts: Reinforcement learning with ltlf/ldlf restraining specifications

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.003421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.446769Z digest=sha256:fd67c870cd012bb01045f79355d2367da25c29a58112da534e8e73d184faebd2

Observation 3560acac-53cd-4c1e-93b4-015dc9c65932 · outbound

This paper cites From language to goals: Inverse reinforcement learning for vision-based instruction following.

Adaptive Reward Design for Reinforcement Learning From language to goals: Inverse reinforcement learning for vision-based instruction following

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.987526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.452848Z digest=sha256:6fdda9b7331daa439e844a294735893be2a6ed5f82c9daed3bed350ee63d5d16

Observation 4c33d730-ac7d-4357-8f29-f7408f49d443 · outbound

This paper cites Reinforcement learning for temporal logic control synthesis with probabilistic satisfaction guarantees.

Adaptive Reward Design for Reinforcement Learning Reinforcement learning for temporal logic control synthesis with probabilistic satisfaction guarantees

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.971633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.458473Z digest=sha256:97f2be18efe8b37ba381d50987b9a23a1f75ce1230006e507b8ae45b03979875

Observation 28d912a4-17f9-4163-96ae-cfc3421f9ada · outbound

This paper cites Deep reinforcement learning with temporal logics.

Adaptive Reward Design for Reinforcement Learning Deep reinforcement learning with temporal logics

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.953462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.464583Z digest=sha256:24e779a0ab140a3f820de035f0d57168d6c0d84c200ee34df6561dd2d814457d

Observation 6c8af446-f2ec-41a3-947d-e3028e860199 · outbound

This paper cites Reward machines: Exploiting reward function structure in reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Reward machines: Exploiting reward function structure in reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.931948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.472017Z digest=sha256:053da11c41cfb449cdc7af5cbae6d38b1677a23d0c999418f281f0f058f86823

Observation 6a025cf4-3809-4f04-952d-b6f39dbbf243 · outbound

This paper cites Temporal-logic-based reward shaping for continuing reinforcement learning tasks.

Adaptive Reward Design for Reinforcement Learning Temporal-logic-based reward shaping for continuing reinforcement learning tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.913149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.481730Z digest=sha256:deb21569b8c855be698ae5757de87721b9c179b77f3082304c86cacc47747d49

Observation 55ee23ac-cc8d-4ced-833f-b07335c6ef3b · outbound

This paper cites A composable specification language for reinforcement learning tasks.

Adaptive Reward Design for Reinforcement Learning A composable specification language for reinforcement learning tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.893885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.493989Z digest=sha256:f5f829b2b5f826a2ab4554926cfdf88fe3e8d831d31a399550217a98b74eece9

Observation 2e1da156-c139-4568-972f-3140e4b28b46 · outbound

This paper cites Compositional reinforcement learning from logical specifications.

Adaptive Reward Design for Reinforcement Learning Compositional reinforcement learning from logical specifications

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.879116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.509070Z digest=sha256:7343c8280979a269a9721587fc70d23f8e82245a676d766b3bccdb8484301056

Observation 990b3a30-54ab-490c-8476-c80992c7bf68 · outbound

This paper cites Model checking of safety properties.

Adaptive Reward Design for Reinforcement Learning Model checking of safety properties

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.862535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.531591Z digest=sha256:b5d0a7059e1ed1f5cff46ab1dcee5b47e8eb653ffdca3eb37468120ec2110780

Observation c8d86e92-4313-4978-9209-df85a676621b · outbound

This paper cites Probabilistic planning with formal performance guarantees for mobile service robots.

Adaptive Reward Design for Reinforcement Learning Probabilistic planning with formal performance guarantees for mobile service robots

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.847530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.537253Z digest=sha256:15381a079d62d6253d349de920b303d989032ed830fdc15d2b0c1850dc1ec17c

Observation e8bef7a2-3962-4188-bddf-15a9f20ee938 · outbound

This paper cites Reinforcement learning with temporal logic rewards.

Adaptive Reward Design for Reinforcement Learning Reinforcement learning with temporal logic rewards

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.833861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.547114Z digest=sha256:d0d05012f59713392cb68957f789cc802e1422569551ee4aff734aaf57ed7cac

Observation 4c2b5c61-fb91-46b8-9fda-e4254a406a0f · outbound

This paper cites Continuous control with deep reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Continuous control with deep reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.552290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.552290Z digest=sha256:6bc451bbd350a2f6020d9a22a46b661b4ba957cabf287faa8042ae99b8f6fc9d

Observation d8fad458-da6e-45ce-b701-5681eb729509 · outbound

This paper cites Human-level control through deep reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Human-level control through deep reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.556926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.556926Z digest=sha256:da7308836f0addacfbaa3c8f7ca0e6582d88850ae958656f892f2b504fdee310

Observation 1a8b3ba6-fda5-4c02-a686-379596ee1bf7 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Asynchronous methods for deep reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.797080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.561606Z digest=sha256:d7fc29294d594fca696a84ebdb7dae408c3f25642c3b2fde7c5f731fab09cefc

Observation c0e50507-434f-4da3-b512-3c9c80236a2b · outbound

This paper cites Algorithms for inverse reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Algorithms for inverse reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.778765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.570702Z digest=sha256:34249235ba5bc1917f2b0b9d6ad89cdd91801628b7cdc5ce4a47c72fb214a3f1

Observation 08ec08ef-dc8d-42a8-80d9-42aabc55c140 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

Adaptive Reward Design for Reinforcement Learning Policy invariance under reward transformations: Theory and application to reward shaping

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.757627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.584332Z digest=sha256:dc52753b67fd8c00772b47acc86383af9517e64d10eae477494a25bb8d8bffe2

Observation 38404a0c-da5b-459f-87bb-38bf87f9a166 · outbound

This paper cites The temporal semantics of concurrent programs.

Adaptive Reward Design for Reinforcement Learning The temporal semantics of concurrent programs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.733684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.588728Z digest=sha256:5e3637776d8b420db722ea735c4432301298db215e72b57583cf0dc9575cd9dd

Observation b79e5b12-0680-40e0-9c64-691dae414074 · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

Adaptive Reward Design for Reinforcement Learning Stable-baselines3: Reliable reinforcement learning implementations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.717040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.593564Z digest=sha256:fb94f481a5aabd1a1baa74f08f43176b575f44ceb9aceb943f335aa609192c97

Observation 74e00a7c-6247-4e1e-9cee-cc5afc606738 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Adaptive Reward Design for Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.597864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.597864Z digest=sha256:f93e61958e0b24eaad15e8845fa7894c3427be9bc8e8df3435463110c894645b

Observation 699c7cc4-085b-4d80-aa89-8ded6c4ed214 · outbound

This paper cites Deep reinforcement learning with double q-learning.

Adaptive Reward Design for Reinforcement Learning Deep reinforcement learning with double q-learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.701426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.602369Z digest=sha256:32d59302327229df266f57ac886ad6dd8275557ac7dcaeb4d11ba6d07bdbb590

Observation f1af1b0f-3017-46e2-9fee-6bf33ef907e7 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

Adaptive Reward Design for Reinforcement Learning A survey of preference-based reinforcement learning methods

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.607767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.607767Z digest=sha256:a72d014efe14566ce866b03eb25f09761dfcff716e803e3588119514c9be4f4b

Pith citing papers

Observation 7def0c9b-206f-4f04-bbc5-6b2628a79826 · inbound

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning cites this paper.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Adaptive Reward Design for Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:30:33.983121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:30:33.390291Z digest=sha256:8d00d5d7ed68b579eb1faafb056682b21c2aa9672f6f41564518d97076a3988b