Pith. sign in

Paper Citation Record · LEDGER

Adaptive Reward Design for Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2412.10917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10917 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:37:31.607767Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:30:33.390291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:30:33.977487Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8b4a402-f74d-4c99-b096-e5f8f7c13658 · outbound

This paper cites Control synthesis from linear temporal logic specifications using model-free reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Control synthesis from linear temporal logic specifications using model-free reinforcement learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.056244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.416772Z digest=sha256:f7eaf23a8ee8e957451bda1f8c8d3ca8ad5f217f354dc62d7312a88e20225736

Observation 5005d4f1-7527-460b-a175-e0e51bcb431c · outbound

This paper cites OpenAI Gym.

Adaptive Reward Design for Reinforcement Learning OpenAI Gym

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.422543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.422543Z digest=sha256:f0a9c48395b76e1e06a10bc05c008c018461ed8baa6b726613708b39ad6af559

Observation 50838048-cec8-4079-b4fe-6631b96abc7d · outbound

This paper cites Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications.

Adaptive Reward Design for Reinforcement Learning Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.042967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.427324Z digest=sha256:516ca32d478a1ddd84ce1a088d716d55a07225d1658af7e35bb53ef5da040383

Observation 6028a52c-44e4-4ea1-9500-054d80fe0e5a · outbound

This paper cites Learning minimally-violating continuous control for infeasible linear temporal logic specifications.

Adaptive Reward Design for Reinforcement Learning Learning minimally-violating continuous control for infeasible linear temporal logic specifications

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.029062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.433079Z digest=sha256:bd5cf6f89295af4e92b4617c4dda7d01b7c8a7a89b9b597fcc626d7034521664

Observation a6937f62-cfca-4933-a5f8-f260ab2aaa7d · outbound

This paper cites Ltl and beyond: Formal languages for reward function specification in reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Ltl and beyond: Formal languages for reward function specification in reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.016510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.440473Z digest=sha256:3d4995e24be7c9e0383eafab8d5a8419a10be91cb223de6c9eece1f78fd58160

Observation ccd9e90b-0801-42ff-9ef7-04b9a0d11362 · outbound

This paper cites Foundations for restraining bolts: Reinforcement learning with ltlf/ldlf restraining specifications.

Adaptive Reward Design for Reinforcement Learning Foundations for restraining bolts: Reinforcement learning with ltlf/ldlf restraining specifications

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.003421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.446769Z digest=sha256:4726ea70f28228bb042b9f30fa5485d6042fe172c92b28e33dc470d55fb96c1f

Observation 3560acac-53cd-4c1e-93b4-015dc9c65932 · outbound

This paper cites From language to goals: Inverse reinforcement learning for vision-based instruction following.

Adaptive Reward Design for Reinforcement Learning From language to goals: Inverse reinforcement learning for vision-based instruction following

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.987526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.452848Z digest=sha256:1af3a7fb3f2c28b136b6db5abd3e44fbbee915042f77262ba215689b6c8c4b0f

Observation 4c33d730-ac7d-4357-8f29-f7408f49d443 · outbound

This paper cites Reinforcement learning for temporal logic control synthesis with probabilistic satisfaction guarantees.

Adaptive Reward Design for Reinforcement Learning Reinforcement learning for temporal logic control synthesis with probabilistic satisfaction guarantees

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.971633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.458473Z digest=sha256:041a4e147e5dcf13e7970268a9796b5479bc52d6d08728075e5267e10e6b83c8

Observation 28d912a4-17f9-4163-96ae-cfc3421f9ada · outbound

This paper cites Deep reinforcement learning with temporal logics.

Adaptive Reward Design for Reinforcement Learning Deep reinforcement learning with temporal logics

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.953462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.464583Z digest=sha256:30428121c5c6a6236bb861da833be4db4424c7db6a351c36a847235b66bcc30c

Observation 6c8af446-f2ec-41a3-947d-e3028e860199 · outbound

This paper cites Reward machines: Exploiting reward function structure in reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Reward machines: Exploiting reward function structure in reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.931948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.472017Z digest=sha256:71b97170ff6ce579d457b64c1a7207186d65e1b7c7e93db0133bdcfa8f8ecc71

Observation 6a025cf4-3809-4f04-952d-b6f39dbbf243 · outbound

This paper cites Temporal-logic-based reward shaping for continuing reinforcement learning tasks.

Adaptive Reward Design for Reinforcement Learning Temporal-logic-based reward shaping for continuing reinforcement learning tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.913149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.481730Z digest=sha256:8d62bf2d6ff538bad2f9dc6d68a2303ac8190e1df06b1f5847594c4d08819a03

Observation 55ee23ac-cc8d-4ced-833f-b07335c6ef3b · outbound

This paper cites A composable specification language for reinforcement learning tasks.

Adaptive Reward Design for Reinforcement Learning A composable specification language for reinforcement learning tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.893885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.493989Z digest=sha256:62c2c655ea69ca4425dc402fc416561e5621332e05076d4264690e2ec720ae63

Observation 2e1da156-c139-4568-972f-3140e4b28b46 · outbound

This paper cites Compositional reinforcement learning from logical specifications.

Adaptive Reward Design for Reinforcement Learning Compositional reinforcement learning from logical specifications

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.879116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.509070Z digest=sha256:0c33d57b0aafb6094e4d670370500519fc6f4a889f1339afc0fe29f787315046

Observation 990b3a30-54ab-490c-8476-c80992c7bf68 · outbound

This paper cites Model checking of safety properties.

Adaptive Reward Design for Reinforcement Learning Model checking of safety properties

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.862535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.531591Z digest=sha256:9d6f31da23b948a1f09bdd3db38c6c5f453887a9a90df88bf10cace217bd6227

Observation c8d86e92-4313-4978-9209-df85a676621b · outbound

This paper cites Probabilistic planning with formal performance guarantees for mobile service robots.

Adaptive Reward Design for Reinforcement Learning Probabilistic planning with formal performance guarantees for mobile service robots

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.847530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.537253Z digest=sha256:452ad39d8e27da76c3bf1a6bad100d14a4abb9410ce97ce7a3d62986b63c93c8

Observation e8bef7a2-3962-4188-bddf-15a9f20ee938 · outbound

This paper cites Reinforcement learning with temporal logic rewards.

Adaptive Reward Design for Reinforcement Learning Reinforcement learning with temporal logic rewards

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.833861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.547114Z digest=sha256:2ae89b63a77af4bf0fa2cc6ac921198ba19b522b8a666abbf98398d4c3c803c3

Observation 4c2b5c61-fb91-46b8-9fda-e4254a406a0f · outbound

This paper cites Continuous control with deep reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Continuous control with deep reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.552290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.552290Z digest=sha256:d2f8a81e3fdfaf795ab17bc295ec295c099f142240c6440a4ebc6ca1777b713d

Observation d8fad458-da6e-45ce-b701-5681eb729509 · outbound

This paper cites Human-level control through deep reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Human-level control through deep reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.556926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.556926Z digest=sha256:93209a52d185b87c38bcf15f0f467b3cdc4d526be8b0f3b6579a0cc9c847334a

Observation 1a8b3ba6-fda5-4c02-a686-379596ee1bf7 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Asynchronous methods for deep reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.797080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.561606Z digest=sha256:6b67caadd81bf1da0e67c0a8e07f1c31f57c33993e90e3bf98048fb5792e4145

Observation c0e50507-434f-4da3-b512-3c9c80236a2b · outbound

This paper cites Algorithms for inverse reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Algorithms for inverse reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.778765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.570702Z digest=sha256:8fc6b67c08d3109eab9d2f16901f23e0474f9540878ded05bf3121430e7005e8

Observation 08ec08ef-dc8d-42a8-80d9-42aabc55c140 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

Adaptive Reward Design for Reinforcement Learning Policy invariance under reward transformations: Theory and application to reward shaping

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.757627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.584332Z digest=sha256:f2190e6f9bcbe626c06c58f432e5a490fd6a3333b6f38ebf4817fbd9c7bfce63

Observation 38404a0c-da5b-459f-87bb-38bf87f9a166 · outbound

This paper cites The temporal semantics of concurrent programs.

Adaptive Reward Design for Reinforcement Learning The temporal semantics of concurrent programs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.733684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.588728Z digest=sha256:900d963b80cab4021cff2f05d5f2917e75b7014376ab34a5f4c3d928dbbda2d6

Observation b79e5b12-0680-40e0-9c64-691dae414074 · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

Adaptive Reward Design for Reinforcement Learning Stable-baselines3: Reliable reinforcement learning implementations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.717040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.593564Z digest=sha256:ebf7fc2601f2dc4b759d321188bb064d511a07be99097a0d746c7e6db0fdcbd2

Observation 74e00a7c-6247-4e1e-9cee-cc5afc606738 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Adaptive Reward Design for Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.597864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.597864Z digest=sha256:e4551c6c7b133e7b658db0f911a61981e8cc1eee7b3f4e65392e717497395494

Observation 699c7cc4-085b-4d80-aa89-8ded6c4ed214 · outbound

This paper cites Deep reinforcement learning with double q-learning.

Adaptive Reward Design for Reinforcement Learning Deep reinforcement learning with double q-learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.701426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.602369Z digest=sha256:c81dbf358e699ebb29d60f33b840e1370ca84f837bfe4f6122dc15adc68b5d9e

Observation f1af1b0f-3017-46e2-9fee-6bf33ef907e7 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

Adaptive Reward Design for Reinforcement Learning A survey of preference-based reinforcement learning methods

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.607767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.607767Z digest=sha256:f77bcfb9190fc56904c5604c008fefae2c5abc7c76ef8d8d979f206b1b876d9f

Pith citing papers

Observation 7def0c9b-206f-4f04-bbc5-6b2628a79826 · inbound

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning cites this paper.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Adaptive Reward Design for Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:30:33.983121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:30:33.390291Z digest=sha256:cae60f8f0bbdce86d7a7f4c51e1e3f393975a3f4db17a48584b1019e4ccc24c0