Pith. sign in

Paper Citation Record · LEDGER

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards

As of 23 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2411.17861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17861 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:52:46.731946Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:28:06.474838Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T20:28:07.394748Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26a01fb4-8144-45d4-b498-45d2635589e9 · outbound

This paper cites Robustness measures and monitors for time window temporal logic.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Robustness measures and monitors for time window temporal logic

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.562833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.484155Z digest=sha256:e64149706e26dc90f9619e16bb486249766cf62872995c811f86b6fc7f48c303

Observation b831f535-cec6-4552-8dd3-5d889cac1b20 · outbound

This paper cites Q-learning for robust satisfaction of signal temporal logic specifications.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Q-learning for robust satisfaction of signal temporal logic specifications

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.535758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.493814Z digest=sha256:e38c34d722d834517a9dc2d0974d8fa728ec7f71136a9752577c10536133f011

Observation 02c8c338-4d29-4487-baa7-c5d32e88ac3b · outbound

This paper cites u diger Ehlers, Bettina K \.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards u diger Ehlers, Bettina K \

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.502104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.501307Z digest=sha256:5f74833092c1a3b0f7e600f46e9cd52bfba5218b03406000d577b9f6c11cde20

Observation 575564f0-34d5-4a2b-a6cb-eb8881e3f4d3 · outbound

This paper cites Temporal-logic-constrained hybrid reinforcement learning to perform optimal aerial monitoring with delivery drones.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Temporal-logic-constrained hybrid reinforcement learning to perform optimal aerial monitoring with delivery drones

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.478714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.512189Z digest=sha256:fb8f38d7a268d8f78db84600dcc9f4b64a92384dd916e35254dbb71b5fe91d26

Observation 4f82c5e7-bdc9-4640-ba75-73677b8c7b4c · outbound

This paper cites Principles of model checking.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Principles of model checking

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.455806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.532382Z digest=sha256:4e40f9176f8e00d784cfadb7d6a41f4454aeb3a8ec23668325b00f5c629344b3

Observation 46279274-41d4-4a85-be1c-63164cd10115 · outbound

This paper cites Structured reward shaping using signal temporal logic specifications.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Structured reward shaping using signal temporal logic specifications

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.431795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.554330Z digest=sha256:880997af3a0c1193b3bfdc1abf7e010c248502110b0ed0142c55ea72003bebad

Observation 507a73d9-7d7b-45a8-961e-1653a957f7cb · outbound

This paper cites Efficient online reinforcement learning with offline data.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Efficient online reinforcement learning with offline data

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.561945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.561945Z digest=sha256:3868caa15d809a902e1221fe3cdb26b5000b66722354520ef6602b90224c1b32

Observation 140370b8-4ab6-4a8c-bfd8-1f734537dc12 · outbound

This paper cites Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.392580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.571214Z digest=sha256:7722e2514a8e6b9036355e388e6d8ff0fbfa01d3b72a6d9859f254a0dcd4cb12

Observation 4f49832c-82b9-4a68-8481-deeebed5c022 · outbound

This paper cites Provably efficient exploration in policy optimization.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Provably efficient exploration in policy optimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.365131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.579436Z digest=sha256:874bd85387025ecda0a5c54d411bbcd064b6631099eb5d2b0c2942628922bef9

Observation 074d7896-4186-4531-a08c-f8cb89d001db · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Decision transformer: Reinforcement learning via sequence modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.586476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.586476Z digest=sha256:154f610d39d9b9131bca7e88162ee431e7788be1e7f37268f803aba220aeada9

Observation a71427fb-82c6-4d19-a888-d51377e037bc · outbound

This paper cites Mirror learning: A unifying framework of policy optimisation.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Mirror learning: A unifying framework of policy optimisation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.325509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.591702Z digest=sha256:48ed1519e82c036571a49eea06eaadf413e064343536f5db791d90a128d7622a

Observation a537d976-ec31-47ec-a8b5-e08437152fc5 · outbound

This paper cites Imitation Bootstrapped Reinforcement Learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Imitation Bootstrapped Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.600129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.600129Z digest=sha256:ab45ba202d14b0e6b805724dce86de6dc98120d0e46905bcf95d80803ecbdbdc

Observation 315a9cc6-2786-4d74-b15d-c7d494085f9f · outbound

This paper cites A tutorial on mm algorithms.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards A tutorial on mm algorithms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.300989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.606014Z digest=sha256:29f95180199777ba46a03b250f5e045e089b08513a976233dd4326d43398c7ab

Observation abb12aef-0a87-4c8d-8847-37421b20b81b · outbound

This paper cites Reward machines: Exploiting reward function structure in reinforcement learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Reward machines: Exploiting reward function structure in reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.612212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.612212Z digest=sha256:5e652a99e22f9ba7154bbd484644f0ffa3686e4ce32d316fcb8e7e35f1df0262

Observation 19539de4-6464-4d0d-917a-110a92845c3b · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Approximately optimal approximate reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.617397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.617397Z digest=sha256:09b571b786358d294b29142fc70c212741bdae42051eb88d3748c47a8f60aef8

Observation c99c3e8f-28c8-4e46-b47c-e89d8c6d131c · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Conservative q-learning for offline reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.624935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.624935Z digest=sha256:e0b5b4ad9a7e87a427b89e82748d3e9298dd5e3a7fe8d3d4f8e67c50ec2869c6

Observation 35a87e66-ffdd-4134-b9d5-5cb559e10415 · outbound

This paper cites Reinforcement learning with temporal logic rewards.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Reinforcement learning with temporal logic rewards

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.229610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.630179Z digest=sha256:b72b05d662b93578d520d27ce7e1cdbd0c81def59cf085d96e5ff60368a78d89

Observation 70a1a7f9-d3d9-4b9e-a01d-af8814f4e334 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.637160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.637160Z digest=sha256:267248ed20cffad4bd6a3c2ff3036aed15c7c548e7c1ef8dde564b2d8a88c05e

Observation 0547adb5-2f2c-43f0-a5c7-d748303d54f0 · outbound

This paper cites Advice-guided reinforcement learning in a non-markovian environment.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Advice-guided reinforcement learning in a non-markovian environment

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.193723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.644641Z digest=sha256:b4561bc1d714ae65c2116c2ab24f8eccbf5b630012a462847bf93c8320b8cc09

Observation ba6f8e0c-4952-4f05-bafb-8bad039f17c5 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Policy invariance under reward transformations: Theory and application to reward shaping

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.163130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.651318Z digest=sha256:f744a14a2123ce3c2291f551668a86e6726bfd124a9e89eb194a65fb4b74c5eb

Observation 9675e615-6219-4b16-8b3d-31c0b35025b3 · outbound

This paper cites Implicit human perception learning in complex and unknown environments.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Implicit human perception learning in complex and unknown environments

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.129012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.658404Z digest=sha256:9f2eb43f6d51ee842d8c37e8d58d615f4decc0948fb235eb0491380201b293cb

Observation d61732d2-32ce-49e4-ab93-009ff1927aae · outbound

This paper cites Agnostic system identification for model-based reinforcement learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Agnostic system identification for model-based reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.106305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.663767Z digest=sha256:0d858c506c972b89e0a2af58017c78a48b338a45af51920d8026f49e16d88ef5

Observation 4932ad35-cd2e-4111-a6f4-7adc7d0841c7 · outbound

This paper cites Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.083880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.669943Z digest=sha256:27a3e62ff74644a8fa5b219dc608cdaeec033eaf19676e3de55ab2e2a9da9899

Observation c8228c20-aed5-48aa-86a2-167f48d9e72f · outbound

This paper cites Kickstarting Deep Reinforcement Learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Kickstarting Deep Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.677141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.677141Z digest=sha256:7fe4743a22ee302cb5855a1f62a2066de99abcf60853dcb7e1e29a5b2b54feba

Observation 5cfa168b-e794-4500-9c49-3ef6bf4068e7 · outbound

This paper cites Trust region policy optimization.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Trust region policy optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.060838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.683744Z digest=sha256:e2c5576109ad329efc8b67efef02eed6c5e4e35e770997969f1528ed84359d37

Observation c9d8d462-38ba-43fd-9fda-13a0da0b570a · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.695180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.695180Z digest=sha256:5868bdae8bfaef377aa9de367b1876e4265c438558d16d19f4e16dd38b706125

Observation 9734904f-ca20-45c8-9a93-1ba29d8df102 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.701697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.701697Z digest=sha256:a8e25749e162d8970f831c27e4286618a846bfe1a4b0e5404543e01ab553fafd

Observation af193ef4-b9b9-412f-aea1-c2643a4c53c6 · outbound

This paper cites Reinforcement learning: An introduction.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Reinforcement learning: An introduction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.714546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.714546Z digest=sha256:698c60068ff843071c9f785af6bbba0a1d1d02ac4870c818009289c4f37e394f

Observation 1c0d99f7-56a0-45ca-af40-07e6a58348ca · outbound

This paper cites an unresolved cited work.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:52:47.005125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.720242Z digest=sha256:5b1582d0084bdc3b90f74f7d6c510ba07dbd0c7f9d9b229a96f49089dd71a66c

Observation 0b05f1c8-2b1c-425d-b52e-2168edbf0018 · outbound

This paper cites Time window temporal logic.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Time window temporal logic

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:46.971335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.726539Z digest=sha256:ebb837d46972aa5811907b9fbaac659f0fab2f7a6a96c24744175b76cbdd9e87

Observation 970abbdd-c3a1-4355-b40e-bdf92c9c81c6 · outbound

This paper cites Joint inference of reward machines and policies for reinforcement learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Joint inference of reward machines and policies for reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:46.925300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.731946Z digest=sha256:c1ec515b4bdd8be811f607e0549a765cb43338443106c6e36b54bc2185a21dc1

Pith citing papers

Observation 8ff85624-1885-40b4-9c98-286e164ae58a · inbound

ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins cites this paper.

ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:28:07.444921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:28:06.474838Z digest=sha256:88034ad5ce342518f6b0e9b44ee06b5dcb7e5048627ae15c478f9da9eedd89bb