Pith. sign in

Paper Citation Record · LEDGER

Proactive Constrained Policy Optimization with Preemptive Penalty

As of 16 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2508.01883.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01883 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:26:41.702511Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 158c1d0a-84f8-4e6d-a6f6-822a4f06fc5f · outbound

This paper cites Safe exploration in continuous action spaces, 2018.

Proactive Constrained Policy Optimization with Preemptive Penalty Safe exploration in continuous action spaces, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:45.597366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:38.755778Z digest=sha256:257db419d924f3c48cf12967eaa2bb210069ad9d1da606145c8e5d197c1c746e

Observation 8c770962-57f9-42a2-a67b-b091b3d94b69 · outbound

This paper cites A lyapunov-based approach to safe reinforcement learning.

Proactive Constrained Policy Optimization with Preemptive Penalty A lyapunov-based approach to safe reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:38.856977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:38.856977Z digest=sha256:ee42f074b95363d2bc814e7b240a422bf995ed2efc1b6114269880317f69c452

Observation 9107fa5c-4e99-473c-8222-a1263edd31d0 · outbound

This paper cites Lyapunov-based safe reinforcement learning for microgrid energy management.IEEE transactions on neural networks and learning systems, 2024.

Proactive Constrained Policy Optimization with Preemptive Penalty Lyapunov-based safe reinforcement learning for microgrid energy management.IEEE transactions on neural networks and learning systems, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:45.400330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:38.949435Z digest=sha256:051ff43bbdac7ecaea2d477d58fe01f9cd760b51b2b92818a374d1c05a86f758

Observation 7cb9a1bb-0a77-456a-93c4-7d1eb72ba55e · outbound

This paper cites Reward Constrained Policy Optimization.

Proactive Constrained Policy Optimization with Preemptive Penalty Reward Constrained Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:39.090493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:39.090493Z digest=sha256:c142e90e20a76706e68adf2d178be7f80cd9e54340961e2ab4e411b344fd6d20

Observation 7d7809c2-a09d-4131-add5-db0c898de8be · outbound

This paper cites A review of safe reinforcement learning: Methods, theories and applications.

Proactive Constrained Policy Optimization with Preemptive Penalty A review of safe reinforcement learning: Methods, theories and applications

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:39.201603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:39.201603Z digest=sha256:803885cca013c3bc1f2b0e91e77c56e062ea985ca4bab9aadea5f887699e3374

Observation dfdbf03e-c4e2-4a24-9ec1-0a616736697b · outbound

This paper cites Cvar-constrained policy optimization for safe reinforcement learning.

Proactive Constrained Policy Optimization with Preemptive Penalty Cvar-constrained policy optimization for safe reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:45.212121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:39.344167Z digest=sha256:3572f8ba34ffcebe36bf32f2fe8b1aaafc8c796e054e8dd32cae0f33426227a1

Observation 25954455-b329-4f15-853e-1ab49c69c1cd · outbound

This paper cites Safe reinforcement learning for multi-agent systems with risk constraints.

Proactive Constrained Policy Optimization with Preemptive Penalty Safe reinforcement learning for multi-agent systems with risk constraints

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:45.017761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:39.493126Z digest=sha256:811dc8f2d075877f7c3b82269c50d7295c771ece37301a073aee552490326f0e

Observation 21b67990-3376-4db2-b984-bbac0bf513a4 · outbound

This paper cites Lyapunov-based safe policy optimization for continuous control, 2019.

Proactive Constrained Policy Optimization with Preemptive Penalty Lyapunov-based safe policy optimization for continuous control, 2019

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:44.769847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:39.574101Z digest=sha256:72b9d2319142d3db2b63be81608505591a73dd8816a45219ce726db040d8ecfd

Observation 17118456-c96d-4c44-b024-1d6d4f6219be · outbound

This paper cites Responsive safety in reinforcement learning by pid lagrangian methods.

Proactive Constrained Policy Optimization with Preemptive Penalty Responsive safety in reinforcement learning by pid lagrangian methods

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:39.672273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:39.672273Z digest=sha256:90bf6a9ee0369dfe1a74fbda5cc116931aeed2054876c91e0557401b2d657b70

Observation 35880f46-6317-4f9a-8698-05ac91fbb800 · outbound

This paper cites Soufi Enayati, Mehran Ghafarian Tamizi, and Homayoun Najjaran.

Proactive Constrained Policy Optimization with Preemptive Penalty Soufi Enayati, Mehran Ghafarian Tamizi, and Homayoun Najjaran

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:44.588775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:39.772541Z digest=sha256:e9d9f28d984bf16b16e9d34f01cc8b8c20143eb9eece45edc107bd76a0133825

Observation dc73e218-5bfb-4075-89f4-7853786c8c85 · outbound

This paper cites A survey of constraint formulations in safe reinforcement learning, 2024.

Proactive Constrained Policy Optimization with Preemptive Penalty A survey of constraint formulations in safe reinforcement learning, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:44.430941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:39.862428Z digest=sha256:77e371950c81976b54d93f3bb2799328f302f98ae4cce40e0cc2d38f2713207e

Observation 69baed02-188c-4f37-84b3-29394e97d553 · outbound

This paper cites Projection-Based Constrained Policy Optimization.

Proactive Constrained Policy Optimization with Preemptive Penalty Projection-Based Constrained Policy Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:39.944282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:39.944282Z digest=sha256:c5483e0c107e41dc8f5dda32965622aeaa9e05d4f97b4fbcd7e5469404bdab07

Observation c184f696-0fc4-4987-91df-61f0d986fd55 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

Proactive Constrained Policy Optimization with Preemptive Penalty Risk-constrained reinforcement learning with percentile risk criteria

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:44.278532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.042380Z digest=sha256:ca5fba964b37c52827957574bbded7e77767e6ea818c18361b720ed0b24633d5

Observation 1d51875b-3f80-43d4-a743-6fa57a6b21ab · outbound

This paper cites Constrained policy optimization.

Proactive Constrained Policy Optimization with Preemptive Penalty Constrained policy optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:40.158865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:40.158865Z digest=sha256:79a7da02a0d87f9708e1a9f82db1276e64daf5c33d9ccbf0663642f4e3cea315

Observation 712f9f21-044c-4653-8c19-b3b69a38efb9 · outbound

This paper cites Safe reinforcement learning using advantage-based intervention.

Proactive Constrained Policy Optimization with Preemptive Penalty Safe reinforcement learning using advantage-based intervention

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:44.152979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.273470Z digest=sha256:36bde86cd591773db102b61af7175893f8b4dae6374a21c2c58187321b431d32

Observation 1223e959-01da-4f0f-95b3-0f492f85ad12 · outbound

This paper cites Provably efficient primal-dual reinforcement learning for cmdps with non-stationary objectives and constraints.

Proactive Constrained Policy Optimization with Preemptive Penalty Provably efficient primal-dual reinforcement learning for cmdps with non-stationary objectives and constraints

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.998691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.366679Z digest=sha256:8656fa19cd00832dc706023417d11af51c756ac33ac1a4ae98129d8809de201c

Observation 90958591-c93f-48bf-acba-838bfe795348 · outbound

This paper cites Scalable primal- dual actor-critic method for safe multi-agent rl with general utilities.

Proactive Constrained Policy Optimization with Preemptive Penalty Scalable primal- dual actor-critic method for safe multi-agent rl with general utilities

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.903161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.479056Z digest=sha256:fae5c67d055e316ecb7b6ba568dabec364d6df3312d6f2f78707fc70a3d0fa48

Observation fbf3189f-7c34-49e7-aee6-a0dcf555708d · outbound

This paper cites Adaptive primal-dual method for safe reinforce- ment learning, 2024.

Proactive Constrained Policy Optimization with Preemptive Penalty Adaptive primal-dual method for safe reinforce- ment learning, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.771070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.594650Z digest=sha256:0c901b9acd997f20f1be1e50b857ed209175c93063440dd6169ce6e8430c2bc7

Observation abb19f79-54c5-445b-a110-6ecf2a944da9 · outbound

This paper cites Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation.

Proactive Constrained Policy Optimization with Preemptive Penalty Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.612791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.692585Z digest=sha256:58c827bf3b105196085e885bb3f76b8c8fdf088fa4e567c080d40334681f02d6

Observation bd75ed25-de71-418e-8265-a3878072853d · outbound

This paper cites Safe cor: A dual-expert approach to integrating imitation learning and safe reinforcement learning using constraint rewards.

Proactive Constrained Policy Optimization with Preemptive Penalty Safe cor: A dual-expert approach to integrating imitation learning and safe reinforcement learning using constraint rewards

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.462053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.768219Z digest=sha256:d49cd1ae932ac98a592ff2e55d659ee3f1d89c86d81731675c47affd1cbff63c

Observation 4f735d69-43ea-4c44-b06e-f32235da32ee · outbound

This paper cites Safe reinforcement learning via episodic control.

Proactive Constrained Policy Optimization with Preemptive Penalty Safe reinforcement learning via episodic control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.327459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.854334Z digest=sha256:bb589828466153b9b3b5b1cee695484a6bd69238ba3c44ae0c4a893abc4d3bb6

Observation 730a5251-5b7c-405b-93a1-3c6824cdfb27 · outbound

This paper cites Comprehensive overview of reward engineering and shaping in advancing reinforcement learning applications.

Proactive Constrained Policy Optimization with Preemptive Penalty Comprehensive overview of reward engineering and shaping in advancing reinforcement learning applications

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.126303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.933393Z digest=sha256:e0195d680aa3fec21ddaf73f0d0ea00345fd6ae89acc2930b6aa8c5804cce02b

Observation 3e5e93af-91a8-42d0-bd39-283a535f8cff · outbound

This paper cites A review of safe reinforcement learning methods for modern power systems.

Proactive Constrained Policy Optimization with Preemptive Penalty A review of safe reinforcement learning methods for modern power systems

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.946267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.025604Z digest=sha256:dda5d70a307cc83b73d0f569c1dc81e0edb13c4a2ce838421c60a23891939cb8

Observation 7f434828-22fa-4faf-aad4-c3659ecf033c · outbound

This paper cites Constrained Markov decision processes.

Proactive Constrained Policy Optimization with Preemptive Penalty Constrained Markov decision processes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:41.098054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:41.098054Z digest=sha256:6a2ae868593f16243357086b8bd335e2688ff177d0216937ecb2926b97c12ea9

Observation 0f0ee32c-7a79-4a4e-b7d0-8b28b8b53bbf · outbound

This paper cites Reinforcement learning, 2015.

Proactive Constrained Policy Optimization with Preemptive Penalty Reinforcement learning, 2015

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.769405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.160613Z digest=sha256:8382adaad8634eefe624eae995c2de6be553ad87bcb7872f3ccc6282442a7790

Observation c2a75454-4b9a-4462-9e1c-b6d578cdd7b6 · outbound

This paper cites Asynchronous methods for deep reinforce- ment learning.

Proactive Constrained Policy Optimization with Preemptive Penalty Asynchronous methods for deep reinforce- ment learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:41.251357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:41.251357Z digest=sha256:b3d59175b90fa17748814312805bf89fe780c77e59d7391429dfe72a03b13c82

Observation 96babdb9-16e3-42da-8e6c-e4d1a9a1d8f9 · outbound

This paper cites Constrained deep networks: Lagrangian optimization via log-barrier extensions.

Proactive Constrained Policy Optimization with Preemptive Penalty Constrained deep networks: Lagrangian optimization via log-barrier extensions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:41.316580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:41.316580Z digest=sha256:8832f6c535b65df808d8dbb903c13e00e9c3d38542ef2a25b026b67add1dfb17

Observation 8189934b-557a-43e5-aa0b-eafd085a65df · outbound

This paper cites A tutorial on mm algorithms.

Proactive Constrained Policy Optimization with Preemptive Penalty A tutorial on mm algorithms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.576444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.379442Z digest=sha256:eae204ec0894795da9d28c352b8036861f5df797f1e1361f903ad254bea51aeb

Observation 2cc1ca94-bbdd-476f-98c4-5d4883ce31e4 · outbound

This paper cites Safety gymnasium: A unified safe reinforcement learning benchmark.

Proactive Constrained Policy Optimization with Preemptive Penalty Safety gymnasium: A unified safe reinforcement learning benchmark

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.435045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.440539Z digest=sha256:583c2db92cda7fff0d0a4759ddca315e7a8d6cb2a8dd33be2472c26e57861300

Observation 9eac2dca-eb20-4493-b0ca-fde350a1a3be · outbound

This paper cites Constrained update projection approach to safe policy optimization.

Proactive Constrained Policy Optimization with Preemptive Penalty Constrained update projection approach to safe policy optimization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.315793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.506850Z digest=sha256:31207a08d4af62e82f5ad01245b940320b6d86d9465fc4ddb3c7f79e1d38e544

Observation 60bb6ac8-e7bf-4a41-8bcc-a64de835a208 · outbound

This paper cites First order constrained optimization in policy space.

Proactive Constrained Policy Optimization with Preemptive Penalty First order constrained optimization in policy space

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.162364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.574762Z digest=sha256:796572452a3a592200ddfabb9ffb848c8d0470c135b58176f78e2958569726bf

Observation 91f14207-99f6-4e88-ad5d-3f74c8fcc513 · outbound

This paper cites Trust region policy optimization.

Proactive Constrained Policy Optimization with Preemptive Penalty Trust region policy optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:41.646373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:41.646373Z digest=sha256:45be53453d5363bf009b8351a877a11fdd14678e955e5242916bb7f4084f0301

Observation b7862b4c-a11e-45be-8c1c-e40d78bea53d · outbound

This paper cites repulsive force.

Proactive Constrained Policy Optimization with Preemptive Penalty repulsive force

Reference 33

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T05:26:41.994351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.702511Z digest=sha256:173e891ef2287277773a482d5e843e1b6ab307950b371bf0455ea070ef49b257

Pith citing papers

No inbound Pith citation observations are available.