Pith. sign in

Paper Citation Record · LEDGER

Proactive Constrained Policy Optimization with Preemptive Penalty

As of 16 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2508.01883.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01883 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:26:41.702511Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 158c1d0a-84f8-4e6d-a6f6-822a4f06fc5f · outbound

This paper cites Safe exploration in continuous action spaces, 2018.

Proactive Constrained Policy Optimization with Preemptive Penalty Safe exploration in continuous action spaces, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:45.597366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:38.755778Z digest=sha256:7461a58bd63dc41cbfb3eadf4619a77290fd8f5e6c6085764fd4df265d6598c8

Observation 8c770962-57f9-42a2-a67b-b091b3d94b69 · outbound

This paper cites A lyapunov-based approach to safe reinforcement learning.

Proactive Constrained Policy Optimization with Preemptive Penalty A lyapunov-based approach to safe reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:38.856977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:38.856977Z digest=sha256:884a21ba4f3600cd0754158e2a407b2ec5ade28a25e34378902240f2112067ad

Observation 9107fa5c-4e99-473c-8222-a1263edd31d0 · outbound

This paper cites Lyapunov-based safe reinforcement learning for microgrid energy management.IEEE transactions on neural networks and learning systems, 2024.

Proactive Constrained Policy Optimization with Preemptive Penalty Lyapunov-based safe reinforcement learning for microgrid energy management.IEEE transactions on neural networks and learning systems, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:45.400330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:38.949435Z digest=sha256:7cdcae10223a886cb95dbdd65f0611bf509070fef79000659cd473c52cdec6df

Observation 7cb9a1bb-0a77-456a-93c4-7d1eb72ba55e · outbound

This paper cites Reward Constrained Policy Optimization.

Proactive Constrained Policy Optimization with Preemptive Penalty Reward Constrained Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:39.090493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:39.090493Z digest=sha256:3e73330401e5cf27d4e5c437861101dd13686e8a712d52076f5dbabba5ba1400

Observation 7d7809c2-a09d-4131-add5-db0c898de8be · outbound

This paper cites A review of safe reinforcement learning: Methods, theories and applications.

Proactive Constrained Policy Optimization with Preemptive Penalty A review of safe reinforcement learning: Methods, theories and applications

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:39.201603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:39.201603Z digest=sha256:96c0bda987f528991082faebe216670aafb780abd04570d14059ff1ac7fee838

Observation dfdbf03e-c4e2-4a24-9ec1-0a616736697b · outbound

This paper cites Cvar-constrained policy optimization for safe reinforcement learning.

Proactive Constrained Policy Optimization with Preemptive Penalty Cvar-constrained policy optimization for safe reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:45.212121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:39.344167Z digest=sha256:3df169e7e9054588257c2960ab65a05b96e1863f8961d6cd8fed723cb8f189f5

Observation 25954455-b329-4f15-853e-1ab49c69c1cd · outbound

This paper cites Safe reinforcement learning for multi-agent systems with risk constraints.

Proactive Constrained Policy Optimization with Preemptive Penalty Safe reinforcement learning for multi-agent systems with risk constraints

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:45.017761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:39.493126Z digest=sha256:f8eb0971a969ea11f309ab8566e563172a72b62db7583329bb3a88fd77fceb9b

Observation 21b67990-3376-4db2-b984-bbac0bf513a4 · outbound

This paper cites Lyapunov-based safe policy optimization for continuous control, 2019.

Proactive Constrained Policy Optimization with Preemptive Penalty Lyapunov-based safe policy optimization for continuous control, 2019

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:44.769847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:39.574101Z digest=sha256:c22a1d94e41310a2c78f5435875a009fedc92dcb2532587f65caf50580a1cf62

Observation 17118456-c96d-4c44-b024-1d6d4f6219be · outbound

This paper cites Responsive safety in reinforcement learning by pid lagrangian methods.

Proactive Constrained Policy Optimization with Preemptive Penalty Responsive safety in reinforcement learning by pid lagrangian methods

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:39.672273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:39.672273Z digest=sha256:3b8ad8092f598948acff55b979dbce6ba355dab575ddf1ec06f53eb43a039e41

Observation 35880f46-6317-4f9a-8698-05ac91fbb800 · outbound

This paper cites Soufi Enayati, Mehran Ghafarian Tamizi, and Homayoun Najjaran.

Proactive Constrained Policy Optimization with Preemptive Penalty Soufi Enayati, Mehran Ghafarian Tamizi, and Homayoun Najjaran

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:44.588775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:39.772541Z digest=sha256:a0c77444c6c926eb8dad410ac0664908d1fbd61fda6e6f83f82b7dea7aa40c86

Observation dc73e218-5bfb-4075-89f4-7853786c8c85 · outbound

This paper cites A survey of constraint formulations in safe reinforcement learning, 2024.

Proactive Constrained Policy Optimization with Preemptive Penalty A survey of constraint formulations in safe reinforcement learning, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:44.430941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:39.862428Z digest=sha256:17c4c802d8d73d1ddc32e35e19d358eeb2bb3c9806c572d0e97cc0adeedf10db

Observation 69baed02-188c-4f37-84b3-29394e97d553 · outbound

This paper cites Projection-Based Constrained Policy Optimization.

Proactive Constrained Policy Optimization with Preemptive Penalty Projection-Based Constrained Policy Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:39.944282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:39.944282Z digest=sha256:c03419ba6e124532f27723bed15af7f649d9b9ef504388e3f4dc7e68e889e2de

Observation c184f696-0fc4-4987-91df-61f0d986fd55 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

Proactive Constrained Policy Optimization with Preemptive Penalty Risk-constrained reinforcement learning with percentile risk criteria

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:44.278532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.042380Z digest=sha256:6a1b45825267e412b64f9438a77bd25e47b525050aded4612bcc2549b1ba131c

Observation 1d51875b-3f80-43d4-a743-6fa57a6b21ab · outbound

This paper cites Constrained policy optimization.

Proactive Constrained Policy Optimization with Preemptive Penalty Constrained policy optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:40.158865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:40.158865Z digest=sha256:52e5f00a8acefcb1b05ef8899003ac8ebab5cad49dcaa43ea8379abb6fe694e8

Observation 712f9f21-044c-4653-8c19-b3b69a38efb9 · outbound

This paper cites Safe reinforcement learning using advantage-based intervention.

Proactive Constrained Policy Optimization with Preemptive Penalty Safe reinforcement learning using advantage-based intervention

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:44.152979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.273470Z digest=sha256:43c1829a712790e59bf746bee0082e343e323284b15f18eb25d305bb03bc07e8

Observation 1223e959-01da-4f0f-95b3-0f492f85ad12 · outbound

This paper cites Provably efficient primal-dual reinforcement learning for cmdps with non-stationary objectives and constraints.

Proactive Constrained Policy Optimization with Preemptive Penalty Provably efficient primal-dual reinforcement learning for cmdps with non-stationary objectives and constraints

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.998691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.366679Z digest=sha256:d7f2cf9d0622e2dd7d9506f72e883c5d8b35b02e985292820c23a64bbe377df2

Observation 90958591-c93f-48bf-acba-838bfe795348 · outbound

This paper cites Scalable primal- dual actor-critic method for safe multi-agent rl with general utilities.

Proactive Constrained Policy Optimization with Preemptive Penalty Scalable primal- dual actor-critic method for safe multi-agent rl with general utilities

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.903161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.479056Z digest=sha256:14f6627599d01d966219a7e9d127eeaff45d6064d3a1e6d24c26275b292c17c5

Observation fbf3189f-7c34-49e7-aee6-a0dcf555708d · outbound

This paper cites Adaptive primal-dual method for safe reinforce- ment learning, 2024.

Proactive Constrained Policy Optimization with Preemptive Penalty Adaptive primal-dual method for safe reinforce- ment learning, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.771070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.594650Z digest=sha256:e69cc891ce72d44e82376e78819b1fe67201b0fd495ea4049e8e5ced00d670c0

Observation abb19f79-54c5-445b-a110-6ecf2a944da9 · outbound

This paper cites Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation.

Proactive Constrained Policy Optimization with Preemptive Penalty Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.612791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.692585Z digest=sha256:9be25774b46278bb1555a59a08e41f05d44d582b845898345e0f3573164c8c79

Observation bd75ed25-de71-418e-8265-a3878072853d · outbound

This paper cites Safe cor: A dual-expert approach to integrating imitation learning and safe reinforcement learning using constraint rewards.

Proactive Constrained Policy Optimization with Preemptive Penalty Safe cor: A dual-expert approach to integrating imitation learning and safe reinforcement learning using constraint rewards

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.462053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.768219Z digest=sha256:0d9951792c02df5977f5ef997b704f11b454e7d345aee71a48c3ea9138dc04a4

Observation 4f735d69-43ea-4c44-b06e-f32235da32ee · outbound

This paper cites Safe reinforcement learning via episodic control.

Proactive Constrained Policy Optimization with Preemptive Penalty Safe reinforcement learning via episodic control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.327459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.854334Z digest=sha256:d770c3caadec9e84b2d4e513073d40a7383b93378b61bacde1ddabc49d3799f8

Observation 730a5251-5b7c-405b-93a1-3c6824cdfb27 · outbound

This paper cites Comprehensive overview of reward engineering and shaping in advancing reinforcement learning applications.

Proactive Constrained Policy Optimization with Preemptive Penalty Comprehensive overview of reward engineering and shaping in advancing reinforcement learning applications

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:43.126303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:40.933393Z digest=sha256:b337de76fb60615afc289007459ffd823c7d5da3e019560ee82b69bb5bae19ef

Observation 3e5e93af-91a8-42d0-bd39-283a535f8cff · outbound

This paper cites A review of safe reinforcement learning methods for modern power systems.

Proactive Constrained Policy Optimization with Preemptive Penalty A review of safe reinforcement learning methods for modern power systems

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.946267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.025604Z digest=sha256:98be85f879627b03856507693683a8aaae1a013ebb2853e289dcda55b582a6a0

Observation 7f434828-22fa-4faf-aad4-c3659ecf033c · outbound

This paper cites Constrained Markov decision processes.

Proactive Constrained Policy Optimization with Preemptive Penalty Constrained Markov decision processes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:41.098054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:41.098054Z digest=sha256:3bd4aebfdd62f3b77347456fb9c5da21e9cae76f0ebd02d05531f2b3c1cc2f9d

Observation 0f0ee32c-7a79-4a4e-b7d0-8b28b8b53bbf · outbound

This paper cites Reinforcement learning, 2015.

Proactive Constrained Policy Optimization with Preemptive Penalty Reinforcement learning, 2015

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.769405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.160613Z digest=sha256:5edebd90017ef4427dd147643a3d24e28b8f21404e29757cffe3ac4d34bc70e2

Observation c2a75454-4b9a-4462-9e1c-b6d578cdd7b6 · outbound

This paper cites Asynchronous methods for deep reinforce- ment learning.

Proactive Constrained Policy Optimization with Preemptive Penalty Asynchronous methods for deep reinforce- ment learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:41.251357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:41.251357Z digest=sha256:acd64c73513353187083dd53a302f29d0a20b3557a5e414332fb2d9ccbba6533

Observation 96babdb9-16e3-42da-8e6c-e4d1a9a1d8f9 · outbound

This paper cites Constrained deep networks: Lagrangian optimization via log-barrier extensions.

Proactive Constrained Policy Optimization with Preemptive Penalty Constrained deep networks: Lagrangian optimization via log-barrier extensions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:41.316580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:41.316580Z digest=sha256:176fcf8b0329c09ab5768aec3ab61b45cf6825d0e2937fcc7a57504fccae9bd5

Observation 8189934b-557a-43e5-aa0b-eafd085a65df · outbound

This paper cites A tutorial on mm algorithms.

Proactive Constrained Policy Optimization with Preemptive Penalty A tutorial on mm algorithms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.576444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.379442Z digest=sha256:dfb4d084496dca4ba705cfa1c167bfd3090263f5294cc50b90e25ed8da9497b9

Observation 2cc1ca94-bbdd-476f-98c4-5d4883ce31e4 · outbound

This paper cites Safety gymnasium: A unified safe reinforcement learning benchmark.

Proactive Constrained Policy Optimization with Preemptive Penalty Safety gymnasium: A unified safe reinforcement learning benchmark

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.435045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.440539Z digest=sha256:984ac0fb7ae761176cdc1a56afa3513fcd01782d3651763db3e07b8e8cad903b

Observation 9eac2dca-eb20-4493-b0ca-fde350a1a3be · outbound

This paper cites Constrained update projection approach to safe policy optimization.

Proactive Constrained Policy Optimization with Preemptive Penalty Constrained update projection approach to safe policy optimization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.315793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.506850Z digest=sha256:dd9ca8311e174343ec780444db154af227a58e3a9ef5a2c6e18d33442a08ec30

Observation 60bb6ac8-e7bf-4a41-8bcc-a64de835a208 · outbound

This paper cites First order constrained optimization in policy space.

Proactive Constrained Policy Optimization with Preemptive Penalty First order constrained optimization in policy space

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:26:42.162364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.574762Z digest=sha256:020d2c57d323bb7e3f194113d1946279808aef56954cd449c748475800e96f6a

Observation 91f14207-99f6-4e88-ad5d-3f74c8fcc513 · outbound

This paper cites Trust region policy optimization.

Proactive Constrained Policy Optimization with Preemptive Penalty Trust region policy optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:26:41.646373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:26:41.646373Z digest=sha256:8b9e91138ef13f4897e5b94e289089105d6318b70e0147234faec243ca9c93af

Observation b7862b4c-a11e-45be-8c1c-e40d78bea53d · outbound

This paper cites repulsive force.

Proactive Constrained Policy Optimization with Preemptive Penalty repulsive force

Reference 33

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T05:26:41.994351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:26:41.702511Z digest=sha256:ac44e023cd55782d37c949e73cde092b3a7fc38baf57be90268fd546459b62cb

Pith citing papers

No inbound Pith citation observations are available.