Pith. sign in

Paper Citation Record · LEDGER

Central Path Proximal Policy Optimization

As of 16 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 3 inbound Pith citation observations for arXiv:2506.00700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00700 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:07:09.566879Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:23:46.697427Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T11:35:19.078046Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 074502be-9dda-46de-81e9-5d3b8d32d5fd · outbound

This paper cites A safe exploration approach to constrained Markov decision processes.

Central Path Proximal Policy Optimization A safe exploration approach to constrained Markov decision processes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:07:10.495244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:07:08.931227Z digest=sha256:438befaa20810eafa66b59226491b73b97024bb5f1316eb96016277ddfe15ecc

Observation 78af1034-67c6-4a9f-8fd3-24376f09bebb · outbound

This paper cites Direct Behavior Specification via Constrained Reinforcement Learning.

Central Path Proximal Policy Optimization Direct Behavior Specification via Constrained Reinforcement Learning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:07:10.187120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:07:09.093230Z digest=sha256:a20588d7dfb73e1186c73c3d58fee872be22a34fdb3f814da1beadb163e05783

Observation 4faecdac-4fa9-411c-a493-11ed459e1bb1 · outbound

This paper cites Responsive Safety in Reinforcement Learning by PID Lagrangian Methods.

Central Path Proximal Policy Optimization Responsive Safety in Reinforcement Learning by PID Lagrangian Methods

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:07:09.904564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:07:09.280854Z digest=sha256:eeb3c6470203bfeef5f0362751913eb4e3bce0e2f0d486110427e0be21277a8d

Observation d55bf51f-2cc4-4f0f-8774-b0fda9159b60 · outbound

This paper cites Constrained Reinforcement Learning with Smoothed Log Barrier Function.

Central Path Proximal Policy Optimization Constrained Reinforcement Learning with Smoothed Log Barrier Function

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:09.385071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:09.385071Z digest=sha256:d4f452a6b741cc507f0622630e5a34dd1b1298703225dde9d3fca87a81ce3936

Observation d3628220-eedf-49dc-a3a3-71f7824ecd61 · outbound

This paper cites Yiming Zhang, Quan Vuong, and Keith Ross.

Central Path Proximal Policy Optimization Yiming Zhang, Quan Vuong, and Keith Ross

Reference 13

Resolution
verified exact
doi, observed 2026-08-07T12:07:09.751998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:07:09.480957Z digest=sha256:28811e659e9448d4bdee4cae28f45f797574df2f54d3a8f3bb6be0f0dff12d75

Observation 5d638274-83d0-44c6-8aef-e4d9f5cc284b · outbound

This paper cites surrogate advantage trick.

Central Path Proximal Policy Optimization surrogate advantage trick

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:07:10.352342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:07:09.566879Z digest=sha256:7e7aa3c469277171a09e03856aaa7d6160635e9466d9b466b84dd4fbe5bbb886

Observation da807287-27f3-458c-99d7-ddad26648efb · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Central Path Proximal Policy Optimization High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:09.141595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:09.141595Z digest=sha256:21e196240cc08c5ab83c0bf58ff300b8a573ae6af0a33fac74f1fd968e58300d

Observation d1d34400-fb63-46a9-b893-1e669deacd40 · outbound

This paper cites Dimitri P Bertsekas.

Central Path Proximal Policy Optimization Dimitri P Bertsekas

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:07:10.664488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:07:08.578032Z digest=sha256:28b02ac8b0f2ff431510af084487ad24e113c07709f521ab9be5a52a0a718744

Observation 3f78ce86-e4d7-4720-8ad6-c368ea7bfc38 · outbound

This paper cites Lyapunov-based Safe Policy Optimization for Continuous Control.

Central Path Proximal Policy Optimization Lyapunov-based Safe Policy Optimization for Continuous Control

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.633497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.633497Z digest=sha256:d5c9eba0d7f8de704423686063dbdd61da54036b0d63edf5e45440177dba1d0a

Observation 2f5a27ad-336f-436e-9782-1670098eb2d9 · outbound

This paper cites Jincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li, Csaba Szepesvári, and Dale Schuurmans.

Central Path Proximal Policy Optimization Jincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li, Csaba Szepesvári, and Dale Schuurmans

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.702350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.702350Z digest=sha256:df6ca4873ecbb36655c8047a4772515d39b7e559b5bb408296ad1057a6ddfb28

Observation 6ba153c4-cc06-4dc8-bf5c-d9d8f999b133 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Central Path Proximal Policy Optimization Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.990567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.990567Z digest=sha256:8dddad479b3d0ab931e94088e99f542d642ca29fc454f01dd8c839857395f7d6

Observation 4b21ca17-83c7-407c-bef8-f3c3fc4823e1 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Central Path Proximal Policy Optimization A unified view of entropy-regularized Markov decision processes

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.844809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.844809Z digest=sha256:807d3a7d8422860ec8b51d2bea345c3959bfde4f8774221d0facf53bd6337f0e

Observation 9eb6c714-6dfd-4a21-b04b-ee6f1af7c67d · outbound

This paper cites On PI Controllers for Updating Lagrange Multipliers in Constrained Optimization.

Central Path Proximal Policy Optimization On PI Controllers for Updating Lagrange Multipliers in Constrained Optimization

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:07:10.017454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:07:09.219280Z digest=sha256:8a838bffdd593233b9010ba87ab4e2aad5d580bdc87c91b7c6764567e1510fd4

Observation fc27e03e-c0c5-4d02-a012-6a3d05bf8902 · outbound

This paper cites Embedding Safety into RL: A New Take on Trust Region Methods.

Central Path Proximal Policy Optimization Embedding Safety into RL: A New Take on Trust Region Methods

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.760132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.760132Z digest=sha256:5250859f3d98626d3443d427d0b8e78e960e992e50c44321473bd9a34c986c12

Pith citing papers

Observation 22fa1a46-a0d1-460c-b1d8-0af62d3c0ded · inbound

Symbiotic Agents: A Novel Paradigm for Trustworthy AGI-driven Networks cites this paper.

Symbiotic Agents: A Novel Paradigm for Trustworthy AGI-driven Networks Central Path Proximal Policy Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T18:23:46.697427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:23:46.697427Z digest=sha256:488f2056ac2b34f5abba7c570182635bf362f40081df44b724c684c50a3a60d0

Observation 7a1ef6a6-4859-4cd6-84de-2a19d6439ba4 · inbound

The Geometry of Nonlinear Reinforcement Learning cites this paper.

The Geometry of Nonlinear Reinforcement Learning Central Path Proximal Policy Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.630698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.630698Z digest=sha256:611cd0ac0bcd1711467d28c5dba8604c03c8fdde83d3df327af284ff79852521

Observation d095a95e-6a3c-48b7-817c-c5acf95f3998 · inbound

Bounded Ratio Reinforcement Learning cites this paper.

Bounded Ratio Reinforcement Learning Central Path Proximal Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:19.079386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T04:50:11.020901Z digest=sha256:8e809072f51ba52465ac650992c48fa54ae81a100fdb12a7431c0a2c333e252b