Pith. sign in

Paper Citation Record · LEDGER

Central Path Proximal Policy Optimization

As of 9 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 2 inbound Pith citation observations for arXiv:2506.00700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00700 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:07:09.566879Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:37:41.630698Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T11:35:19.078046Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 074502be-9dda-46de-81e9-5d3b8d32d5fd · outbound

This paper cites A safe exploration approach to constrained Markov decision processes.

Central Path Proximal Policy Optimization A safe exploration approach to constrained Markov decision processes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:07:10.495244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:07:08.931227Z digest=sha256:9dedc67af33c71e871d46db45d2ec5073d059a20a8e68c614541c1500471111c

Observation 78af1034-67c6-4a9f-8fd3-24376f09bebb · outbound

This paper cites Direct Behavior Specification via Constrained Reinforcement Learning.

Central Path Proximal Policy Optimization Direct Behavior Specification via Constrained Reinforcement Learning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:07:10.187120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:07:09.093230Z digest=sha256:dca07c666ee91468429aa6ae5bfca1546004ac1f982144b22e0eb4ec66e06bb1

Observation 4faecdac-4fa9-411c-a493-11ed459e1bb1 · outbound

This paper cites Responsive Safety in Reinforcement Learning by PID Lagrangian Methods.

Central Path Proximal Policy Optimization Responsive Safety in Reinforcement Learning by PID Lagrangian Methods

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:07:09.904564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:07:09.280854Z digest=sha256:c909d3c0fe834b70e405bd05c4e1a22ac24e9fc47d27265ec4e5e67e4ab2a79e

Observation d55bf51f-2cc4-4f0f-8774-b0fda9159b60 · outbound

This paper cites Constrained Reinforcement Learning with Smoothed Log Barrier Function.

Central Path Proximal Policy Optimization Constrained Reinforcement Learning with Smoothed Log Barrier Function

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:09.385071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:09.385071Z digest=sha256:fec7d214fab1edb5608007c12bf6f89bd6b42d7ae15e172a274db07d139c35f6

Observation d3628220-eedf-49dc-a3a3-71f7824ecd61 · outbound

This paper cites Yiming Zhang, Quan Vuong, and Keith Ross.

Central Path Proximal Policy Optimization Yiming Zhang, Quan Vuong, and Keith Ross

Reference 13

Resolution
verified exact
doi, observed 2026-08-07T12:07:09.751998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:07:09.480957Z digest=sha256:e98e49dec5bf1f602e968577aed75a0f39031a26a5f3f2da83400493387bf466

Observation 5d638274-83d0-44c6-8aef-e4d9f5cc284b · outbound

This paper cites surrogate advantage trick.

Central Path Proximal Policy Optimization surrogate advantage trick

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:07:10.352342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:07:09.566879Z digest=sha256:2543f791bde45650290b4526169a9222f8eac7ddecbbe2b4f90c013773143319

Observation da807287-27f3-458c-99d7-ddad26648efb · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Central Path Proximal Policy Optimization High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:09.141595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:09.141595Z digest=sha256:6ad0f8df5e820cbcf1f57329398d340e7664f0bd3f0530a5397e29d1b98332bf

Observation d1d34400-fb63-46a9-b893-1e669deacd40 · outbound

This paper cites Dimitri P Bertsekas.

Central Path Proximal Policy Optimization Dimitri P Bertsekas

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:07:10.664488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:07:08.578032Z digest=sha256:a073621ac5bf283124afe570ede8059960f74755c6384f1c1bcd5f273f843e2c

Observation 3f78ce86-e4d7-4720-8ad6-c368ea7bfc38 · outbound

This paper cites Lyapunov-based Safe Policy Optimization for Continuous Control.

Central Path Proximal Policy Optimization Lyapunov-based Safe Policy Optimization for Continuous Control

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.633497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.633497Z digest=sha256:cac8c99bc623155d1c6d8462970d7bb9ea1eaed32844aef9637923e86d0befae

Observation 2f5a27ad-336f-436e-9782-1670098eb2d9 · outbound

This paper cites Jincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li, Csaba Szepesvári, and Dale Schuurmans.

Central Path Proximal Policy Optimization Jincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li, Csaba Szepesvári, and Dale Schuurmans

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.702350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.702350Z digest=sha256:fd9dfee0def550ed3599701faafed24b4d9ae10c8de255734bdc6c51d05f18a2

Observation 6ba153c4-cc06-4dc8-bf5c-d9d8f999b133 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Central Path Proximal Policy Optimization Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.990567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.990567Z digest=sha256:cf684329abb4c5ff227d5fc74ea6f102c752719971a302c6de50e332192b9c34

Observation 4b21ca17-83c7-407c-bef8-f3c3fc4823e1 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Central Path Proximal Policy Optimization A unified view of entropy-regularized Markov decision processes

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.844809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.844809Z digest=sha256:958d4dfccd725d775901a95e9bd7d54ac168380a3585f62c649b766b24bf0ccb

Observation 9eb6c714-6dfd-4a21-b04b-ee6f1af7c67d · outbound

This paper cites On PI Controllers for Updating Lagrange Multipliers in Constrained Optimization.

Central Path Proximal Policy Optimization On PI Controllers for Updating Lagrange Multipliers in Constrained Optimization

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:07:10.017454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:07:09.219280Z digest=sha256:ed4f30a24cf82fb2640339ab69a3b81d9c1082867f01b0a6938a9953329ef7f8

Observation fc27e03e-c0c5-4d02-a012-6a3d05bf8902 · outbound

This paper cites Embedding Safety into RL: A New Take on Trust Region Methods.

Central Path Proximal Policy Optimization Embedding Safety into RL: A New Take on Trust Region Methods

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.760132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.760132Z digest=sha256:ccf311a88d702fe8e61c4f05a10ef31ad04d01ccad12f211f642d77679c32a14

Pith citing papers

Observation 7a1ef6a6-4859-4cd6-84de-2a19d6439ba4 · inbound

The Geometry of Nonlinear Reinforcement Learning cites this paper.

The Geometry of Nonlinear Reinforcement Learning Central Path Proximal Policy Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:41.630698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:41.630698Z digest=sha256:b23a96363d10d2fb35c684c696aafe1ac5daa4ff1b9f4e071ff481ca315fdd29

Observation d095a95e-6a3c-48b7-817c-c5acf95f3998 · inbound

Bounded Ratio Reinforcement Learning cites this paper.

Bounded Ratio Reinforcement Learning Central Path Proximal Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:19.079386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:50:11.020901Z digest=sha256:a3674095411b6bbe52cd1ca59d23399b77dc8bb8fb1d6eb129bc4c58f38b1572