Pith. sign in

Paper Citation Record · LEDGER

Behavior Proximal Policy Optimization

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2302.11312.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.11312 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:56.424024Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5068b278-5f7a-45ef-94e6-75e76d93dd16 · inbound

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL cites this paper.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Behavior Proximal Policy Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:56.424024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:56.424024Z digest=sha256:68de992859ba2454c2992487d971f0ee36e7b93723b256fc029c59c543296c1f

Observation 2de92e6a-ec70-4159-9a24-cf51af5f8e47 · inbound

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization cites this paper.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Behavior Proximal Policy Optimization

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.085023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.085023Z digest=sha256:859bc6a10d5895e4af06bea80e970eb053ff98f836389cc405668e76513d62c3

Observation ae06fb0f-9b78-4722-af00-df34d5711fd4 · inbound

Mitigating Data Scarcity in Spaceflight Applications for Offline Reinforcement Learning Using Physics-Informed Deep Generative Models cites this paper.

Mitigating Data Scarcity in Spaceflight Applications for Offline Reinforcement Learning Using Physics-Informed Deep Generative Models Behavior Proximal Policy Optimization

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:38:18.462367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T21:35:52.012244Z digest=sha256:e7b202c5e8f344fa84223934446abbdd45c8f2764aea9e9d8d4421f276866eaf

Observation b8b72fd1-6758-4061-b1cc-1b21a4a45727 · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Behavior Proximal Policy Optimization

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.036585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:07c640b78921f3239490f314295dc06978d49d186e645a547905237d09d3e7d4