Pith. sign in

Paper Citation Record · LEDGER

Process-Supervised Reinforcement Learning for Code Generation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2502.01715.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01715 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:31:54.286136Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T12:44:40.230064Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 18e4674b-e365-46e7-8ec7-e59fadbb27af · inbound

Improving LLM-Generated Code Quality with GRPO cites this paper.

Improving LLM-Generated Code Quality with GRPO Process-Supervised Reinforcement Learning for Code Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:54.286136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:54.286136Z digest=sha256:bfc04304ab30462118dd36ec38724d772bde3893fc1229826a54ad71d5e16d54

Observation aaffbf05-6329-479d-bf10-f10098639611 · inbound

Reinforcement Learning in hyperbolic space for multi-step reasoning cites this paper.

Reinforcement Learning in hyperbolic space for multi-step reasoning Process-Supervised Reinforcement Learning for Code Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.819837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.819837Z digest=sha256:4b14b27f54215ec39d4fb13d7db7efe90e49414725473daa1d7ac34399cecbad

Observation 1ebf6c85-9541-44f4-ba5d-45ae9cf8dc64 · inbound

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs cites this paper.

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs Process-Supervised Reinforcement Learning for Code Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:38.089096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:27:38.089096Z digest=sha256:ef77f089f27ab72c1d28f2c91b2fa5e28355fb40fda374dee4f613326d6f20e4

Observation 862d71d8-3201-45e2-a169-e6611ec11551 · inbound

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation cites this paper.

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation Process-Supervised Reinforcement Learning for Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:35.177823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:19:35.177823Z digest=sha256:e69d498ac1fd9939a0a19f54b9cd535f8ac8e46006cc83a4732976bf016e6af7

Observation e0374b31-8d81-4623-8e1e-6d44d37a3bfb · inbound

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning cites this paper.

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning Process-Supervised Reinforcement Learning for Code Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:38:18.788744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T21:36:14.478007Z digest=sha256:13f548eacc76739cd3f76c9d719f4e03ecf19ea3ef53542291316a1b0836bdcb

Observation 6a3feea1-6dfe-49be-aa00-6b9d94158540 · inbound

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs cites this paper.

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Process-Supervised Reinforcement Learning for Code Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:45:59.984342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:58:10.013475Z digest=sha256:618b028a46f06a98e090e31e03923a0478a38eba25dde291a5b66af2cf7bdce9

Observation cd0ae5ae-5a2c-43c2-abbf-20c4a72c161e · inbound

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning cites this paper.

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning Process-Supervised Reinforcement Learning for Code Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:01:00.233468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:51:19.555272Z digest=sha256:95fb4404d136f5aa2c4a3eb30eac42dc6d4feab87da94d022c12849c4cb72326

Observation 1d03fe18-37d5-47bd-8db7-482b1e88f803 · inbound

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling cites this paper.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Process-Supervised Reinforcement Learning for Code Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.288936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:65e333c1c14bf0c24b462d194068a97383ab4339a4d65105196827ebedf8bc62

Observation fae0bdb1-6d17-4048-8adb-07a787463b99 · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Process-Supervised Reinforcement Learning for Code Generation

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:40.231715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:de69be60fa5497493912c164c0256107c3b3b54117b2ab900db8617226397b7e