Pith. sign in

Paper Citation Record · LEDGER

Process Supervision-Guided Policy Optimization for Code Generation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2410.17621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.17621 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:28.131857Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T14:43:30.588160Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d5201cf0-82e2-40c3-9b48-2f0f7f9a2223 · inbound

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review cites this paper.

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review Process Supervision-Guided Policy Optimization for Code Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:28.131857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:28.131857Z digest=sha256:4b96b337af9743fd5aa92ee667ef95eb39c9460d779a18f7722bca5242c07faa

Observation 70f3149e-36fb-4255-b48f-bdcf2e5f911e · inbound

Improving LLM-Generated Code Quality with GRPO cites this paper.

Improving LLM-Generated Code Quality with GRPO Process Supervision-Guided Policy Optimization for Code Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.155905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.155905Z digest=sha256:f52d249d6d7c1fe268ed0241dd166864ac2d518c1b219eea1e72313145adf39a

Observation ad068b07-0870-4b6b-bb7b-5f3be6dc076e · inbound

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation cites this paper.

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation Process Supervision-Guided Policy Optimization for Code Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:31.060590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:19:31.060590Z digest=sha256:71ae6022c047dccb76612f272d0d9f08faaf3720c4e55606803a0813b076a493

Observation c8489a72-52e3-43f3-804e-e60e61562a26 · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning Process Supervision-Guided Policy Optimization for Code Generation

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:53.874104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:02581c452994e9154b4c92851eb6d3d19719b35c4a2039e0f619863472cf9766

Observation 6fe221c3-b9fd-452c-8fef-c02ef0aaaf71 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Process Supervision-Guided Policy Optimization for Code Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:29.990844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:f82376a108a5b53852b28ed8b0b081e4049c2b05f865665c9e64de44d7a55825

Observation 2d2dfbc7-05d3-4832-a588-d3cf1a98149c · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Process Supervision-Guided Policy Optimization for Code Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:16.491340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:788ae479542e308f65bbc841d4a0f8df4ccc4d4cf47955185e3013da5cd29909

Observation 17bf7669-c6b4-422c-8398-920bf9b1bfa3 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Process Supervision-Guided Policy Optimization for Code Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:02:40.813856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:443a7d94436d01036bbb04ca7471f2e20c2f2ff8996b4e2d0c13af04e2dba639

Observation a85312dc-57a2-4f30-842f-1112c4493d91 · inbound

Process Rewards with Learned Reliability cites this paper.

Process Rewards with Learned Reliability Process Supervision-Guided Policy Optimization for Code Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:53:06.948837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T14:51:13.966538Z digest=sha256:104ef7ba07ec134ab5f7083a3e0d4f4d8118be188136363c3cc7b34f2fbce5d7

Observation 97fb297e-54e8-4f69-b0cc-ea5613685052 · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding Process Supervision-Guided Policy Optimization for Code Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.831367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:025ae24641a592994418c034f67350a075e332cea44fa98abaeb4700630bf8e0

Observation 558d2a0e-7a17-4c1b-8eb0-2cb63a278b6e · inbound

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL cites this paper.

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL Process Supervision-Guided Policy Optimization for Code Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:43:30.589913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T14:41:13.191919Z digest=sha256:46493978b1f586eba4a99814fbdb46e99c76b6559976d92c4ffbd080287e379e