Pith. sign in

Paper Citation Record · LEDGER

Process-Supervised Reinforcement Learning for Code Generation

As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 9 inbound Pith citation observations for arXiv:2502.01715.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01715 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:10:39.001050Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:31:54.286136Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T12:44:40.230064Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44db55d9-7fb6-4803-94ee-e4ec4ec07176 · outbound

This paper cites GPT-4 Technical Report.

Process-Supervised Reinforcement Learning for Code Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.746645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.746645Z digest=sha256:dce865145d4c7fb42c06f691685d7d0040486544e7ecae112ad9665cd7920a22

Observation 0056b658-88b8-42cb-8aff-613fa78adb8a · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.751842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.751842Z digest=sha256:179ea1c671ca4a8cfd5450b2eea71b877b9261c7df5f0b73a28fcafcca8c0a5f

Observation 6c32569f-3a9d-4705-8abf-cf0b3e87ab48 · outbound

This paper cites Concrete Problems in AI Safety.

Process-Supervised Reinforcement Learning for Code Generation Concrete Problems in AI Safety

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.757164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.757164Z digest=sha256:958bc0b3fab74855ca445c33505450e36b151c9c495ffeeb5bab031d7ac23a17

Observation 20a2211f-d9c2-46ed-8157-fb84d8f6dae9 · outbound

This paper cites Program Synthesis with Large Language Models.

Process-Supervised Reinforcement Learning for Code Generation Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.762277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.762277Z digest=sha256:251a8cbf23a6416794e24f5e28c4b3f0ad23e9cb9a0c4e5f40ffb781e7f24320

Observation c367e2ca-b73d-4664-a06f-075c3c6e5282 · outbound

This paper cites Language Models are Few-Shot Learners.

Process-Supervised Reinforcement Learning for Code Generation Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.767129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.767129Z digest=sha256:85e3c4ddb4775a5e55ce6729277656058f7e8996ab31ea449600cedce22a2fca

Observation 81b336dd-1a86-4be9-b198-ad2b53e51433 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Process-Supervised Reinforcement Learning for Code Generation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.772031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.772031Z digest=sha256:29ff33c52d2811f5e97e48fc519761a728cc4d002b4ecac9cea0db766fbd7688

Observation 11ed19ae-6dd5-4a38-9bfb-088af0c68af3 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Process-Supervised Reinforcement Learning for Code Generation Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.777654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.777654Z digest=sha256:60f03f8c863d1ac4df25b32c973697c1b0162ec9bbd5188947202925dbef9094

Observation 9cb63d96-72bc-49c9-bfb8-cbba843e0a92 · outbound

This paper cites Process Supervision-Guided Policy Optimization for Code Generation.

Process-Supervised Reinforcement Learning for Code Generation Process Supervision-Guided Policy Optimization for Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.782468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.782468Z digest=sha256:d7478ddb5c9b94d4dca9cd3d1abbaa9c2dd037cba9bec7452692da3cd5207b15

Observation 763fd09c-214a-4f75-8455-fa3a916adb48 · outbound

This paper cites Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access.

Process-Supervised Reinforcement Learning for Code Generation Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.787613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.787613Z digest=sha256:6b263b15d242c5610916c0026428d5ab501de032b7e3aadca4f6179d1e2b8e70

Observation 012cf20e-210f-4178-b988-d7b3b102023c · outbound

This paper cites The Llama 3 Herd of Models.

Process-Supervised Reinforcement Learning for Code Generation The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.792717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.792717Z digest=sha256:3e166f4abc448dd1a2d1194641512b718f42779869800da8cb5c86c6798b4da5

Observation 7b976b9b-1e7e-4b0a-b036-cab8504a3c38 · outbound

This paper cites CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

Process-Supervised Reinforcement Learning for Code Generation CodeBERT: A Pre-Trained Model for Programming and Natural Languages

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.797764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.797764Z digest=sha256:214b784ff5b9b8872dc02479b189869b751d5945fe837394eca999ed1556668e

Observation 477ad882-2a87-40b3-bdf6-9882e2eac1a4 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.802256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.802256Z digest=sha256:51267dd56d068e9cd649a4000b6ec85ca5339431ca0fceccee69b7be9b90926e

Observation 38f56300-3a0c-407c-9728-b45cdc92ef5e · outbound

This paper cites UniXcoder: Unified Cross-Modal Pre-training for Code Representation.

Process-Supervised Reinforcement Learning for Code Generation UniXcoder: Unified Cross-Modal Pre-training for Code Representation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.807462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.807462Z digest=sha256:be7a93532c861a2e9c4a9f08df54a9c435b73c381d6c0dc419df1d7b169d2115

Observation 06359d2f-5ce0-4aab-8cf5-1c94af8c2df3 · outbound

This paper cites GraphCodeBERT: Pre-training Code Representations with Data Flow.

Process-Supervised Reinforcement Learning for Code Generation GraphCodeBERT: Pre-training Code Representations with Data Flow

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.812077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.812077Z digest=sha256:e5f642518cff81ebc7144a61a25c566603792568813933953a1c60929f211ad5

Observation bbab7549-6628-4783-8e7e-9fd8ec298cdd · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Process-Supervised Reinforcement Learning for Code Generation DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.816815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.816815Z digest=sha256:9d13c376deb6e4903427f74b5b26f2868bbdabc11af037debd8acf3afc0b9273

Observation 15a9ec98-f86f-4d15-b564-d810cf708beb · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-09T15:10:39.783815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T15:10:38.821485Z digest=sha256:f2648c6ec17bf25657fc290b0a056d227c5fc4d5acf22407c3264fbe300c890e

Observation e7c9a691-87fa-4cf0-b645-0c8c88e73e3c · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-09T15:10:39.767454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T15:10:38.825981Z digest=sha256:6a19bb324b8aafdff8789c0eb8dd45a6933f35419d6492b8a34aaef20bfe1e6b

Observation 7d60ab94-c0e6-4cd8-ba32-9e9591c08c3b · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Process-Supervised Reinforcement Learning for Code Generation A Survey on Large Language Models for Code Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.830590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.830590Z digest=sha256:01dd24fef176827691b874a83ed193ea2146fa7752ebd4a41507a70bef4423d9

Observation 2d0c8187-abce-4049-8641-ef4c0fcd7ff7 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.835716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.835716Z digest=sha256:190f5c9c7da5a8730188c4a04088eab380c6c84bb579be0b01acdffe93904222

Observation ec13d256-b9ae-4e69-b018-ec4e62afcfd8 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-09T15:10:39.741845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T15:10:38.840331Z digest=sha256:20d85bd0cec2d7151adce75110326bdcdbc87d5cd1ada1fb63dc1aebe96562c1

Observation 8dbd8ac3-f10b-4ffe-a237-a13a61f4b112 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Process-Supervised Reinforcement Learning for Code Generation RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.844627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.844627Z digest=sha256:a3c81570437c406dcd68a9ea257ab66e7ae01e3b141a7afd10e5d5d8867e4e8d

Observation d0aa659f-55d3-41eb-a76e-edc90acd3175 · outbound

This paper cites StarCoder: may the source be with you!.

Process-Supervised Reinforcement Learning for Code Generation StarCoder: may the source be with you!

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.849701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.849701Z digest=sha256:fea82751550af352bf9cb6a240b127f799f03ac5c029bee3f45bbca39be40363

Observation 54c0398b-cf32-4f36-a34d-92907d7de799 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.854560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.854560Z digest=sha256:ec0bbc5bffbdd75d526421712c72984ce57a684c76561c4922d7cec234928982

Observation 777d1184-b482-4e03-8384-2d3dddb99d23 · outbound

This paper cites Let's Verify Step by Step.

Process-Supervised Reinforcement Learning for Code Generation Let's Verify Step by Step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.859562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.859562Z digest=sha256:e087dcf9865155622bf5918cceee91c2450bfc8419d1f9ba7285f80f54412780

Observation bf02f726-c214-480a-b365-4b18e1a0c4ea · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-09T15:10:39.716141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T15:10:38.864580Z digest=sha256:8cf58468009e63f2bdc277a05270f956cd2dabcabfda6b8830d40b49635875ae

Observation 399b8bb8-37a4-4c0a-aa2f-376e220713c2 · outbound

This paper cites RLTF: Reinforcement Learning from Unit Test Feedback.

Process-Supervised Reinforcement Learning for Code Generation RLTF: Reinforcement Learning from Unit Test Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.868972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.868972Z digest=sha256:fc63f0a8ce9d0341b3b5afc3718d3271aabdbd46468fb95a7d1d2f200cb06bc3

Observation da135484-570d-4fc5-b66b-a249d7f9f50f · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Process-Supervised Reinforcement Learning for Code Generation Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.874401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.874401Z digest=sha256:b2aad76d3db85d3b250616fa458ce165ed47567196e494b9ec87cc239fc02218

Observation 54fc7869-4a4c-4613-8f6a-c6c38907ce13 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

Process-Supervised Reinforcement Learning for Code Generation Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.879367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.879367Z digest=sha256:77f47b7c468b2fb5b7f040419a0fec0b3d7316230e6582284cf5bced507bd497

Observation 0aedf4ac-34bd-41ce-b874-7a9d7fb118cc · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Process-Supervised Reinforcement Learning for Code Generation Playing Atari with Deep Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.884186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.884186Z digest=sha256:50663e26376a35c642829ef842c9307c10bd2ffaa004f89c3daa0683c147f99a

Observation bd164991-ce4a-49c2-b077-7b6446eb505f · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.889088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.889088Z digest=sha256:250a00c30e035f88b7c832dd6f3b757ef4fc67dabca6850dd21f62485b40c3dc

Observation a5cf2432-338a-4442-be1d-5965b8d2d9a3 · outbound

This paper cites CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis.

Process-Supervised Reinforcement Learning for Code Generation CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.893567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.893567Z digest=sha256:6a629742cd7371a45f94c0dbc1a5ce90b3c095d807b5f9c2ad230b6fc2863a14

Observation b48f425f-3241-492f-b801-20288909d7d8 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.898254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.898254Z digest=sha256:bcfdcf3d123a73820079832b847f84cd97dc3a1ed3a77843499f6b7f05e37f7e

Observation 9534e56e-ce12-4d14-934f-789adc462854 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.903595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.903595Z digest=sha256:221498ce6669ac5e0cdf897bc1fc641a1422bd11ef25a6840e57e6e7f1598415

Observation 7519ed64-db05-4ee2-9168-3ce1d56b98da · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

Process-Supervised Reinforcement Learning for Code Generation CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.908142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.908142Z digest=sha256:dc5bf05c726b1996cb2ad7941410f03638ffb1b411f1aa1dfabe6c7455dd4481

Observation 9dd5e846-051b-45f2-ae80-93322471ba29 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.913160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.913160Z digest=sha256:7f37fda0084e46e1d772af5b2e5c1c0c3df719277dff7f632913c92815ff7cac

Observation b22b0fb4-2ea6-45f7-9200-0afb4b0bd28a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Process-Supervised Reinforcement Learning for Code Generation Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.917519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.917519Z digest=sha256:4475673a9724e14c832e0ffa07b200404ffc709f25d825794c90b414ccc7df57

Observation 3e62896b-58e7-403e-bc94-5d8f11da6227 · outbound

This paper cites Execution-based Code Generation using Deep Reinforcement Learning.

Process-Supervised Reinforcement Learning for Code Generation Execution-based Code Generation using Deep Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.922030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.922030Z digest=sha256:18b6c91a1ded240aa666ecc39c18ea99436dba197b0b80711847095fc6348444

Observation 54915dd4-887f-4690-9423-85276bffe2ee · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Process-Supervised Reinforcement Learning for Code Generation LLaMA: Open and Efficient Foundation Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.926569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.926569Z digest=sha256:77d091f3518d2650cb7d63b1abda6d85fef556181f415192d673edc06fee7321

Observation 4625abeb-b736-42cf-8643-24031b2b9917 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Process-Supervised Reinforcement Learning for Code Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.930919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.930919Z digest=sha256:b42cf0c6ee6072cd963c129899929afcb1d163be9203520a5200a461dabf404c

Observation f9f01a97-994a-4618-afa8-76c2534bdf8e · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Process-Supervised Reinforcement Learning for Code Generation Solving math word problems with process- and outcome-based feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.935579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.935579Z digest=sha256:65670a354e03d2a524177c0def8d11c930d5c60f7a1aa5fc7deb3438df21997c

Observation 81c24362-4f0a-4cdb-9134-116180b7db97 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-09T15:10:39.661393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T15:10:38.940441Z digest=sha256:a769dbee34c0d3b15d94f1b5d1267d1976cc6e52592cff2574ea261645be1057

Observation 8454ba4b-efd8-41b4-840c-51b2d9f2ff55 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.944659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.944659Z digest=sha256:1f5ac1c05f093316b297974f1fe94281f18aac517d79e14ffe8a36f210dbdb87

Observation 247ed4be-1511-457e-b6a0-261ed9b59708 · outbound

This paper cites Compilable Neural Code Generation with Compiler Feedback.

Process-Supervised Reinforcement Learning for Code Generation Compilable Neural Code Generation with Compiler Feedback

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-09T15:10:39.157183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T15:10:38.949102Z digest=sha256:a0cae53bdc9e5e346a2451b6153151b1a83cab6a353e893a3af58cf6385c58f1

Observation 0cc9f4bb-5586-44e6-9fab-b4b98adca092 · outbound

This paper cites CodeT5+: Open Code Large Language Models for Code Understanding and Generation.

Process-Supervised Reinforcement Learning for Code Generation CodeT5+: Open Code Large Language Models for Code Understanding and Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.953669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.953669Z digest=sha256:8fe700f4da800c04b4603dfee28afa055679a37d7c8a03442c9808158c045524

Observation 626d8ee4-88f9-42ef-92a6-01ac4b2928b1 · outbound

This paper cites CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation.

Process-Supervised Reinforcement Learning for Code Generation CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.958717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.958717Z digest=sha256:854f8c36ffcf45fe9f681d0fb4fff65d597f6e5a4dcfca86240f1ec9f27c9727

Observation 054dac79-5ce3-4e3c-8023-ddb7455b472f · outbound

This paper cites Hybrid Code Networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning.

Process-Supervised Reinforcement Learning for Code Generation Hybrid Code Networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-09T15:10:39.104213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T15:10:38.963219Z digest=sha256:66baf74b609270d990376fb1c8d69ee8becf0284324bfe70a4eccdddbb13309e

Observation b102dee4-1e04-41d4-8eb8-727ec912214f · outbound

This paper cites Recursively Summarizing Books with Human Feedback.

Process-Supervised Reinforcement Learning for Code Generation Recursively Summarizing Books with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.967859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.967859Z digest=sha256:8acfc606ec8a1113ed175b11688fdfb26d5bd6bbf3585fecf10bef86e6d8cd2d

Observation 068badc7-02e3-4791-a5a0-d8c04c9de420 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-09T15:10:39.635553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T15:10:38.972239Z digest=sha256:2ec55fc46ab2f14784d1f39cb70d990d3e22cf999dcf882e01d2a6f8e8657a63

Observation 88e70d61-df5f-4499-8ba5-0301212479d5 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-09T15:10:39.618283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T15:10:38.976377Z digest=sha256:e46eee5a4e4490e958cc7df6eccc8b334b0efd020c60916a6c514d77e6f24a40

Observation a0fe9872-cd95-4a06-aca8-11e1f8c28c74 · outbound

This paper cites Large Language Models Meet NL2Code: A Survey.

Process-Supervised Reinforcement Learning for Code Generation Large Language Models Meet NL2Code: A Survey

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.981032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.981032Z digest=sha256:3ff93d6b863607e0871c92c47f668acba680eecb515b6e9c2a9ed7c7b7f15551

Observation fb166f4f-8517-4990-9242-8b043d6b6f61 · outbound

This paper cites an unresolved cited work.

Process-Supervised Reinforcement Learning for Code Generation Unresolved cited work

Reference 51

Resolution
verified exact
doi, observed 2026-08-09T15:10:39.039045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T15:10:38.986407Z digest=sha256:1a1fa6b5af259140e340c514cb6bcbba1e2262911abf72c0de2f6061599ac590

Observation a443e119-8987-407d-b456-34cff9e5173a · outbound

This paper cites Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code.

Process-Supervised Reinforcement Learning for Code Generation Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.991652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.991652Z digest=sha256:d78d0776fed7b4407579085ab54f9cceca37ca021495d5db37ccf640d299bed2

Observation 2c9d9588-16f3-47b8-af7f-3226fbbdbd79 · outbound

This paper cites online" 'onlinestring :=.

Process-Supervised Reinforcement Learning for Code Generation online" 'onlinestring :=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.996261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.996261Z digest=sha256:5d47c7b58ba4542e3fbe74a99a270f3316ffbc5e1c7b7968f8fc00b76bafc6d2

Observation 9d8e089c-8dda-4b85-9f4c-c38fd3e6e9a2 · outbound

This paper cites write newline.

Process-Supervised Reinforcement Learning for Code Generation write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:39.001050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:39.001050Z digest=sha256:681b1cf081f6bb1ed766718e428e94c243aca161156480b475fd44d2fe23be18

Pith citing papers

Observation 18e4674b-e365-46e7-8ec7-e59fadbb27af · inbound

Improving LLM-Generated Code Quality with GRPO cites this paper.

Improving LLM-Generated Code Quality with GRPO Process-Supervised Reinforcement Learning for Code Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:54.286136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:54.286136Z digest=sha256:a2826cf721c946335cdb8db10f42500dc3c075e4475f19d8c41f801f1360263a

Observation aaffbf05-6329-479d-bf10-f10098639611 · inbound

Reinforcement Learning in hyperbolic space for multi-step reasoning cites this paper.

Reinforcement Learning in hyperbolic space for multi-step reasoning Process-Supervised Reinforcement Learning for Code Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.819837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.819837Z digest=sha256:12d30339e6eaac91e43eccc00e972900a74b7d6f836ecde3c59b2a5e78b95f02

Observation 1ebf6c85-9541-44f4-ba5d-45ae9cf8dc64 · inbound

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs cites this paper.

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs Process-Supervised Reinforcement Learning for Code Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:38.089096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:27:38.089096Z digest=sha256:6fffbfba4eac526c3fbc6549f4a0cdaff2e46193cb31b88fd45ece7b614d4156

Observation 862d71d8-3201-45e2-a169-e6611ec11551 · inbound

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation cites this paper.

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation Process-Supervised Reinforcement Learning for Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:35.177823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:19:35.177823Z digest=sha256:eb798615f0a9dc504e8a7c7f568b0f764c66a368d3f1119045c8181eb03862b4

Observation e0374b31-8d81-4623-8e1e-6d44d37a3bfb · inbound

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning cites this paper.

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning Process-Supervised Reinforcement Learning for Code Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:38:18.788744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T21:36:14.478007Z digest=sha256:dbbbf2af77b922084c87dcb41c99bbac97b7499630643bf2048b5a3272b6158c

Observation 6a3feea1-6dfe-49be-aa00-6b9d94158540 · inbound

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs cites this paper.

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Process-Supervised Reinforcement Learning for Code Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:45:59.984342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:58:10.013475Z digest=sha256:1f3b89cb77bbdf44ad0f0ef9c022b6c05ae755ead6299d86eef507a1e4b93487

Observation cd0ae5ae-5a2c-43c2-abbf-20c4a72c161e · inbound

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning cites this paper.

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning Process-Supervised Reinforcement Learning for Code Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:01:00.233468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:51:19.555272Z digest=sha256:47a7031cd79127ad038c6f569c8b4f4ea4cbdb8ff793311358d64e06e33f4bac

Observation 1d03fe18-37d5-47bd-8db7-482b1e88f803 · inbound

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling cites this paper.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Process-Supervised Reinforcement Learning for Code Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.288936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:946ead190a23433ee97c342f9ec36ba863d4f24b9f9b44eb080c2468a5af77fc

Observation fae0bdb1-6d17-4048-8adb-07a787463b99 · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Process-Supervised Reinforcement Learning for Code Generation

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:40.231715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:568c8bb035efb9f060c6afe18d1d21e9ea1596d35ccb243f2a3e04e61c4b2ab7