Pith. sign in

Paper Citation Record · LEDGER

Improving LLM-Generated Code Quality with GRPO

As of 22 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 2 inbound Pith citation observations for arXiv:2506.02211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02211 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:31:54.358241Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:14:04.101113Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T17:26:04.783434Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7d8571f-bed6-4320-94eb-c67e61bc8125 · outbound

This paper cites Program Synthesis with Large Language Models.

Improving LLM-Generated Code Quality with GRPO Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.092772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.092772Z digest=sha256:1e186f799f7d52d540960a588cb47dd833d565c78c3a79ca3d230fde6f7bee9b

Observation 27ef6c89-a7dc-41ec-84a4-cefc545ee0fd · outbound

This paper cites StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback.

Improving LLM-Generated Code Quality with GRPO StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.249493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.249493Z digest=sha256:9bf11409d35eddfe2729143000aed415376ba3f2ab935bfa2f06ed0a92e5417f

Observation 42e8fd50-6085-47b2-9879-b8b77ad352ed · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

Improving LLM-Generated Code Quality with GRPO RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.421299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.421299Z digest=sha256:805b898a3941facf0482cb784ea8972efe9a43151582cac5c8456c812e8252a4

Observation e9bfec68-2cfd-40d0-a16a-60fa32f7e1e5 · outbound

This paper cites 2 OLMo 2 Furious.

Improving LLM-Generated Code Quality with GRPO 2 OLMo 2 Furious

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.700853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.700853Z digest=sha256:d903ee41f027614ed9746cd9296a75244474c886f4aeb56958d83de95d05350e

Observation 4e323f52-8cda-464e-850c-c18facdd882c · outbound

This paper cites Qwen2.5 Technical Report.

Improving LLM-Generated Code Quality with GRPO Qwen2.5 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.801015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.801015Z digest=sha256:479adf12469327910d66734b32be44691918148af577e6d88b690dc662e292a9

Observation db786141-1b76-41c9-9c68-35409a456dcc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Improving LLM-Generated Code Quality with GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.979789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.979789Z digest=sha256:336f687b822eae9ec1a6c51a38b8195ff9bfe87c21e96317ee9bc5f063b50a14

Observation 18e4674b-e365-46e7-8ec7-e59fadbb27af · outbound

This paper cites Process-Supervised Reinforcement Learning for Code Generation.

Improving LLM-Generated Code Quality with GRPO Process-Supervised Reinforcement Learning for Code Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:54.286136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:54.286136Z digest=sha256:3a7127ec9a083e7d8e2f3c6c2bc3017f42d8476b89c3b2a1487fe703530d764b

Observation 378849d3-2eeb-4410-9d09-7f42e5337343 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Improving LLM-Generated Code Quality with GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:54.358241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:54.358241Z digest=sha256:3456a359d967b935c8a23a6c030d89559110ccd92cf2f71318d964e6073c74c7

Observation 16b8eb3b-5ed2-474b-9cc0-dc6b5efd9bd6 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

Improving LLM-Generated Code Quality with GRPO Measuring Coding Challenge Competence With APPS

Reference 1977

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.495008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.495008Z digest=sha256:ef1d11509ba928192df2de85fb98dd4073ccce6197ba5d1f8a7b5de7850d70c5

Observation f03c37c0-91bf-4042-876b-76f091be84d8 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

Improving LLM-Generated Code Quality with GRPO CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.592291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.592291Z digest=sha256:66bcfa18f78e8c00f3b85f6c8a67dfb26a70069fa1a9ac85a3647fa6a2a7888d

Observation 6ca98cb9-9959-472b-be3d-e676d2059e9d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Improving LLM-Generated Code Quality with GRPO Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.908094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.908094Z digest=sha256:4b6300a9d1d039f73fd987e35eb42258a91ea0c46450be3bb06905af2fb1eb69

Observation 7aa5ae01-ac28-44cb-aa26-bbf9606695d6 · outbound

This paper cites Iterative Self-Training for Code Generation via Reinforced Re-Ranking.

Improving LLM-Generated Code Quality with GRPO Iterative Self-Training for Code Generation via Reinforced Re-Ranking

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:31:54.623757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:31:54.067709Z digest=sha256:093b32b77edb1e38849640f34fb8b21b0bd5ee25043b1eb931f9cff84755bf2b

Observation 6bbf3403-b156-47ef-9a44-b9cca6cd9fd6 · outbound

This paper cites The Llama 3 Herd of Models.

Improving LLM-Generated Code Quality with GRPO The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:54.201555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:54.201555Z digest=sha256:9d5abceb145bfebb24f2811ddd629523869475c782cb6d73c65cf94913dafa66

Observation 70f3149e-36fb-4255-b48f-bdcf2e5f911e · outbound

This paper cites Process Supervision-Guided Policy Optimization for Code Generation.

Improving LLM-Generated Code Quality with GRPO Process Supervision-Guided Policy Optimization for Code Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.155905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.155905Z digest=sha256:cfbed82bc3785fa1bc062b4c44f570e57fd9bf1ea7b5b5aa0561799ad89aea2d

Pith citing papers

Observation c4a76982-bab2-4716-b3a1-911f1a06ad74 · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models Improving LLM-Generated Code Quality with GRPO

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:04.101113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:04.101113Z digest=sha256:3f05877679bf9493c94e656e8ec590e93e9e1c87c69c0f3d503253f9b95f4501

Observation 053b9dab-34da-4aa9-834e-b03b938dc901 · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code Improving LLM-Generated Code Quality with GRPO

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:26:04.786091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:b9879eeecefca2fae05e8cc1dc20ae1b166b8ad16f5e7b1df07d3dbd1df1625b