Pith. sign in

Paper Citation Record · LEDGER

StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2402.01391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01391 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:39:09.585165Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 493ee7ce-f59d-4a86-87b1-2c8bae5e5a44 · inbound

A Survey on Large Language Models for Code Generation cites this paper.

A Survey on Large Language Models for Code Generation StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:18:06.739448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:18:06.304134Z digest=sha256:da87a04a6decf4db58f651e1c120e7925e25cf6ef4ec61bdfac7ad55a0727897

Observation a6fe8069-4af6-42a9-b8a4-475274dab7e2 · inbound

MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic Testing cites this paper.

MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic Testing StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:18:31.317120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T22:17:02.681066Z digest=sha256:f97e672c40996ec05f34907df7f85c0053f1ebfebc1b08884dee7d0cd1ec0e10

Observation dae15ffe-6ac2-4676-832a-1328af38d70a · inbound

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs cites this paper.

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:05:15.099853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:05:15.099853Z digest=sha256:4ff20bbc5c78be7e1454864aba0ea336751ac6054152e0a111e606aef11a4edd

Observation 6bbc4a7e-b8cc-4b9e-86d0-5d060bdf148d · inbound

Preference Optimization for Reasoning with Pseudo Feedback cites this paper.

Preference Optimization for Reasoning with Pseudo Feedback StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.560207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.560207Z digest=sha256:a37e6988c718803fdb6bbb80e126280360cc9af38ed81712a2c537ce10ccf876

Observation 7f1aaa35-7be3-47ad-a19d-409ca0333f57 · inbound

Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning cites this paper.

Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:11.364925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:11.364925Z digest=sha256:0e030e6aaa55176cc0264184389da934cb6ef113028489eaf21b4384c2fb1be8

Observation 9d034c76-dece-451c-96ae-c69ff824e8a2 · inbound

Distilling Desired Comments for Enhanced Code Review with Large Language Models cites this paper.

Distilling Desired Comments for Enhanced Code Review with Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:28:51.280416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:28:51.280416Z digest=sha256:5ac4694826f332d8e50c71f759b5402814cefc7414d1450849957c1e092be2a7

Observation 7a48100f-4cdd-435d-8ebb-f657497a38e3 · inbound

ACECODER: Acing Coder RL via Automated Test-Case Synthesis cites this paper.

ACECODER: Acing Coder RL via Automated Test-Case Synthesis StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:53:49.686606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:53:49.686606Z digest=sha256:b9a9b2471ff07cfe0cc7ec03dc66c0c3b5d9792d4b013cbdda55e2e840768607

Observation 1c172a57-6eb3-497a-9cd2-7e8bdbb84290 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 164

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:23.489044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:bd4b6ab30f614ee4a74627153184e2920624a9b62c3f493ada7f1410515c2c2a

Observation 3403b63f-cad3-4b34-a036-c1401269f0ff · inbound

Themisto: Jupyter-Based Runtime Benchmark cites this paper.

Themisto: Jupyter-Based Runtime Benchmark StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:39:09.585165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:39:09.585165Z digest=sha256:c46e57126eb0fb07b4f18027d34221497793974ffac63f5ad9f02689e7f9a959

Observation ea30a76c-e77e-40e4-b4e9-06aaaf99fb45 · inbound

Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs cites this paper.

Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:34:01.054883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:34:01.054883Z digest=sha256:eefd43226f232429f1bd240f8a7b88afc2ccbbc4f0bc84e026ccecbaa89e258f

Observation 44bd3ffe-f9ed-4004-ad79-63854f7f6591 · inbound

CRPE: Expanding The Reasoning Capability of Large Language Model for Code Generation cites this paper.

CRPE: Expanding The Reasoning Capability of Large Language Model for Code Generation StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:22:19.042890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:22:19.042890Z digest=sha256:15c8e35c01d3ee109180cfcd54506b285a6d9766c6247ca656178b865277989f

Observation 59474258-9f83-4571-bd6e-e0f7e94a8363 · inbound

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution cites this paper.

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:03.712435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:03.712435Z digest=sha256:495c883cdc5bbe47a3ad6f1144eb2ae5ac9c3b357c5f1e709694875e9065e1ad

Observation 6bb10acb-174c-40f0-81d2-ae3bcbc23691 · inbound

Training Language Models to Generate Quality Code with Program Analysis Feedback cites this paper.

Training Language Models to Generate Quality Code with Program Analysis Feedback StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:04:58.439514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:04:58.439514Z digest=sha256:4d8e89f41381f5c3e0fc8cb30c7dd385f2620e2de6da12675f6bc997dc336747

Observation 457bd8db-44ed-4218-8753-a5053c9b1667 · inbound

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization cites this paper.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:24.562628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:24.562628Z digest=sha256:df9782d43ba0456da9c549ea42ddf8cdd7518a76577e469807e015cd7546851d

Observation 4887bd46-fc39-4dd5-b312-36bb30a7e0e8 · inbound

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review cites this paper.

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:28.236569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:28.236569Z digest=sha256:8a175b8f7d2377540e3cb963465faf68bb039f10c1954e9011f41e2026215bbd

Observation 27ef6c89-a7dc-41ec-84a4-cefc545ee0fd · inbound

Improving LLM-Generated Code Quality with GRPO cites this paper.

Improving LLM-Generated Code Quality with GRPO StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.249493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.249493Z digest=sha256:ae5f9525673207f89f0903f1d79005395ab68c699a3fca01d6b9cfb3c82ebeb7

Observation 14c18af4-3176-4b4b-a3dc-f9fdb4b4e445 · inbound

D-LiFT: Improving LLM-based Decompiler Backend via Code Quality-driven Fine-tuning cites this paper.

D-LiFT: Improving LLM-based Decompiler Backend via Code Quality-driven Fine-tuning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:41:56.947388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:41:56.947388Z digest=sha256:9e00581bc4dd9db86d97e5b81d21b8a9b452fd21382684328be5575de33cf8f1

Observation 7c073181-d46a-45d0-908e-9845c553223b · inbound

SysTemp: A Multi-Agent System for Template-Based Generation of SysML v2 cites this paper.

SysTemp: A Multi-Agent System for Template-Based Generation of SysML v2 StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:10.403693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:10.403693Z digest=sha256:bc703113691eb8479e5cdf3e9f3f97d8b3478b7410ceba639007c4039891a409

Observation 7a511d87-afd3-4702-83c6-9d560eabbca4 · inbound

ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle cites this paper.

ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:47.513069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:47.513069Z digest=sha256:792816a449ab08d772b67ac3eaa81e47ec709a04ed683e3efb52e8896291739c

Observation e11cf801-ef24-4569-aa48-584de858c0cc · inbound

ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge cites this paper.

ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:32:01.539614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T03:29:23.464348Z digest=sha256:e7c07bf2b0a0997762b630444707334d508b5ecf63ccd5c6b3dd9614583318d2

Observation 1182d1ba-ef9a-4b20-a296-7ae0aea84ed5 · inbound

Repair-R1: Better Test Before Repair cites this paper.

Repair-R1: Better Test Before Repair StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:19:19.040967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:19:19.040967Z digest=sha256:217bd43458f55d5eb539ca90aea62c053649b2c7de06711c75511d2bdf544b2a

Observation 983fec0b-a38e-41d8-aba5-5a554a3e9642 · inbound

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention cites this paper.

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:56:50.471275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T20:54:30.449792Z digest=sha256:c0f044ed63bdd12dd5946b135a3817e7bf8bd09d533f66e684d0a8add3229ca6

Observation bbe70dc4-8e78-4d4e-88a8-b072a15b3a7e · inbound

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models cites this paper.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.749999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.749999Z digest=sha256:60838a0545844f3c762934f9c2d7ad240ca94be04ea80172e73f8a43b82f7059

Observation 26cfeec9-3740-40b3-b940-d4cd569333d8 · inbound

Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data cites this paper.

Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T09:02:26.930703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:02:26.930703Z digest=sha256:e1b1716fc6eaff499990002aea4729117ed6961488d24717a44dd025b98424d6

Observation b21c6ace-e1f4-4553-8d50-6b21cb945071 · inbound

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping cites this paper.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.826796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:b0771609100ca6d9b6d3b6aed0b61b8c6f0b46ea9bf6117f54545dc744ac82a4

Observation 9d74a1e8-b880-4c2e-bf55-8761bab3834c · inbound

Towards Enabling An Artificial Self-Construction Software Life-cycle via Autopoietic Architectures cites this paper.

Towards Enabling An Artificial Self-Construction Software Life-cycle via Autopoietic Architectures StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:45:23.468601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T12:43:14.903173Z digest=sha256:b52d9c6d1b0d34e630312ba01d6a514ddf11542d39d3c9691f74421a559c85d6

Observation ae70bde9-1c87-4c90-aff3-4b911f425f8f · inbound

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems cites this paper.

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.660929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T07:02:02.992871Z digest=sha256:3e85740c18216e55bad799ceee905bfe31096ba48831a4bb730b9beb741ab641

Observation 52d9851c-dd6b-438e-b8a0-86ca5b61a069 · inbound

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning cites this paper.

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:19:46.690201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T00:18:18.340391Z digest=sha256:a2690e1d3a93d3540957eaf4b1ebcb3faf4ee60ec96670feeb37802c6cbe5fd7

Observation 81432182-5677-4745-bd85-ee10950b3dc0 · inbound

BoostLoRA: Growing Effective Rank by Boosting Adapters cites this paper.

BoostLoRA: Growing Effective Rank by Boosting Adapters StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.696129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T08:00:21.048408Z digest=sha256:0a2280bf2b312c952af36b938557177e7acbf4c220a881e031ca8b6fd6729e62

Observation c8fd815c-4d19-4ff3-9422-04b899c04fca · inbound

Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning cites this paper.

Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:25.357376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T19:21:30.956682Z digest=sha256:af7445056472115ce0df138e38d02fe7c38017d207ab073c437d57c535f6b198

Observation 2836c5b3-c6a5-4726-9aa4-7c5e83278da3 · inbound

Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation cites this paper.

Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:08:21.267584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T14:03:45.869373Z digest=sha256:43fe6105fb44abadf6c2aba8df18a5d5fa6158e4ac2eb5683eda3db87a6c7edd

Observation f96a4253-1842-4991-b3d9-ad54b92d3e15 · inbound

Distilling Game Code World Model Generation into Lightweight Large Language Models cites this paper.

Distilling Game Code World Model Generation into Lightweight Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:04:44.656018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T13:58:37.956333Z digest=sha256:7f8f411b0a882bbe2fbe95d931901b5cf836532783c5df471c36ff73e547c483

Observation c3ffb77b-e6c1-4d7a-89fc-d70816408c83 · inbound

Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning cites this paper.

Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:03:24.427771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T11:54:07.211895Z digest=sha256:55959383d40cffb6bdd8735b65d45f45f5e48f2849c985bb5c28794f122297d7

Observation 819bc428-11f8-4236-a84a-3f8b2dd030b1 · inbound

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback cites this paper.

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:53:32.268045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T06:11:06.960389Z digest=sha256:fffa61463f89442d425f8eec6002d3eb9069c4cee0ac57c739b83910cf7ca09f

Observation 1e40299e-9fa3-4d54-a36f-a848fbcf333f · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:10:55.862963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:c14b9f8250b00a34d0a0b9ec3ebb03659718ab757b567245e6db237b2167e186

Observation 8b70a479-b2a7-40a8-b68c-7f5172e8d7ab · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:00:19.817888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:d7a970b6ea1c313c349dadd2cd24a020d083456543fed5ff1f21849734ab07bf

Observation a8185407-e9a5-4fd2-906a-d6a1d7386175 · inbound

MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning cites this paper.

MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:35.169574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T09:31:56.712360Z digest=sha256:5128e3ff24a5ba118c8a64d9b84596d949e394e783a3dd37664802a44594e8bf

Observation e5986654-96d5-413b-8126-212f1c8c847e · inbound

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study cites this paper.

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:33.953898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T17:07:21.486960Z digest=sha256:de6d0ad9c539bbb1e7e6f2cac5abe07495988e74f3e1d024cecfbc0df0297444

Observation 396d5770-455f-4181-9e4a-5c07690b25d9 · inbound

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning cites this paper.

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:53.759973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:53.759973Z digest=sha256:2fdcadc9baee229cf8cd52d6fa593209cab7b6f078d26d800b5b20434c42d886

Observation ab486f05-3228-4ee3-9670-56dd3dd2c8af · inbound

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning cites this paper.

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:42.958248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:27:42.958248Z digest=sha256:42d19c01f3ef3bef53b6b61bff49d4492fa7866f9e8bb0a817fe2de1d69c3237