Pith. sign in

Paper Citation Record · LEDGER

RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2410.02089.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.02089 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:50.613928Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 91e1c7b4-1331-4b0d-8ee6-10cdad3b9da5 · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:20:59.399537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:4955443664ab86c03004fce14231266e632c270b48242df00ee4ef724af48bce

Observation 4f86d814-6f5b-4ad7-a836-828d61c88d09 · inbound

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution cites this paper.

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T10:27:56.341144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T10:27:56.185943Z digest=sha256:5d8e3b4a5d1cdcb81a175913e1541f5db19dd91ab4a389460ee8069515fc0901

Observation dfd6e967-1c5b-4d72-be45-fe8404b61f48 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 206

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:23.564056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:0cc811d18fd21d7b9be495995b0c69c3c784ace60566b35de1cfddf307c2348d

Observation 43bae483-6515-488d-813f-dedbbcfb33ce · inbound

Gemma 3 Technical Report cites this paper.

Gemma 3 Technical Report RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:22:12.170943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T22:18:55.976503Z digest=sha256:40e758a89bdee814260d5826d52d06f1bada1320b220d6f77cc8ab769e1142c7

Observation c56626ec-d521-44db-b591-9947b9a4ba79 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T19:32:00.981090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:3fd2c7cc04bd01c60ae5e34fad28776ccd7efb41584652f13ccdf3e3136ec984

Observation 41cd47ac-0822-4c76-8982-68592cfe2bbc · inbound

Reinforcing General Reasoning without Verifiers cites this paper.

Reinforcing General Reasoning without Verifiers RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.613928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:50.613928Z digest=sha256:a10937519ccab7d34716bad6bf4ce8f8abff2bb71827685adb1b023ab5998aa1

Observation 484deb15-1a48-4b08-a34d-273bc18125e1 · inbound

Training Language Models to Generate Quality Code with Program Analysis Feedback cites this paper.

Training Language Models to Generate Quality Code with Program Analysis Feedback RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:04:58.623446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:04:58.623446Z digest=sha256:f9465a127e6fed037309e853eda88d46e7cffd9685584d3717dde119c7f69cfc

Observation 5badb4da-5bed-4308-8c0b-fd622e9b4717 · inbound

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization cites this paper.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:24.815311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:24.815311Z digest=sha256:02092aaa98d9f510f3310bbdf478d61c5b52182f7d51aa505cff094f22e693f5

Observation 1aff602b-b69d-462f-8792-7d66865a4e5d · inbound

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training cites this paper.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.658504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.658504Z digest=sha256:e1e77df19cef250caf14f58c5f094cf95fdfd216355cf477c463c942710da59c

Observation 42e8fd50-6085-47b2-9879-b8b77ad352ed · inbound

Improving LLM-Generated Code Quality with GRPO cites this paper.

Improving LLM-Generated Code Quality with GRPO RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:53.421299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:53.421299Z digest=sha256:e0c06e3d1acbf86f589c0743715aba771af81e419370eec734fcf6e284ffa724

Observation 43b36156-2531-46f0-b9dd-afe892cd943c · inbound

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models cites this paper.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.562765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.562765Z digest=sha256:989b95d6e24761099e8cc1dfac4c61b5d382ca58057c5814186a3da6d23b0818

Observation 3b9a162a-60ac-446a-8ddc-30c8f0d508da · inbound

Reinforcement learning fine-tuning of language model for instruction following and math reasoning cites this paper.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.805270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.805270Z digest=sha256:16431a3a8b859fad269227fdee9d81780e8bc5216f0a2c86f794b1905bdf9af5

Observation 9b939a93-5116-451a-b01f-0489329ea5b0 · inbound

CodeGrad: Integrating Multi-Step Verification with Gradient-Based LLM Refinement cites this paper.

CodeGrad: Integrating Multi-Step Verification with Gradient-Based LLM Refinement RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:41.213688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:12:41.213688Z digest=sha256:ea2503c483a00b17be0d943678abba7c235acd94527dadb7d173db6a0878e888

Observation ea83a352-454f-448a-810f-8d2cc060d73a · inbound

VERIRL: Boosting the LLM-based Verilog Code Generation via Reinforcement Learning cites this paper.

VERIRL: Boosting the LLM-based Verilog Code Generation via Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:28:23.056538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:28:23.056538Z digest=sha256:8f14079f582cfcc5b1f700a0092a6cb39bef711ed2450c7bd6128279e0c0fe5a

Observation 97650b8c-9389-4fae-9456-c96ce892e39f · inbound

Short window attention enables long-term memorization cites this paper.

Short window attention enables long-term memorization RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:11:21.813433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T12:10:42.646127Z digest=sha256:14fe654b8b481b7f0d4dc23c864f17105e12f20b3f884c0a2aa8c22b67a4b7e7

Observation 9651214a-fb95-4701-9ae9-a09d438c2181 · inbound

KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta cites this paper.

KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T13:45:23.514060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:45:23.514060Z digest=sha256:ed6732a60e096c1ea446feb34d10775b47e06566371d45e162b68b16fddea29a

Observation ba509e48-0377-4279-bbfb-229312551200 · inbound

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation cites this paper.

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:31.452883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:19:31.452883Z digest=sha256:ae457462990c393abc8361c88b7e85011419df81257c42d180730413d3c7e8d4

Observation 6f6b0ecb-48a8-4d63-a242-e6c291cdc467 · inbound

NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable Architectures cites this paper.

NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable Architectures RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T10:38:11.981408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:38:11.981408Z digest=sha256:499ebf14f44af4d4659de0c7696ad2144ddb767646e17d43f66a8324fa4ddf13

Observation b38a1b66-17ac-4be1-844c-5a99239250b9 · inbound

Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints cites this paper.

Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:10:47.874733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:12:53.062214Z digest=sha256:ab6562868ce018aa59d2637396e50c9666ce37a92ee65e0ba197f8e54c8e6b99

Observation f5e72726-7d9c-4e30-9322-09cbe9cfdc61 · inbound

An Iterative Test-and-Repair Framework for Competitive Code Generation cites this paper.

An Iterative Test-and-Repair Framework for Competitive Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:48.523439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:44:32.950977Z digest=sha256:cb89cd34f1cabd9a3ea61500f84986df4921412181f376162fd3dd790b289a82

Observation 98fd10ca-0dbe-4f9b-bc8d-ac1fcde25d4e · inbound

An Iterative Test-and-Repair Framework for Competitive Code Generation cites this paper.

An Iterative Test-and-Repair Framework for Competitive Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T09:22:44.557413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:22:44.557413Z digest=sha256:70a3ba7904e51e8efeb04bd0c4501eb0731d446bc00dd59d2a6441dcfd905153

Observation b6d7d84f-326b-455e-8077-d709f2e0fac7 · inbound

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning cites this paper.

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:30:53.249113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:45:40.915428Z digest=sha256:53e6a8429db42d402400600480aaff79f9b55c5b9163fa0fd5e740cd6b309d35

Observation 1027fdd9-2fe5-4d0a-8b08-6c6c3423f5e0 · inbound

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning cites this paper.

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:58.374374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:42:57.596073Z digest=sha256:3b9338be610064c563c5ab419a4d1d397d505c6acf890173303fc4259d327830

Observation d4a371a1-8bf5-49fd-8ab3-d17134a5225b · inbound

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents cites this paper.

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:25:35.836993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:22:13.551709Z digest=sha256:f9bb18e0d3bb55779aa6d25749aa4dd6cb42cb1048d8d8c6a2beef999734cdf5

Observation 5335f5e0-a2ef-4c60-aca3-0809b19ee7ab · inbound

CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora cites this paper.

CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:29:25.217935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:59:44.880102Z digest=sha256:b4bd389e2411a0470cd1620e1d6c42c745284c09d25a279da2ec986000c49892

Observation de5616c2-4498-479a-80ef-866762313403 · inbound

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents cites this paper.

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:26.080342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:12:37.970638Z digest=sha256:ac250f64912d4a30ae31196f4ab6a6d8e62fb1f9472d7c7f7d9f8533071a103b

Observation 1e6959cd-4a2f-4088-9d8f-8fa9aee28c31 · inbound

BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models cites this paper.

BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:11:18.979788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:09:40.321500Z digest=sha256:69ff7743ae1714d2ba8f6340e0a808f84cbfb79d66d976b4da0762de40553544

Observation 71e6a669-defc-4543-bd0c-571fbabf2cb5 · inbound

BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models cites this paper.

BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:07:22.367744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T06:03:32.270553Z digest=sha256:733dc2da7a5e5634202c1cd1a0d737f4902acebe913f6d8a63bd50b7901133f8

Observation 61dd3dc0-56a7-474c-93e4-1a540b0357bf · inbound

Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards cites this paper.

Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:43:27.601951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T01:41:31.631505Z digest=sha256:30fb7db9225f4f0255a380fab3690b2e440b5fe7f0fd40d63ec2742bc5510954

Observation 706ce023-1d4b-400f-9605-f8ed0fb8259a · inbound

Self-Supervised On-Policy Distillation for Reasoning Language Models cites this paper.

Self-Supervised On-Policy Distillation for Reasoning Language Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:43:22.222344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T14:42:55.368104Z digest=sha256:6ed4652111f522fd712c31b7c386b0b9283ebab023a840d807ea0d1e89c3dc13

Observation cbbfb1d5-2ad0-4a09-851a-d943557595ac · inbound

HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL cites this paper.

HydroAgent: Closing the Gap Between Frontier LLMs and Human Experts in Hydrologic Model Calibration via Simulator-Grounded RL RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:28:18.834772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T13:28:03.965754Z digest=sha256:167710a1118c9bd419394fc61a3a96f73c2e1abbcacc425a28e5857bc9683267

Observation 4b8c36f2-35bd-48e7-b0fe-ebdaa2ae94e0 · inbound

Code as Agent Harness cites this paper.

Code as Agent Harness RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 104

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:58:14.301565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:54:54.558241Z digest=sha256:433aaab8064912268af76ab28cc927263cb09c8883d23b51eb5ff981685ed818

Observation e60c5963-68fb-4953-90eb-5fc7254f05f9 · inbound

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models cites this paper.

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:13:58.422663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:11:35.785227Z digest=sha256:b9cee7ed333cbeeecfa34c0cf844ae77ca5ccb282ffdfec17c7ec8304167797c

Observation 59c618b0-de54-4a45-99fb-e3b65c89d6f4 · inbound

Self-Policy Distillation via Capability-Selective Subspace Projection cites this paper.

Self-Policy Distillation via Capability-Selective Subspace Projection RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T05:34:40.386519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T05:31:42.803465Z digest=sha256:f8cc5eaf8629e195866cf274eb6f7cbe7a956b64b98d45a85c378cb6e789a706

Observation 86301b5d-5ca7-45a6-93ff-f7022016e5a4 · inbound

Learning the Error Patterns of Language Models cites this paper.

Learning the Error Patterns of Language Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:03:29.670798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:56:04.187007Z digest=sha256:309a8db1447a82cee14948090bfff20eab7d266956b6461b44ddec4670e30f7c

Observation cf5dea75-95b2-4735-be05-ef2454e75c9d · inbound

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs cites this paper.

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:46:33.093345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T09:40:35.258818Z digest=sha256:b89cf32e7fd2febe2d9dc781900b903a8dcc660bcd2707b4c1057de696fc2cf3

Observation af5d5b50-2658-4b97-b35b-aa63b36d48df · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 298

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:38.270687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:baf9e027382a4a5bc4c97078369333b07275ea6f1f199a120ee2dd872a43003b

Observation 0db63301-0fce-4dd1-a679-b74a2b8679f3 · inbound

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs cites this paper.

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-06-30T01:34:09.489245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T01:29:42.919461Z digest=sha256:d7bf5b72c10952c1306a922994bac40cec43c2b5f7bf6634d97fe41f2896916c

Observation 94718f18-42ee-4043-b317-adf8860d58f3 · inbound

DecompRL: Solving Harder Problems by Learning Modular Code Generation cites this paper.

DecompRL: Solving Harder Problems by Learning Modular Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.791771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-03T16:30:34.793328Z digest=sha256:63d1b61a4b6eed2f286f82b1fb03161d1b7b0b729ede541dd7ff09bfe420e27d

Observation 16ed8299-fb9d-4606-9c31-af71c070ade7 · inbound

Beyond the Need for Speed: Energy-Aware Code Generation via Simulation-Guided Reinforcement Learning cites this paper.

Beyond the Need for Speed: Energy-Aware Code Generation via Simulation-Guided Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T17:00:48.664985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:00:48.664985Z digest=sha256:bdc70e2879b58e4e950657d75c42603111ff66ff7c604b2ff7357d2fb9c34ac0

Observation dfd47062-a949-4b64-969c-22bb38ce2cf8 · inbound

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards cites this paper.

Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T11:28:23.511747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:28:23.511747Z digest=sha256:ad851bc65734dab4ecbad9b3c17d1cc52f08f3b28d6345c78d19aceb231b8cc5

Observation ae5926b7-33cb-4f8d-93b0-fb5d51422def · inbound

Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation cites this paper.

Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T13:39:15.458226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:39:15.458226Z digest=sha256:f5b9cde324119d4b34de21bbcf8f24c2ca10e50005aee16521b4945f4fa91e6e

Observation 873b42e5-cbfb-4569-9dae-a9a5bf5f35a1 · inbound

CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation cites this paper.

CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T09:12:22.385849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:12:22.385849Z digest=sha256:02686c80b86e8c0781a3d1296de2a9ec536172e097cd6084ae2f8a8ffc56c051

Observation 8597521d-8618-40b5-9d1a-cfa2c439d041 · inbound

Training Large Language Models for Self-Explanation Faithfulness cites this paper.

Training Large Language Models for Self-Explanation Faithfulness RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-01T08:36:30.457193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:36:30.457193Z digest=sha256:73a50843311f41db271ccde007990298784436da6a9a5d894954f72a97363f94

Observation bba3d182-65cb-4f0c-a6a5-429435763b84 · inbound

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills cites this paper.

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T04:28:45.552110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:28:45.552110Z digest=sha256:9ef94627e9b1131807afe5ccf15d2f647eae0c06dd29a58774d331094a247ce3

Observation 9989336a-ed70-49c6-9097-e3ccdfc965fc · inbound

RLPF: Reinforcement Learning from Performance Feedback for Code Generation cites this paper.

RLPF: Reinforcement Learning from Performance Feedback for Code Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T10:53:08.880515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:53:08.880515Z digest=sha256:d5f1d6de7428be55188dba21b8eb04d464c76cc4fb4ae6c2a8352dab1eab3c33

Observation 4b0e5296-46c0-4996-9f78-40f3af77103c · inbound

LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation cites this paper.

LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:38.260775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:38.260775Z digest=sha256:645c784a82bfc4882d5596c3e892aa0319b38136679437c956243d8072130382

Observation 6b017c0d-f03d-4ccb-8932-104c8caa8a47 · inbound

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning cites this paper.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.219331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.219331Z digest=sha256:4a3390a1506384f26b0da9eb673fd5320317781d5ba66a83b95fe40a3b1f302d