Pith. sign in

Paper Citation Record · LEDGER

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

As of 17 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 10 inbound Pith citation observations for arXiv:2505.23387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23387 v3

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:29.507664Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:27:35.889180Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved46
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 19bb2ec8-78db-4d56-8cfd-77d0d0bda6c0 · outbound

This paper cites GPT-4 Technical Report.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.483646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.483646Z digest=sha256:37c4ec5f435602b098e0c498485972bc79001adf208d3842a994bd412fbef314

Observation 468da172-88cd-4971-b37d-8c5a53c6bfc9 · outbound

This paper cites SantaCoder: don't reach for the stars!.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization SantaCoder: don't reach for the stars!

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.535332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.535332Z digest=sha256:f60d82bd83f2265294bd6a8a78db76e9fe9acc1c936742b7f46069f63cefe8ab

Observation 0f266989-003f-473f-bf97-4dbd7d2645d2 · outbound

This paper cites An orchestrated survey of methodologies for automated software test case generation.Journal of systems and software, 86(8):1978–2001, 2013.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization An orchestrated survey of methodologies for automated software test case generation.Journal of systems and software, 86(8):1978–2001, 2013

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:34.630886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:23.612845Z digest=sha256:78cf18e69dbd0cb432e5d17c4da06dd7122c91781ca68f9f96667d7b06f3f03d

Observation 9eb01a86-3448-44a7-9bbb-77b90512f6c9 · outbound

This paper cites Introducing claude 3.5 sonnet, 6 2024.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Introducing claude 3.5 sonnet, 6 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:34.460316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:23.694238Z digest=sha256:fb2c386175508a06cb5374ebb7942dc0eb75e0f07fc53b56c211e2a3e48aac58

Observation f4771a24-b08c-4463-a240-b05cf17fb19d · outbound

This paper cites Claude 3.7 sonnet and claude code, 2 2025.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Claude 3.7 sonnet and claude code, 2 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:34.313485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:23.775023Z digest=sha256:9d36d0c4f7f63f6492a467d06e80f050f6e1f3d3d79ff04e4dc759e3b045632d

Observation 03503e04-0ef7-4ff0-b052-418442383f6f · outbound

This paper cites Program Synthesis with Large Language Models.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Program Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.855659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.855659Z digest=sha256:eb04482694552d114bc74aa278f47e29b0f3d07a5049ad5b6e6f4bd4dae42f6f

Observation 9fcb4b02-dc24-4bf0-a9d1-56a22cd1799d · outbound

This paper cites Code alpaca: An instruction-following llama model for code generation.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Code alpaca: An instruction-following llama model for code generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:24.037931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:24.037931Z digest=sha256:40b5d4aa72bd26e1825969c1f7daf9722eb00d178be2c8701682ea15cd8f15cd

Observation 73ac9bfd-78db-4414-a0c4-8a553ef074e6 · outbound

This paper cites an unresolved cited work.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:34.158212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:24.120147Z digest=sha256:c44cc1f941b18c2ac2a855afb679c62507c2ee78bc5ecdb9a3d371c554c4b33f

Observation 9c38a33c-8575-4f14-9e31-570569b308fc · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Evaluating Large Language Models Trained on Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:24.180001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:24.180001Z digest=sha256:7909d5f987315186f4491b342008f32d65c20d6658adc139f898f811427f2228

Observation ebecbb5f-4933-4c80-a794-f472fc657670 · outbound

This paper cites An introduction to algorithms and the big o notation.Introduction to Programming with Fortran: With Coverage of Fortran 90, 95, 2003, 2008 and 77, pages 359–364, 2015.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization An introduction to algorithms and the big o notation.Introduction to Programming with Fortran: With Coverage of Fortran 90, 95, 2003, 2008 and 77, pages 359–364, 2015

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:33.953988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:24.233906Z digest=sha256:63078b63b1defd60d3a63d48523f290abd739415481f9d74a0a0737c91f606fd

Observation f0523b29-292a-4f91-9e29-160329a78769 · outbound

This paper cites Mhpp: Exploring the capabilities and limitations of language models beyond basic code generation, 2024.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Mhpp: Exploring the capabilities and limitations of language models beyond basic code generation, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:33.772239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:24.356750Z digest=sha256:538901db925c34f444a96e67a28eb509e548b3bf1fea10cd5518f68a3d9fe027

Observation b258fdba-ecdd-4cda-acbc-bd4bf3c12972 · outbound

This paper cites Docker.lınea].[Junio de 2017].

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Docker.lınea].[Junio de 2017]

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:33.580132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:24.485638Z digest=sha256:acaa6058a597fc0c3c6ace4aced998bb727827752a8c3960b5385b14a46968e8

Observation 457bd8db-44ed-4218-8753-a5053c9b1667 · outbound

This paper cites StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:24.562628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:24.562628Z digest=sha256:1643f2c6bce8945d0576088ba82fb8d4459ec3b32b881d3a55714f6aadd017a9

Observation 0ac43d93-2f3e-493a-a312-8bf844b68a09 · outbound

This paper cites Mercury: A code efficiency benchmark for code large language models.Advances in Neural Information Processing Systems, 37, 2024.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Mercury: A code efficiency benchmark for code large language models.Advances in Neural Information Processing Systems, 37, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:33.390117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:24.645596Z digest=sha256:72225588775ce5e7f3afcc0203dfe165e0f0533d58252edb3fb4d8decc78ce66

Observation 45663742-0ae8-4af7-800d-af78d13af267 · outbound

This paper cites Chapman and Hall/CRC, 1994.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Chapman and Hall/CRC, 1994

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:24.724108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:24.724108Z digest=sha256:3751b157c471973a30ac28ed7e795cc3e76add93f13892415ede0a67341bca3b

Observation 5badb4da-5bed-4308-8c0b-fd622e9b4717 · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:24.815311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:24.815311Z digest=sha256:80ca7bdc05090b3b7a20494c77476c428d3fac357628685d884bb00786c3d632

Observation 722ebdaf-510b-4de9-a384-35e57fd0cebb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:24.899820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:24.899820Z digest=sha256:04d18b392c5ade9915f5bc77491334be38a37fe1bad5b71e0ba7e90b365e272a

Observation f36078d5-eae7-49a9-b5e7-19be210922b6 · outbound

This paper cites Measuring coding challenge competence with apps.NeurIPS, 2021.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Measuring coding challenge competence with apps.NeurIPS, 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:33.219909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:24.978032Z digest=sha256:fbcf8236bcaf385a45eecc8642f268c87ecfeaa3e6aa1be20c943dff41240476

Observation 08b14e2d-1020-4ba0-b9b3-75ecd3c05244 · outbound

This paper cites Codecot: Tackling code syntax errors in cot reasoning for code generation, 2024.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Codecot: Tackling code syntax errors in cot reasoning for code generation, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:33.019551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:25.021399Z digest=sha256:870dca0f2f68c2f2614d347dcb6d4f8a30f0b8f23396fce7075fe8ce04f019f2

Observation b478f50e-0d9c-4771-b248-cec0d4d2b89b · outbound

This paper cites Effilearner: Enhancing efficiency of generated code via self-optimization.Advances in Neural Information Processing Systems, 37:84482–84522, 2024.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Effilearner: Enhancing efficiency of generated code via self-optimization.Advances in Neural Information Processing Systems, 37:84482–84522, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:32.790748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:25.112854Z digest=sha256:72eebbe9a256079387c486c72d887b2c3472bc42f6043b5720cad4416620fdf4

Observation 446d9409-f8a4-4237-aab9-80fbe517bf5f · outbound

This paper cites Effibench: Benchmarking the efficiency of automatically generated code.Advances in Neural Information Processing Systems, 37:11506–11544, 2024.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Effibench: Benchmarking the efficiency of automatically generated code.Advances in Neural Information Processing Systems, 37:11506–11544, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:32.549057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:25.160863Z digest=sha256:4df9b9720507036140a0590cf7c7a275cf6ebd1952b7a0dfc5de935b94f08a16

Observation 8d701659-174a-4c4d-afeb-08c29b63f765 · outbound

This paper cites EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:25.227030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:25.227030Z digest=sha256:4f5d39386e24baf6ae0f8fd8c6b8fc568635124d0604a6400b250d7ac22a730a

Observation 92b7a3da-f89a-4ff5-aa5d-51097c8b5401 · outbound

This paper cites Bias testing and mitigation in llm-based code generation.ACM Transactions on Software Engineering and Methodology, 2024.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Bias testing and mitigation in llm-based code generation.ACM Transactions on Software Engineering and Methodology, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:32.310975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:25.294749Z digest=sha256:fcaafd4530b1c32e4ae4a1cd767b7bff2fb55e7689a2c1456a23746b805e0e03

Observation e47b6424-dece-40b9-be72-358d58537a21 · outbound

This paper cites Measuring the Influence of Incorrect Code on Test Generation.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Measuring the Influence of Incorrect Code on Test Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:25.350738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:25.350738Z digest=sha256:1ba87aaa2cf67e782490c37aaacc22c02df9a37a60a9195e3652a5998709722b

Observation 28372f51-1e5a-44a1-a159-8e2af24c6d33 · outbound

This paper cites Zhang, Michael Luck, Qingwen Bu, Yuhao Qing, and Heming Cui.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Zhang, Michael Luck, Qingwen Bu, Yuhao Qing, and Heming Cui

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:25.414753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:25.414753Z digest=sha256:aa8ae830a6fe918df07cb39406f575a4f41108611be3d57ec2afd0db6da294d2

Observation 6e9830e7-59f7-4d76-b684-647962f480f7 · outbound

This paper cites OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:25.476308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:25.476308Z digest=sha256:aaea25acdd5f7ee7de78d8649f0d5d8f4f8d3a599a5ec63de53af5d1cab0bbfb

Observation 85dcd90d-1106-43d6-be37-fce6f7f671e7 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Qwen2.5-Coder Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:25.619468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:25.619468Z digest=sha256:239081eb5563882c3799a54b794040c8a2630264d9942591b34cd5b18c855c93

Observation 494b4c79-dec5-4857-a27c-146e6ac0731c · outbound

This paper cites Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:31.958871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:25.703099Z digest=sha256:09091b6ccce447965b045f41190f0aa3256f7e865c8c362f03d3bd9ab5e54b67

Observation bdd1d97f-5a85-4f00-bf56-e344fc4e7c4c · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:25.816411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:25.816411Z digest=sha256:3506118c66216bf644fee613431b6ae0b81402688f8d736b480fc97cf47376ed

Observation a57ee1a0-903a-49fb-8d93-adc1c0903f55 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization A Survey on Large Language Models for Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:25.949940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:25.949940Z digest=sha256:c7e1370f783e46ddb892beeccbbeb93cb95e72f9bd910dfd2e1e148ce9d4dae3

Observation 5f6a845a-08fe-4409-936e-3cd8fcac0ce8 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Gonzalez, Hao Zhang, and Ion Stoica

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:26.056841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:26.056841Z digest=sha256:6bdee65b200cb75357502b2d710204c78553c917bb3292437d35985a362ce6b7

Observation 1149084b-f88c-4659-ab68-623086d0a859 · outbound

This paper cites Coderl: Mastering code generation through pretrained models and deep reinforcement learning.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Coderl: Mastering code generation through pretrained models and deep reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:31.646333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:26.174145Z digest=sha256:927d5dd5e86aad4fa6c176f993368470b81637eda48a9afa1ae1bd8a41b547dd

Observation d1335fe5-2825-46e4-9764-ee4e5a722cf1 · outbound

This paper cites Competition-level code generation with alphacode.Science, 378(6624):1092–1097, 2022.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Competition-level code generation with alphacode.Science, 378(6624):1092–1097, 2022

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:26.277047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:26.277047Z digest=sha256:087f0420f59bf99fff3e662e1f75429fe3a44bed5d222a4eefc8b1d2bd3afc64

Observation 2fbd8b7c-5ee1-4003-a086-1a9d9256e68c · outbound

This paper cites DeepSeek-V3 Technical Report.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization DeepSeek-V3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:26.377727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:26.377727Z digest=sha256:42aceb42bb5da9c539500f800199bdf9f5df31621177be4b445118c0cf8febe8

Observation 1dcf8880-cb30-428f-a2c4-c71d6bf5fce9 · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:26.394353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:26.394353Z digest=sha256:143af4873ad5c583d334163d97068ab619fcefae4211b78aceb080e96f7f3445

Observation 901b1cf3-73e3-462e-8e60-18cfe99e9c15 · outbound

This paper cites Evaluating Language Models for Efficient Code Generation.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Evaluating Language Models for Efficient Code Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:26.404596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:26.404596Z digest=sha256:988482f24434f0720e04ee34eb75c27202a4ba99314da453b602ba00f5794daf

Observation 3d2f7533-58d2-4b83-b9c0-79448fbafd5d · outbound

This paper cites Refining chatgpt-generated code: Characterizing and mitigating code quality issues.ACM Transactions on Software Engineering and Methodology, 33(5):1–26, 2024.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Refining chatgpt-generated code: Characterizing and mitigating code quality issues.ACM Transactions on Software Engineering and Methodology, 33(5):1–26, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:31.355615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:26.564217Z digest=sha256:bbe1101c220cd038fe4df9dd7144af250ee3fa516ad1f64248fcaba6f8f2df02

Observation b645b98e-207d-455b-a7fa-b67c1182f7a9 · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization StarCoder 2 and The Stack v2: The Next Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:26.687342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:26.687342Z digest=sha256:b6ec469c93b7fdfa96298f837e30682ac3fa79af23bf80f0d3f0f75f3cde6664

Observation eb8d02a6-cfb5-4e96-b0ad-45eb984f4414 · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:26.792726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:26.792726Z digest=sha256:1f896b12a5066ebbf2ad6c800aadd1a1ad171c714ad99482bb23e5da53c90a5c

Observation f35a0c77-24bc-44db-b4fa-10f686a0fd94 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:31.081251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:26.908398Z digest=sha256:e8e65be84f64e3f66df441493f34675004a14f5a033aae7278cb275aac6b53f7

Observation 93a71e4a-fe8b-4a52-82b8-9f69abb06b3c · outbound

This paper cites OctoPack: Instruction Tuning Code Large Language Models.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization OctoPack: Instruction Tuning Code Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:27.026954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:27.026954Z digest=sha256:880afd9e2ccf8879fb7a1e8e912bc9a329646bb71ca69a7520117a521059e09e

Observation b028048c-f4f5-4732-91dd-22fbbe098333 · outbound

This paper cites CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:27.094115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:27.094115Z digest=sha256:b6947dfb896d6ed3fe59b38f51fde0c2fd0f666887620842a243966be52dca22

Observation 064be610-7300-4050-b626-fd9dd388fd70 · outbound

This paper cites Introducing openai o3 and o4-mini, 4 2025.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Introducing openai o3 and o4-mini, 4 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:30.800178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:27.149254Z digest=sha256:4c170f243803dfdd1215df9c2949715d430ca790d1be8645aed5a3d20572cc56

Observation 45870295-1a6a-4563-b2c4-6584df34f3b8 · outbound

This paper cites Zhang, Heming Cui, Siu-Ming Yiu, Dong Huang, See-Kiong Ng, and Luu Anh Tuan.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Zhang, Heming Cui, Siu-Ming Yiu, Dong Huang, See-Kiong Ng, and Luu Anh Tuan

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:30.584280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:27.204810Z digest=sha256:5cb5ccced8d0c58a2c38051a640816b2c7facb88102a1ad48a8631ffc2d24577

Observation d768de9e-9a5d-4dd6-91bd-2962dc171164 · outbound

This paper cites How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:27.276056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:27.276056Z digest=sha256:437692540ee445978128b1b0536209f63a5b89d87a3937a09fc90031a15de46a

Observation 43e114bd-4d8f-457f-9ca7-bfe336ed532c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:27.345353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:27.345353Z digest=sha256:d5ecdc8b410d5c6197f04db8ae09c2f17f24a6d2e2a6e46da24bc926c3bc4066

Observation 77eeb386-b4ad-4930-96a9-263eb0ef08b3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:27.475345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:27.475345Z digest=sha256:16535b377210f877c8af68863047b09ad19a1f5a265356a526b1fc2ffec1f1d0

Observation 1957a6dd-8bff-4a8e-809e-b216b82fbce4 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization HybridFlow: A Flexible and Efficient RLHF Framework

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:27.598252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:27.598252Z digest=sha256:d22da7ef19f21b98c5b4a142e15d3a9f29e47bf492deb25dd46f3a0ad4e1f2df

Observation 564cc66c-78cb-4f7c-9a6d-36d52c42892d · outbound

This paper cites Efficient and Green Large Language Models for Software Engineering: Literature Review, Vision, and the Road Ahead.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Efficient and Green Large Language Models for Software Engineering: Literature Review, Vision, and the Road Ahead

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:27.702240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:27.702240Z digest=sha256:070ca3728c9688a43a6081904cd13b4744d8fef0d3aa5897f666f2984834ff09

Observation ae1fa2d6-8fdb-48d9-8e30-51fbb49bd72a · outbound

This paper cites Learning Performance-Improving Code Edits.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Learning Performance-Improving Code Edits

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:27.836357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:27.836357Z digest=sha256:0f9f2cacb49da4baa6b206a55750975851ef7282135bb506e06647a3033d7cc8

Observation e1bb270e-7f0f-4abc-a70d-79a771fb7f03 · outbound

This paper cites CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:27.939356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:27.939356Z digest=sha256:1c27f39802b627bacf2ec75c53f8e1d61a46935c4116fa6974b78084ab628e99

Observation acf1e616-757f-40d2-8eac-a58cebcbb587 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:27.997213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:27.997213Z digest=sha256:16e7ac50da97694111ca9c4f027345208ed4a67c0a87be2c8461ca213370b089

Observation c1ec2e52-8874-46a8-a3f7-314142942e33 · outbound

This paper cites ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:28.085835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:28.085835Z digest=sha256:87854ebccb31d1c53dd5678c76f4cc7fd7ce7d2a2e4b63bee5d885181396b911

Observation 11b8eb8f-4ef8-49a9-9824-7dd480fb5620 · outbound

This paper cites Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:28.141142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:28.141142Z digest=sha256:867b2a58dfbd3064af8a2c8a4b4ab8604dbe00c2148fedafb7cc9e191b07f77d

Observation ad6dd3da-fa85-4285-8449-befd7caddbd3 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Finetuned Language Models Are Zero-Shot Learners

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:28.229504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:28.229504Z digest=sha256:a359e0f6253209628f6a40c04c936ea15eaca8c4463826313c582d2f98372ffc

Observation 6158f97a-520d-47e9-a578-c453f4f2ba4d · outbound

This paper cites Magicoder: Empow- ering code generation with OSS-instruct.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Magicoder: Empow- ering code generation with OSS-instruct

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:30.396424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:28.315645Z digest=sha256:b2eee292d08006600e593e5e980d94aa010b78aa0636e2329e6ed1687e888766

Observation 2eb1c552-0b1d-46b8-8c6d-933607e2e3e2 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:28.424341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:28.424341Z digest=sha256:eea08189e2278487a054b93c3643866048a0659d7734bb5cafc8aa9958681f30

Observation dba9e3d8-c622-4c48-900b-ffb25c58fa41 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:28.568252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:28.568252Z digest=sha256:017c87a5c527b13fb98918a40b701d14834ed249ccdad7a9389bcfff7a3774f0

Observation 8ae38b73-4d33-4c53-aa58-bd40f5be8991 · outbound

This paper cites Qwen2.5 Technical Report.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Qwen2.5 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:28.668207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:28.668207Z digest=sha256:5346a43b71f8e068119c4f61101b8191d39e933bda64c9f68c63619db6073c21

Observation b7813645-9f34-4a7b-b5ea-f2014113db0f · outbound

This paper cites LLM4EFFI: Leveraging Large Language Models to Enhance Code Efficiency and Correctness.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization LLM4EFFI: Leveraging Large Language Models to Enhance Code Efficiency and Correctness

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:28.777021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:28.777021Z digest=sha256:e37ee79a5655b7b9bc9fb83cfe0c79e9e4ce2b20a0bae575e2d074d0e34d416c

Observation 1479b453-fd6a-4059-9697-2fae65a8813c · outbound

This paper cites Focused-dpo: Enhancing code generation through focused preference optimization on error-prone points.arXiv preprint arXiv:2502.11475, 2025.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Focused-dpo: Enhancing code generation through focused preference optimization on error-prone points.arXiv preprint arXiv:2502.11475, 2025

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:28.872627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:28.872627Z digest=sha256:edb4322b4c1569c11ebc40893c4f5ef3ee8b265967a95af26f19b0f46a1c3a43

Observation 1a1d3a14-8e2a-4b1e-bc73-9b1a235d957f · outbound

This paper cites A systematic literature review on large language models for automated program repair.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization A systematic literature review on large language models for automated program repair

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:28.952622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:28.952622Z digest=sha256:6148008953186e37c762df6876183564c658d080af8e8bfb4c403937381d3717

Observation 53c2939c-7ca9-4d28-967d-42f849ce694c · outbound

This paper cites Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:29.052783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:29.052783Z digest=sha256:a766461bb3b0fb0a97be628e963ec6637bcb8de31c95fa9213f57553a2fb4c2f

Observation 9e139f44-2a97-4edf-a4a0-a2c171e0a289 · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:29.173199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:29.173199Z digest=sha256:85ff5abbfa0295caf72491b24aff88a946bd2ee8075e019d0407fba57264dcb4

Observation 6737dbb1-562f-43f3-b232-1010bf84302a · outbound

This paper cites Debug like a human: A large language model debugger via verifying runtime execution step-by-step, 2024.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization Debug like a human: A large language model debugger via verifying runtime execution step-by-step, 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:30.261687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:29.282465Z digest=sha256:13c7a521aad10f7f33a3e570c3c77a9c879992b50e1105cc134720d09bea2aac

Observation 53fa0cad-94a7-426b-8f78-94bf78001534 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:29.403751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:29.403751Z digest=sha256:103442916e990256a5e3fd65f055f018ae0bb7b0c767f2060d0953b357497957

Observation 3d7a45dd-2500-4667-b35e-b9e700dca317 · outbound

This paper cites <thinking> thing_content </thinking> <solution> solution_content </solution>.

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization <thinking> thing_content </thinking> <solution> solution_content </solution>

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:50:30.109450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:50:29.507664Z digest=sha256:88b753854c9871efc29f9b930080242c63c542ac358ed66fd4f43800c59167a2

Pith citing papers

Observation d24ad08c-0077-4fc8-bcb5-6bd98235db2a · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.299885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:0cc5819aa72cb972286d6010c79555ce3380f2ad57e02dd73f271a766f873220

Observation 63ce1b29-3b42-47d7-9ac3-da77138d62bc · inbound

ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution cites this paper.

ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T19:27:35.889180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:27:35.889180Z digest=sha256:580961cee68e6d6e6e5b76778053c61d6dc9f3b3683fc1562599156c43391ac1

Observation 8091cb00-e0a4-413e-8745-630dca012d56 · inbound

An Initial Exploration of Contrastive Prompt Tuning to Generate Energy-Efficient Code cites this paper.

An Initial Exploration of Contrastive Prompt Tuning to Generate Energy-Efficient Code Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:40:10.533028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T16:37:16.968173Z digest=sha256:78ac0efd0761651848fd228e6b1f8c7fe8af39e5cc6f054c3058cfd1fe0587c4

Observation 0ab1904a-5441-4146-b456-9d111cebdfb7 · inbound

Paper Espresso: From Paper Overload to Research Insight cites this paper.

Paper Espresso: From Paper Overload to Research Insight Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:45:47.882211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:37:25.753942Z digest=sha256:ecb2e21d61ed06c431373d4e7a7214efcb2102666377eddab8cbe36e9155d11c

Observation f1c30b62-e7be-45ca-865c-b5e7fc4d2ce5 · inbound

AutoVecCoder: Teaching LLMs to Generate Explicitly Vectorized Code cites this paper.

AutoVecCoder: Teaching LLMs to Generate Explicitly Vectorized Code Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:43:14.954223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T11:41:58.923091Z digest=sha256:c5b80cd5b694ae805cad5437cb16446055a20e6812e47b87515cfc31c6a44c9e

Observation cf593b53-fee9-4d4c-8b60-9aa8ef710dc3 · inbound

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL cites this paper.

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:43:30.585461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T14:41:13.191919Z digest=sha256:68f54770b99772f3fe18f275d36c264a8cdd1227b78e54f2eb63e125775bd74e

Observation 929a6b79-6e5d-4392-9ea1-d87fc9f941b7 · inbound

Chiseling Out Efficiency: Structured Skeleton Supervision for Efficient Code Generation cites this paper.

Chiseling Out Efficiency: Structured Skeleton Supervision for Efficient Code Generation Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T18:57:17.261414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T21:43:14.839335Z digest=sha256:2f710af47940a5319da3330b427b05f860afe6a182ee85b45458dd5ebe00abb7

Observation 340d2a8d-ac0f-4077-82c6-24e66db3da73 · inbound

SkelDPO: A Skeleton-Guided Direct Preference Optimization Framework for Efficient Code Generation cites this paper.

SkelDPO: A Skeleton-Guided Direct Preference Optimization Framework for Efficient Code Generation Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.720297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T21:41:02.035059Z digest=sha256:3847e813c8fda071467e3e475e9aff44d56a4103423d2656012e8c40a585eec3

Observation 4ab10d0a-cf07-46ae-bac0-8c78fb288242 · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:10:55.920584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:7de73fccda9f9f1f110a42858a4424b330f50694906bc39e30cf892eb8fae8d4

Observation f75767b9-bece-4bef-92df-e6872303254f · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:00:19.847467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:d71f4c78abcc8dd638c8615a23dc19b06d29a21f58def31070c4e8c339dad62d