Pith. sign in

Paper Citation Record · LEDGER

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation

As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2511.07833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.07833 v3

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T23:13:43.754235Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:54:52.398402Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T00:54:53.394757Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact19
  • verified fuzzy2
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78f8feb9-63a0-4407-aab6-7a9673cbd448 · outbound

This paper cites online" 'onlinestring :=.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation online" 'onlinestring :=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:15:27.077451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:863e8ad1687bfb532a164ca8f00d72f759016f0ed2f6d718500818bd30471d73

Observation 0dda1b3c-ec3d-453f-862e-506693063364 · outbound

This paper cites write newline.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation write newline

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:15:27.073870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:404096c2ea4f8d225b15df6a1451b5d2d57ace0240c0a80df03841f709cac801

Observation b8fbbeec-10c0-443b-948c-61f774d9dc34 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 3

Resolution
verified exact
doi, observed 2026-05-17T23:15:26.628592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:49b4462cd5e6aebcbd191569a1969d99acf3bb2349cb99fa1c0c815b3b2f734b

Observation 9af2921c-36e5-4d64-b5b9-26c3390a575c · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.100254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:5c10d3d1e6f56abb81ca34796c69d9b196c238e5c772a3f0b20fac3c4463810b

Observation ff021806-6cbc-4c4e-9f76-2040cafb1f6d · outbound

This paper cites Program Synthesis with Large Language Models.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Program Synthesis with Large Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.709803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:175569ad880f720ed1a8104b03af538027dc2b122e62965c1176b20bddfa16b7

Observation 1732b6aa-d882-4e1a-82a8-01be7729ace3 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.704346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:a9d085060e7c3cef54c7cef258582e4215523ad7067673868c74d929e7ea2720

Observation f88ba77c-3ad0-423d-b2a4-0c85c4a8f729 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:15:26.720782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:c523602a5f842bc7373da0bbb1409c108d34d5563843422612d942a5a8395be3

Observation b1eda87f-922b-4c70-8774-9d6a501c5b47 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.687761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:5a27ae05a47559dc24ace9daf58e34ad3847c620a926915d49a1cac981cd5ed4

Observation 65231ea9-88c8-4882-99e6-313aa18a55cb · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.109445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:308cfc60ea50ffb4ac81a7179dca4ddfe1282d7a6bc7c95b663cb0c1964cd4eb

Observation 4722bb41-a52e-45c6-8b49-523a937285b8 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.106174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:aab9d0096c75aa0f1c81dc918129bbddb053069cf90c6c5cfb6cf5993964b149

Observation f1abe81b-2b82-4b9a-b184-5cf7b69e997e · outbound

This paper cites A survey on large language models for code generation.ACM Transactions on Software Engineering and Methodology, 35 (2):1–72.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation A survey on large language models for code generation.ACM Transactions on Software Engineering and Methodology, 35 (2):1–72

Reference 11

Resolution
verified exact
doi, observed 2026-05-17T23:15:26.622996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:3721c47ce99cf69177ca255a4603ae52bdcf31799194f49a6e98acb8b8752ae3

Observation 8b8e813d-2af0-465f-8e07-a0f2d67932be · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.067389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:52939789ef6307b54e4392e8c0535200c876780af96668f2da265b70a1857cd9

Observation df5a381d-ad51-42b3-9a13-c7d69b8bf89b · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:15:26.729709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:3116817b1632eacdcd3717a7f709b0829863921f5c593a9182bf54ff892e5157

Observation 8ad6cdf3-0ee1-4abf-895f-66c145ce1fe6 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.070238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:2fbb599b3706bf1b8446b40840de9c8417d87b4750b0c0e0aba499d7a6d81c6f

Observation f130e6de-927b-40dc-a66f-41c6705c35da · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.064608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:a63b87a43a570e2fe8f0911c8f83883dd0879d2d5db0f1a7fc1c508f243ae387

Observation d253d772-7e72-419a-be9a-2a5655fdfa45 · outbound

This paper cites 2 OLMo 2 Furious.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation 2 OLMo 2 Furious

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.681922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:9ffcd518121ec9793ca9f8d8287c3133d9dd1e836fb0be1c77edb36478fc7e8b

Observation 74db933c-5edc-4773-95cc-5a5dea52f840 · outbound

This paper cites OpenAI o1 System Card.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation OpenAI o1 System Card

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.739295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:41516a508ff00ae5ff04b84dcac1da60e9f9f0ebc1acecbf97904b691e767712

Observation 0128674f-5ac9-4ba2-b120-789082f92ea8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Proximal Policy Optimization Algorithms

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.715339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:fb746810993d6ceff8cb8201b9597ac426abebbf41cbe2293126e5dea5ce7a64

Observation 49832b57-ef4f-4347-9ac6-c76e8f00b508 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.768199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:4d74139956af4c791da2a535a82e6f711a34f29cb2baf2fe3790c47957c74076

Observation 587966da-50c4-4cc3-83bd-94e80bf9dd04 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.097267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:6d2315c011359fff0b494020eafae67edfc6ffe9349ef4774d806b3355d0e985

Observation e32b2ea2-1a58-457f-979f-3b3d0eeae31a · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.675857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:befdee8d547bc169c544d5018f2fafd90bc20f53fad26ea9dfbb92f06fcd6a33

Observation 088eedf2-c4f2-4838-b9e4-7ee7292faf8c · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.103118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:fe906d7b2d9ddc20b6d1974bd206f2768193e056bf9ab33e671abc8fdc6cb522

Observation b5332958-03b3-4694-8bc9-a08f8aa574a7 · outbound

This paper cites Demystifying llm-based software engineering agents.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Demystifying llm-based software engineering agents

Reference 23

Resolution
verified exact
doi, observed 2026-05-17T23:15:26.616947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:43bcedb5025576667897146e64cae0973f68fbd7d7f591da3a578365715b4254

Observation 37f2ec39-82a2-48c0-a334-6b26acdf5d34 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.698984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:50667a65141580711f46fbf42175ead642d2d814fdaebf8542ebcf6b365b63c4

Observation 3c6e2350-a71e-4b38-b8a7-d1f6d650143f · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:15:26.757913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:f53f08c88b6a27c628c0ebf0d28d7ee3d6ba192865dd69651397c22081f7fbf7

Observation 7d664539-6eb5-49a5-888c-b6feb604fda2 · outbound

This paper cites Qwen3 Technical Report.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Qwen3 Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.693481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:f85f7ff1b327afe6f55b972c9c5416e306d1c54483c430f54942f5c460ce5162

Observation 07f74694-db5e-4fad-a0f2-5248e4aac6fd · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.089038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:7d194e8bb0684b7f52b972d5f1ee9d9996d4e7d59b10c66efa835b21a8b3bc50

Observation 22249bd6-35ac-4837-870e-f8c1fba5d130 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.091879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:bdaaddeb9548d3597797cdb853c3858d100148229876e3fe1b019490bc36d80b

Observation 1831d8ae-812e-4df3-a114-7d67837e9418 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.725203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:ce63fba7e2c449eb88e13eab9c28b9c9bddec89bb0423d145ff769bf004854d4

Observation 04809f1f-fea8-4b58-a2c3-8888db5b6d5a · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:15:26.747000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:a6984df40a592694804d2804f9f8809feb6316d038e516185e422d6031c048e0

Observation 99842e8b-521f-4a83-b48a-a42e6b4d92f8 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.734331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:7a5c174ae5fe26895ae049c7bad4ab1f8bedf8e5953a70d532d0363e528d0409

Observation bb5db8fd-4415-4d2c-bb58-9cc652e33de5 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:15:26.752887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:dc21541cac9ddab7ecc4418470439a4aef9fb871177b0d3292632d27d5f20f8d

Observation 7f5e4ec9-72af-4c30-bd9b-54387848f990 · outbound

This paper cites Group Sequence Policy Optimization.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Group Sequence Policy Optimization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.763421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:0020f13ec41f83c7752d6e9b9c2a1f9d2704e7fe9c0207f2ac6bd80ceb0188d0

Observation bce6996a-dd7c-4530-8e04-81d35c727c3e · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.086126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:2948f17eb6c5fa5ecbec9146e24a561b0c66a3de48940aa8a7394806fbf8b842

Observation a5897ef3-213c-4350-8834-62f414bfaa3b · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.080401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:8671c79ec927c0fb4fb94bc0155f33af9d79dd4395d8289fc73f1ef4f2baad20

Observation e9b48f81-5e85-4f64-8170-38d6f713ba79 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.094592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:7a83a357daeb225fe824fb269a3f2639e86399745d87cc8a1ed0db26c2a068bc

Observation e0f11304-120d-49c9-ba97-261317bcfb35 · outbound

This paper cites an unresolved cited work.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:15:27.083288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:5f5c2dfadccc1e8d0172d584bf656775371e1d9d38538622e41664658f741816

Pith citing papers

Observation 7307b0b4-e442-4238-bc34-fa070b144c78 · inbound

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute cites this paper.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.262444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.262444Z digest=sha256:8e662ed387185997dc5837d15fabbe3c877d2a1d55087c476c88fe7bd05f96e0

Observation 5a218909-b531-48f3-9c29-d41cf980f411 · inbound

TaPR: Test-Aware Policy Refinement for Feedback-Conditioned Code Generation cites this paper.

TaPR: Test-Aware Policy Refinement for Feedback-Conditioned Code Generation MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:54:53.401490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:54:52.398402Z digest=sha256:ffa5a47daddd0b4ea2807be6190d5c2a92ee0933326998dbc4b78e31c70f13d7