Pith. sign in

Paper Citation Record · LEDGER

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 18 inbound Pith citation observations for arXiv:2506.11425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11425 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:16:08.504844Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:47:59.075366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:55.557008Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e8a4e8c-3519-4492-9a04-355aaf229ffb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.416963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.416963Z digest=sha256:19a11cdc06592de803922b91263adcc3ee3b78e9f8ef74073dedbca1965101f5

Observation 2be21a07-10c6-4a7e-8a2c-de4d7ffe1fbd · outbound

This paper cites Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:16:08.792345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:16:08.420827Z digest=sha256:e91f33d06cd0a6c1201803dce724f0ae686e8e6bca9922d9fb34410e14cc3b18

Observation 26e9f587-880d-4a73-af48-7edce2138133 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Francis Christiano.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Francis Christiano

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:16:08.782440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:16:08.424091Z digest=sha256:72c91f9f7ac735b3b695f2f463fd3ff6dd320c8db3322bfb0a0e94bd4485564b

Observation 9ef3b7a6-547a-4159-8202-4cc7b3f3955b · outbound

This paper cites Training language models to follow instructions with human feedback.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Training language models to follow instructions with human feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.427531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.427531Z digest=sha256:577ec0178fea081e730a588cb78a8b3218aed95c8af025cfab296c74818ad4cd

Observation 52a0bac0-0a79-48b3-9954-97d235ffae56 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.431125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.431125Z digest=sha256:208120a9c3d5d6be6c6cccb11bbbe3cc204690aca86eaf36899c3fb7fec5980f

Observation 648ece32-822b-46ac-bbb7-c8210ea004b3 · outbound

This paper cites Program Synthesis with Large Language Models.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Program Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.434730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.434730Z digest=sha256:7cba898f038768ba704229196dbc5ca25bf43ed4d4f163355bbb0fdaec673a66

Observation 296b8740-ec45-45a5-ac84-fbdb4c8ad79e · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Measuring Coding Challenge Competence With APPS

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.438286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.438286Z digest=sha256:18ecd7fb3080421a5f790b725535462129437ff5fc0177caf892e19a81f78681

Observation d0d8936d-8db1-40fc-b6db-a3e511e0c6b0 · outbound

This paper cites Planning In Natural Language Improves LLM Search For Code Generation.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Planning In Natural Language Improves LLM Search For Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.441362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.441362Z digest=sha256:8f315ab6a9fa00ccca46d5fcb8e3a6d21f6b2682ee9bd8c6d4d5977828413f69

Observation 4315de77-c1db-4f60-8992-5305968fe354 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.444325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.444325Z digest=sha256:94442731a4fc94463328950aaa49775b6b4689d906d3c4eb0555d84c445157a1

Observation d9d2ea16-8ced-45e5-b3b9-8b4c7e20e7a4 · outbound

This paper cites Training Software Engineering Agents and Verifiers with SWE-Gym.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Training Software Engineering Agents and Verifiers with SWE-Gym

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.447400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.447400Z digest=sha256:5dfee7b440fb317c81b742640b8997cc909367766a080e3e6382b8535640cde0

Observation ca8a7a1b-1699-4885-a4f5-e1e89984aa7a · outbound

This paper cites STar: Bootstrapping reasoning with reasoning.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards STar: Bootstrapping reasoning with reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.450383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.450383Z digest=sha256:ca023eb4aa7233996e4211ad89466ba2e132aafbb45a32aa93082f29d4aba183

Observation 3a744ce8-b809-4fce-8d52-db70b358bad4 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Reinforced Self-Training (ReST) for Language Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.453021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.453021Z digest=sha256:2e00508665e4cb711445ea1066e0d26ab340bb04b5c1125f8cf12b159a2ec519

Observation beb20196-bca5-4497-ad80-500001be350b · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Agentless: Demystifying LLM-based Software Engineering Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.456160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.456160Z digest=sha256:bf79f1d985148aeef6b2f6427d08b1e12ce9ea06267a1aea706bd87857240b65

Observation 1326b1c1-5622-47b3-b55f-f8146d15cacb · outbound

This paper cites Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:16:08.767942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:16:08.459059Z digest=sha256:69983c50cfa99ed276df3c0e29d27e1a79f2eb336952d6b104bf9486dbf0daaf

Observation 0d47d712-a232-45d9-8689-ce8afc3db309 · outbound

This paper cites SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.461930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.461930Z digest=sha256:89ed879cc8ec0a616fc9ddbbedcc15999b5aa092510bc9bcc341e680f6d5a332

Observation 23dfffdb-63d4-418f-819d-d61496ea7326 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.464749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.464749Z digest=sha256:c0496ee030ef2c97d9f2f72bee706741dafd78f3fb52ba047f308dbccdd171e8

Observation 1925b323-2750-4627-87d1-38c89723b7df · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.467524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.467524Z digest=sha256:cc49997709f431ed34695692f5a1856c41ddd926d434de7f7f275201c7a94636

Observation 1e6cc14a-eafc-44b5-8a91-32cb06d268a9 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.471002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.471002Z digest=sha256:6359faba1abc9f230df4604e864c6094702865b58a84c3a23cfa174fc424d479

Observation 28a93968-05ef-4510-b354-23379f10466e · outbound

This paper cites DeepSeek-V3 Technical Report.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeek-V3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.473951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.473951Z digest=sha256:0ef7e4f1f669965d059f4a6419ef28e1f332f5be2c621e6c79566ab2ed5213fa

Observation 9c9ef027-3deb-4379-bb19-94e1e6950ccf · outbound

This paper cites Commit0: Library Generation from Scratch.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Commit0: Library Generation from Scratch

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.477293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.477293Z digest=sha256:e5a0169c87a262d8f0989e55fc9a343f8e5ea8d14cb0e73557c812363ebb9668

Observation ccac6345-da87-4ac5-a2ed-e3325ea68349 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.480495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.480495Z digest=sha256:cf19ff2884fe9bb0f6609bf5cc43794b12cbd05ff63d8111aeac88b0f8cb6172

Observation 76698c99-b65f-464a-9946-d155b06aff8d · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards ReAct: Synergizing Reasoning and Acting in Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.483469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.483469Z digest=sha256:f73db36dfb5368a2d9d3f6660b2a3b4c41824b41ab147c799e0c5da3b17b69db

Observation 110172c0-7cf8-4510-8de3-9f3642ec23ee · outbound

This paper cites Proximal Policy Optimization Algorithms.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.486283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.486283Z digest=sha256:341de4868532d450ce0ec768a54d4b9d55719d730213168bbb6bd360f5468942

Observation 228809fc-4eca-4d16-96c0-1a1a8348cfe5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.489630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.489630Z digest=sha256:56a702fb368346e45843067772a85743230b4f1f3cd1f72143357987a1631bd1

Observation e082b267-d4f3-424e-adf6-f4a1d9132bfd · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.492855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.492855Z digest=sha256:61c0f1c97988e4c131c959ea73eacb370395e6d9c4b2cbd4ead58d83ea916ce3

Observation c4f109a2-c72f-48c7-b4d2-63befd7472aa · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Mind2Web: Towards a Generalist Agent for the Web

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.495825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.495825Z digest=sha256:6b159255fcb7884fea90e325de029f0bd53aa7f83e1af4fa72b43e7079acd128

Observation 05cff3f7-a161-4d6c-9da5-72e6cb9872cf · outbound

This paper cites ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.498785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.498785Z digest=sha256:a88cf4f8b6aaf568c4ea7bf54c9885acd709a30abc0c84fc66c99adc42cc7906

Observation 83c58c56-16f3-4c58-97a3-5084691e6531 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.501801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.501801Z digest=sha256:476f606ebb423e7d8b9aa276c3e9546bebacef63e1424e43a1a985cbdf182c36

Observation e6aefe98-2f29-4ad0-b8e9-c6a42f5efe51 · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.504844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.504844Z digest=sha256:a10ce0d43e276622a5a16b34d7ae64ab8b6d0b5db7397df1e244700d7be7f633

Pith citing papers

Observation 781e070c-9bf1-4920-951e-e58502d4508d · inbound

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? cites this paper.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.748215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:419a5a3a72636743c7e65c847781f2da26dd8e5a82c1be153b7bbebbefc4c514

Observation 5b423429-604d-4cd8-81b3-3549e256206a · inbound

SWE-IF: Aligning Code Evaluation with Human Preference cites this paper.

SWE-IF: Aligning Code Evaluation with Human Preference Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:02:07.533516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:02:07.533516Z digest=sha256:9bfcfa98494e995413c6408c4c4282a5ce83dbd165751d3289700a4de3d41212

Observation be23bde6-0887-4ca3-8224-7e9394f1323a · inbound

SERA: Soft-Verified Efficient Repository Agents cites this paper.

SERA: Soft-Verified Efficient Repository Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T07:17:13.394276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:17:13.394276Z digest=sha256:30f77aa37da1f691fc85d4c9f2d3b543f07d222fd109e922d3f1a99ccf501c54

Observation e5d51157-6f35-401b-8ef8-9bbce2b3ba2b · inbound

Fate of Secondary Droplets Produced by High-speed Raindrops Interacting with a Liquid Pool cites this paper.

Fate of Secondary Droplets Produced by High-speed Raindrops Interacting with a Liquid Pool Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T22:37:45.905230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:37:45.905230Z digest=sha256:a116192ec8d5f6a5dfadf715d343f60570dfc1853d5ca88ed7d43d9c5f7ef3aa

Observation 185f9dde-b76d-43cc-aa4a-48b40b711c9d · inbound

SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents cites this paper.

SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:00:59.385102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:21:12.961613Z digest=sha256:527838dd7050121b43396dd48e470b6e32df5798b4f81eca965217f3cf6bc212

Observation b3de5738-ffac-4b63-a2f8-55a9dce22fc5 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.121327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:e0936c89189391f82a9f99ab04b7a6a89b69b4507cfd6234948f27b3a38fad26

Observation 740bad67-0768-4658-9cd9-47c50619a692 · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:22:54.783701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:21:26.640493Z digest=sha256:07813c47a5ae83520e2005644a56c076674ec95f5a58c4ef082ae1524f6c5fa0

Observation 7bfd02ae-5fcf-4231-b1ec-fc33107928ff · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:43:45.447084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T21:42:49.452347Z digest=sha256:b9b3a1977d4ca705b2265435a386dcf0cb4b1022b699daee9126b3b6d096e616

Observation 3a167a8b-7dcc-41da-90bf-ab85eb4d7060 · inbound

ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL cites this paper.

ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.921429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:49:40.239199Z digest=sha256:39624cdc427e476e37ef0ba5a8474c725fe8952d4e6f1b4cc1e9bed5bfb2ea5a

Observation 076142ad-da01-4de0-a46d-bc4db2f109f8 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.212323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:10:06.685591Z digest=sha256:17629a9b80223ced377b7da75ecc46004275f09acbe5dc0eff1658fd1e1d9574

Observation 9c15bb12-a2e7-4b9f-8422-7f29e39538a1 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T15:15:09.675778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:15:09.675778Z digest=sha256:432a52920c95ad4e17801f464eaa7a37a900c8ca8ad2784b158720ba9173fc48

Observation b55f1ba7-cf7d-4b55-b1b0-60b03a905c3b · inbound

AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation cites this paper.

AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.320110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:03:54.185019Z digest=sha256:d671129b759f0bf8f7046b2549b6d45d0e63c012632efd75c34b2917fc7c98c3

Observation cdbe7401-d4c8-4938-b932-ec10e1591b6f · inbound

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents cites this paper.

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:58:33.180739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T06:49:12.070481Z digest=sha256:6ed56a24d6c5ec71632b1b5d9597408c8617f4aac6d7813da672685a6fda6cf8

Observation 8f126582-5c6b-4f46-86c0-ef9ee3d9f93d · inbound

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows cites this paper.

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:55.558939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:18:07.134598Z digest=sha256:3d3278cdf0fd8a827dbdaba6aa03c5782a6cb7b990c28587b5326af4040ff3a1

Observation daced041-f2dc-4f63-8d6b-cc3dd8ae344b · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:cd23ad3069f72280bf72cf1252a1187cd5e604a026e2eeb64be6462a58d9d649

Observation 6bd6e6ce-b523-4888-a300-1909ca46bcc4 · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:10.280348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:10.280348Z digest=sha256:935407ed8e1f4127e923de4934c59729f1eb9fba5b2ce7dbbf4cebcd92096df2

Observation 003a6523-95fe-4860-8329-fc4f1bbafd50 · inbound

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning cites this paper.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.759049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.759049Z digest=sha256:c023f001397992ebb7fc980f918d4dd58b6d0434b394faa57a50e033aed400a7

Observation a67c7176-23ad-40d9-a1f1-c3825501dc7b · inbound

Self-Evolving Coding Agents cites this paper.

Self-Evolving Coding Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T19:47:59.075366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:47:59.075366Z digest=sha256:a7def2aeb14f97fabd14ae9f1625da75e5cb7257d5f5f63c00fae5a07e186890