Pith. sign in

Paper Citation Record · LEDGER

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

As of 20 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 19 inbound Pith citation observations for arXiv:2506.11425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11425 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:16:08.504844Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:39:03.647525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:55.557008Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e8a4e8c-3519-4492-9a04-355aaf229ffb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.416963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.416963Z digest=sha256:f6ae125965c3bf9a5a0e94590b58a439604644331f8fb245360613e96813b058

Observation 2be21a07-10c6-4a7e-8a2c-de4d7ffe1fbd · outbound

This paper cites Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:16:08.792345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:16:08.420827Z digest=sha256:e7fb4a99cf4ea45377b7f2cfb5bec8a8a9e6c00fca228d5369c66953ff3cb127

Observation 26e9f587-880d-4a73-af48-7edce2138133 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Francis Christiano.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Francis Christiano

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:16:08.782440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:16:08.424091Z digest=sha256:c1822b5aecb94f8c3bd8ec63252b4d8bf400ffe3b81fd86746abdc4bb3960668

Observation 9ef3b7a6-547a-4159-8202-4cc7b3f3955b · outbound

This paper cites Training language models to follow instructions with human feedback.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Training language models to follow instructions with human feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.427531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.427531Z digest=sha256:eac973a825c413cf8bd593cbfc5f27671e846378e19aeff017a438f277a13af0

Observation 52a0bac0-0a79-48b3-9954-97d235ffae56 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.431125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.431125Z digest=sha256:40f815abfe1f055b0c61d32bc80a6712ecafcce8b53c8259135794618243438c

Observation 648ece32-822b-46ac-bbb7-c8210ea004b3 · outbound

This paper cites Program Synthesis with Large Language Models.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Program Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.434730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.434730Z digest=sha256:704fb9ac24c24810c2d76172251993aa66d3580a4bffe8c992354eceacec5a30

Observation 296b8740-ec45-45a5-ac84-fbdb4c8ad79e · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Measuring Coding Challenge Competence With APPS

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.438286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.438286Z digest=sha256:33202444ac3e1ec3893c8b0d34c4e4f73b0fa7e0b0a6c3b6f438f8cd40aebf4c

Observation d0d8936d-8db1-40fc-b6db-a3e511e0c6b0 · outbound

This paper cites Planning In Natural Language Improves LLM Search For Code Generation.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Planning In Natural Language Improves LLM Search For Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.441362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.441362Z digest=sha256:f66e611eecfe54413fdce844f893b80f3e4f0148b29a726311fc057739311673

Observation 4315de77-c1db-4f60-8992-5305968fe354 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.444325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.444325Z digest=sha256:06576af54bfeda9631e3115cc3183f0264b43730e1e070ac6168d78f3ff83261

Observation d9d2ea16-8ced-45e5-b3b9-8b4c7e20e7a4 · outbound

This paper cites Training Software Engineering Agents and Verifiers with SWE-Gym.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Training Software Engineering Agents and Verifiers with SWE-Gym

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.447400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.447400Z digest=sha256:48154684a026a19482c88bcffedb6762f51609e678d1ac7a1369fd12b416478c

Observation ca8a7a1b-1699-4885-a4f5-e1e89984aa7a · outbound

This paper cites STar: Bootstrapping reasoning with reasoning.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards STar: Bootstrapping reasoning with reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.450383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.450383Z digest=sha256:02c24ee73219a308d4aff90a21b40b50a3200c00e73064b408c76dd6c3c6b1d5

Observation 3a744ce8-b809-4fce-8d52-db70b358bad4 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Reinforced Self-Training (ReST) for Language Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.453021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.453021Z digest=sha256:a081f5bd47a6de10eb33e0275b9c6091166d6fe51d18d7f96318b9ea6543960e

Observation beb20196-bca5-4497-ad80-500001be350b · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Agentless: Demystifying LLM-based Software Engineering Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.456160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.456160Z digest=sha256:ed4096a0420216629665f485d33e32fddbc77728e2a0bb9fc80c4d7124c02547

Observation 1326b1c1-5622-47b3-b55f-f8146d15cacb · outbound

This paper cites Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:16:08.767942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:16:08.459059Z digest=sha256:42927c0c17d0dc0038984e352b35488cbfc8c52b481c894eaa08257d257a10a9

Observation 0d47d712-a232-45d9-8689-ce8afc3db309 · outbound

This paper cites SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.461930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.461930Z digest=sha256:f94820bc57e755f90c058f486cc8d30a0b701b8e0cfae929381ed293e585b65c

Observation 23dfffdb-63d4-418f-819d-d61496ea7326 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.464749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.464749Z digest=sha256:0c049afafac2ef81c0030ed5d2f40d6a9bc142af85ac0c96866b005a7c171905

Observation 1925b323-2750-4627-87d1-38c89723b7df · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.467524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.467524Z digest=sha256:b058b9fae8a07ba745b063cadeea731b95dcac7a4b38023886db57b27a389880

Observation 1e6cc14a-eafc-44b5-8a91-32cb06d268a9 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.471002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.471002Z digest=sha256:c6fb7e5afbcf70fab845a0bc4ab86f4b66aa2d9fe57e8982a5c791f1f872d6cb

Observation 28a93968-05ef-4510-b354-23379f10466e · outbound

This paper cites DeepSeek-V3 Technical Report.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeek-V3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.473951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.473951Z digest=sha256:584242924058ab0c6fa2ce3c70e34e7583b9b103a533dcad61464f27229f7543

Observation 9c9ef027-3deb-4379-bb19-94e1e6950ccf · outbound

This paper cites Commit0: Library Generation from Scratch.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Commit0: Library Generation from Scratch

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.477293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.477293Z digest=sha256:7284935677e88f3b1b336a8135b0f897c9e2a7305d72ecf4449512b2cc675d4f

Observation ccac6345-da87-4ac5-a2ed-e3325ea68349 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.480495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.480495Z digest=sha256:f02b70c9ad25a2d57c6b8467a6c2b1c05094cca296d7bdbfc97219e86f02943f

Observation 76698c99-b65f-464a-9946-d155b06aff8d · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards ReAct: Synergizing Reasoning and Acting in Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.483469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.483469Z digest=sha256:06b74c2353d4fe29e9b8fd6d1a8f2f15b83525f54754fbe8dafcb4daebf2a376

Observation 110172c0-7cf8-4510-8de3-9f3642ec23ee · outbound

This paper cites Proximal Policy Optimization Algorithms.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.486283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.486283Z digest=sha256:a4d5d3bac7fcf8a39a2944bab24b94036ec952d0f6b347ab1533db440cff7a7d

Observation 228809fc-4eca-4d16-96c0-1a1a8348cfe5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.489630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.489630Z digest=sha256:91b8c38bae776e94830b8b0e922fc9c69fb1e5436f4ece60aa33475c962462ea

Observation e082b267-d4f3-424e-adf6-f4a1d9132bfd · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.492855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.492855Z digest=sha256:ef0a2684ec9096a7b52814d725b6ac0384bb165723d6ed33a9bd4a903e43e676

Observation c4f109a2-c72f-48c7-b4d2-63befd7472aa · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Mind2Web: Towards a Generalist Agent for the Web

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.495825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.495825Z digest=sha256:e83d2c8ebca8c92de9af98f54fea37cfeea11203fc1e1ac6d1fd034e597b91a6

Observation 05cff3f7-a161-4d6c-9da5-72e6cb9872cf · outbound

This paper cites ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.498785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.498785Z digest=sha256:56ed0ac9f3326f312c64044a7e953f580249e6605810e21b02787eb82d41a883

Observation 83c58c56-16f3-4c58-97a3-5084691e6531 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.501801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.501801Z digest=sha256:48bafd14459dc2b028c823c1b53e6c0139263e40034b5bb050c835bbd79411fb

Observation e6aefe98-2f29-4ad0-b8e9-c6a42f5efe51 · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.504844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.504844Z digest=sha256:6afa13338ed3033ba985b3d83d36915c0a4e0ac5a651f5ca0c4cc9b7b6a447e5

Pith citing papers

Observation 781e070c-9bf1-4920-951e-e58502d4508d · inbound

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? cites this paper.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.748215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:7184669bb1741467473d48227d809b56be02cd3d9651b114352022edfc432eb5

Observation 5b423429-604d-4cd8-81b3-3549e256206a · inbound

SWE-IF: Aligning Code Evaluation with Human Preference cites this paper.

SWE-IF: Aligning Code Evaluation with Human Preference Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:02:07.533516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:02:07.533516Z digest=sha256:6aa1c64591fc77040cf8097410aee764035131f33b221778d11e46fb5a357caf

Observation be23bde6-0887-4ca3-8224-7e9394f1323a · inbound

SERA: Soft-Verified Efficient Repository Agents cites this paper.

SERA: Soft-Verified Efficient Repository Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T07:17:13.394276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:17:13.394276Z digest=sha256:b151bc06227d50c3d922f56eaa10c325891e74917c731f1647a2bbd0a5f96f78

Observation e5d51157-6f35-401b-8ef8-9bbce2b3ba2b · inbound

Fate of Secondary Droplets Produced by High-speed Raindrops Interacting with a Liquid Pool cites this paper.

Fate of Secondary Droplets Produced by High-speed Raindrops Interacting with a Liquid Pool Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T22:37:45.905230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:37:45.905230Z digest=sha256:47b7523138c5022120520c79490b2ac20aa57c618e43af92ce6eaddf7ee82fb5

Observation 185f9dde-b76d-43cc-aa4a-48b40b711c9d · inbound

SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents cites this paper.

SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:00:59.385102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:21:12.961613Z digest=sha256:66e7cb8eb04dfb3574df28ee2e4b8dea3caf392c10febe9f00802b4dcef24a05

Observation b3de5738-ffac-4b63-a2f8-55a9dce22fc5 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.121327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:5eac83f6ed19da410dd91f374dbb752ebed0acb6ef0d6a6cef7182e3d9a5f9ec

Observation 740bad67-0768-4658-9cd9-47c50619a692 · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:22:54.783701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:21:26.640493Z digest=sha256:c0139ad5ffba562c12a73028efe5ef5d8d00efe7260ee469db71c8778e517941

Observation 7bfd02ae-5fcf-4231-b1ec-fc33107928ff · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:43:45.447084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T21:42:49.452347Z digest=sha256:41b0078d9dc5fc1f6bcc74a3c8630b4a3470ab5422762645d344d0f14d6fdea5

Observation 3a167a8b-7dcc-41da-90bf-ab85eb4d7060 · inbound

ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL cites this paper.

ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.921429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T14:49:40.239199Z digest=sha256:a1bd82b891e6c205a8e37a8d62191836e451238cf845dfcffc099f25e5407ed3

Observation 076142ad-da01-4de0-a46d-bc4db2f109f8 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.212323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T11:10:06.685591Z digest=sha256:ddebe96c32ff03b5f39c942209f6033f87be3bb37290002a5c73c2ed81c9fe0e

Observation 9c15bb12-a2e7-4b9f-8422-7f29e39538a1 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T15:15:09.675778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:15:09.675778Z digest=sha256:5fd0dba3f48ac7328eda9b243a8e9fa39ca663530900d2366200c861def5b723

Observation b55f1ba7-cf7d-4b55-b1b0-60b03a905c3b · inbound

AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation cites this paper.

AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.320110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T10:03:54.185019Z digest=sha256:578b37c246b53758482acf31257a4066b75c0bc809dceaf9e4f5ea8e4b9494d0

Observation cdbe7401-d4c8-4938-b932-ec10e1591b6f · inbound

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents cites this paper.

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:58:33.180739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T06:49:12.070481Z digest=sha256:942ac874bee806bab10b88eaf470cf5199919ef1c4d5c899c8d3f83215a09c89

Observation 8f126582-5c6b-4f46-86c0-ef9ee3d9f93d · inbound

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows cites this paper.

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:55.558939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T20:18:07.134598Z digest=sha256:132c6f5ab79de60ac9e209348b36cab5d0b5805ef5c8633b9b7aa437670fd27b

Observation daced041-f2dc-4f63-8d6b-cc3dd8ae344b · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:7a7c7ecae5d1ad99fe4c9179d7329b75414f0a42f5ae7135f6fd1dbaa367b413

Observation 6bd6e6ce-b523-4888-a300-1909ca46bcc4 · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:10.280348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:10.280348Z digest=sha256:4b929e03f67bd77816901c03ab61f1fa8597b7c2407e01b9dffcea0a333984f8

Observation 003a6523-95fe-4860-8329-fc4f1bbafd50 · inbound

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning cites this paper.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.759049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.759049Z digest=sha256:2b6a757b4cf6e7a46593c1e0f0e1cf5ae5c47adf2182f50f098aba5b4d8cb9a6

Observation a67c7176-23ad-40d9-a1f1-c3825501dc7b · inbound

Self-Evolving Coding Agents cites this paper.

Self-Evolving Coding Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T19:47:59.075366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:47:59.075366Z digest=sha256:6c705e657660c2d00dfca16606a12c58b8b69122661ecaad9cec46fe8b566360

Observation ca4ee6f6-5d92-472b-95e6-2c5810b604f1 · inbound

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing cites this paper.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.647525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.647525Z digest=sha256:c89e974c2ab4e753bf4b0fe4ab5bdf1087599fc3f71bea860ecb553f0cd975b5