Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:16:08.504844Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 18 inbound Pith citation observations for arXiv:2506.11425.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:16:08.504844Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T19:47:59.075366Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:18:55.557008Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2e8a4e8c-3519-4492-9a04-355aaf229ffb · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2be21a07-10c6-4a7e-8a2c-de4d7ffe1fbd · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26e9f587-880d-4a73-af48-7edce2138133 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Francis Christiano
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ef3b7a6-547a-4159-8202-4cc7b3f3955b · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Training language models to follow instructions with human feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52a0bac0-0a79-48b3-9954-97d235ffae56 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648ece32-822b-46ac-bbb7-c8210ea004b3 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Program Synthesis with Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 296b8740-ec45-45a5-ac84-fbdb4c8ad79e · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Measuring Coding Challenge Competence With APPS
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d8936d-8db1-40fc-b6db-a3e511e0c6b0 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Planning In Natural Language Improves LLM Search For Code Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4315de77-c1db-4f60-8992-5305968fe354 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9d2ea16-8ced-45e5-b3b9-8b4c7e20e7a4 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Training Software Engineering Agents and Verifiers with SWE-Gym
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca8a7a1b-1699-4885-a4f5-e1e89984aa7a · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards STar: Bootstrapping reasoning with reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a744ce8-b809-4fce-8d52-db70b358bad4 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Reinforced Self-Training (ReST) for Language Modeling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beb20196-bca5-4497-ad80-500001be350b · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Agentless: Demystifying LLM-based Software Engineering Agents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1326b1c1-5622-47b3-b55f-f8146d15cacb · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d47d712-a232-45d9-8689-ce8afc3db309 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23dfffdb-63d4-418f-819d-d61496ea7326 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1925b323-2750-4627-87d1-38c89723b7df · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e6cc14a-eafc-44b5-8a91-32cb06d268a9 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28a93968-05ef-4510-b354-23379f10466e · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeek-V3 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c9ef027-3deb-4379-bb19-94e1e6950ccf · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Commit0: Library Generation from Scratch
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccac6345-da87-4ac5-a2ed-e3325ea68349 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76698c99-b65f-464a-9946-d155b06aff8d · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards ReAct: Synergizing Reasoning and Acting in Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 110172c0-7cf8-4510-8de3-9f3642ec23ee · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Proximal Policy Optimization Algorithms
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 228809fc-4eca-4d16-96c0-1a1a8348cfe5 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e082b267-d4f3-424e-adf6-f4a1d9132bfd · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4f109a2-c72f-48c7-b4d2-63befd7472aa · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards Mind2Web: Towards a Generalist Agent for the Web
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05cff3f7-a161-4d6c-9da5-72e6cb9872cf · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83c58c56-16f3-4c58-97a3-5084691e6531 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6aefe98-2f29-4ad0-b8e9-c6a42f5efe51 · outbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 781e070c-9bf1-4920-951e-e58502d4508d · inbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b423429-604d-4cd8-81b3-3549e256206a · inbound
SWE-IF: Aligning Code Evaluation with Human Preference Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be23bde6-0887-4ca3-8224-7e9394f1323a · inbound
SERA: Soft-Verified Efficient Repository Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5d51157-6f35-401b-8ef8-9bbce2b3ba2b · inbound
Fate of Secondary Droplets Produced by High-speed Raindrops Interacting with a Liquid Pool Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 185f9dde-b76d-43cc-aa4a-48b40b711c9d · inbound
SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3de5738-ffac-4b63-a2f8-55a9dce22fc5 · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 740bad67-0768-4658-9cd9-47c50619a692 · inbound
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7bfd02ae-5fcf-4231-b1ec-fc33107928ff · inbound
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a167a8b-7dcc-41da-90bf-ab85eb4d7060 · inbound
ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 076142ad-da01-4de0-a46d-bc4db2f109f8 · inbound
Trading Human Curation for Synthetic Augmentation in RLVR Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c15bb12-a2e7-4b9f-8422-7f29e39538a1 · inbound
Trading Human Curation for Synthetic Augmentation in RLVR Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b55f1ba7-cf7d-4b55-b1b0-60b03a905c3b · inbound
AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdbe7401-d4c8-4938-b932-ec10e1591b6f · inbound
SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f126582-5c6b-4f46-86c0-ef9ee3d9f93d · inbound
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation daced041-f2dc-4f63-8d6b-cc3dd8ae344b · inbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd6e6ce-b523-4888-a300-1909ca46bcc4 · inbound
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 003a6523-95fe-4860-8329-fc4f1bbafd50 · inbound
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a67c7176-23ad-40d9-a1f1-c3825501dc7b · inbound
Self-Evolving Coding Agents Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.