{"work":{"id":"d6614c64-acca-4134-bd4e-80ec66ffdf31","openalex_id":"https://openalex.org/W6966659450","doi":"10.48550/arxiv.2504.07164","arxiv_id":"2504.07164","raw_key":null,"title":"R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents","authors":null,"authors_text":"Naman Jain, Jaskirat Singh, Manish Shetty, Liang Zheng, Koushik Sen, and Ion Stoica","year":2025,"venue":"cs.SE","abstract":"Improving open-source models on real-world SWE tasks (solving GITHUB issues) faces two key challenges: 1) scalable curation of execution environments to train these models, and, 2) optimal scaling of test-time compute. We introduce AgentGym, the largest procedurally-curated executable gym environment for training real-world SWE-agents, consisting of more than 8.7K tasks. AgentGym is powered by two main contributions: 1) SYNGEN: a synthetic data curation recipe that enables scalable curation of executable environments using test-generation and back-translation directly from commits, thereby reducing reliance on human-written issues or unit tests. We show that this enables more scalable training leading to pass@1 performance of 34.4% on SWE-Bench Verified benchmark with our 32B model. 2) Hybrid Test-time Scaling: we provide an in-depth analysis of two test-time scaling axes; execution-based and execution-free verifiers, demonstrating that they exhibit complementary strengths and limitations. Test-based verifiers suffer from low distinguishability, while execution-free verifiers are biased and often rely on stylistic features. Surprisingly, we find that while each approach individually saturates around 42-43%, significantly higher gains can be obtained by leveraging their complementary strengths. Overall, our approach achieves 51% on the SWE-Bench Verified benchmark, reflecting a new state-of-the-art for open-weight SWE-agents and for the first time showing competitive performance with proprietary models such as o1, o1-preview and sonnet-3.5-v2 (with tools). We will open-source our environments, models, and agent trajectories.","external_url":"https://arxiv.org/abs/2504.07164","cited_by_count":0,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2504.07164","created_at":"2026-05-09T22:13:58.368435+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"R2E-gym: Procedural environments and hybrid verifiers for scaling open-weights SWE agents","render_title":"R2E-gym: Procedural environments and hybrid verifiers for scaling open-weights SWE agents"},"hub":{"state":{"work_id":"d6614c64-acca-4134-bd4e-80ec66ffdf31","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":25,"external_cited_by_count":0,"distinct_field_count":4,"first_pith_cited_at":"2025-04-28T11:08:22+00:00","last_pith_cited_at":"2026-07-07T09:37:45+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T15:49:38.512863+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":4},{"context_role":"dataset","n":2}],"polarity_counts":[{"context_polarity":"background","n":5},{"context_polarity":"use_dataset","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}