SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

Adnan El Assadi; Albert Liu; Christopher Settles; Daniel Wang; Derek Chen; Erik Quintanilla; Fenil Faldu; Ishan Gupta; Ivan Bercovich; Jesse Hu

arxiv: 2606.07682 · v1 · pith:BJQH5Z2Knew · submitted 2026-06-05 · 💻 cs.SE · cs.AI

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

Rishi Desai , Jesse Hu , Joan Cabezas , Neel Harsola , Pratyush Shukla , Roey Ben Chaim , Adnan El Assadi , Omkaar Mukund Kamath

show 18 more authors

Fenil Faldu Prannay Hebbar Jiankai Sun Yiyuan Li Pramod Srinivasan Ishan Gupta Christopher Settles Daniel Wang Derek Chen Pranav Raja Albert Liu Marek \v{S}uppa Nevasini Sasikumar Luyang Kong Erik Quintanilla Xiangyi Li Ivan Bercovich Steven Dillmann

This is my paper

classification 💻 cs.SE cs.AI

keywords swe-marathonagentsagenttasksbenchmarkscompletecurrentenvironment

0 comments

read the original abstract

AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent benchmarks largely evaluate short-form tasks, such as single pull requests, small tickets, or 5-10 minute exercises, limiting our ability to measure agents' capabilities in planning, long-context understanding, and memory use. We introduce SWE-Marathon, a benchmark of 20 long-horizon tasks spanning software engineering and adjacent technical domains. Each task consists of a unique executable environment, a human-written reference solution, and a multi-layer verification suite. Logged agent attempts average 27.2M total tokens, making SWE-Marathon substantially longer-horizon than existing SWE and command-line agent benchmarks. Current frontier coding agents solve fewer than 30% of tasks. Failures often arise from poor self-verification, self-reported infeasibility, and premature termination. We also observe reward-hacking behavior in 13.8% of rollouts, where agents attempt to exploit the environment or verifier to bypass the intended workflow. SWE-Marathon includes adversarial review of test suites and execution environments, as well as multi-layer checks designed to prevent shortcut solutions. We release SWE-Marathon, evaluation code, and agent trajectories at https://swe-marathon.org/.

This paper has not been read by Pith yet.

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

discussion (0)