REVIEW 3 major objections 5 minor 31 references
The paper argues that a declining end-to-end success rate on longer tasks proves deployment limits, not failure mechanisms, and that only a pre-specified compositional baseline plus a horizon residual can license claims about long-horizon m
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 10:25 UTC pith:WNSH2SFN
load-bearing objection A disciplined position paper: the horizon-residual protocol gives long-horizon evaluation a pre-registered way to distinguish compounding from mechanism, and its biggest unknown—checkpoint composability—is already named as the paper's own falsifier. the 3 major comments →
Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that 'longer is harder' conflates distinct phenomena: more required work, harder local decisions, and trajectory-induced degradation, in which earlier execution makes later steps harder (context rot being the special case where the visible text is the culprit). To claim a long-horizon failure mechanism, a benchmark must pre-register a decomposition of the task into verifiable semantic stages, declare a checkpoint protocol specifying environment state, visible history, revealed information, and budget, run the same deployed agent from those checkpoints to estimate conditional stage success probabilities q_i, form the product baseline P_expected = ∏ q_i, and measur
What carries the argument
The load-bearing object is the product baseline: decompose a long task into semantic stages, estimate q_i = Pr(stage i success | prior stages complete, declared checkpoint protocol) using the same agent configuration, and predict P_expected = ∏ q_i. The horizon residual Γ_H = ln(P_expected / P_observed) is the log-ratio of that prediction to natural end-to-end success, measured in nats. It is a protocol-specific contrast, not a causal quantity; its role is to make a mismatch visible so that follow-up interventions can be targeted. The paper also proposes two task annotations—work exposure (number of required verifiable stages) and dependency depth (length of the deepest required chain)—to se
Load-bearing premise
The load-bearing premise is that checkpoint states can be composed into a single coherent rollout; if the checkpoints used to estimate stage probabilities are states that never occur along any natural trajectory, the residual compares two arbitrary protocols rather than diagnosing long-horizon failure.
What would settle it
Run a single persistent-state benchmark under two pre-registered, admissible checkpoint protocols: if the sign of the horizon residual flips between them on the same tasks and no declared admissibility rule excludes either protocol, the residual is not stable enough to serve as a disciplined diagnostic starting point.
If this is right
- End-to-end success curves alone should be treated as deployment results; a mechanism claim requires a pre-specified compositional prediction and the horizon residual.
- A low full-task success rate is compatible with no long-horizon mechanism: for example, four independent stages at 80% local success predict 41% overall success, so the residual, not the raw curve, is what carries diagnostic information.
- Benchmarks that adopt the protocol would pre-register stage decompositions, checkpoint definitions, and budget rules, and would report raw counts, bootstrap intervals, and sensitivity across admissible checkpoint choices.
- The stage-wise decomposition of the residual shows where the natural rollout falls behind the checkpoint baseline, pointing follow-up interventions to the right stage (though not to the origin of the degradation).
- The proposed work-exposure and dependency-depth annotations would let benchmarks report task structure separately from local difficulty, enabling matched comparisons across task families.
Where Pith is reading between the lines
- If residual reporting becomes common, the practical meaning of a leaderboard changes: a model with a near-zero residual on a long task would be understood as having its long-task behavior fully explained by local competence, while a large residual would identify history/state interaction as the bottleneck to target.
- The same counterfactual logic could be applied to historical results: where published benchmarks report both oracle-reset and persistent-state scores, one can compose the reset-state stage scores into a P_expected and compute the residual post hoc, turning old comparisons into a first test of the proposal.
- The dependence of the residual on the chosen decomposition suggests a natural meta-experiment: vary the stage decomposition of a fixed task family and measure how much Γ_H moves; large movement would indicate that decomposition choices, not agent behavior, dominate the diagnostic.
- Outside coding and terminal tasks, replayable checkpoints are harder to construct, so the framework's extension to web, desktop, or embodied agents would likely require approximate checkpoints or learned resets; testing that transfer is an open empirical question.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that declining end-to-end success on long-horizon benchmarks is a deployment observation that does not, by itself, license claims about long-horizon failure mechanisms. It proposes a diagnostic protocol: decompose the task into verifiable stages, run the same agent from declared checkpoints to estimate conditional stage success probabilities qi, form the product baseline P_expected = Π qi (Eq. 1), measure P_observed under a natural end-to-end rollout, and report the horizon residual ΓH = log(P_expected/P_observed) (Eq. 2). The paper emphasizes that ΓH is protocol-specific, not causal; gives a stage-wise decomposition (Eq. 3); provides a delta-method variance approximation (Eq. 4) with caveats recommending bootstrap; discusses interpretation, baseline selection, task annotations (work exposure, dependency depth); proposes a benchmark design protocol; reviews related benchmarks; states scope conditions, limitations, and falsifiers; and closes with a strong claim that mechanism claims require pre-specified compositional prediction plus residual. The paper is a position paper with no new empirical data.
Significance. If the proposal is adopted, it would raise the evidentiary bar for interpreting long-horizon benchmark results: papers reporting only end-to-end curves would be read as deployment evidence, not mechanism evidence. The core mathematical identities (Eqs. 1–3) are elementary but correctly framed, and the paper is commendably explicit about its own scope conditions: the cross-world composition problem, verifier error, protocol sensitivity, zero-probability handling, and the need for bootstrap intervals. The paper also supplies a concrete, auditable null model and a falsifiability list. The main value is as a methodological position piece: it systematizes and disseminates a 'compose local success' check that is often absent in the long-horizon evaluation literature. The paper's strength is its disciplinary clarity about what an evaluative protocol can and cannot establish. However, as the paper itself acknowledges, the practical validity of the whole framework rests on the ability to construct compatible checkpoints at manageable cost; this is not yet demonstrated. The paper has no empirical validation, so its contribution is methodological advocacy rather than established scientific fi
major comments (3)
- [§3.2, Eq. (1) and §7, falsifier 2] The load-bearing premise of the product baseline is that the checkpoint-estimated qi values can be composed into a counterfactual for the natural end-to-end rollout. The paper itself names the 'cross-world composition problem' in §3.2: if the checkpoint states (oracle resets, repaired states, fresh contexts) are not reachable or representative of states along successful natural rollouts, then ΓH is an arbitrary protocol contrast rather than a diagnosis of long-horizon failure. The paper lists as falsifier 2 the possibility that 'compatible checkpoints cannot be constructed at manageable cost for representative coding benchmarks' but provides no evidence that they can be; the cited examples (ChainSWE, SWE-Milestone) use canonical oracle states, not agent-achieved trajectory states. This is not an internal inconsistency, but it means the strong 'only by' claim in §8 is conditional on an op
- [§5, protocol step 7] The robustness criterion says conclusions should be reported only when the sign of ΓH is stable across the pre-declared admissible set. This is sensible but potentially too permissive: sign stability alone ignores magnitude and can be satisfied when all residuals are near zero. The paper should specify a minimal effect size or a resolution threshold (e.g., requiring the bootstrap interval for ΓH to exclude a chosen δ, say |ΓH| > 0.2 nats, and to not cross zero). Without such a threshold, the criterion is not fully auditable and cannot distinguish 'robustly no effect' from 'no evidence'.
- [§3.2, 'For a heterogeneous task family'] The statement that predictions should be composed within each task and then aggregated using the same task weights as P_observed is correct but underdeveloped. The paper should specify the formal aggregation rule: e.g., P_expected = (1/T) Σ_t Π_i q_{t,i} when tasks are equally weighted, and clarify how ΓH should be interpreted when task-level residuals are heterogeneous (e.g., a positive aggregate residual may be driven by a few hard tasks). Without this, the protocol in §5 step 4 is ambiguous for multi-task benchmarks.
minor comments (5)
- [§3.1, Eq. (1)] The notation S<i is not defined; it presumably means S1=...=S_{i-1}=1. Please define explicitly at first use.
- [§3.3, Eq. (4)] The delta-method variance approximation is presented with the correct caveat that independence is violated and bootstrap is recommended. However, the paper does not give a variance expression that accounts for the nested nature of the estimates: later ri are estimated only from rollouts with successful prefixes, and the same task units appear in both checkpoint and natural evaluations. Consider adding an explicit statement that in realistic settings with shared tasks, the effective sample size for the residual is closer to the number of complete task-level units, not the sum of per-stage trials.
- [§4, Table 2 and 'Local difficulty should likewise be measured'] The paper proposes matching measured local difficulty across shallow and deep conditions, but does not explain how to match local difficulty when the stage distributions differ, nor how to detect a failure of matching. The protocol in §5 should include an explicit test of local-difficulty balance (e.g., comparing per-stage success distributions with compatibility intervals) before drawing conclusions from the residual.
- [§6, Table 3] The table is informative but dense; consider adding a column for whether the study reports any composite/expected prediction, which directly relates to the paper's thesis and would make the table's message more immediate.
- [References] Several references are dated 2026 and appear to be preprints (e.g., SWE-Milestone, ChainSWE, SWE-Marathon). The paper does not note which have been peer-reviewed. Please add an explicit statement about the status of each cited benchmark.
Circularity Check
No significant circularity: the horizon residual is a declared protocol contrast, not a prediction fitted to its own target.
full rationale
The paper's central quantity is ΓH = log(Pexpected / Pobserved), where Pexpected is the product of checkpoint-measured conditional stage probabilities (Eq. 1) and Pobserved is an independently measured end-to-end success rate. Neither quantity is fitted to the other; the product baseline is explicitly presented as an auditable null model, not as an output tuned to reproduce the natural rollout. The paper repeatedly states that the residual is not a causal answer and does not identify a mechanism, so it does not rename an input as a discovered explanation. The chain-rule decomposition is definitional rather than a substantive derivation, and the main argument is normative: it proposes a reporting standard. The one load-bearing premise—that checkpoint states can be composed into a coherent counterfactual for the natural rollout—is explicitly acknowledged in Section 3.2 ('otherwise the qi values may describe stages that cannot be combined in one coherent rollout') and is listed as falsifier #2 in Section 7. This makes the practical applicability conditional, but it is a stated scope condition rather than a circularity. I also find no self-citations by the present authors and no fitted parameter renamed as a prediction; cited external benchmarks are used as illustrative evidence, not as the source of the proposed definition. Accordingly, no circular step is exhibited, and the appropriate score is 0.
Axiom & Free-Parameter Ledger
axioms (4)
- standard math Probability chain rule: P(S1∩...∩Sn) = ∏ P(Si | S1∩...∩S_{i-1})
- domain assumption Tasks can be decomposed into semantically verifiable stages with replayable checkpoints and acceptance conditions.
- domain assumption The same deployed agent can be run from declared checkpoint states with a meaningful correspondence to natural rollouts.
- domain assumption Estimation uncertainty can be approximated by bootstrap over complete task units.
read the original abstract
Long-horizon benchmarks often show that agents fail more as tasks become longer. This observation is useful for deployment, but it does not by itself explain why failure occurs. More stages create more opportunities for ordinary errors to compound; longer tasks may also contain harder individual decisions or become harder as conversation history, tool outputs, and environment changes accumulate. We use trajectory-induced degradation to mean this last possibility: earlier execution makes later work harder. When the harmful accumulation is specifically the text visible to the model, it is often called context rot. In this position paper, we argue that to claim a "long-horizon failure", benchmarks must compare actual full-task success against a baseline prediction built from short, individual stages. We call the log-ratio between this prediction and actual success the horizon residual. The comparison must use the same agent configuration and specify in advance how stages, checkpoints, information, and budgets will be chosen. The residual shows that the full rollout differs from the chosen baseline; targeted experiments are still needed to explain why.
Figures
Reference graph
Works this paper leans on
-
[1]
Deng, Gangda and Chen, Zhaoling and Yu, Zhongming and Fan, Haoyang and Liu, Yuhong and Yang, Yuxin and Parikh, Dhruv and Kannan, Rajgopal and Cong, Le and Wang, Mengdi and Zhang, Qian and Prasanna, Viktor and Tang, Xiangru and Wang, Xingyao , journal =
-
[2]
The Illusion of Diminishing Returns: Measuring Long Horizon Execution in
Sinha, Akshit and Arun, Arvindh and Goel, Shashwat and Staab, Steffen and Geiping, Jonas , booktitle =. The Illusion of Diminishing Returns: Measuring Long Horizon Execution in. 2026 , note =
2026
-
[3]
arXiv preprint arXiv:2604.11978 , year =
The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break , author =. arXiv preprint arXiv:2604.11978 , year =
-
[4]
Preprint , year =
Towards Long-Horizon Agents: A Survey---Foundation, Evolution, Harness, Optimization, Application, and Frontier , author =. Preprint , year =
-
[5]
Proceedings of the 43rd International Conference on Machine Learning , series =
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length , author =. Proceedings of the 43rd International Conference on Machine Learning , series =. 2026 , note =
2026
-
[6]
arXiv preprint arXiv:2607.08964 , year =
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading , author =. arXiv preprint arXiv:2607.08964 , year =
-
[7]
, journal =
Lam, Man Ho and Wang, Chaozheng and Liu, Hange and Xiao, Jingyu and Li, Haau-sing and Huang, Jen-tse and Zhuo, Terry Yue and Lyu, Michael R. , journal =. 2026 , url =
2026
-
[8]
2026 , url =
Jin, Qirui and Tung, Lingching and Li, Kenan and Shi, Qiyang and She, Yushi and Jia, Huanzhong and Zhao, Harrison and Xia, Kejing and Du, Zhenbang and Zhang, Yikai and Pei, Jiaxin and Zhang, Zhenyu and Qi, Zhen and Duan, Yuyan and Lee, Wenke and Jin, Zijian , journal =. 2026 , url =
2026
-
[9]
2026 , url =
Orlanski, Gabriel and Roy, Devjeet and Yun, Alexander and Shin, Changho and Gu, Alex and Ge, Albert and Adila, Dyah and Roberts, Nicholas and Sala, Frederic and Albarghouthi, Aws , journal =. 2026 , url =
2026
-
[10]
Le, Tue and Thai, Minh V. T. and Manh, Dung Nguyen and Nhat, Huy Phan and Bui, Nghi D. Q. , journal =. 2025 , url =
2025
-
[11]
2025 , url =
Deng, Xiang and Da, Jeff and Pan, Edwin and He, Yannis Yiming and Ide, Charles and Garg, Kanak and Lauffer, Niklas and Park, Andrew and Pasari, Nitin and Rane, Chetan and Sampath, Karmini and Krishnan, Maya and Kundurthy, Srivatsa and Hendryx, Sean and Wang, Zifan and Bharadwaj, Vijay and Holm, Jeff and Aluri, Raja and Zhang, Chen Bo Calvin and Jacobson, ...
2025
-
[12]
2025 , url =
Ding, Jingzhe and Long, Shengda and Pu, Changxin and Zhou, Huan and Gao, Hongwan and Gao, Xiang and He, Chao and Hou, Yue and Hu, Fei and Li, Zhaojian and Shi, Weiran and Wang, Zaiyuan and Zan, Daoguang and Zhang, Chenchen and Zhang, Xiaoxu and Chen, Qizhi and Cheng, Xianfu and Deng, Bo and Gu, Qingshui and Hua, Kai and Lin, Juntao and Liu, Pai and Li, Mi...
2025
-
[13]
Training Software Engineering Agents and Verifiers with
Pan, Jiayi and Wang, Xingyao and Neubig, Graham and Jaitly, Navdeep and Ji, Heng and Suhr, Alane and Zhang, Yizhe , booktitle =. Training Software Engineering Agents and Verifiers with. 2025 , note =
2025
-
[14]
and Tang, Xiangru and Zhuge, Mingchen and Pan, Jiayi and Song, Yueqi and Li, Bowen and Singh, Jaskirat and Tran, Hoang H
Wang, Xingyao and Li, Boxuan and Song, Yufan and Xu, Frank F. and Tang, Xiangru and Zhuge, Mingchen and Pan, Jiayi and Song, Yueqi and Li, Bowen and Singh, Jaskirat and Tran, Hoang H. and Li, Fuqiang and Ma, Ren and Zheng, Mingzhang and Qian, Bill and Shao, Yanjun and Muennighoff, Niklas and Zhang, Yizhe and Hui, Binyuan and Lin, Junyang and Brennan, Robe...
2025
-
[15]
2025 , url =
Miserendino, Samuel and Wang, Michele and Patwardhan, Tejal and Heidecke, Johannes , journal =. 2025 , url =
2025
-
[16]
and Barnes, Elizabeth and Chan, Lawrence , booktitle =
Kwa, Thomas and West, Ben and Becker, Joel and Deng, Amy and Garcia, Katharyn and Hasin, Max and Jawhar, Sami and Kinniment, Megan and Rush, Nate and Von Arx, Sydney and Bloom, Ryan and Broadley, Thomas and Du, Haoxing and Goodrich, Brian and Jurkovic, Nikola and Miles, Luke Harold and Nix, Seraphina and Lin, Tao and Painter, Chris and Parikh, Neev and Re...
2025
-
[17]
and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =
Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =. 2024 , note =
2024
-
[18]
and Wettig, Alexander and Lieret, Kilian and Yao, Shunyu and Narasimhan, Karthik and Press, Ofir , booktitle =
Yang, John and Jimenez, Carlos E. and Wettig, Alexander and Lieret, Kilian and Yao, Shunyu and Narasimhan, Karthik and Press, Ofir , booktitle =. 2024 , note =
2024
-
[19]
Agentless: Demystifying
Xia, Chunqiu Steven and Deng, Yinlin and Dunn, Soren and Zhang, Lingming , journal =. Agentless: Demystifying. 2024 , url =
2024
-
[20]
2024 , note =
Zhang, Yuntong and Ruan, Haifeng and Fan, Zhiyu and Roychoudhury, Abhik , booktitle =. 2024 , note =
2024
-
[21]
2024 , url =
Arora, Daman and Sonwane, Atharv and Wadhwa, Nalin and Mehrotra, Abhav and Utpala, Saiteja and Bairi, Ramakrishna and Kanade, Aditya and Natarajan, Nagarajan , journal =. 2024 , url =
2024
-
[22]
and Zhu, Hao and Zhou, Xuhui and Lo, Robert and Sridhar, Abishek and Cheng, Xianyi and Ou, Tianyue and Bisk, Yonatan and Fried, Daniel and Alon, Uri and Neubig, Graham , journal =
Zhou, Shuyan and Xu, Frank F. and Zhu, Hao and Zhou, Xuhui and Lo, Robert and Sridhar, Abishek and Cheng, Xianyi and Ou, Tianyue and Bisk, Yonatan and Fried, Daniel and Alon, Uri and Neubig, Graham , journal =. 2023 , url =
2023
-
[23]
2024 , url =
Xie, Tianbao and Zhang, Danyang and Chen, Jixuan and Li, Xiaochuan and Zhao, Siheng and Cao, Ruisheng and Hua, Toh Jing and Cheng, Zhoujun and Shin, Dongchan and Lei, Fangyu and Liu, Yitao and Xu, Yiheng and Zhou, Shuyan and Savarese, Silvio and Xiong, Caiming and Zhong, Victor and Yu, Tao , journal =. 2024 , url =
2024
-
[24]
2024 , note =
Trivedi, Harsh and Khot, Tushar and Hartmann, Mareike and Manku, Ruskin and Dong, Vinty and Li, Edward and Gupta, Shashank and Sabharwal, Ashish and Balasubramanian, Niranjan , booktitle =. 2024 , note =
2024
-
[25]
2024 , url =
Wijk, Hjalmar and Lin, Tao and Becker, Joel and Jawhar, Sami and Parikh, Neev and Broadley, Thomas and Chan, Lawrence and Chen, Michael and Clymer, Josh and Dhyani, Jai and Ericheva, Elena and Garcia, Katharyn and Goodrich, Brian and Jurkovic, Nikola and Karnofsky, Holden and Kinniment, Megan and Lajko, Aron and Nix, Seraphina and Sato, Lucas and Saunders...
2024
-
[26]
Transactions of the Association for Computational Linguistics , volume =
Lost in the Middle: How Language Models Use Long Contexts , author =. Transactions of the Association for Computational Linguistics , volume =
-
[27]
Context Rot: How Increasing Input Tokens Impacts
Hong, Kelly and Troynikov, Anton and Huber, Jeff , institution =. Context Rot: How Increasing Input Tokens Impacts. 2025 , url =
2025
-
[28]
The Memory Curse: How Expanded Recall Erodes Cooperative Intent in
Liu, Jiayuan and Li, Tianqin and Du, Shiyi and Luo, Xin and Zeng, Haoxuan and Tewolde, Emanuel and Lee, Tai Sing and Wang, Tonghan and Kingsford, Carl and Conitzer, Vincent , journal =. The Memory Curse: How Expanded Recall Erodes Cooperative Intent in. 2026 , url =
2026
-
[29]
2026 , url =
Huang, Wenqi and Lee, Charley and Tng, Leonard and Ge, Serena , journal =. 2026 , url =
2026
-
[30]
2026 , url =
Desai, Rishi and Hu, Jesse and Cabezas, Joan and Harsola, Neel and Shukla, Pratyush and others , journal =. 2026 , url =
2026
-
[31]
International Conference on Learning Representations , year =
Let's Verify Step by Step , author =. International Conference on Learning Representations , year =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.