{"id":"a699e8e4-39cd-40c0-8950-ce320e8a0df7","arxiv_id":"2608.06663","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A survey of 1,547 papers defines the 'horizon gap' and documents that long-horizon agent research is converging on trajectory-level process signals instead of outcome-only scores.","lead":"A survey of 1,547 arXiv papers argues that LLM agents fail on multi-hour tasks because single outcome scores stop being informative as task length grows, and the field is responding by building denser step-level signals. It separates long-horizon, long-context, and long-term memory, maps six research categories, and names unsolved measurement problems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Synthesis rests on single-annotator corpus labels; systematic misclassification could manufacture the very pattern the paper reports.","rationale":"The reader's weakest_assumption is the corpus being representative and labels reliable, and I agree that this is the single most load-bearing concern. The paper explicitly discloses its single-annotator, rule-based classification and the small validation sample, so the concern is not that the paper hides a flaw but that its central synthesis leans on unverified empirical base. Everything else — the six-section framing, the harness-versus-model problem, the correlated-measurement-bias argument — is argued carefully and hedged appropriately, and the paper's own language calibration and limitation statements are unusually thorough. I cannot find a more pressing internal inconsistency. The taxonomy's self-organization from the corpus it then describes is a real circularity risk, but the paper partly mitigates it by releasing rules as scripts and inviting recomputation; the decisive missing piece is a second annotator or an independent label audit. Given that the empirical base is not yet independently checkable, the appropriate verdict remains CONDITIONAL rather than ACCEPT. If the release and an inter-annotator study confirm the labels, the verdict could move to ACCEPT without further substantive changes.","tokens_in":33221,"tokens_out":2406,"duration_ms":24573,"concrete_test":"Conduct a blind second-annotator study: give a second annotator the published category definitions and a stratified sample of at least 200 papers (oversampling memory, execution, and foundations), without revealing the author's labels. Measure Cohen's kappa and per-category agreement. Separately, recompute Table 1 counts and Figure 6 using a held-out relabeling of the full corpus by scripts alone, with the override list applied automatically and documented; if the external-to-context ratio (294:103) or the execution-to-recovery ratio (338:245) moves by more than a few percentage points, the synthesis-level ratios are not robust.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is an empirical generalization: as horizon grows, outcome-only signals become uninformative and the field responds by manufacturing denser step-level signal (§1, §9). This claim is substantiated largely through corpus statistics rather than through any single experiment: execution is the largest category (n=584), memory's external subcategory dwarfs context (294 vs 103), and Figure 6's growth timeline shows execution leading early and memory surging in 2026. All of these numbers trace back to one annotator's pipeline: harvest-time topic filters, TF-IDF + SVD + k-means cluster defaults (k=18), keyword rules, an LLM/agent-signal gate with cluster exemptions, and an explicit override list (§2.2). The validation sample is 85 papers (30 kept, 55 excluded); the estimated residual misclassification rate 'on the order of one in twenty' is a single unquantified point estimate with no confidence interval and no per-category breakdown. If the errors are systematic — e.g., if the cluster scoping or the override list embodies the author's prior that execution/memory dominate — the taxonomy's shape, and hence the synthesized narrative, could be an artifact of the classification rather than a discovery about the literature. The paper is transparent about many limitations, but transparency does not by itself establish that the labels are reliable enough to carry the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey defines 'the horizon gap' as the distance between single-step model capability and reliable completion of tasks spanning many steps, and maps the 2024-2026 arXiv literature responding to it. The corpus (1,547 papers) is built via an eight-thread seed harvest with a disclosed two-stage bleed filter (26.8% excluded), supplemented by 128 targeted papers for under-covered foundations/safety work. The paper disambiguates long-horizon (task property), long-context (model property), and long-term memory (system property); organizes the corpus into six lifecycle categories (planning, memory, execution, training, evaluation, foundations) crossed with a horizon-locus axis (within-context, within-task-beyond-context, cross-task-persistent); and advances a cross-cutting thesis that as horizon grows, outcome-only signals become uninformative and the field responds by manufacturing denser step-level signals, visible in process reward models and credit assignment (training) and in trajectory-level diagnostics and benchmark-critique work (evaluation). It closes by naming three open measurement problems: model-versus-harness attribution, correlated measurement bias between training and evaluation signals, and whether long-horizon reliability admits a general predictive theory.","tokens_in":33402,"tokens_out":16815,"duration_ms":160748,"significance":"Should the synthesis hold, the central contribution is a reframing: six sub-literatures of long-horizon agent research read as one coordinated response to signal sparsity, together with a clean tripartite disambiguation (long-horizon vs long-context vs long-term memory) that is itself genuinely useful, and two open measurement problems (model-versus-harness attribution; correlated measurement bias in process-level signals) concrete enough to guide future experiments. The paper is internally consistent: I verified the corpus arithmetic (1,939 unique raw hits; 520 dropped over two stages = 26.8%; 1,419 seed + 128 supplement = 1,547; Table 1's categories and subcategories sum exactly to the corpus total, and the supplement distribution across categories sums to 128). The methodology is unusually transparent for the genre — disclosed bleed rates with quantified per-stage drops, an explicit rule-plus-override classification procedure, hedged 'hypothesis, not finding' language for corpus-shape observations, and one genuinely falsifiable prediction (the verifiability-gradient delay).","major_comments":[{"comment":"The quantitative texture of the survey — per-category counts, the external:context (294:103), orchestration:recovery (338:245), and rl:supervision (130:37) ratios used in the section Assessments, and the Figure 6 growth timeline — rests on a single annotator's pipeline validated by a manual re-check of only 85 papers (30 kept, 55 excluded). The reported residual error ('on the order of one in twenty') is a point estimate with no confidence interval, no per-category breakdown, and no second annotator; because the re-check was performed by the same annotator against their own criterion, it establishes rule-consistency rather than construct validity against an independent expert. The headline ratios are robust to uniform ±5% noise, but systematic directional bias (e.g., a broadened 'external' tag absorbing in-context work, or an over-broad 'orchestration' tag) is exactly the failure mode the validation cannot detect at this scale, and it would change the conclusions the Assessments draw. The skeptic's worry that misclassification could manufacture the reported pattern is, on reading the paper, only partially warranted: the paper's own hedging and the independently cited literature (the PRM800K result, the SWE-bench+ critique, the trajectory-level diagnostics of §7) would survive substantial label noise. But the per-category ratios and the §9 corpus-shape hypotheses are as fragile as the skeptic says. Because the corpus is the paper's primary empirical artifact, I ask that the corpus and scripts be released at submission rather than at camera-ready, and that the validation be extended to per-category residual-error estimates with intervals and, ideally, an independent second-annotator pass on a stratified sample.","section":"§2.2, Table 1, Figure 6"},{"comment":"The abstract's central claim — 'Across all categories, we find the same pattern: outcome-only signals grow uninformative as horizons lengthen' — is stated as a universal, and §9 repeats the universal form ('Across every category, the object being stored, scored, attributed over, audited, and debugged is shifting from the outcome to the execution trace'). The body supports the signal-density version of the claim strongly only for training (§6) and evaluation (§7). The §3 assessment describes a commitment-versus-robustness trade-off, the §4 assessment a persistence-versus-fidelity trade-off, and the §5 assessment a harness-versus-model attribution pattern; these are structural responses to horizon growth, but they are not the same claim about outcome-only signals growing uninformative, and §10 itself narrows the outcome-to-process shift to training and evaluation. To make the headline claim match the evidence, either restrict the abstract/conclusion formulation to the two well-supported instantiations plus a clearly-labeled broader 'densification of structure' reading, or supply per-category evidence that signal density specifically is the cross-cutting response.","section":"Abstract, §9, §10"},{"comment":"§9's second interpretive hypothesis (the 'verifiability-gradient' account) argues from Table 1's 2026-share column — foundations at 77%, memory at 72% — but that column includes the 128 supplement papers, which §2.2 declares 'recency-biased by construction' and which are excluded from Figure 6 for precisely that reason. For foundations, 93 of 103 papers are supplement papers harvested with dedicated 2024+ queries, so the 77% figure is substantially an artifact of the supplement's construction; using it as evidence for a substantive claim about the field's recent priorities is inconsistent with the paper's own exclusion of the supplement from the growth analysis. The 2026-share column should be recomputed for seed-only papers, or the hypothesis should be argued from seed-only shares.","section":"Table 1, §9"},{"comment":"The taxonomy and the counts are entangled: the exploratory k-means clustering over the seed pool supplied the 'cluster-based category defaults' that the classification then used, so the category boundaries — and hence Table 1's counts and Figures 2 and 6 — were fit to the same data they are then used to summarize. The paper discloses this in §2.2, which is to its credit, but the section Assessments nevertheless treat several counts as findings (e.g., §4: 'the external-to-context size ratio... is itself informative'). Because the clustering defaults, keyword rules, and override list jointly determine the counts, the paper should (i) state at each point where a count is used as evidence that the count inherits the taxonomy's construction choices, and (ii) release the cluster assignments alongside the labels so the dependence can be audited.","section":"§2.2, §2.3"}],"minor_comments":[{"comment":"The text contains ligature artifacts ('eﬀicient', 'eﬀiciency', 'diﬀiculty', 'suﬀicient' in §§2-8) and spacing artifacts ('Data A vailability', 'W ALL-E') that should be cleaned in the published version; they appear to be rendering artifacts rather than intentional formatting.","section":"Throughout"},{"comment":"The phrase 'rather than force the corpus down to a pre-registered size' is ambiguous: the earlier mention of 'our initial ~1,100–1,400 planning estimate' suggests an internal planning estimate, and readers should not infer a formal preregistration; please clarify the status of the estimate.","section":"§2.2"},{"comment":"The caption states that the figure is 'Sourced entirely from primary papers' while the plotted bars mix 'reported or illustrative ranges'; the two classes of ranges should be distinguished visually and in the caption, and the cited artifact file (artifacts/benchmark_durations.csv) should be released with the corpus so the point estimates can be traced to their sources.","section":"Figure 5"},{"comment":"The prose synthesizes roughly 130 exemplar papers out of a 1,547-paper corpus, but the criterion for which papers become prose exemplars rather than corpus rows is not stated; one sentence describing the exemplar-selection principle (representativeness, influence, topical span, or recency) would help readers judge whether the narrative is anchored to a purposive sample.","section":"§§3-9"},{"comment":"Several load-bearing diagnostic findings cited in support of the synthesis are 2026 arXiv preprints without peer-review status (e.g., [76] on the accuracy-correction paradox, [92] on GRPO as an implicit process reward model, [108] on coherence collapse); the paper's language calibration is exemplary for its own claims, and the same discipline should be extended to first-mention status flags for cited preprints whose findings have not yet been independently replicated.","section":"§§6-8"}],"recommendation":"major_revision","confidential_remarks":"For the editor: this is a large, well-crafted survey whose main empirical artifact is the corpus; I would want to see the corpus and scripts actually released before publication, since several quantitative claims (Figure 6, per-category ratios) are not otherwise checkable. The survey is closely tied to the authors' companion survey [15], sharing harvest infrastructure and taxonomy methodology; the manuscript discloses this clearly, but the editor may wish to confirm that the two papers' division of labor is communicated to readers and that the shared methodology is not double-counted as two independent empirical contributions. A number of load-bearing citations are 2026 arXiv preprints; I do not read this as a defect given the field's pace, but it is a stability risk worth flagging. Finally, the paper is a pure survey with no new experiments; its suitability depends on the journal's policy on surveys, since the contribution is synthesis and taxonomy — substantial and useful, but descriptive rather than experimentally falsifiable, with the exception of the paper's own explicitly falsifiable predictions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you work on LLM agents and want a map of where the long-horizon subfield actually is. The paper's main contribution is conceptual: separating long-horizon (task property), long-context (model property), and long-term memory (system property), then organizing 1,547 papers along a horizon-locus axis. That framing is genuinely useful and not something the planning, memory, GUI-agent, or agentic-RL surveys I know provide. The synthesis—that outcome-only signals get uninformative as horizons grow and the field responds by manufacturing denser step-level signal—is a defensible reading of the literature, and the paper is careful to label corpus-shape observations as hypotheses rather than findings. The internal arithmetic checks out, the bleed-filter numbers are disclosed, and the paper is candid about its own limits, including the missing theory in §8 and the recency-biased supplement. Credit where due: this is an honest, well-scoped survey, not a dressed-up blog post.\n\nThe soft spot is exactly where the reader put it. The corpus labels come from a single annotator, the validation sample is 85 papers, the residual misclassification rate is an unquantified point estimate, and the taxonomy was built from the same corpus it then organizes. If the errors are systematic—say, the cluster gates or override list encode a prior that execution and memory dominate—then ratios like external-to-context memory 294:103 and the Figure 6 growth timeline could partly reflect the classification pipeline rather than the literature. That concern is real but not fatal, because the paper's central thesis does not stand or fall on any single count, and because the authors explicitly flag the limits. What would move it from conditional to accept is straightforward: release the corpus and scripts with a commit hash, add a second annotator or a formal reliability measure, and publish the exact query strings and override rules. The stress-test worry is plausible and worth taking seriously, but it is an addressable methodological gap, not evidence of fabrication.\n\nFor whom: grad students and researchers entering long-horizon agent work will get the most value; practitioners may skim §§3–7 for the taxonomy and the measurement-problem discussion. I'd bring it to a reading group and I'd cite it if I wrote in this area. Recommendation: send it to serious peer review. The referee should push on data availability and label reliability before acceptance, but the paper deserves the engagement.","headline":"A serious, unusually transparent survey that names a real gap and builds a useful taxonomy; the corpus behind it needs to be public and independently labeled before the empirical claims carry full weight.","tokens_in":758,"tokens_out":801,"would_cite":true,"duration_ms":25789,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims the field's scattered long-horizon agent research is one coordinated response to signal sparsity: as horizons grow, outcome-only signals become uninformative, and training and evaluation both manufacture denser step-level…","keywords":["horizon gap","long-horizon agents","long-context","long-term memory","process reward models","trajectory-level evaluation","harness vs model","correlated measurement bias"],"falsifier":"Re-run the full 1,547-paper classification with a second independent annotator using the released scripts and compare per-category ratios; if, say, the external-to-context memory ratio of 294:103 or the Figure 6 growth timeline changes materially, the survey's structural pattern is an artifact of labeling. Alternatively, the harness-versus-model claim would be settled by a controlled experiment holding the model fixed across harnesses and holding the harness fixed across models and measuring achievable task horizon at a fixed reliability: if harness changes move the horizon little, the paper's central attribution problem dissolves.","tokens_in":32919,"feed_emoji":"🤖","tokens_out":4753,"duration_ms":60922,"temperature":0.7,"pith_summary":"The paper identifies a \"horizon gap\": large language models solve single-step problems well but fail at tasks spanning hours, losing track of earlier decisions, declaring unfinished work done, and drifting from goals. It surveys 1,547 papers from 2024 to 2026 and argues that the field's many strands are one coordinated response to a single pressure: as horizons grow, outcome-only signals (one final reward, one pass/fail check) become uninformative, so researchers manufacture denser step-level signals for training and evaluation alike. The paper disambiguates long-horizon (a task property), long-context (a model property), and long-term memory (a system property), then organizes the corpus into six lifecycle categories crossed with where the horizon is carried. A sympathetic reader would care because the survey turns scattered agent research into a testable structural claim and names two concrete measurement problems that block progress: separating model capability from harness capability, and avoiding correlated bias between training and evaluation signals.","feed_headline":"1,547 papers show long-horizon AI fails as signals thin out","feed_subtitle":"Outcome-only rewards go uninformative as tasks lengthen; the field responds by building denser step-level signals.","key_machinery":"The organizing device is a two-axis taxonomy: six lifecycle categories (planning, memory, execution, training, evaluation, foundations) crossed with where the horizon is carried (within-context, within-task-beyond-context, cross-task-persistent). The load-bearing analytical identity is the structural analogy between training and evaluation: both need a reliable signal of partial-trajectory progress, and both manufacture it at the step level, which creates the risk of correlated measurement bias. The paper's evidence base is a systematically harvested corpus with a disclosed bleed filter; the ratio external:context memory (294:103) and the category growth timeline are the concrete observations carrying the synthesis.","core_discovery":"The central claim is that outcome-only supervision and outcome-only evaluation stop working as task horizon grows, and the field's response is the same everywhere: replace the single terminal signal with denser, step-level signal. In training this appears as process reward models and credit assignment; in evaluation as trajectory-level diagnostics and benchmark-purification work; in memory and execution as traceable, recoverable trajectories. The paper further claims the execution trajectory is becoming the shared unit of analysis across all six categories, and that this convergence makes trajectory logging a precondition for answering the field's two open measurement problems.","pith_inferences":["If correlated measurement bias is real, published long-horizon progress may be systematically overstated, and a decisive test would compare process signals built from disjoint assumptions.","The taxonomy predicts research attention follows visible failure; a testable extension is whether foundations-paper counts lag capability demonstrations by a roughly constant delay.","Because the corpus was labeled by a single annotator with rule-based defaults, all category ratios should be treated as hypotheses; an independent re-labeling of the same corpus would be a cheap check.","If trajectory retention becomes the norm, privacy and auditability of stored traces become load-bearing, connecting the memory-security thread to the oversight thread."],"forward_implications":["If the thesis is right, improving long-horizon reliability means investing in step-level signal design for both training and evaluation, not just scaling models or contexts.","The harness-versus-model question must be answered before long-horizon capability can be rationally pursued; a controlled same-model/different-harness comparison is the implied decisive experiment.","The convergence on trajectory as the unit of analysis implies trajectory logging becomes infrastructure, not an implementation detail.","Benchmark critiques (leakage, weak tests, coherence collapse) are not side literature; they are first-class evidence that outcome scores overstate capability.","Training algorithms described as outcome-based, such as GRPO, may already perform implicit process-level credit assignment, so the outcome/process distinction is not architectural."],"supporting_citations":[{"why":"Supplies the measured horizon-growth trend that the survey treats as measured-and-reported rather than as a settled law.","marker":"[1]"},{"why":"Anchor finding that intrinsic self-correction does not reliably improve reasoning, carried through the recovery literature.","marker":"[2]"},{"why":"Anchor for benchmark-construction critiques showing agentic benchmarks can over- or under-estimate performance by large margins.","marker":"[5]"},{"why":"Establishes the operating-system metaphor for external memory that organizes the largest memory subcategory.","marker":"[45]"},{"why":"Reference result that process supervision beats outcome supervision, at a step-level annotation cost no agentic setting can pay per task.","marker":"[89]"},{"why":"Supports the claim that outcome-based GRPO implicitly performs process-reward-like credit attribution.","marker":"[92]"},{"why":"Supplies SWE-bench as the common-currency benchmark whose evolution and critique anchor the evaluation section.","marker":"[98]"},{"why":"Documents solution leakage and weak-test inflation in SWE-bench, dropping a top system's resolution rate from 12.47% to 3.97%.","marker":"[103]"},{"why":"Trajectory-level diagnostic behind the coherence-collapse finding and the claim that Pass@1 misdiagnoses failure.","marker":"[108]"}],"fun_headline_variants":["Outcome-only rewards fail as tasks lengthen: the horizon gap","Long-horizon LLMs need denser signals, 1,547 papers show","Horizon gap: why long tasks break LLM agents' tracking","Step-level signals are the answer to long-horizon AI failures","Survey of 1,547 papers reveals the horizon gap in LLM agents"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The corpus is representative and its labels are trustworthy; classification rested on a single annotator's rules, with residual misclassification estimated from only 85 manually re-checked papers, so systematic labeling errors would invalidate the category ratios and growth timeline on which the synthesis rests.","fun_headline_variants_meta":{"raw":{"variants":["Outcome-only rewards fail as tasks lengthen: the horizon gap","Long-horizon LLMs need denser signals, 1,547 papers show","Horizon gap: why long tasks break LLM agents' tracking","Step-level signals are the answer to long-horizon AI failures","Survey of 1,547 papers reveals the horizon gap in LLM agents"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000353,"raw_usage":{"total_tokens":1923,"prompt_tokens":950,"completion_tokens":973,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":890}},"tokens_in":566,"tokens_out":973,"duration_ms":7723,"temperature":1.0,"reasoning_tokens":890,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:56:23.263099+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the full 1,547-paper classification with a second independent annotator using the released scripts and compare per-category ratios; if, say, the external-to-context memory ratio of 294:103 or the Figure 6 growth timeline changes materially, the survey's structural pattern is an artifact of labeling. Alternatively, the harness-versus-model claim would be settled by a controlled experiment holding the model fixed across harnesses and holding the harness fixed across models and measuring achievable task horizon at a fixed reliability: if harness changes move the horizon little, the paper's central attribution problem dissolves.","supporting_citations":[],"review_version":1}