{"id":"8b58fde7-705c-473c-8c21-cba8f87e8af5","arxiv_id":"2608.09799","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A coding agent's success on a consolidated specification does not guarantee the same tested behavior when the identical contract is reached through a different revision history.","lead":"SpecPath is a new evaluation that tests whether AI coding agents that succeed on a direct task description still succeed when the same final requirements are described through different revision histories. Across five tasks and fourteen agent configurations, aggregate accuracy stays flat while 35 of 100 directly successful runs fail on at least one equivalent history.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing direct-direct rerun control leaves the 36.4% any-CPV estimate unseparated from run-to-run stochasticity; a four-run direct-direct null baseline is required before the path-sensitivity claim can be accepted.","rationale":"The reader's weakest-assumption analysis and my independent reading converge on the same inferential gap: the paper conditions on direct success and interprets subsequent alternative-history failures as specification-path effects, but it never measures how often a direct run fails after another direct run succeeds. The duplicate condition is not an adequate control because it is itself one of the path conditions and already shows substantial flip risk. The paper is otherwise strong: verifier calibration, visible-history review, transparent attrition accounting, and independent audit are real supports, and the Discussion already narrows the claim to \"can change which directly competent blocks remain correct.\" The concern does not invalidate the diagnostic method; it changes what the headline number can assert. The recommended fix is a direct-direct null baseline, after which the claim boundary can be stated cleanly. Verdict remains CONDITIONAL, matching the reader's assessment.","tokens_in":14057,"tokens_out":6982,"duration_ms":69919,"concrete_test":"Add a direct-direct control: for each of the 127 complete blocks (or at least the 100 direct-success blocks), schedule four additional direct-condition runs under the same task, model deployment, scaffold, and budget, with fresh seeds. Compute a direct-direct any-CPV analog: original direct succeeds and at least one of the four extra direct runs fails. Aggregate by task-macro and bootstrap 95% intervals as in Equation (8). If this null any-CPV interval contains 36.4%, or its lower bound falls below 25.6%, the observed effect is consistent with stochastic flip; if the null point estimate is clearly below the primary interval, the path-sensitivity claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on a conditional comparison: among blocks where direct succeeds (D=1), failure under at least one of four alternative histories (R=1) is counted as specification-path sensitivity. But every run is sampled independently, and a second direct run can fail for stochastic reasons. The duplicate condition, which restates the same final contract, already produces an 18.3% task-macro CPV, and the four variant-specific estimates are 18.3%, 10.8%, 14.1%, and 12.2%. Under a simple independence model, these four per-history flip rates imply an expected any-failure rate of roughly 1 - (0.817 * 0.892 * 0.859 * 0.878) = 45%, comparable to the observed 36.4%. No direct-direct rerun control exists, and the paper explicitly concedes that \"separately sampled agent runs remain stochastic\" (Discussion) and that CPV \"is not a common-randomness estimate\" (Threats to Validity). Without a direct-direct null baseline, the headline 35/100 result does not separate path sensitivity from run-to-run noise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"SpecPath is a controlled evaluation that varies only the revision history leading to a fixed final software contract, holding the repository, verifier, agent system, and execution budget constant. Across five task families and fourteen coding-agent configurations, the paper reports that aggregate final-contract realization is nearly unchanged (78.8% direct vs. 78.7% average over alternative histories), yet 35 of 100 complete blocks that succeed on the direct specification fail on at least one contract-equivalent history. The paper interprets this as specification-path sensitivity and proposes this as a distinct evaluation dimension for coding agents. It includes a detailed construction pipeline, verifier calibration against gold and mutant patches, independent human audits of contract equivalence, explicit attrition accounting, and cluster-bootstrap inference.","tokens_in":14293,"tokens_out":6906,"duration_ms":59421,"significance":"If the central claim holds, the paper makes a valuable methodological contribution: it demonstrates that canonical success on a consolidated requirement does not imply robustness to the history by which the requirement became final, and that aggregate accuracy can hide which blocks are actually succeeding. The construction pipeline is careful and unusually thorough for this area: the verifier is validated against base, gold, and mutant patches before agent outcomes are used; contract equivalence is checked by hidden-trace replay and independent review; and the inference respects the small number of task families via cluster bootstrap. The main risk is that the headline 35/100 result is not separated from run-to-run stochasticity, because every execution is an independent sample and no direct-direct rerun control is reported.","major_comments":[{"comment":"The any-CPV estimate of 36.4% (95% CI 25.6–45.1%) is computed on blocks where direct succeeds (D=1) and at least one of duplicate, override, cancellation, or split fails (R=1). Because every execution is an independent sample, a second direct run can fail for stochastic reasons, and the manuscript reports no direct-direct rerun control. The duplicate condition, which restates the same final contract in a near-identical way, already produces an 18.3% task-macro variant-specific CPV; under a simple independence model using the four variant-specific rates (18.3%, 10.8%, 14.1%, 12.2%), the expected any-failure rate is 1 − (0.817 × 0.892 × 0.859 × 0.878) ≈ 45%, close to the observed 36.4%. The Discussion states that \"separately sampled agent runs remain stochastic\" and Threats to Validity states that CPV \"is not a common-randomness estimate,\" but no null baseline is provided. Without a direct-direct CPV estimate (e.g., from the existing three repeats of the direct condition), the headline \"35 of 100\" cannot be separated from run-to-run noise and the central claim of specification-path sensitivity is not established. This is load-bearing for RQ2 and the abstract.","section":""},{"comment":"The abstract claims that \"requirement histories that are equivalent in their final meaning lead the same agent system to produce behaviorally different programs.\" This is a causal attribution, but the design is a single paired sample per block with no common-randomness control; the paper itself says in Threats to Validity that CPV \"is not a common-randomness estimate of a deterministic prompt treatment.\" At most, the data show that a block that succeeds on direct can fail on an alternative history in one sampled run. The causal wording should be tempered to \"can coincide with\" or \"is associated with\" unless the direct-direct baseline isolates the path effect. This wording affects the central claim and should be corrected in revision.","section":""},{"comment":"Duplicate is simultaneously a core history in the any-CPV definition and a control for \"inert repetition.\" The 18.3% duplicate-specific CPV, being the largest of the four variant-specific estimates, indicates that even a history with no change in active obligations produces frequent failures. This weakens the interpretation that the observed sensitivity is specific to revision structure (override/cancellation/split); the result may reflect sensitivity to any textual variation or sheer run-to-run noise. The paper acknowledges that the ordering is exploratory and that the controls do not identify a unique mechanism, but the introduction's framing of \"specification-path sensitivity\" as a distinct failure mode of active-contract resolution needs to be reconciled with the size of the duplicate effect. A direct-direct baseline would clarify whether duplicate's effect exceeds the noise floor.","section":""}],"minor_comments":[{"comment":"The phrase \"Grok-4.20-Nonreasoningdeployments\" appears to be a typo; if it refers to non-reasoning deployments, please insert a space or use parentheses.","section":""},{"comment":"The dashed line marking direct FCR is not labeled in the legend; add an annotation to make the reference visible.","section":""},{"comment":"The raster uses a green/red pass/fail palette; for colorblind accessibility, consider a diverging palette or additional cell patterns.","section":""},{"comment":"The phrase \"RepoBenchinstead\" should read \"RepoBench instead\".","section":""},{"comment":"The paper cites both \"35 of 100\" and the 36.4% task-macro any-CPV. Please clarify explicitly that the former is the unweighted block count and the latter is the task-macro estimate, so readers do not conflate the two.","section":""},{"comment":"The notation bPh in Eq. (7) is defined for a history h, while Eq. (8) reuses the macro-averaging notation for CPV; consider adding a sentence distinguishing the two macro definitions to prevent confusion.","section":""}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well-executed methodologically: the verifier calibration, independent audits, attrition accounting, and cluster-bootstrap inference are strengths. The key gap is the missing direct-direct rerun baseline, which is required to separate the headline 35/100 observation from run-to-run stochasticity. Because the data already contain three repeats of the direct condition per block, the authors should be able to compute such a baseline without new experiments. If the direct-direct flip rate is substantially below the observed any-CPV, the paper would be a strong contribution. I did not see a circularity problem: the evaluation pipeline is external and the verifier is validated before agent outcomes are used. The five-family scale is small, but the paper's honest treatment of that limitation is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, SpecPath is the first evaluation I've seen that holds repository, final contract, verifier, and budget fixed while varying only the history that leads to the contract. That's a genuinely useful counterfactual, and the construction pipeline—verifier calibration against gold and mutant patches, hidden-trace replay for contract equivalence, independent audits, attrition accounting—is careful and transparent. The distinction between average FCR and conditional path violation is a good idea, and the raster of executable signatures is a nice way to show what changes. Credit where due: this is a serious attempt to measure something benchmarks don't.\n\nSecond, the central empirical claim—35 of 100 direct-success blocks fail on at least one equivalent history—does not yet separate specification-path sensitivity from run-to-run stochasticity. The paper's own duplicate condition, which is the same final contract stated twice, already produces an 18.3% task-macro CPV. That's your noise floor. If you assume the four history conditions fail independently at the per-history rates reported (18.3%, 10.8%, 14.1%, 12.2%), the expected any-CPV is about 45%, which is right on top of the observed 36.4% (95% CI 25.6–45.1%). So the headline number is statistically indistinguishable from what you'd get from pure sampling variation. The paper acknowledges this in the Discussion—'separately sampled agent runs remain stochastic'—and in Threats to Validity, but the abstract and conclusion lean on the 35/100 as evidence of path sensitivity. That's a mismatch between the strength of the claim and the design.\n\nThe fix is straightforward: add a direct-direct rerun condition, where the same consolidated history is run twice in the same block. That gives you a null distribution for CPV under identical specification. If the any-CPV from alternative histories is substantially higher than the direct-direct CPV, the path effect is real. Without it, the result is compatible with noise.\n\nThe paper is still worth reading. The controlled counterfactual, the history operators, and the measurement framework are solid, and the negative-control logic (paraphrase, length-matched) shows the authors know what competing explanations look like. But the headline result is undersupported. With a rerun control, this becomes a strong contribution.\n\nRecommendation: send to peer review, but make the rerun baseline a requirement. The method deserves referee time; the current inference does not.","headline":"A well-built diagnostic for a real problem, but the headline 35/100 needs a rerun control before you can call it path sensitivity.","tokens_in":14764,"tokens_out":2594,"would_cite":true,"duration_ms":23034,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Coding agents that pass a direct specification often fail on an equivalent history that reaches the same final contract, despite stable aggregate accuracy.","keywords":["specification-path sensitivity","coding agents","contract-equivalent histories","conditional path violation","active-contract resolution","benchmark methodology","requirements evolution","metamorphic testing"],"falsifier":"Re-run each direct-condition block multiple times (say ten times) under identical configuration and count how often a direct success flips to failure between identical executions; if that direct-versus-direct flip rate matches or exceeds the observed 36.4% conditional path violation rate, the path-sensitivity conclusion would be an artifact of run-to-run noise rather than history structure.","tokens_in":13773,"feed_emoji":"🔀","tokens_out":8108,"duration_ms":61021,"temperature":0.7,"pith_summary":"This paper argues that coding agents can be sensitive to the path by which a final specification is reached, even when every path leaves the same contract in force. Using SpecPath, a controlled evaluation that fixes the repository, final contract, verifier, agent system, and execution budget while varying only the revision history, the authors show that aggregate accuracy stays nearly constant (direct 78.8% versus an average of 78.7% across equivalent histories) yet 35 of 100 blocks that succeed directly fail on at least one equivalent history. The paper concludes that implementation success on a consolidated request does not guarantee specification-path invariance, and that evaluation of evolving requirements should measure such invariance separately. The finding matters because real requirements evolve through edits, overrides, and cancellations, so a coding agent must determine which requirements still count before writing code.","feed_headline":"35 of 100 direct successes fail when spec history changes","feed_subtitle":"Average accuracy barely moves, yet which blocks succeed flips across contract-equivalent histories.","key_machinery":"The load-bearing machinery is the contract-equivalent history family built from a hidden trace of requirement atoms. Each atom is the smallest testable behavior with an explicit scope, polarity, and observation; replay applies deactivations and activations, and normalization removes presentation order to yield one final active contract. SpecPath then forms blocks that fix task, model deployment, scaffold, and repeat, and defines the conditional path violation: a block where direct execution succeeds but at least one equivalent history fails. Control histories (duplicate, split, override, cancellation, paraphrase-direct, and length-matched) separate revision structure from wording, repetition, and added context, and the executable verifier, not patch identity, is the oracle.","core_discovery":"The central discovery is that stable aggregate performance can coexist with systematic path-conditioned failures. SpecPath defines contract-equivalent histories as histories whose hidden replay traces normalize to the same final contract, then executes each agent configuration on fresh copies of the same repository under each history. Across five task families and fourteen agent configurations, direct-condition final-contract realization is 78.8%, while the average over the four equivalent histories is 78.7%; however, among 100 complete blocks with direct success, 35 fail under at least one alternative history, giving a task-macro conditional-path-violation estimate of 36.4%. The authors take this as evidence that the path to a specification, not only its final text, changes which programs an agent produces.","pith_inferences":["A natural extension is to adapt the paired-history design to documentation updates, API migrations, or translated requirements, treating specification-path sensitivity as a general axis of instruction-following evaluation.","The 35-of-100 figure is a lower bound: only 127 of 210 possible complete blocks were scored, so missing executions could shift the violation rate; a replication with denser scoring would settle how much attrition matters.","Because the duplicate condition shows the largest variant-specific estimate, downstream work could isolate the role of repetition and salience by varying the number and position of repeated atoms while holding the resolved contract fixed.","If a short explicit recap of the final contract is prepended to each history and CPV drops without lowering direct accuracy, that would support the paper's view that the problem is active-contract resolution rather than raw implementation skill."],"forward_implications":["An agent that succeeds on a consolidated issue has not thereby demonstrated that it would succeed when the same contract is reached through edits, overrides, or split turns.","Benchmark reports should pair direct final-contract accuracy with a conditional path-violation rate, because near-identical means can hide which blocks succeed.","Contract-equivalent history pairs act as metamorphic tests: same repository, verifier, and budget, with only the route to the final contract changed.","The SpecPath design is portable to any task family with a validated executable verifier and multiple histories that resolve to one contract.","Equal average accuracy across prompts should not be read as equal behavior without a paired analysis of which executions change."],"supporting_citations":[{"why":"Establishes the single-issue repository benchmark whose consolidated specification is the direct condition SpecPath contrasts with equivalent histories.","marker":"Jimenez et al. 2024"},{"why":"Provides semantics-preserving prompt and history transformations that SpecPath adapts to the contract-equivalence relation.","marker":"Wang et al. 2023"},{"why":"Supplies the metamorphic-testing framing in which equivalent inputs are expected to yield equivalent outputs.","marker":"Segura et al. 2016"},{"why":"Introduces contrast-set behavioral testing, the template for SpecPath's matched counterfactual executions.","marker":"Ribeiro et al. 2020"},{"why":"Defines the oracle problem that bounds SpecPath's claims to tested behavior rather than program equivalence.","marker":"Barr et al. 2015"},{"why":"Demonstrates length and position effects in long-context use, motivating the length-matched and position controls.","marker":"Liu et al. 2024"},{"why":"Supplies the consistency and completeness view of requirements evolution that underlies the active-contract formulation.","marker":"Zowghi and Gervasi 2003"}],"fun_headline_variants":["35 of 100 direct wins fail when spec history changes","Stable average, but 35% of coding wins flip with spec path","Spec-path sensitivity: 35 of 100 successes flip on equivalent histories","Same final spec, different path: 35% of agents fail","Aggregate scores hide 35% path-dependent coding failures"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The decisive assumption is that a single failure on an alternative history, after the same block succeeds directly, is evidence of path sensitivity rather than ordinary run-to-run randomness; the paper's own duplicate near-repeat condition already shows a similar flip rate, and the authors note that sampled runs remain stochastic.","fun_headline_variants_meta":{"raw":{"variants":["35 of 100 direct wins fail when spec history changes","Stable average, but 35% of coding wins flip with spec path","Spec-path sensitivity: 35 of 100 successes flip on equivalent histories","Same final spec, different path: 35% of agents fail","Aggregate scores hide 35% path-dependent coding failures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000704,"raw_usage":{"total_tokens":3160,"prompt_tokens":917,"completion_tokens":2243,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":2153}},"tokens_in":533,"tokens_out":2243,"duration_ms":15607,"temperature":1.0,"reasoning_tokens":2153,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:30:40.031537+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run each direct-condition block multiple times (say ten times) under identical configuration and count how often a direct success flips to failure between identical executions; if that direct-versus-direct flip rate matches or exceeds the observed 36.4% conditional path violation rate, the path-sensitivity conclusion would be an artifact of run-to-run noise rather than history structure.","supporting_citations":[],"review_version":1}