Pith. sign in

REVIEW 3 major objections 6 minor 99 references

SpecPath: Testing Coding Agents Across Contract-Equivalent Specification Histories

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Coding agents that pass a direct specification often fail on an equivalent history that reaches the same final contract, despite stable aggregate accuracy.

desk verdict A well-built diagnostic for a real problem, but the headline 35/100 needs a rerun control before you can call it path sensitivity. read the letter →

arxiv 2608.09799 v1 pith:NC5CQFQ5 submitted 2026-08-10 cs.SE

classification cs.SE
keywords specification-pathsensitivitycodingagentscontract-equivalenthistoriesconditionalpathviolationactive-contractresolutionbenchmarkmethodologyrequirementsevolutionmetamorphictesting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that coding agents can be sensitive to the path by which a final specification is reached, even when every path leaves the same contract in force. Using SpecPath, a controlled evaluation that fixes the repository, final contract, verifier, agent system, and execution budget while varying only the revision history, the authors show that aggregate accuracy stays nearly constant (direct 78.8% versus an average of 78.7% across equivalent histories) yet 35 of 100 blocks that succeed directly fail on at least one equivalent history. The paper concludes that implementation success on a consolidated request does not guarantee specification-path invariance, and that evaluation of evolving requirements should measure such invariance separately. The finding matters because real requirements evolve through edits, overrides, and cancellations, so a coding agent must determine which requirements still count before writing code.

What carries the argument

The load-bearing machinery is the contract-equivalent history family built from a hidden trace of requirement atoms. Each atom is the smallest testable behavior with an explicit scope, polarity, and observation; replay applies deactivations and activations, and normalization removes presentation order to yield one final active contract. SpecPath then forms blocks that fix task, model deployment, scaffold, and repeat, and defines the conditional path violation: a block where direct execution succeeds but at least one equivalent history fails. Control histories (duplicate, split, override, cancellation, paraphrase-direct, and length-matched) separate revision structure from wording, repetition, and added context, and the executable verifier, not patch identity, is the oracle.

What would settle it

Re-run each direct-condition block multiple times (say ten times) under identical configuration and count how often a direct success flips to failure between identical executions; if that direct-versus-direct flip rate matches or exceeds the observed 36.4% conditional path violation rate, the path-sensitivity conclusion would be an artifact of run-to-run noise rather than history structure.

Watch

Extended reading notes

Core claim

The central discovery is that stable aggregate performance can coexist with systematic path-conditioned failures. SpecPath defines contract-equivalent histories as histories whose hidden replay traces normalize to the same final contract, then executes each agent configuration on fresh copies of the same repository under each history. Across five task families and fourteen agent configurations, direct-condition final-contract realization is 78.8%, while the average over the four equivalent histories is 78.7%; however, among 100 complete blocks with direct success, 35 fail under at least one alternative history, giving a task-macro conditional-path-violation estimate of 36.4%. The authors take this as evidence that the path to a specification, not only its final text, changes which programs an agent produces.

Load-bearing premise

The decisive assumption is that a single failure on an alternative history, after the same block succeeds directly, is evidence of path sensitivity rather than ordinary run-to-run randomness; the paper's own duplicate near-repeat condition already shows a similar flip rate, and the authors note that sampled runs remain stochastic.

Editorial extensions

If this is right

  • An agent that succeeds on a consolidated issue has not thereby demonstrated that it would succeed when the same contract is reached through edits, overrides, or split turns.
  • Benchmark reports should pair direct final-contract accuracy with a conditional path-violation rate, because near-identical means can hide which blocks succeed.
  • Contract-equivalent history pairs act as metamorphic tests: same repository, verifier, and budget, with only the route to the final contract changed.
  • The SpecPath design is portable to any task family with a validated executable verifier and multiple histories that resolve to one contract.
  • Equal average accuracy across prompts should not be read as equal behavior without a paired analysis of which executions change.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to adapt the paired-history design to documentation updates, API migrations, or translated requirements, treating specification-path sensitivity as a general axis of instruction-following evaluation.
  • The 35-of-100 figure is a lower bound: only 127 of 210 possible complete blocks were scored, so missing executions could shift the violation rate; a replication with denser scoring would settle how much attrition matters.
  • Because the duplicate condition shows the largest variant-specific estimate, downstream work could isolate the role of repetition and salience by varying the number and position of repeated atoms while holding the resolved contract fixed.
  • If a short explicit recap of the final contract is prepended to each history and CPV drops without lowering direct accuracy, that would support the paper's view that the problem is active-contract resolution rather than raw implementation skill.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. SpecPath is a controlled evaluation that varies only the revision history leading to a fixed final software contract, holding the repository, verifier, agent system, and execution budget constant. Across five task families and fourteen coding-agent configurations, the paper reports that aggregate final-contract realization is nearly unchanged (78.8% direct vs. 78.7% average over alternative histories), yet 35 of 100 complete blocks that succeed on the direct specification fail on at least one contract-equivalent history. The paper interprets this as specification-path sensitivity and proposes this as a distinct evaluation dimension for coding agents. It includes a detailed construction pipeline, verifier calibration against gold and mutant patches, independent human audits of contract equivalence, explicit attrition accounting, and cluster-bootstrap inference.

Significance. If the central claim holds, the paper makes a valuable methodological contribution: it demonstrates that canonical success on a consolidated requirement does not imply robustness to the history by which the requirement became final, and that aggregate accuracy can hide which blocks are actually succeeding. The construction pipeline is careful and unusually thorough for this area: the verifier is validated against base, gold, and mutant patches before agent outcomes are used; contract equivalence is checked by hidden-trace replay and independent review; and the inference respects the small number of task families via cluster bootstrap. The main risk is that the headline 35/100 result is not separated from run-to-run stochasticity, because every execution is an independent sample and no direct-direct rerun control is reported.

major comments (3)
  1. The any-CPV estimate of 36.4% (95% CI 25.6–45.1%) is computed on blocks where direct succeeds (D=1) and at least one of duplicate, override, cancellation, or split fails (R=1). Because every execution is an independent sample, a second direct run can fail for stochastic reasons, and the manuscript reports no direct-direct rerun control. The duplicate condition, which restates the same final contract in a near-identical way, already produces an 18.3% task-macro variant-specific CPV; under a simple independence model using the four variant-specific rates (18.3%, 10.8%, 14.1%, 12.2%), the expected any-failure rate is 1 − (0.817 × 0.892 × 0.859 × 0.878) ≈ 45%, close to the observed 36.4%. The Discussion states that "separately sampled agent runs remain stochastic" and Threats to Validity states that CPV "is not a common-randomness estimate," but no null baseline is provided. Without a direct-direct CPV estimate (e.g., from the existing three repeats of the direct condition), the headline "35 of 100" cannot be separated from run-to-run noise and the central claim of specification-path sensitivity is not established. This is load-bearing for RQ2 and the abstract.
  2. The abstract claims that "requirement histories that are equivalent in their final meaning lead the same agent system to produce behaviorally different programs." This is a causal attribution, but the design is a single paired sample per block with no common-randomness control; the paper itself says in Threats to Validity that CPV "is not a common-randomness estimate of a deterministic prompt treatment." At most, the data show that a block that succeeds on direct can fail on an alternative history in one sampled run. The causal wording should be tempered to "can coincide with" or "is associated with" unless the direct-direct baseline isolates the path effect. This wording affects the central claim and should be corrected in revision.
  3. Duplicate is simultaneously a core history in the any-CPV definition and a control for "inert repetition." The 18.3% duplicate-specific CPV, being the largest of the four variant-specific estimates, indicates that even a history with no change in active obligations produces frequent failures. This weakens the interpretation that the observed sensitivity is specific to revision structure (override/cancellation/split); the result may reflect sensitivity to any textual variation or sheer run-to-run noise. The paper acknowledges that the ordering is exploratory and that the controls do not identify a unique mechanism, but the introduction's framing of "specification-path sensitivity" as a distinct failure mode of active-contract resolution needs to be reconciled with the size of the duplicate effect. A direct-direct baseline would clarify whether duplicate's effect exceeds the noise floor.
minor comments (6)
  1. The phrase "Grok-4.20-Nonreasoningdeployments" appears to be a typo; if it refers to non-reasoning deployments, please insert a space or use parentheses.
  2. The dashed line marking direct FCR is not labeled in the legend; add an annotation to make the reference visible.
  3. The raster uses a green/red pass/fail palette; for colorblind accessibility, consider a diverging palette or additional cell patterns.
  4. The phrase "RepoBenchinstead" should read "RepoBench instead".
  5. The paper cites both "35 of 100" and the 36.4% task-macro any-CPV. Please clarify explicitly that the former is the unweighted block count and the latter is the task-macro estimate, so readers do not conflate the two.
  6. The notation bPh in Eq. (7) is defined for a history h, while Eq. (8) reuses the macro-averaging notation for CPV; consider adding a sentence distinguishing the two macro definitions to prevent confusion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: SpecPath's central CPV result is a measured conditional proportion on externally validated histories, not a fitted or self-citation-derived quantity.

full rationale

The paper's load-bearing claim is empirical, not derivational. Contract equivalence is defined by the replay and normalization equations (Eqs. 2-3) and is checked by hidden-trace replay plus independent human review, not by the agents' successes; the verifier is validated against base, gold, and mutant patches before agent outcomes are scored. FCR and any-CPV are direct observed proportions (Eqs. 7-8) computed over frozen histories and fixed configurations, and no parameter is fitted to agent outcomes and then renamed as a prediction. The 35-of-100 result is a conditional count from measured runs, not a quantity forced by a definition or by a self-citation chain. Self-citations in the reference list are related-work context, not load-bearing evidence for the central result. The paper's own acknowledged limitations, such as the statement that 'separately sampled agent runs remain stochastic' (Discussion) and that CPV 'is not a common-randomness estimate' (Threats to Validity), weaken the causal interpretation of the measured effect but do not make the measurement circular; they are validity threats, not reductions of the result to its inputs. No circular step matching the enumerated patterns is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on the atom decomposition model and on the validity of contract-equivalent history construction; no numerical free parameters are fitted to produce the result, and no new physical or conceptual entities beyond the benchmark itself are introduced.

assumptions (3)
  • domain assumption A requirement atom is the smallest behavior that can be activated or revoked and tested independently (Problem Formulation).
    The entire atom-replay model assumes requirements decompose into independent atomic behaviors with scope, polarity, and observation; real PRs may not decompose cleanly.
  • domain assumption Replay equation (2) correctly models how a history changes the active contract: M_t = (M_{t-1} \ D_t) U A_t.
    Assumes every revision event is a set-theoretic delete/add on atoms, ignoring partial modifications and ambiguous scope changes.
  • domain assumption Independent human reviewers can recover the same final contract from the visible histories as the hidden trace replay (Visible-History Review and Freeze).
    The equivalence result rests on unanimous reviewer agreement on 25 histories; human judgment is treated as ground truth.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SpecPath: Testing Coding Agents Across Contract-Equivalent Specification Histories." pith.science (2026). https://pith.science/paper/NC5CQFQ5

@misc{pith2026260809799,
  author       = {Pith},
  title        = {Pith review of: SpecPath: Testing Coding Agents Across Contract-Equivalent Specification Histories},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NC5CQFQ5}},
  note         = {Machine review of arXiv:2608.09799}
}
read the original abstract

Modern coding agents increasingly appear capable of following complex software requirements, yet their success leaves a critical ambiguity: do they resolve the active specification, or merely follow the most salient path by which it was stated? We identify specification-path sensitivity, a failure mode in which requirement histories that are equivalent in their final meaning lead the same agent system to produce behaviorally different programs. This reframes evolving-requirement evaluation as active-contract resolution: before writing code, an agent must determine which requirements still count. Building on this view, we introduce SpecPath, a diagnostic evaluation that holds the repository, final contract, verifier, agent system, and execution budget fixed while changing only the revision path that leads to the contract. Rather than treating each patch as an isolated pass or failure, SpecPath uses paired executable outcomes to reveal whether an agent realizes the same tested behavior across contract-equivalent histories. Across five calibrated software tasks and fourteen coding-agent configurations, aggregate direct and revision-history accuracy is nearly unchanged; nevertheless, 35 of 100 complete blocks that succeed on the direct specification fail on at least one equivalent history. These results show that implementation success on a consolidated request does not guarantee specification-path invariance. Evaluating evolving requirements therefore calls for controlled tests of whether agents are robust to the path by which a specification becomes final.

Figures

Figures reproduced from arXiv: 2608.09799 by the authors.

Figure 1
Figure 1. A specification-path counterfactual from Tracecat PR #1245. The direct and override histories resolve to the same [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Construction and validation of each SpecPath task family. Every stage emits a named artifact and must pass its validity gate; a failed audit returns the affected stage for remediation and re-review. Only the frozen package crosses into agent evaluation, whose outcomes never influence task repair or selection. to-test coverage is explicit; gold passes all five histories; base passes none; selective negatives are none… view at source ↗
Figure 3
Figure 3. Stable aggregate accuracy conceals path-conditioned failures. Panel (a) reports task-macro FCR and 95% task-cluster [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

99 extracted references · 61 canonical work pages

  1. [1]

    and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik R

    Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik R. , title =. The Twelfth International Conference on Learning Representations , year =

  2. [2]

    and Wettig, Alexander and Lieret, Kilian and Yao, Shunyu and Narasimhan, Karthik and Press, Ofir , title =

    Yang, John and Jimenez, Carlos E. and Wettig, Alexander and Lieret, Kilian and Yao, Shunyu and Narasimhan, Karthik and Press, Ofir , title =. Advances in Neural Information Processing Systems , volume =. 2024 , url =

  3. [3]

    Proceedings of the ACM on Software Engineering , volume =

    Xia, Chunqiu Steven and Deng, Yinlin and Dunn, Soren and Zhang, Lingming , title =. Proceedings of the ACM on Software Engineering , volume =. 2025 , doi =

  4. [4]

    The Twelfth International Conference on Learning Representations , year =

    Liu, Tianyang and Xu, Canwen and McAuley, Julian , title =. The Twelfth International Conference on Learning Representations , year =

  5. [5]

    2026 , eprint =

    Shen, Haiyang and Chen, Xuanzhong and Xu, Wendong and Ma, Yun and Chen, Liang and Li, Kuan , title =. 2026 , eprint =. doi:10.48550/arXiv.2605.24110 , url =

  6. [6]

    2026 , eprint =

    Deng, Gangda and Chen, Zhaoling and Yu, Zhongming and Fan, Haoyang and Liu, Yuhong and Yang, Yuxin and Parikh, Dhruv and Kannan, Rajgopal and Cong, Le and Wang, Mengdi and Zhang, Qian and Prasanna, Viktor and Tang, Xiangru and Wang, Xingyao , title =. 2026 , eprint =. doi:10.48550/arXiv.2603.13428 , url =

  7. [8]

    Proceedings of the 43rd International Conference on Machine Learning , year=

    Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning , author=. Proceedings of the 43rd International Conference on Machine Learning , year=

  8. [10]

    International Conference on Learning Representations (ICLR) , year=

    Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning , author=. International Conference on Learning Representations (ICLR) , year=

Show all 99 references
  1. [11]

    Findings of the Association for Computational Linguistics: ACL 2025 , pages=

    To code or not to code? adaptive tool integration for math language models via expectation-maximization , author=. Findings of the Association for Computational Linguistics: ACL 2025 , pages=

  2. [12]

    International Conference on Learning Representations (ICLR) , year=

    Reverse-Engineered Reasoning for Open-Ended Generation , author=. International Conference on Learning Representations (ICLR) , year=

  3. [13]

    , title =

    Lam, Man Ho and Wang, Chaozheng and Liu, Hange and Xiao, Jingyu and Li, Haau-sing and Huang, Jen-tse and Zhuo, Terry Yue and Lyu, Michael R. , title =. 2026 , eprint =. doi:10.48550/arXiv.2605.14415 , url =

  4. [14]

    2026 , eprint =

    Huang, Yonghui (Andie) and Ma, Lin and Tahir, Amjed and Zhang, Qian and Xiao, Liwen and Xiao, Lysa , title =. 2026 , eprint =. doi:10.48550/arXiv.2607.01855 , url =

  5. [15]

    Advances in Neural Information Processing Systems , volume =

    Wang, Zhengren and Ling, Rui and Wang, Chufan and Yu, Yongan and Wang, Sizhe and Li, Zhiyu and Xiong, Feiyu and Zhang, Wentao , title =. Advances in Neural Information Processing Systems , volume =. 2025 , url =

  6. [16]

    Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , publisher =

    Wang, Sizhe and Wang, Zhengren and Ma, Dongsheng and Yu, Yongan and Ling, Rui and Li, Zhiyu and Xiong, Feiyu and Zhang, Wentao , title =. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , publisher =. 2026 , doi =

  7. [17]

    2025 , eprint =

    Rawal, Ruchit and Chiang, Jeffrey Yang Fan and Shen, Chihao and Tian, Jeffery Siyuan and Mahajan, Aastha and Goldstein, Tom and Chen, Yizheng , title =. 2025 , eprint =. doi:10.48550/arXiv.2510.13859 , url =

  8. [18]

    2026 , eprint =

    Sobal, Vlad and Yang, Shuo and Zhang, Yuting and Xia, Wei and Soatto, Stefano , title =. 2026 , eprint =. doi:10.48550/arXiv.2606.19613 , url =

  9. [19]

    2026 , eprint =

    Yan, Lu and Chen, Xuan and Zhang, Xiangyu , title =. 2026 , eprint =. doi:10.48550/arXiv.2603.17104 , url =

  10. [20]

    2026 , eprint =

    Raghavendra, Mohit and Gunjal, Anisha and Sabharwal, Aakash and He, Yunzhong , title =. 2026 , eprint =. doi:10.48550/arXiv.2606.30573 , url =

  11. [21]

    2025 , eprint =

    Zhan, Zexun and Gao, Shuzheng and Hu, Ruida and Gao, Cuiyun , title =. 2025 , eprint =. doi:10.48550/arXiv.2509.18808 , url =

  12. [22]

    2025 , eprint =

    Wang, Peiding and Zhang, Li and Liu, Fang and Shi, Lin and Li, Minxiao and Shen, Bo and Fu, An , title =. 2025 , eprint =. doi:10.48550/arXiv.2503.22688 , url =

  13. [23]

    From Tools to Teammates: Evaluating

    Rakotonirina, Nathana. From Tools to Teammates: Evaluating. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , publisher =. 2025 , doi =

  14. [24]

    The Fourteenth International Conference on Learning Representations , year =

    Laban, Philippe and Hayashi, Hiroaki and Zhou, Yingbo and Neville, Jennifer , title =. The Fourteenth International Conference on Learning Representations , year =

  15. [25]

    and Pombal, Jos

    Canaverde, Beatriz and Alves, Duarte M. and Pombal, Jos. 2026 , eprint =. doi:10.48550/arXiv.2605.06353 , url =

  16. [26]

    Findings of the Association for Computational Linguistics: EMNLP 2025 , publisher =

    Han, Chi and Liu, Xin and Wang, Haodong and Li, Shiyang and Yang, Jingfeng and Jiang, Haoming and Wang, Zhengyang and Yin, Qingyu and Qiu, Liang and Yu, Changlong and Gao, Yifan and Li, Zheng and Yin, Bing and Shang, Jingbo and Ji, Heng , title =. Findings of the Association f...

  17. [27]

    Findings of the Association for Computational Linguistics: ACL 2024 , publisher =

    Yan, Jianhao and Luo, Yun and Zhang, Yue , title =. Findings of the Association for Computational Linguistics: ACL 2024 , publisher =. 2024 , doi =

  18. [28]

    2026 , eprint =

    Zhai, Zhiyuan and Li, Ming and Wang, Xin , title =. 2026 , eprint =. doi:10.48550/arXiv.2604.23283 , url =

  19. [29]

    and Wang, Kevin and Li, Junbo and Vikalo, Haris and Akella, Aditya and Wang, Zhangyang , title =

    Zhu, Jianing and Ro, Yeonju and Robertson, John T. and Wang, Kevin and Li, Junbo and Vikalo, Haris and Akella, Aditya and Wang, Zhangyang , title =. 2026 , eprint =. doi:10.48550/arXiv.2605.26302 , url =

  20. [30]

    Findings of the Association for Computational Linguistics: NAACL 2025 , publisher =

    Lu, Jiarui and Holleis, Thomas and Zhang, Yizhe and Aumayer, Bernhard and Nan, Feng and Bai, Haoping and Ma, Shuang and Ma, Shen and Li, Mengyu and Yin, Guoli and Wang, Zirui and Pang, Ruoming , title =. Findings of the Association for Computational Linguistics: NAACL 2025 , p...

  21. [31]

    The Thirteenth International Conference on Learning Representations , year =

    Yao, Shunyu and Shinn, Noah and Razavi, Pedram and Narasimhan, Karthik , title =. The Thirteenth International Conference on Learning Representations , year =

  22. [32]

    IEEE Transactions on Software Engineering , year =

    Guo, Guoxiang (Aaron) and Aleti, Aldeida and Neelofar, Neelofar and Tantithamthavorn, Chakkrit and Qi, Yuanyuan and Chen, Tsong Yueh , title =. IEEE Transactions on Software Engineering , year =. doi:10.1109/TSE.2026.3701230 , url =

  23. [33]

    Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , publisher =

    Wang, Shiqi and Li, Zheng and Qian, Haifeng and Yang, Chenghao and Wang, Zijian and Shang, Mingyue and Kumar, Varun and Tan, Samson and Ray, Baishakhi and Bhatia, Parminder and Nallapati, Ramesh and Ramanathan, Murali Krishna and Roth, Dan and Xiang, Bing , title =. Proceeding...

  24. [34]

    2024 , eprint =

    Wang, Xiaoyin and Zhu, Dakai , title =. 2024 , eprint =. doi:10.48550/arXiv.2406.06864 , url =

  25. [35]

    2026 IEEE/ACM 48th International Conference on Software Engineering , publisher =

    Wang, You and Pradel, Michael and Liu, Zhongxin , title =. 2026 IEEE/ACM 48th International Conference on Software Engineering , publisher =. 2026 , numpages =. doi:10.1145/3744916.3764576 , url =

  26. [36]

    Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , publisher =

    Ribeiro, Marco Tulio and Wu, Tongshuang and Guestrin, Carlos and Singh, Sameer , title =. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , publisher =. 2020 , doi =

  27. [37]

    and Mulcaire, Phoebe and Ning, Qiang and Singh, Sameer and Smith, Noah A

    Gardner, Matt and Artzi, Yoav and Basmov, Victoria and Berant, Jonathan and Bogin, Ben and Chen, Sihao and Dasigi, Pradeep and Dua, Dheeru and Elazar, Yanai and Gottumukkala, Ananth and Gupta, Nitish and Hajishirzi, Hannaneh and Ilharco, Gabriel and Khashabi, Daniel and Lin, K...

  28. [38]

    A Survey on Metamorphic Testing , journal =

    Segura, Sergio and Fraser, Gordon and S. A Survey on Metamorphic Testing , journal =. 2016 , doi =

  29. [39]

    and Harman, Mark and McMinn, Phil and Shahbaz, Muzammil and Yoo, Shin , title =

    Barr, Earl T. and Harman, Mark and McMinn, Phil and Shahbaz, Muzammil and Yoo, Shin , title =. IEEE Transactions on Software Engineering , volume =. 2015 , doi =

  30. [40]

    Advances in Neural Information Processing Systems , volume =

    Liu, Jiawei and Xia, Chunqiu Steven and Wang, Yuyao and Zhang, Lingming , title =. Advances in Neural Information Processing Systems , volume =. 2023 , doi =

  31. [41]

    and Lin, Kevin and Hewitt, John and Paranjape, Ashwin and Bevilacqua, Michele and Petroni, Fabio and Liang, Percy , title =

    Liu, Nelson F. and Lin, Kevin and Hewitt, John and Paranjape, Ashwin and Bevilacqua, Michele and Petroni, Fabio and Liang, Percy , title =. Transactions of the Association for Computational Linguistics , volume =. 2024 , doi =

  32. [42]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , publisher =

    Bai, Yushi and Lv, Xin and Zhang, Jiajie and Lyu, Hongchang and Tang, Jiankai and Huang, Zhidian and Du, Zhengxiao and Liu, Xiao and Zeng, Aohan and Hou, Lei and Dong, Yuxiao and Tang, Jie and Li, Juanzi , title =. Proceedings of the 62nd Annual Meeting of the Association for ...

  33. [43]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , publisher =

    Jiang, Yuxin and Wang, Yufei and Zeng, Xingshan and Zhong, Wanjun and Li, Liangyou and Mi, Fei and Shang, Lifeng and Jiang, Xin and Liu, Qun and Wang, Wei , title =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)...

  34. [44]

    2023 , eprint =

    Zhou, Jeffrey and Lu, Tianjian and Mishra, Swaroop and Brahma, Siddhartha and Basu, Sujoy and Luan, Yi and Zhou, Denny and Hou, Le , title =. 2023 , eprint =. doi:10.48550/arXiv.2311.07911 , url =

  35. [45]

    Information and Software Technology , volume =

    Zowghi, Didar and Gervasi, Vincenzo , title =. Information and Software Technology , volume =. 2003 , doi =

  36. [46]

    IEEE Transactions on Software Engineering , volume =

    Madampe, Kashumi and Hoda, Rashina and Grundy, John , title =. IEEE Transactions on Software Engineering , volume =. 2022 , doi =

  37. [47]

    Findings of the Association for Computational Linguistics: ACL 2025 , publisher =

    Pan, Jane and Shar, Ryan and Pfau, Jacob and Talwalkar, Ameet and He, He and Chen, Valerie , title =. Findings of the Association for Computational Linguistics: ACL 2025 , publisher =. 2025 , doi =

  38. [48]

    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , publisher =

    Kwan, Wai-Chung and Zeng, Xingshan and Jiang, Yuxin and Wang, Yufei and Li, Liangyou and Shang, Lifeng and Jiang, Xin and Liu, Qun and Wong, Kam-Fai , title =. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , publisher =. 2024 , doi =

  39. [49]

    Advances in Neural Information Processing Systems , volume =

    Kim, Myeongsoo and Garg, Shweta and Ray, Baishakhi and Kumar, Varun and Deoras, Anoop , title =. Advances in Neural Information Processing Systems , volume =. 2025 , url =

  40. [50]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume =

    Yang, Weibin and Xie, Liangru and Cai, Jieyun and Yan, Yuxiang and Dai, Hong-Ning and Wang, Hao , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2026 , doi =

  41. [51]

    The Thirteenth International Conference on Learning Representations , year =

    Han, Hojae and Hwang, Seung-won and Samdani, Rajhans and He, Yuxiong , title =. The Thirteenth International Conference on Learning Representations , year =

  42. [52]

    Proceedings of the Fifth ACM SIGPLAN International Conference on Functional Programming , publisher =

    Claessen, Koen and Hughes, John , title =. Proceedings of the Fifth ACM SIGPLAN International Conference on Functional Programming , publisher =. 2000 , doi =

  43. [53]

    , title =

    Jaffe, Andrew and Reicin, Noah and Choi, Jinho D. , title =. 2026 , eprint =. doi:10.48550/arXiv.2601.18924 , url =

  44. [54]

    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , publisher =

    Gupta, Akash and Sheth, Ivaxi and Raina, Vyas and Gales, Mark and Fritz, Mario , title =. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , publisher =. 2024 , doi =

  45. [55]

    Zhang, Zhihan and Li, Shiyang and Zhang, Zixuan and Liu, Xin and Jiang, Haoming and Tang, Xianfeng and Gao, Yifan and Li, Zheng and Wang, Haodong and Tan, Zhaoxuan and Li, Yichuan and Yin, Qingyu and Yin, Bing and Jiang, Meng , title =. Proceedings of the 2025 Conference of th...

  46. [56]

    2026 , eprint =

    Chen, Jialong and Xu, Xander and Wei, Hu and Chen, Chuan and Zhao, Bing , title =. 2026 , eprint =. doi:10.48550/arXiv.2603.03823 , url =

  47. [57]

    2026 , eprint =

    Orlanski, Gabriel and Roy, Devjeet and Yun, Alexander and Shin, Changho and Gu, Alex and Ge, Albert and Adila, Dyah and Roberts, Nicholas and Sala, Frederic and Albarghouthi, Aws , title =. 2026 , eprint =. doi:10.48550/arXiv.2603.24755 , url =

  48. [58]

    2026 , eprint =

    Guo, Guoxiang and Tantithamthavorn, Chakkrit and Neelofar, Neelofar and Qi, Yuanyuan and Aleti, Aldeida , title =. 2026 , eprint =. doi:10.48550/arXiv.2606.25747 , url =

  49. [59]

    Bai, Y.; Lv, X.; Zhang, J.; Lyu, H.; Tang, J.; Huang, Z.; Du, Z.; Liu, X.; Zeng, A.; Hou, L.; Dong, Y.; Tang, J.; and Li, J. 2024. LongBench : A Bilingual, Multitask Benchmark for Long Context Understanding. In Proceedings of the 62nd Annual Meeting of the Association for Comp...

  50. [60]

    T.; Harman, M.; McMinn, P.; Shahbaz, M.; and Yoo, S

    Barr, E. T.; Harman, M.; McMinn, P.; Shahbaz, M.; and Yoo, S. 2015. The Oracle Problem in Software Testing: A Survey. IEEE Transactions on Software Engineering, 41(5): 507--525

  51. [61]

    Deng, G.; Chen, Z.; Yu, Z.; Fan, H.; Liu, Y.; Yang, Y.; Parikh, D.; Kannan, R.; Cong, L.; Wang, M.; Zhang, Q.; Prasanna, V.; Tang, X.; and Wang, X. 2026. SWE-Milestone : Evaluating AI Agents on Continuous Software Evolution. Version 4; ICML 2026, arXiv:2603.13428

  52. [62]

    F.; Mulcaire, P.; Ning, Q.; Singh, S.; Smith, N

    Gardner, M.; Artzi, Y.; Basmov, V.; Berant, J.; Bogin, B.; Chen, S.; Dasigi, P.; Dua, D.; Elazar, Y.; Gottumukkala, A.; Gupta, N.; Hajishirzi, H.; Ilharco, G.; Khashabi, D.; Lin, K.; Liu, J.; Liu, N. F.; Mulcaire, P.; Ning, Q.; Singh, S.; Smith, N. A.; Subramanian, S.; Tsarfat...

  53. [63]

    A.; Aleti, A.; Neelofar, N.; Tantithamthavorn, C.; Qi, Y.; and Chen, T

    Guo, G. A.; Aleti, A.; Neelofar, N.; Tantithamthavorn, C.; Qi, Y.; and Chen, T. Y. 2026. MORTAR : Multi-turn Metamorphic Testing for LLM -based Dialogue Systems. IEEE Transactions on Software Engineering, 1--18

  54. [64]

    Gupta, A.; Sheth, I.; Raina, V.; Gales, M.; and Fritz, M. 2024. LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 14633--14652. Miami, Flori...

  55. [65]

    Han, C.; Liu, X.; Wang, H.; Li, S.; Yang, J.; Jiang, H.; Wang, Z.; Yin, Q.; Qiu, L.; Yu, C.; Gao, Y.; Li, Z.; Yin, B.; Shang, J.; and Ji, H. 2025. Can Language Models Follow Multiple Turns of Entangled Instructions? In Findings of the Association for Computational Linguistics:...

  56. [66]

    A.; Ma, L.; Tahir, A.; Zhang, Q.; Xiao, L.; and Xiao, L

    Huang, Y. A.; Ma, L.; Tahir, A.; Zhang, Q.; Xiao, L.; and Xiao, L. 2026. Regression Accumulation in Multi-Turn LLM Programming Conversations. Accepted to ASE 2026; formal proceedings metadata not yet available, arXiv:2607.01855

  57. [67]

    Jaffe, A.; Reicin, N.; and Choi, J. D. 2026. RIFT : Reordered Instruction Following Testbed To Evaluate Instruction Following in Singular Multistep Prompt Structures. arXiv:2601.18924

  58. [68]

    E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K

    Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K. R. 2024. SWE-bench : Can Language Models Resolve Real-World GitHub Issues? In The Twelfth International Conference on Learning Representations

  59. [69]

    Laban, P.; Hayashi, H.; Zhou, Y.; and Neville, J. 2026. LLM s Get Lost in Multi-Turn Conversation. In The Fourteenth International Conference on Learning Representations

  60. [70]

    H.; Wang, C.; Liu, H.; Xiao, J.; Li, H.-s.; Huang, J.-t.; Zhuo, T

    Lam, M. H.; Wang, C.; Liu, H.; Xiao, J.; Li, H.-s.; Huang, J.-t.; Zhuo, T. Y.; and Lyu, M. R. 2026. SWE-Chain : Benchmarking Coding Agents on Chained Release-Level Package Upgrades. arXiv:2605.14415

  61. [71]

    S.; Wang, Y.; and Zhang, L

    Liu, J.; Xia, C. S.; Wang, Y.; and Zhang, L. 2023. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. In Advances in Neural Information Processing Systems, volume 36, 21558--21572

  62. [72]

    F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P

    Liu, N. F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12: 157--173

  63. [73]

    Liu, T.; Xu, C.; and McAuley, J. 2024. RepoBench : Benchmarking Repository-Level Code Auto-Completion Systems. In The Twelfth International Conference on Learning Representations

  64. [74]

    Madampe, K.; Hoda, R.; and Grundy, J. 2022. A Faceted Taxonomy of Requirements Changes in Agile Contexts. IEEE Transactions on Software Engineering, 48(10): 3737--3752

  65. [75]

    Raghavendra, M.; Gunjal, A.; Sabharwal, A.; and He, Y. 2026. SWE-INTERACT : Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions. arXiv:2606.30573

  66. [76]

    C.; Hamdy, M.; Campos, J

    Rakotonirina, N. C.; Hamdy, M.; Campos, J. A.; Weber, L.; Testoni, A.; Fadaee, M.; Pezzelle, S.; and Del Tredici, M. 2025. From Tools to Teammates: Evaluating LLM s in Multi-Session Coding Interactions. In Proceedings of the 63rd Annual Meeting of the Association for Computati...

  67. [77]

    Rawal, R.; Chiang, J. Y. F.; Shen, C.; Tian, J. S.; Mahajan, A.; Goldstein, T.; and Chen, Y. 2025. Benchmarking Correctness and Security in Multi-Turn Code Generation. arXiv:2510.13859

  68. [78]

    T.; Wu, T.; Guestrin, C.; and Singh, S

    Ribeiro, M. T.; Wu, T.; Guestrin, C.; and Singh, S. 2020. Beyond Accuracy: Behavioral Testing of NLP Models with CheckList . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 4902--4912. Online: Association for Computational Linguistics

  69. [79]

    B.; and Ruiz-Cort \'e s, A

    Segura, S.; Fraser, G.; S \'a nchez, A. B.; and Ruiz-Cort \'e s, A. 2016. A Survey on Metamorphic Testing. IEEE Transactions on Software Engineering, 42(9): 805--824

  70. [80]

    Shen, H.; Chen, X.; Xu, W.; Ma, Y.; Chen, L.; and Li, K. 2026. EvoCode-Bench : Evaluating Coding Agents in Multi-Turn Iterative Interactions. arXiv:2605.24110

  71. [81]

    Sobal, V.; Yang, S.; Zhang, Y.; Xia, W.; and Soatto, S. 2026. StaminaBench : Stress-Testing Coding Agents over 100 Interaction Turns. arXiv:2606.19613

  72. [82]

    Wang, H.; Feng, W.; Yu, J.; Liu, C.; Nie, P.; Lin, F.; Liu, J.; Huang, R.; Lin, J.; Chen, W.; et al. 2026 a . Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation. arXiv preprint arXiv:2607.05382

  73. [83]

    Wang, H.; Li, L.; Qu, C.; Xu, W.; Zhu, F.; Chu, W.; and Lin, F. 2025 a . To code or not to code? adaptive tool integration for math language models via expectation-maximization. In Findings of the Association for Computational Linguistics: ACL 2025, 3060--3075

  74. [84]

    Wang, H.; Que, H.; Xu, Q.; Liu, M.; Zhou, W.; Feng, J.; Zhong, W.; Ye, W.; Yang, T.; Huang, W.; et al. 2026 b . Reverse-Engineered Reasoning for Open-Ended Generation. In International Conference on Learning Representations (ICLR)

  75. [85]

    Wang, H.; Wei, C.; Ren, W.; Liu, J.; Lin, F.; and Chen, W. 2026 c . RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time. arXiv preprint arXiv:2604.11626

  76. [86]

    Wang, H.; Xu, Q.; Liu, C.; Wu, J.; Lin, F.; and Chen, W. 2026 d . Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning. In International Conference on Learning Representations (ICLR)

  77. [87]

    Wang, H.; Xu, Q.; Wang, C.; Xue, T.; Peng, C.; Chen, W.; and Lin, F. 2026 e . Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning. In Proceedings of the 43rd International Conference on Machine Learning

  78. [88]

    Wang, P.; Zhang, L.; Liu, F.; Shi, L.; Li, M.; Shen, B.; and Fu, A. 2025 b . CodeIF-Bench : Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation. Version 4, revised 23 November 2025, arXiv:2503.22688

  79. [89]

    K.; Roth, D.; and Xiang, B

    Wang, S.; Li, Z.; Qian, H.; Yang, C.; Wang, Z.; Shang, M.; Kumar, V.; Tan, S.; Ray, B.; Bhatia, P.; Nallapati, R.; Ramanathan, M. K.; Roth, D.; and Xiang, B. 2023. ReCode : Robustness Evaluation of Code Generation Models. In Proceedings of the 61st Annual Meeting of the Associ...

  80. [90]

    Wang, S.; Wang, Z.; Ma, D.; Yu, Y.; Ling, R.; Li, Z.; Xiong, F.; and Zhang, W. 2026 f . CodeFlowBench : A Multi-turn, Iterative Benchmark for Complex Code Generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pap...

  81. [91]

    Wang, X.; and Zhu, D. 2024. Validating LLM -Generated Programs with Metamorphic Prompt Testing. arXiv:2406.06864

  82. [92]

    Wang, Y.; Pradel, M.; and Liu, Z. 2026. Are ``Solved Issues'' in SWE-bench Really Solved Correctly? An Empirical Study. In 2026 IEEE/ACM 48th International Conference on Software Engineering. New York, NY, USA: Association for Computing Machinery. ISBN 979-8-4007-2025-3

  83. [93]

    Wang, Z.; Ling, R.; Wang, C.; Yu, Y.; Wang, S.; Li, Z.; Xiong, F.; and Zhang, W. 2025 c . MaintainCoder : Maintainable Code Generation Under Dynamic Requirements. In Advances in Neural Information Processing Systems, volume 38

  84. [94]

    S.; Deng, Y.; Dunn, S.; and Zhang, L

    Xia, C. S.; Deng, Y.; Dunn, S.; and Zhang, L. 2025. Demystifying LLM -Based Software Engineering Agents. Proceedings of the ACM on Software Engineering, 2(FSE): 801--824

  85. [95]

    Yan, J.; Luo, Y.; and Zhang, Y. 2024. RefuteBench : Evaluating Refuting Instruction Following for Large Language Models. In Findings of the Association for Computational Linguistics: ACL 2024, 13775--13791. Bangkok, Thailand: Association for Computational Linguistics

  86. [96]

    Yan, L.; Chen, X.; and Zhang, X. 2026. When the Specification Emerges: Benchmarking Faithfulness Loss in Long-Horizon Coding Agents. arXiv:2603.17104

  87. [97]

    E.; Wettig, A.; Lieret, K.; Yao, S.; Narasimhan, K.; and Press, O

    Yang, J.; Jimenez, C. E.; Wettig, A.; Lieret, K.; Yao, S.; Narasimhan, K.; and Press, O. 2024. SWE-agent : Agent-Computer Interfaces Enable Automated Software Engineering. In Advances in Neural Information Processing Systems, volume 37, 50528--50652

  88. [98]

    Zhai, Z.; Li, M.; and Wang, X. 2026. Revisable by Design: A Theory of Streaming LLM Agent Execution. arXiv:2604.23283

  89. [99]

    Zhan, Z.; Gao, S.; Hu, R.; and Gao, C. 2025. SR-Eval : Evaluating LLM s on Code Generation under Stepwise Requirement Refinement. Version 2, revised 17 March 2026, arXiv:2509.18808

  90. [100]

    T.; Wang, K.; Li, J.; Vikalo, H.; Akella, A.; and Wang, Z

    Zhu, J.; Ro, Y.; Robertson, J. T.; Wang, K.; Li, J.; Vikalo, H.; Akella, A.; and Wang, Z. 2026. Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems. arXiv:2605.26302

  91. [101]

    Zowghi, D.; and Gervasi, V. 2003. On the Interplay Between Consistency, Completeness, and Correctness in Requirements Evolution. Information and Software Technology, 45(14): 993--1009

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.