REVIEW 3 major objections 6 minor 55 references
Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Under an enforced enterprise-style policy, every tested coding agent lands on a worse success–cost point, and the best choice depends on the policy.
desk verdict A genuinely new benchmark contribution with a transparent audit, but the headline magnitudes rest on three-trial point estimates and author-written solvability witnesses that need independent validation before the numbers are taken as precise. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is Boundary-Bench, a hardening plugin that layers a configurable, operating-system-enforced policy onto Terminal-Bench 2.1 and measures each bundle's success–cost shift relative to its own control. Policies are built from three axes—network egress (N), filesystem scope (F), and privilege (P)—and the paper evaluates three nested levels: control (root, open egress, writable filesystem), non-root (privilege drop only), and NIST-derived high (a fixed 205-domain egress allowlist, read-only system trees outside a small writable set, and a no-new-privileges lockdown), each enforced by native Linux mechanisms such as a default-deny proxy, read-only bind mounts, and an unprivileged user with an empty capability set. The argument is carried by the per-bundle change metrics $\Delta SR_m(\ell)$ and $\Delta C_m(\ell)$ of Eq. (1) plotted as a success–cost Pareto frontier, plus the solvability audit that replays reference solutions under the strict policy, authors policy-compliant witnesses where needed, and additively repairs over-specified verifiers so that valid non-root solutions are credited. This audit is what lets the paper attribute failures to models rather than to tasks the policy forecloses.
What would settle it
Run the 32 witnessed tasks and the seven blocked-by-design tasks from a clean Terminal-Bench 2.1 image under NIST-derived high with the repaired verifiers, without access to the paper's witness trajectories, and have independent implementers attempt each task; then inspect whether each witness passes the verifier by performing the task's intended computation rather than by matching a hardcoded constant or a non-standard route (protein-assembly's SNAP/FLAG values, dna-assembly's coreutils slicing at canonical offsets, or configure-git-webserver's custom SSH server with no UNIX account). If a substantial share of the blocked-by-design tasks turn out to be solvable, or if the witnesses do not reproduce on fresh sandboxes, the measured policy penalties and the model-versus-policy attribution fail.
Extended reading notes
Core claim
The paper's own claim is that hardening does not uniformly degrade coding agents: under NIST-derived high, all twelve bundles shift to strictly worse success–cost operating points, but by amounts and in currencies that differ across models, so the ranking of models is policy-dependent. Success losses reach 18.3 percentage points and cost inflation reaches 167.3 percent, and the two axes disagree—the bundle that gives up the least success (Grok 4.5, −7.1 points) pays the largest cost inflation (+167.3 percent), while the bundle that loses the most success (Claude Sonnet 5, −18.3 points) does so at a modest +21.4 percent cost rise. The additional failures are timeouts and completed-but-wrong solutions rather than early stops, and on the two deep-sampled tasks the median per-run cost roughly triples even where the success rate holds. To ground the comparison, the paper establishes solvability witnesses for 82 of the 89 tasks under the strict policy (50 via the unchanged reference solution, 32 via authored policy-compliant witnesses), identifies seven tasks blocked by design, and repairs five over-specified verifiers after showing their original checks rejected valid non-root solutions. The paper explicitly treats all numbers as describing a spread across bundles, not a ranking.
Load-bearing premise
The load-bearing premise is that the 32 study-authored solvability witnesses are genuine proofs that their tasks remain solvable under NIST-derived high, and that a task with no witness really is foreclosed by the policy; if those witnesses are contrived—for example, hardcoded expected outputs, canonical-offset file slicing, or custom servers that bypass the task's intended architecture—then the paper's separation of model failures from policy-foreclosed tasks, and every reported success penalty, is distorted.
Editorial extensions
If this is right
- Unconstrained leaderboards misrepresent deployed performance: a model that looks best in a permissive sandbox can be the wrong choice under a specific enterprise policy, so model selection should be repeated under the deployment's own policy.
- Deployment budgets must price hardening: because failed runs under policy are timeouts and wrong solutions rather than early stops, wall-clock and token budgets need headroom, and matched passing runs already cost 13% more wall-clock time, 14% more tool calls, and 26% more tokens.
- Reference-solution compatibility predicts where policy will cost money and success before any agent runs, giving benchmark authors a cheap way to flag tasks that will inflate measured penalties.
- Benchmarks should report a policy axis as standard practice, since the paper shows that a common restricted environment restores discriminative headroom that permissive leaderboards have lost.
Reading between the lines
- The 97% attribution of added cost to workaround construction suggests a concrete design target: agents that detect a denial early and plan a policy-compliant route in one pass could recover most of the efficiency loss, which is a testable engineering claim beyond what the paper measures.
- Because the seven blocked-by-design tasks are tied to the specific 205-domain allowlist and read-only mount set, the measured penalties would shift if enterprise policies admit more package mirrors or writable system paths; the paper's policy axis makes such sensitivity explicit rather than hidden.
- A policy-conditional leaderboard would reorder models relative to today's unconstrained rankings, and a natural extension is to check whether the policy-conditional ordering is stable across benchmark families or task pools beyond Terminal-Bench.
- The cost inflation that appears even on tasks where success holds (median per-run cost roughly tripling on both deep-sampled tasks) implies that organizations should price policy not just in lost tasks but in inference spend on successful work.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces Boundary-Bench, a hardening plugin that layers operating-system-enforced security policies onto Terminal-Bench 2.1 and evaluates 12 frozen model-harness bundles under three nested policy levels: control, non-root, and NIST-derived high. The central empirical claim is that hardening is non-uniform: under the strictest policy every bundle's point estimate of success decreases and mean cost per task increases, with success losses up to 18.3 percentage points and cost inflation up to 167.3%, so that model choice is policy-dependent. The paper also decomposes policy-induced failures into timeouts, wrong solutions, and early stops; attributes cost inflation to workaround construction; measures blocked actions at the enforcement boundary; and audits task solvability under the strictest policy, including authoring 32 policy-compliant solvability witnesses and additively repairing five over-specified verifiers. Full per-bundle tables, infrastructure provenance, provider-side intervention documentation, and code are released.
Significance. If the central result holds, the paper makes a valuable contribution: it introduces the security policy level as an explicit independent variable for coding-agent evaluation, a dimension missing from standard leaderboards. The strengths are concrete: native Linux enforcement with pre-flight probes, transparent verifier-audit methodology, explicit documentation of provider-side safety interventions and of a benchmark-integrity artifact (Appendix C), and full per-bundle result tables. The claim that reference-solution compatibility anticipates where a policy will cost success and money is a falsifiable and practically useful prediction. The main risk is the validity of the 32 study-authored solvability witnesses, which carry the separation between model failures and policy-foreclosed tasks; because the paper itself shows the same verifier-award-without-computation failure mode in count-dataset-tokens, this risk is concrete rather than hypothetical.
major comments (3)
- [Appendix B, Table 7; Appendix C] The 32 study-authored solvability witnesses in Table 7 are load-bearing: the paper's own decomposition shows that the large policy penalties are concentrated on the 32 affected tasks (83.5% of additional failures; cost multiplier 2.59x versus 1.14x on unaffected tasks), so the validity of these witnesses determines whether the measured success losses are model failures or policy-foreclosed tasks. Several entries are suspect: protein-assembly hardcodes the SNAP/FLAG values; dna-assembly and dna-insert build primers by pure-coreutils slicing at canonical offsets with no primer-design computation; and configure-git-webserver passes via a custom paramiko SSH server that authenticates a login with no corresponding UNIX account. Appendix C demonstrates that the same verifier-award-without-computation failure mode is real in this benchmark (count-dataset-tokens), so the burden is on the authors to show that each witness performs the substantive task work rather than merely reaching the official passing state. Please provide verifier-level evidence or an independent audit for each of the 32 witnesses, and rerun the affected/unaffected analysis after removing any witness that cannot be defended.
- [Results: Restriction Sensitivity; Tables 12-13] The central claim that under NIST-derived high every bundle's point estimate worsens on both axes is universal, but the supporting numbers are three-trial point estimates with no intervals on costs. In Table 13, for example, MiniMax M3 moves from $18.98 to $19.87 and GPT-5.6 Luna from $12.54 to $13.63 on the 82-task pool, both within plausible run-to-run variability. The paper disclaims ranking and notes the absence of intervals, but the universal direction claim is the paper's headline; please add bootstrap or other confidence intervals for the per-bundle deltas, and a sign test or a sensitivity analysis that excludes the smallest changes, to show that the pattern is not driven by noise.
- [Cost Inflation under Policy; Figure 4] The predictive claim that reference-solution compatibility anticipates where a policy will cost both money and success is partly operationalized by the authors' own witness-authoring effort: a task is affected exactly when the shipped reference solution fails and the authors can write a policy-compliant substitute. If the witness pool is contaminated, the affected/unaffected distinction is not an independent task property and the prediction is circular. Please report how the task-level verifier checks the substantive product for each affected task, and ideally show that the same predictive split holds when the affected set is defined by a property independent of the witness-authoring process, for example whether any canonical toolchain step is blocked by a specific policy axis.
minor comments (6)
- [Experimental Setup; Table 3] The trial protocol says invalid attempts are rerun until the cell holds three valid trials, so it is unclear how Table 3 can report a residual 0.4% of runs per condition as provider- and verifier-side errors excluded; please reconcile.
- [Benchmark Artifacts under Policy] The sentence 'removing a roughly common five to seven points from every measured penalty' is not supported by Tables 12 and 13; the per-bundle differences between the 89-task and 82-task success drops range from about 4.5 to 7.5 points.
- [Figure 2] The dotted line labeled 'robustness Pareto frontier' is not defined in the caption or the surrounding text; please define the dominance relation used.
- [Table 14] The caption does not state the policy level; make explicit that the frontier is for NIST-derived high on the 82-task pool.
- [Results: Cost Inflation under Policy] The 97% workaround-construction attribution is based on the authors' reading of 200 trajectories; please state whether this coding was done with a protocol and whether any inter-rater reliability check was performed.
- [Appendix G] The relationship between the 6,194 verified-complete trials and the 9,612 total trials is only given in the appendix; consider stating in the main text that blocked-action counts are lower bounds due to trace truncation.
Circularity Check
No significant circularity: measured success and cost are external; one minor definitional component in the blocked-by-design classification.
-
self definitional
[Boundary-Bench, Benchmark Audit; Results, Benchmark Artifacts under Policy]
"A task is blocked by design when the policy forbids an action the task itself requires, so no policy-compliant solution exists. ... Under hardening the seven blocked-by-design tasks are guaranteed failures for every bundle alike"
The 'blocked by design' category is defined as 'no policy-compliant solution exists', so the statement that such tasks are guaranteed failures is true by construction. Because the full-pool headline success drops include these seven tasks, a portion of the measured penalty is definitional rather than empirically discovered. The paper mitigates this by separately reporting the 82-task witnessed pool and showing that non-uniform degradation persists there, so this step is not load-bearing for the central claim.
full rationale
The paper's central results are direct measurements against external Terminal-Bench verifiers: success rates and dollar costs are recorded per trial under each policy, and the headline shifts are computed from those measurements via Eq. (1). No parameter is fitted to the data and then renamed as a prediction; the cost and success penalties are not derived from the model's own outputs in a circular way. The only definitional component is the blocked-by-design classification, which is disclosed and robustness-checked by restricting to the 82 witnessed tasks. The solvability-witness process is a correctness assumption about task validity, not a circular reduction: the authors' inability to author a witness does not mathematically force the measured penalties on the witnessed pool. There is also no load-bearing self-citation chain or imported uniqueness theorem. Overall, the derivation chain is self-contained, and the minor definitional aspect is appropriately separated from the main empirical comparison.
Assumptions & free parameters
free parameters (3)
- timeout_budget_fraction =
0.95
- early_stop_bottom_decile =
0.10 (bottom decile of task runtimes)
- egress_allowlist =
205-domain fixed list
assumptions (4)
- domain assumption Terminal-Bench official verifiers, after additive repairs, correctly measure task success.
- domain assumption Native Linux enforcement mechanisms cannot be bypassed by agents and are correctly inherited by child processes.
- ad hoc to paper Study-authored solvability witnesses are genuine policy-compliant solutions.
- domain assumption The egress allowlist is representative of enterprise restrictions and is not fitted to the evaluated tasks.
Cite this review
Pith. "Pith review of Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments." pith.science (2026). https://pith.science/paper/I4ZTPQDU
@misc{pith2026260802670,
author = {Pith},
title = {Pith review of: Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments},
year = {2026},
howpublished = {\url{https://pith.science/paper/I4ZTPQDU}},
note = {Machine review of arXiv:2608.02670}
}
read the original abstract
Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems, non-root execution) constrain them like any other software. Existing benchmarks, however, evaluate agents almost exclusively in permissive sandboxes, so it is unknown how performance changes when policy is enforced. In this work, we evaluate 12 coding agents on Terminal-Bench 2.1 across nested security policy levels derived from common real-world enterprise restrictions. Hardening is never free but far from uniform: under the strictest policy, success losses reach 18.3 points and cost inflation 167.3\%, and the two axes disagree; the model that best preserves success is also the one that loses the most efficiency, so model choice is policy-dependent. Beyond aggregate scores, we characterize how agents behave when policy blocks their actions and decompose the failures hardening induces: runs grind into timeouts or wrong solutions rather than stopping early, in a mix that differs by model. To ground comparisons, we verify task solvability under the strictest policy, separating model failures from tasks the policy forecloses. We release Boundary-Bench, an open-source hardening plugin enabling policy-constrained evaluation of coding agents on Terminal-Bench and compatible benchmarks.
Figures
Reference graph
Works this paper leans on
-
[1]
Merrill, Mike A. and Shaw, Alexander G. and Carlini, Nicholas and Li, Boxuan and Raj, Harsh and Bercovich, Ivan and Shi, Lin and Shin, Jeong Yeon and Walshe, Thomas and Buchanan, E. Kelly and Shen, Junhong and Ye, Guanghao and Lin, Haowei and Poulos, Jason and Wang, Maoyu and Nezhurina, Marianna and Jitsev, Jenia and Lu, Di and Mastromichalakis, Orfeas Me...
-
[2]
doi:10.48550/ARXIV.2306.04528 , abstract =
Zhu, Kaijie and Wang, Jindong and Zhou, Jiaheng and Wang, Zichen and Chen, Hao and Wang, Yidong and Yang, Linyi and Ye, Wei and Zhang, Yue and Gong, Neil Zhenqiang and Xie, Xing , year =. doi:10.48550/ARXIV.2306.04528 , abstract =
-
[3]
Wang, Boxin and Xu, Chejian and Wang, Shuohang and Gan, Zhe and Cheng, Yu and Gao, Jianfeng and Awadallah, Ahmed Hassan and Li, Bo , year =. Adversarial. doi:10.48550/ARXIV.2111.02840 , abstract =
-
[4]
doi:10.48550/ARXIV.2510.03285 , abstract =
Kara, Su and Faisal, Fazle and Nath, Suman , year =. doi:10.48550/ARXIV.2510.03285 , abstract =
-
[5]
Taori, Rohan and Dave, Achal and Shankar, Vaishaal and Carlini, Nicholas and Recht, Benjamin and Schmidt, Ludwig , year =. Measuring. doi:10.48550/ARXIV.2007.00644 , abstract =
-
[6]
and Hashimoto, Tatsunori , year =
Ruan, Yangjun and Dong, Honghua and Wang, Andrew and Pitis, Silviu and Zhou, Yongchao and Ba, Jimmy and Dubois, Yann and Maddison, Chris J. and Hashimoto, Tatsunori , year =. Identifying the. doi:10.48550/ARXIV.2309.15817 , abstract =
-
[7]
Transactions on Machine Learning Research , publisher =
Chen, Lingjiao and Zaharia, Matei and Zou, James , year =. Transactions on Machine Learning Research , publisher =. doi:10.48550/ARXIV.2305.05176 , abstract =
-
[8]
and Nadgir, Nitya and Narayanan, Arvind , year =
Kapoor, Sayash and Stroebl, Benedikt and Siegel, Zachary S. and Nadgir, Nitya and Narayanan, Arvind , year =. Transactions on Machine Learning Research , publisher =. doi:10.48550/ARXIV.2407.01502 , abstract =
Show all 55 references
- [9]
-
[10]
Sah, Tanmay and Srivastava, Vishal and Sah, Dolly and Jordan, Kayden , month = may, year =. The. Proceedings of the. doi:10.1145/3786335.3813160 , language =
- [11]
- [12]
-
[13]
doi:10.48550/ARXIV.2601.06112 , abstract =
Gupta, Aayush , year =. doi:10.48550/ARXIV.2601.06112 , abstract =
-
[14]
doi:10.48550/ARXIV.2410.09024 , abstract =
Andriushchenko, Maksym and Souly, Alexandra and Dziemian, Mateusz and Duenas, Derek and Lin, Maxwell and Wang, Justin and Hendrycks, Dan and Zou, Andy and Kolter, Zico and Fredrikson, Matt and Winsor, Eric and Wynne, Jerome and Gal, Yarin and Davies, Xander , year =. doi:10.48...
- [15]
- [16]
-
[17]
doi:10.48550/ARXIV.2404.07972 , abstract =
Xie, Tianbao and Zhang, Danyang and Chen, Jixuan and Li, Xiaochuan and Zhao, Siheng and Cao, Ruisheng and Hua, Toh Jing and Cheng, Zhoujun and Shin, Dongchan and Lei, Fangyu and Liu, Yitao and Xu, Yiheng and Zhou, Shuyan and Savarese, Silvio and Xiong, Caiming and Zhong, Victo...
-
[18]
and Song, Yufan and Li, Boxuan and Tang, Yuxuan and Jain, Kritanjali and Bao, Mengxue and Wang, Zora Z
Xu, Frank F. and Song, Yufan and Li, Boxuan and Tang, Yuxuan and Jain, Kritanjali and Bao, Mengxue and Wang, Zora Z. and Zhou, Xuhui and Guo, Zhitong and Cao, Murong and Yang, Mingyang and Lu, Hao Yang and Martin, Amaad and Su, Zhe and Maben, Leander and Mehta, Raj and Chi, Wa...
- [19]
-
[20]
and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , year =
Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , year =. doi:10.48550/ARXIV.2310.06770 , abstract =
- [21]
-
[22]
arXiv.org , author =
Authenticated. arXiv.org , author =. 2025 , annote =
2025
-
[23]
doi:10.48550/ARXIV.2602.10975 , abstract =
Zhou, Qixing and Zhang, Jiacheng and Wang, Haiyang and Hao, Rui and Wang, Jiahe and Han, Minghao and Yang, Yuxue and Wu, Shuzhe and Pan, Feiyang and Fan, Lue and Tu, Dandan and Zhang, Zhaoxiang , year =. doi:10.48550/ARXIV.2602.10975 , abstract =
- [24]
- [25]
- [26]
-
[27]
Yehudai, Asaf and Eden, Lilach and Li, Alan and Uziel, Guy and Zhao, Yilun and Bar-Haim, Roy and Cohan, Arman and Shmueli-Scheuer, Michal , year =. A. Findings of the. doi:10.18653/v1/2026.findings-acl.1330 , language =
2026 doi
-
[28]
Sharifloo, Amir Molzam and Heydari, Maedeh and Kazerooni, Parsa and Maninger, Daniel and Mezini, Mira , month = nov, year =. Where. 2025 2nd. doi:10.1109/AIWARE69974.2025.00035 , abstract =
2025
-
[29]
Andriushchenko, M.; Souly, A.; Dziemian, M.; Duenas, D.; Lin, M.; Wang, J.; Hendrycks, D.; Zou, A.; Kolter, Z.; Fredrikson, M.; Winsor, E.; Wynne, J.; Gal, Y.; and Davies, X. 2024. AgentHarm : A Benchmark for Measuring Harmfulness of LLM Agents . Version Number: 3
2024
-
[30]
Chen, L.; Zaharia, M.; and Zou, J. 2024. FrugalGPT : How to Use Large Language Models While Reducing Cost and Improving Performance . Transactions on Machine Learning Research. Version Number: 1
2024
-
[31]
Debenedetti, E.; Shumailov, I.; Fan, T.; Hayes, J.; Carlini, N.; Fabian, D.; Kern, C.; Shi, C.; Terzis, A.; and Tramèr, F. 2025. Defeating Prompt Injections by Design . Version Number: 2
2025
-
[32]
Gupta, A. 2026. ReliabilityBench : Evaluating LLM Agent Reliability Under Production - Like Stress Conditions . Version Number: 1
2026
-
[33]
E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K
Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K. 2023. SWE -bench: Can Language Models Resolve Real - World GitHub Issues ? Version Number: 3
2023
-
[34]
Joint Task Force Interagency Working Group . 2020. Security and Privacy Controls for Information Systems and Organizations . Technical report, National Institute of Standards and Technology. Edition: Revision 5
2020
-
[35]
S.; Nadgir, N.; and Narayanan, A
Kapoor, S.; Stroebl, B.; Siegel, Z. S.; Nadgir, N.; and Narayanan, A. 2025. AI Agents That Matter . Transactions on Machine Learning Research. Version Number: 1
2025
-
[36]
Kara, S.; Faisal, F.; and Nath, S. 2025. WAREX : Web Agent Reliability Evaluation on Existing Benchmarks . Version Number: 1
2025
-
[37]
Levy, I.; Wiesel, B.; Marreed, S.; Oved, A.; Yaeli, A.; Mashkif, N.; and Shlomov, S. 2026. ST - WebAgentBench : A Benchmark for Evaluating Safety and Trustworthiness in Web Agents . ArXiv:2410.06703 [cs.AI]
2026 arXiv
-
[38]
Li, H.; Wen, R.; Shi, S.; Zhang, N.; Vorobeychik, Y.; and Xiao, C. 2026. AgentDyn : Are Your Agent Security Defenses Deployable in Real - World Dynamic Environments ? Version Number: 3
2026
-
[39]
Liu, J.; Qian, C.; Su, Z.; Zong, Q.; Huang, S.; He, B.; and Fung, Y. R. 2026. CostBench : Evaluating Multi - Turn Cost - Optimal Planning and Adaptation in Dynamic Environments for LLM Tool - Use Agents . Version Number: 3
2026
-
[40]
Ma, X.; Wang, Y.; Yao, Y.; Yuan, T.; Zhang, A.; Zhang, Z.; and Zhao, H. 2024. Caution for the Environment : Multimodal LLM Agents are Susceptible to Environmental Distractions . Version Number: 3
2024
-
[41]
A.; Shaw, A
Merrill, M. A.; Shaw, A. G.; Carlini, N.; Li, B.; Raj, H.; Bercovich, I.; Shi, L.; Shin, J. Y.; Walshe, T.; Buchanan, E. K.; Shen, J.; Ye, G.; Lin, H.; Poulos, J.; Wang, M.; Nezhurina, M.; Jitsev, J.; Lu, D.; Mastromichalakis, O. M.; Xu, Z.; Chen, Z.; Liu, Y.; Zhang, R.; Chen,...
2026 arXiv
-
[42]
E.; Kadous, M
Ong, I.; Almahairi, A.; Wu, V.; Chiang, W.-L.; Wu, T.; Gonzalez, J. E.; Kadous, M. W.; and Stoica, I. 2024. RouteLLM : Learning to Route LLMs with Preference Data . Version Number: 4
2024
-
[43]
J.; and Hashimoto, T
Ruan, Y.; Dong, H.; Wang, A.; Pitis, S.; Zhou, Y.; Ba, J.; Dubois, Y.; Maddison, C. J.; and Hashimoto, T. 2023. Identifying the Risks of LM Agents with an LM - Emulated Sandbox . Version Number: 2
2023
-
[44]
Sah, T.; Srivastava, V.; Sah, D.; and Jordan, K. 2026. The Verifier Tax : Horizon Dependent Safety -- Success Tradeoffs in Tool Using LLM Agents . In Proceedings of the ACM Conference on AI and Agentic Systems , 785--799. San Jose CA USA: ACM. ISBN 979-8-4007-2415-2
2026
-
[45]
M.; Heydari, M.; Kazerooni, P.; Maninger, D.; and Mezini, M
Sharifloo, A. M.; Heydari, M.; Kazerooni, P.; Maninger, D.; and Mezini, M. 2025. Where Do LLMs Still Struggle ? An In - Depth Analysis of Code Generation Benchmarks . In 2025 2nd IEEE / ACM International Conference on AI -powered Software ( AIware ) , 249--253. ArXiv:2511.0435...
2025 arXiv
-
[46]
Shi, T.; He, J.; Wang, Z.; Li, H.; Wu, L.; Guo, W.; and Song, D. 2025. Progent: Securing AI Agents with Privilege Control . Version Number: 3
2025
-
[47]
D.; Greenwood, D.; Chan, A.; and Pentland, A
South, T.; Marro, S.; Hardjono, T.; Mahari, R.; Whitney, C. D.; Greenwood, D.; Chan, A.; and Pentland, A. 2025. Authenticated Delegation and Authorized AI Agents
2025
-
[48]
Taori, R.; Dave, A.; Shankar, V.; Carlini, N.; Recht, B.; and Schmidt, L. 2020. Measuring Robustness to Natural Distribution Shifts in Image Classification . Version Number: 2
2020
-
[49]
H.; and Li, B
Wang, B.; Xu, C.; Wang, S.; Gan, Z.; Cheng, Y.; Gao, J.; Awadallah, A. H.; and Li, B. 2021. Adversarial GLUE : A Multi - Task Benchmark for Robustness Evaluation of Language Models . Version Number: 2
2021
-
[50]
J.; Cheng, Z.; Shin, D.; Lei, F.; Liu, Y.; Xu, Y.; Zhou, S.; Savarese, S.; Xiong, C.; Zhong, V.; and Yu, T
Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua, T. J.; Cheng, Z.; Shin, D.; Lei, F.; Liu, Y.; Xu, Y.; Zhou, S.; Savarese, S.; Xiong, C.; Zhong, V.; and Yu, T. 2024. OSWorld : Benchmarking Multimodal Agents for Open - Ended Tasks in Real Computer Environments . Ve...
2024
-
[51]
Xiong, W.; Wang, K.; Song, Y.; Liu, H.; Zhou, S.; Peng, W.; and Li, S. 2025. More Vulnerable than You Think : On the Stability of Tool - Integrated LLM Agents . ArXiv:2506.21967 [cs.CL]
2025 arXiv
-
[52]
Yao, S.; Shinn, N.; Razavi, P.; and Narasimhan, K. 2024. \ τ\ -bench: A Benchmark for Tool - Agent - User Interaction in Real - World Domains . ArXiv:2406.12045 [cs.AI]
2024 arXiv
-
[53]
Yehudai, A.; Eden, L.; Li, A.; Uziel, G.; Zhao, Y.; Bar-Haim, R.; Cohan, A.; and Shmueli-Scheuer, M. 2026. A Survey on Evaluation of LLM -based Agents . In Findings of the Association for Computational Linguistics : ACL 2026 , 26690--26714. San Diego, California, United States...
2026
-
[54]
Zhou, Q.; Zhang, J.; Wang, H.; Hao, R.; Wang, J.; Han, M.; Yang, Y.; Wu, S.; Pan, F.; Fan, L.; Tu, D.; and Zhang, Z. 2026. FeatureBench : Benchmarking Agentic Coding for Complex Feature Development . Version Number: 1
2026
-
[55]
Z.; and Xie, X
Zhu, K.; Wang, J.; Zhou, J.; Wang, Z.; Chen, H.; Wang, Y.; Yang, L.; Ye, W.; Zhang, Y.; Gong, N. Z.; and Xie, X. 2023. PromptRobust : Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts . Version Number: 5
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.