Pith. sign in

REVIEW 3 major objections 6 minor 55 references

Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Under an enforced enterprise-style policy, every tested coding agent lands on a worse success–cost point, and the best choice depends on the policy.

desk verdict A genuinely new benchmark contribution with a transparent audit, but the headline magnitudes rest on three-trial point estimates and author-written solvability witnesses that need independent validation before the numbers are taken as precise. read the letter →

arxiv 2608.02670 v1 pith:I4ZTPQDU submitted 2026-08-02 cs.CR cs.AI

classification cs.CRcs.AI
keywords codingagentssecuritypolicyhardenedenvironmentssuccess-costtrade-offTerminal-BenchsolvabilityauditNISTSP800-53Paretofrontier
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper's central claim is that an enforced security policy is a first-order evaluation axis for coding agents: when agents must run under scoped credentials, restricted egress, and read-only filesystems, their success rates and per-task costs change in ways that permissive-sandbox leaderboards cannot predict. Evaluating twelve frozen model–harness bundles on Terminal-Bench 2.1 under three nested policies, the authors find that under the strictest NIST-derived level every bundle loses success (up to 18.3 percentage points) and pays more per task (up to 167.3% cost inflation), but bundles pay in different currencies: the model that best preserves success absorbs the largest cost increase, while the largest success loss comes with a comparatively modest cost rise. The paper introduces Boundary-Bench, a hardening plugin that enforces policy through native operating-system mechanisms, and pairs it with a solvability audit that separates model failures from tasks the policy forecloses. If the central claim is right, model selection has to be done under the deployment's own policy, and benchmark reporting should treat the policy level as a variable rather than a fixed permissive default.

What carries the argument

The load-bearing mechanism is Boundary-Bench, a hardening plugin that layers a configurable, operating-system-enforced policy onto Terminal-Bench 2.1 and measures each bundle's success–cost shift relative to its own control. Policies are built from three axes—network egress (N), filesystem scope (F), and privilege (P)—and the paper evaluates three nested levels: control (root, open egress, writable filesystem), non-root (privilege drop only), and NIST-derived high (a fixed 205-domain egress allowlist, read-only system trees outside a small writable set, and a no-new-privileges lockdown), each enforced by native Linux mechanisms such as a default-deny proxy, read-only bind mounts, and an unprivileged user with an empty capability set. The argument is carried by the per-bundle change metrics $\Delta SR_m(\ell)$ and $\Delta C_m(\ell)$ of Eq. (1) plotted as a success–cost Pareto frontier, plus the solvability audit that replays reference solutions under the strict policy, authors policy-compliant witnesses where needed, and additively repairs over-specified verifiers so that valid non-root solutions are credited. This audit is what lets the paper attribute failures to models rather than to tasks the policy forecloses.

What would settle it

Run the 32 witnessed tasks and the seven blocked-by-design tasks from a clean Terminal-Bench 2.1 image under NIST-derived high with the repaired verifiers, without access to the paper's witness trajectories, and have independent implementers attempt each task; then inspect whether each witness passes the verifier by performing the task's intended computation rather than by matching a hardcoded constant or a non-standard route (protein-assembly's SNAP/FLAG values, dna-assembly's coreutils slicing at canonical offsets, or configure-git-webserver's custom SSH server with no UNIX account). If a substantial share of the blocked-by-design tasks turn out to be solvable, or if the witnesses do not reproduce on fresh sandboxes, the measured policy penalties and the model-versus-policy attribution fail.

Watch

Extended reading notes

Core claim

The paper's own claim is that hardening does not uniformly degrade coding agents: under NIST-derived high, all twelve bundles shift to strictly worse success–cost operating points, but by amounts and in currencies that differ across models, so the ranking of models is policy-dependent. Success losses reach 18.3 percentage points and cost inflation reaches 167.3 percent, and the two axes disagree—the bundle that gives up the least success (Grok 4.5, −7.1 points) pays the largest cost inflation (+167.3 percent), while the bundle that loses the most success (Claude Sonnet 5, −18.3 points) does so at a modest +21.4 percent cost rise. The additional failures are timeouts and completed-but-wrong solutions rather than early stops, and on the two deep-sampled tasks the median per-run cost roughly triples even where the success rate holds. To ground the comparison, the paper establishes solvability witnesses for 82 of the 89 tasks under the strict policy (50 via the unchanged reference solution, 32 via authored policy-compliant witnesses), identifies seven tasks blocked by design, and repairs five over-specified verifiers after showing their original checks rejected valid non-root solutions. The paper explicitly treats all numbers as describing a spread across bundles, not a ranking.

Load-bearing premise

The load-bearing premise is that the 32 study-authored solvability witnesses are genuine proofs that their tasks remain solvable under NIST-derived high, and that a task with no witness really is foreclosed by the policy; if those witnesses are contrived—for example, hardcoded expected outputs, canonical-offset file slicing, or custom servers that bypass the task's intended architecture—then the paper's separation of model failures from policy-foreclosed tasks, and every reported success penalty, is distorted.

Editorial extensions

If this is right

  • Unconstrained leaderboards misrepresent deployed performance: a model that looks best in a permissive sandbox can be the wrong choice under a specific enterprise policy, so model selection should be repeated under the deployment's own policy.
  • Deployment budgets must price hardening: because failed runs under policy are timeouts and wrong solutions rather than early stops, wall-clock and token budgets need headroom, and matched passing runs already cost 13% more wall-clock time, 14% more tool calls, and 26% more tokens.
  • Reference-solution compatibility predicts where policy will cost money and success before any agent runs, giving benchmark authors a cheap way to flag tasks that will inflate measured penalties.
  • Benchmarks should report a policy axis as standard practice, since the paper shows that a common restricted environment restores discriminative headroom that permissive leaderboards have lost.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 97% attribution of added cost to workaround construction suggests a concrete design target: agents that detect a denial early and plan a policy-compliant route in one pass could recover most of the efficiency loss, which is a testable engineering claim beyond what the paper measures.
  • Because the seven blocked-by-design tasks are tied to the specific 205-domain allowlist and read-only mount set, the measured penalties would shift if enterprise policies admit more package mirrors or writable system paths; the paper's policy axis makes such sensitivity explicit rather than hidden.
  • A policy-conditional leaderboard would reorder models relative to today's unconstrained rankings, and a natural extension is to check whether the policy-conditional ordering is stable across benchmark families or task pools beyond Terminal-Bench.
  • The cost inflation that appears even on tasks where success holds (median per-run cost roughly tripling on both deep-sampled tasks) implies that organizations should price policy not just in lost tasks but in inference spend on successful work.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper introduces Boundary-Bench, a hardening plugin that layers operating-system-enforced security policies onto Terminal-Bench 2.1 and evaluates 12 frozen model-harness bundles under three nested policy levels: control, non-root, and NIST-derived high. The central empirical claim is that hardening is non-uniform: under the strictest policy every bundle's point estimate of success decreases and mean cost per task increases, with success losses up to 18.3 percentage points and cost inflation up to 167.3%, so that model choice is policy-dependent. The paper also decomposes policy-induced failures into timeouts, wrong solutions, and early stops; attributes cost inflation to workaround construction; measures blocked actions at the enforcement boundary; and audits task solvability under the strictest policy, including authoring 32 policy-compliant solvability witnesses and additively repairing five over-specified verifiers. Full per-bundle tables, infrastructure provenance, provider-side intervention documentation, and code are released.

Significance. If the central result holds, the paper makes a valuable contribution: it introduces the security policy level as an explicit independent variable for coding-agent evaluation, a dimension missing from standard leaderboards. The strengths are concrete: native Linux enforcement with pre-flight probes, transparent verifier-audit methodology, explicit documentation of provider-side safety interventions and of a benchmark-integrity artifact (Appendix C), and full per-bundle result tables. The claim that reference-solution compatibility anticipates where a policy will cost success and money is a falsifiable and practically useful prediction. The main risk is the validity of the 32 study-authored solvability witnesses, which carry the separation between model failures and policy-foreclosed tasks; because the paper itself shows the same verifier-award-without-computation failure mode in count-dataset-tokens, this risk is concrete rather than hypothetical.

major comments (3)
  1. [Appendix B, Table 7; Appendix C] The 32 study-authored solvability witnesses in Table 7 are load-bearing: the paper's own decomposition shows that the large policy penalties are concentrated on the 32 affected tasks (83.5% of additional failures; cost multiplier 2.59x versus 1.14x on unaffected tasks), so the validity of these witnesses determines whether the measured success losses are model failures or policy-foreclosed tasks. Several entries are suspect: protein-assembly hardcodes the SNAP/FLAG values; dna-assembly and dna-insert build primers by pure-coreutils slicing at canonical offsets with no primer-design computation; and configure-git-webserver passes via a custom paramiko SSH server that authenticates a login with no corresponding UNIX account. Appendix C demonstrates that the same verifier-award-without-computation failure mode is real in this benchmark (count-dataset-tokens), so the burden is on the authors to show that each witness performs the substantive task work rather than merely reaching the official passing state. Please provide verifier-level evidence or an independent audit for each of the 32 witnesses, and rerun the affected/unaffected analysis after removing any witness that cannot be defended.
  2. [Results: Restriction Sensitivity; Tables 12-13] The central claim that under NIST-derived high every bundle's point estimate worsens on both axes is universal, but the supporting numbers are three-trial point estimates with no intervals on costs. In Table 13, for example, MiniMax M3 moves from $18.98 to $19.87 and GPT-5.6 Luna from $12.54 to $13.63 on the 82-task pool, both within plausible run-to-run variability. The paper disclaims ranking and notes the absence of intervals, but the universal direction claim is the paper's headline; please add bootstrap or other confidence intervals for the per-bundle deltas, and a sign test or a sensitivity analysis that excludes the smallest changes, to show that the pattern is not driven by noise.
  3. [Cost Inflation under Policy; Figure 4] The predictive claim that reference-solution compatibility anticipates where a policy will cost both money and success is partly operationalized by the authors' own witness-authoring effort: a task is affected exactly when the shipped reference solution fails and the authors can write a policy-compliant substitute. If the witness pool is contaminated, the affected/unaffected distinction is not an independent task property and the prediction is circular. Please report how the task-level verifier checks the substantive product for each affected task, and ideally show that the same predictive split holds when the affected set is defined by a property independent of the witness-authoring process, for example whether any canonical toolchain step is blocked by a specific policy axis.
minor comments (6)
  1. [Experimental Setup; Table 3] The trial protocol says invalid attempts are rerun until the cell holds three valid trials, so it is unclear how Table 3 can report a residual 0.4% of runs per condition as provider- and verifier-side errors excluded; please reconcile.
  2. [Benchmark Artifacts under Policy] The sentence 'removing a roughly common five to seven points from every measured penalty' is not supported by Tables 12 and 13; the per-bundle differences between the 89-task and 82-task success drops range from about 4.5 to 7.5 points.
  3. [Figure 2] The dotted line labeled 'robustness Pareto frontier' is not defined in the caption or the surrounding text; please define the dominance relation used.
  4. [Table 14] The caption does not state the policy level; make explicit that the frontier is for NIST-derived high on the 82-task pool.
  5. [Results: Cost Inflation under Policy] The 97% workaround-construction attribution is based on the authors' reading of 200 trajectories; please state whether this coding was done with a protocol and whether any inter-rater reliability check was performed.
  6. [Appendix G] The relationship between the 6,194 verified-complete trials and the 9,612 total trials is only given in the appendix; consider stating in the main text that blocked-action counts are lower bounds due to trace truncation.

Circularity Check

1 steps flagged · score 1.0 of 10

No significant circularity: measured success and cost are external; one minor definitional component in the blocked-by-design classification.

  1. self definitional [Boundary-Bench, Benchmark Audit; Results, Benchmark Artifacts under Policy]
    "A task is blocked by design when the policy forbids an action the task itself requires, so no policy-compliant solution exists. ... Under hardening the seven blocked-by-design tasks are guaranteed failures for every bundle alike"

    The 'blocked by design' category is defined as 'no policy-compliant solution exists', so the statement that such tasks are guaranteed failures is true by construction. Because the full-pool headline success drops include these seven tasks, a portion of the measured penalty is definitional rather than empirically discovered. The paper mitigates this by separately reporting the 82-task witnessed pool and showing that non-uniform degradation persists there, so this step is not load-bearing for the central claim.

full rationale

The paper's central results are direct measurements against external Terminal-Bench verifiers: success rates and dollar costs are recorded per trial under each policy, and the headline shifts are computed from those measurements via Eq. (1). No parameter is fitted to the data and then renamed as a prediction; the cost and success penalties are not derived from the model's own outputs in a circular way. The only definitional component is the blocked-by-design classification, which is disclosed and robustness-checked by restricting to the 82 witnessed tasks. The solvability-witness process is a correctness assumption about task validity, not a circular reduction: the authors' inability to author a witness does not mathematically force the measured penalties on the witnessed pool. There is also no load-bearing self-citation chain or imported uniqueness theorem. Overall, the derivation chain is self-contained, and the minor definitional aspect is appropriately separated from the main empirical comparison.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central empirical comparison rests on the Terminal-Bench verifiers, the enforcement stack, and the author-written solvability witnesses. The first two are domain assumptions with partial validation from pre-flight probes and the verifier audit; the witnesses are the most fragile item because their validity is not independently verified.

free parameters (3)
  • timeout_budget_fraction = 0.95
    Runs reaching 95% of the task's wall-clock budget are classified as timeouts; this threshold is chosen by hand and directly shapes the failure-mode decomposition in Table 3.
  • early_stop_bottom_decile = 0.10 (bottom decile of task runtimes)
    A run ending in the bottom decile of its task's runtime distribution is counted as an early stop; the decile cutoff is arbitrary and influences the claim that agents do not give up more under policy.
  • egress_allowlist = 205-domain fixed list
    The strict-policy egress allowlist is adapted by hand from Claude Code's default and is deliberately not task-fitted, but its composition determines which tasks are blocked or require workarounds and therefore shapes the measured cost inflation.
assumptions (4)
  • domain assumption Terminal-Bench official verifiers, after additive repairs, correctly measure task success.
    All success rates are scored by the benchmark's verifiers; the audit repairs five over-specified verifiers but does not independently validate the rest, so a verifier defect would change measured success.
  • domain assumption Native Linux enforcement mechanisms cannot be bypassed by agents and are correctly inherited by child processes.
    The policy is enforced via proxy, read-only mounts, no_new_privs, and capability drops; pre-flight probes check the mechanisms, but the claim that agents cannot evade them is an operational assumption.
  • ad hoc to paper Study-authored solvability witnesses are genuine policy-compliant solutions.
    The authors define solvability by their own success in writing witnesses that pass the official verifier; this is not independently grounded and is the load-bearing premise for separating model failures from policy-foreclosed tasks.
  • domain assumption The egress allowlist is representative of enterprise restrictions and is not fitted to the evaluated tasks.
    The authors extracted task network dependencies and deliberately excluded 21 domains, assuming an equivalent artifact is reachable through allowlisted hosts for every remaining task; this shapes the policy's difficulty.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments." pith.science (2026). https://pith.science/paper/I4ZTPQDU

@misc{pith2026260802670,
  author       = {Pith},
  title        = {Pith review of: Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/I4ZTPQDU}},
  note         = {Machine review of arXiv:2608.02670}
}
read the original abstract

Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems, non-root execution) constrain them like any other software. Existing benchmarks, however, evaluate agents almost exclusively in permissive sandboxes, so it is unknown how performance changes when policy is enforced. In this work, we evaluate 12 coding agents on Terminal-Bench 2.1 across nested security policy levels derived from common real-world enterprise restrictions. Hardening is never free but far from uniform: under the strictest policy, success losses reach 18.3 points and cost inflation 167.3\%, and the two axes disagree; the model that best preserves success is also the one that loses the most efficiency, so model choice is policy-dependent. Beyond aggregate scores, we characterize how agents behave when policy blocks their actions and decompose the failures hardening induces: runs grind into timeouts or wrong solutions rather than stopping early, in a mix that differs by model. To ground comparisons, we verify task solvability under the strictest policy, separating model failures from tasks the policy forecloses. We release Boundary-Bench, an open-source hardening plugin enabling policy-constrained evaluation of coding agents on Terminal-Bench and compatible benchmarks.

Figures

Figures reproduced from arXiv: 2608.02670 by the authors.

Figure 1
Figure 1. Hardening moves the Pareto frontier down and to the right (lower success, higher cost), by amounts that differ across [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Every bundle loses success and gains cost non [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Policy enforcement raises cost even where suc [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Policy-induced cost inflation concentrates on tasks [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

55 extracted references · 25 canonical work pages

  1. [1]

    and Shaw, Alexander G

    Merrill, Mike A. and Shaw, Alexander G. and Carlini, Nicholas and Li, Boxuan and Raj, Harsh and Bercovich, Ivan and Shi, Lin and Shin, Jeong Yeon and Walshe, Thomas and Buchanan, E. Kelly and Shen, Junhong and Ye, Guanghao and Lin, Haowei and Poulos, Jason and Wang, Maoyu and Nezhurina, Marianna and Jitsev, Jenia and Lu, Di and Mastromichalakis, Orfeas Me...

  2. [2]

    doi:10.48550/ARXIV.2306.04528 , abstract =

    Zhu, Kaijie and Wang, Jindong and Zhou, Jiaheng and Wang, Zichen and Chen, Hao and Wang, Yidong and Yang, Linyi and Ye, Wei and Zhang, Yue and Gong, Neil Zhenqiang and Xie, Xing , year =. doi:10.48550/ARXIV.2306.04528 , abstract =

  3. [3]

    Adversarial

    Wang, Boxin and Xu, Chejian and Wang, Shuohang and Gan, Zhe and Cheng, Yu and Gao, Jianfeng and Awadallah, Ahmed Hassan and Li, Bo , year =. Adversarial. doi:10.48550/ARXIV.2111.02840 , abstract =

  4. [4]

    doi:10.48550/ARXIV.2510.03285 , abstract =

    Kara, Su and Faisal, Fazle and Nath, Suman , year =. doi:10.48550/ARXIV.2510.03285 , abstract =

  5. [5]

    Measuring

    Taori, Rohan and Dave, Achal and Shankar, Vaishaal and Carlini, Nicholas and Recht, Benjamin and Schmidt, Ludwig , year =. Measuring. doi:10.48550/ARXIV.2007.00644 , abstract =

  6. [6]

    and Hashimoto, Tatsunori , year =

    Ruan, Yangjun and Dong, Honghua and Wang, Andrew and Pitis, Silviu and Zhou, Yongchao and Ba, Jimmy and Dubois, Yann and Maddison, Chris J. and Hashimoto, Tatsunori , year =. Identifying the. doi:10.48550/ARXIV.2309.15817 , abstract =

  7. [7]

    Transactions on Machine Learning Research , publisher =

    Chen, Lingjiao and Zaharia, Matei and Zou, James , year =. Transactions on Machine Learning Research , publisher =. doi:10.48550/ARXIV.2305.05176 , abstract =

  8. [8]

    and Nadgir, Nitya and Narayanan, Arvind , year =

    Kapoor, Sayash and Stroebl, Benedikt and Siegel, Zachary S. and Nadgir, Nitya and Narayanan, Arvind , year =. Transactions on Machine Learning Research , publisher =. doi:10.48550/ARXIV.2407.01502 , abstract =

Show all 55 references
  1. [9]

    and Kadous, M Waleed and Stoica, Ion , year =

    Ong, Isaac and Almahairi, Amjad and Wu, Vincent and Chiang, Wei-Lin and Wu, Tianhao and Gonzalez, Joseph E. and Kadous, M Waleed and Stoica, Ion , year =. doi:10.48550/ARXIV.2406.18665 , abstract =

  2. [10]

    Sah, Tanmay and Srivastava, Vishal and Sah, Dolly and Jordan, Kayden , month = may, year =. The. Proceedings of the. doi:10.1145/3786335.3813160 , language =

  3. [11]

    doi:10.48550/arXiv.2410.06703 , abstract =

    Levy, Ido and Wiesel, Ben and Marreed, Sami and Oved, Alon and Yaeli, Avi and Mashkif, Nir and Shlomov, Segev , month = jun, year =. doi:10.48550/arXiv.2410.06703 , abstract =

  4. [12]

    Caution for the

    Ma, Xinbei and Wang, Yiting and Yao, Yao and Yuan, Tongxin and Zhang, Aston and Zhang, Zhuosheng and Zhao, Hai , year =. Caution for the. doi:10.48550/ARXIV.2408.02544 , abstract =

  5. [13]

    doi:10.48550/ARXIV.2601.06112 , abstract =

    Gupta, Aayush , year =. doi:10.48550/ARXIV.2601.06112 , abstract =

  6. [14]

    doi:10.48550/ARXIV.2410.09024 , abstract =

    Andriushchenko, Maksym and Souly, Alexandra and Dziemian, Mateusz and Duenas, Derek and Lin, Maxwell and Wang, Justin and Hendrycks, Dan and Zou, Andy and Kolter, Zico and Fredrikson, Matt and Winsor, Eric and Wynne, Jerome and Gal, Yarin and Davies, Xander , year =. doi:10.48...

  7. [15]

    \ τ\ -bench:

    Yao, Shunyu and Shinn, Noah and Razavi, Pedram and Narasimhan, Karthik , month = jun, year =. \ τ\ -bench:. doi:10.48550/arXiv.2406.12045 , abstract =

  8. [16]

    Xiong, Weimin and Wang, Ke and Song, Yifan and Liu, Hanchao and Zhou, Sai and Peng, Wei and Li, Sujian , month = jun, year =. More. doi:10.48550/arXiv.2506.21967 , abstract =

  9. [17]

    doi:10.48550/ARXIV.2404.07972 , abstract =

    Xie, Tianbao and Zhang, Danyang and Chen, Jixuan and Li, Xiaochuan and Zhao, Siheng and Cao, Ruisheng and Hua, Toh Jing and Cheng, Zhoujun and Shin, Dongchan and Lei, Fangyu and Liu, Yitao and Xu, Yiheng and Zhou, Shuyan and Savarese, Silvio and Xiong, Caiming and Zhong, Victo...

  10. [18]

    and Song, Yufan and Li, Boxuan and Tang, Yuxuan and Jain, Kritanjali and Bao, Mengxue and Wang, Zora Z

    Xu, Frank F. and Song, Yufan and Li, Boxuan and Tang, Yuxuan and Jain, Kritanjali and Bao, Mengxue and Wang, Zora Z. and Zhou, Xuhui and Guo, Zhitong and Cao, Murong and Yang, Mingyang and Lu, Hao Yang and Martin, Amaad and Su, Zhe and Maben, Leander and Mehta, Raj and Chi, Wa...

  11. [19]

    , month = jun, year =

    Liu, Jiayu and Qian, Cheng and Su, Zhaochen and Zong, Qing and Huang, Shijue and He, Bingxiang and Fung, Yi R. , month = jun, year =. doi:10.48550/ARXIV.2511.02734 , abstract =

  12. [20]

    and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , year =

    Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , year =. doi:10.48550/ARXIV.2310.06770 , abstract =

  13. [21]

    2020 , note =

    Security and. 2020 , note =. doi:10.6028/NIST.SP.800-53r5 , urldate =

  14. [22]

    arXiv.org , author =

    Authenticated. arXiv.org , author =. 2025 , annote =

  15. [23]

    doi:10.48550/ARXIV.2602.10975 , abstract =

    Zhou, Qixing and Zhang, Jiacheng and Wang, Haiyang and Hao, Rui and Wang, Jiahe and Han, Minghao and Yang, Yuxue and Wu, Shuzhe and Pan, Feiyang and Fan, Lue and Tu, Dandan and Zhang, Zhaoxiang , year =. doi:10.48550/ARXIV.2602.10975 , abstract =

  16. [24]

    Defeating

    Debenedetti, Edoardo and Shumailov, Ilia and Fan, Tianqi and Hayes, Jamie and Carlini, Nicholas and Fabian, Daniel and Kern, Christoph and Shi, Chongyang and Terzis, Andreas and Tramèr, Florian , year =. Defeating. doi:10.48550/ARXIV.2503.18813 , abstract =

  17. [25]

    Progent:

    Shi, Tianneng and He, Jingxuan and Wang, Zhun and Li, Hongwei and Wu, Linyu and Guo, Wenbo and Song, Dawn , year =. Progent:. doi:10.48550/ARXIV.2504.11703 , abstract =

  18. [26]

    doi:10.48550/ARXIV.2602.03117 , abstract =

    Li, Hao and Wen, Ruoyao and Shi, Shanghao and Zhang, Ning and Vorobeychik, Yevgeniy and Xiao, Chaowei , year =. doi:10.48550/ARXIV.2602.03117 , abstract =

  19. [27]

    Yehudai, Asaf and Eden, Lilach and Li, Alan and Uziel, Guy and Zhao, Yilun and Bar-Haim, Roy and Cohan, Arman and Shmueli-Scheuer, Michal , year =. A. Findings of the. doi:10.18653/v1/2026.findings-acl.1330 , language =

  20. [28]

    Sharifloo, Amir Molzam and Heydari, Maedeh and Kazerooni, Parsa and Maninger, Daniel and Mezini, Mira , month = nov, year =. Where. 2025 2nd. doi:10.1109/AIWARE69974.2025.00035 , abstract =

  21. [29]

    Andriushchenko, M.; Souly, A.; Dziemian, M.; Duenas, D.; Lin, M.; Wang, J.; Hendrycks, D.; Zou, A.; Kolter, Z.; Fredrikson, M.; Winsor, E.; Wynne, J.; Gal, Y.; and Davies, X. 2024. AgentHarm : A Benchmark for Measuring Harmfulness of LLM Agents . Version Number: 3

  22. [30]

    Chen, L.; Zaharia, M.; and Zou, J. 2024. FrugalGPT : How to Use Large Language Models While Reducing Cost and Improving Performance . Transactions on Machine Learning Research. Version Number: 1

  23. [31]

    Debenedetti, E.; Shumailov, I.; Fan, T.; Hayes, J.; Carlini, N.; Fabian, D.; Kern, C.; Shi, C.; Terzis, A.; and Tramèr, F. 2025. Defeating Prompt Injections by Design . Version Number: 2

  24. [32]

    Gupta, A. 2026. ReliabilityBench : Evaluating LLM Agent Reliability Under Production - Like Stress Conditions . Version Number: 1

  25. [33]

    E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K

    Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K. 2023. SWE -bench: Can Language Models Resolve Real - World GitHub Issues ? Version Number: 3

  26. [34]

    Joint Task Force Interagency Working Group . 2020. Security and Privacy Controls for Information Systems and Organizations . Technical report, National Institute of Standards and Technology. Edition: Revision 5

  27. [35]

    S.; Nadgir, N.; and Narayanan, A

    Kapoor, S.; Stroebl, B.; Siegel, Z. S.; Nadgir, N.; and Narayanan, A. 2025. AI Agents That Matter . Transactions on Machine Learning Research. Version Number: 1

  28. [36]

    Kara, S.; Faisal, F.; and Nath, S. 2025. WAREX : Web Agent Reliability Evaluation on Existing Benchmarks . Version Number: 1

  29. [37]

    Levy, I.; Wiesel, B.; Marreed, S.; Oved, A.; Yaeli, A.; Mashkif, N.; and Shlomov, S. 2026. ST - WebAgentBench : A Benchmark for Evaluating Safety and Trustworthiness in Web Agents . ArXiv:2410.06703 [cs.AI]

  30. [38]

    Li, H.; Wen, R.; Shi, S.; Zhang, N.; Vorobeychik, Y.; and Xiao, C. 2026. AgentDyn : Are Your Agent Security Defenses Deployable in Real - World Dynamic Environments ? Version Number: 3

  31. [39]

    Liu, J.; Qian, C.; Su, Z.; Zong, Q.; Huang, S.; He, B.; and Fung, Y. R. 2026. CostBench : Evaluating Multi - Turn Cost - Optimal Planning and Adaptation in Dynamic Environments for LLM Tool - Use Agents . Version Number: 3

  32. [40]

    Ma, X.; Wang, Y.; Yao, Y.; Yuan, T.; Zhang, A.; Zhang, Z.; and Zhao, H. 2024. Caution for the Environment : Multimodal LLM Agents are Susceptible to Environmental Distractions . Version Number: 3

  33. [41]

    A.; Shaw, A

    Merrill, M. A.; Shaw, A. G.; Carlini, N.; Li, B.; Raj, H.; Bercovich, I.; Shi, L.; Shin, J. Y.; Walshe, T.; Buchanan, E. K.; Shen, J.; Ye, G.; Lin, H.; Poulos, J.; Wang, M.; Nezhurina, M.; Jitsev, J.; Lu, D.; Mastromichalakis, O. M.; Xu, Z.; Chen, Z.; Liu, Y.; Zhang, R.; Chen,...

  34. [42]

    E.; Kadous, M

    Ong, I.; Almahairi, A.; Wu, V.; Chiang, W.-L.; Wu, T.; Gonzalez, J. E.; Kadous, M. W.; and Stoica, I. 2024. RouteLLM : Learning to Route LLMs with Preference Data . Version Number: 4

  35. [43]

    J.; and Hashimoto, T

    Ruan, Y.; Dong, H.; Wang, A.; Pitis, S.; Zhou, Y.; Ba, J.; Dubois, Y.; Maddison, C. J.; and Hashimoto, T. 2023. Identifying the Risks of LM Agents with an LM - Emulated Sandbox . Version Number: 2

  36. [44]

    Sah, T.; Srivastava, V.; Sah, D.; and Jordan, K. 2026. The Verifier Tax : Horizon Dependent Safety -- Success Tradeoffs in Tool Using LLM Agents . In Proceedings of the ACM Conference on AI and Agentic Systems , 785--799. San Jose CA USA: ACM. ISBN 979-8-4007-2415-2

  37. [45]

    M.; Heydari, M.; Kazerooni, P.; Maninger, D.; and Mezini, M

    Sharifloo, A. M.; Heydari, M.; Kazerooni, P.; Maninger, D.; and Mezini, M. 2025. Where Do LLMs Still Struggle ? An In - Depth Analysis of Code Generation Benchmarks . In 2025 2nd IEEE / ACM International Conference on AI -powered Software ( AIware ) , 249--253. ArXiv:2511.0435...

  38. [46]

    Shi, T.; He, J.; Wang, Z.; Li, H.; Wu, L.; Guo, W.; and Song, D. 2025. Progent: Securing AI Agents with Privilege Control . Version Number: 3

  39. [47]

    D.; Greenwood, D.; Chan, A.; and Pentland, A

    South, T.; Marro, S.; Hardjono, T.; Mahari, R.; Whitney, C. D.; Greenwood, D.; Chan, A.; and Pentland, A. 2025. Authenticated Delegation and Authorized AI Agents

  40. [48]

    Taori, R.; Dave, A.; Shankar, V.; Carlini, N.; Recht, B.; and Schmidt, L. 2020. Measuring Robustness to Natural Distribution Shifts in Image Classification . Version Number: 2

  41. [49]

    H.; and Li, B

    Wang, B.; Xu, C.; Wang, S.; Gan, Z.; Cheng, Y.; Gao, J.; Awadallah, A. H.; and Li, B. 2021. Adversarial GLUE : A Multi - Task Benchmark for Robustness Evaluation of Language Models . Version Number: 2

  42. [50]

    J.; Cheng, Z.; Shin, D.; Lei, F.; Liu, Y.; Xu, Y.; Zhou, S.; Savarese, S.; Xiong, C.; Zhong, V.; and Yu, T

    Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua, T. J.; Cheng, Z.; Shin, D.; Lei, F.; Liu, Y.; Xu, Y.; Zhou, S.; Savarese, S.; Xiong, C.; Zhong, V.; and Yu, T. 2024. OSWorld : Benchmarking Multimodal Agents for Open - Ended Tasks in Real Computer Environments . Ve...

  43. [51]

    Xiong, W.; Wang, K.; Song, Y.; Liu, H.; Zhou, S.; Peng, W.; and Li, S. 2025. More Vulnerable than You Think : On the Stability of Tool - Integrated LLM Agents . ArXiv:2506.21967 [cs.CL]

  44. [52]

    Yao, S.; Shinn, N.; Razavi, P.; and Narasimhan, K. 2024. \ τ\ -bench: A Benchmark for Tool - Agent - User Interaction in Real - World Domains . ArXiv:2406.12045 [cs.AI]

  45. [53]

    Yehudai, A.; Eden, L.; Li, A.; Uziel, G.; Zhao, Y.; Bar-Haim, R.; Cohan, A.; and Shmueli-Scheuer, M. 2026. A Survey on Evaluation of LLM -based Agents . In Findings of the Association for Computational Linguistics : ACL 2026 , 26690--26714. San Diego, California, United States...

  46. [54]

    Zhou, Q.; Zhang, J.; Wang, H.; Hao, R.; Wang, J.; Han, M.; Yang, Y.; Wu, S.; Pan, F.; Fan, L.; Tu, D.; and Zhang, Z. 2026. FeatureBench : Benchmarking Agentic Coding for Complex Feature Development . Version Number: 1

  47. [55]

    Z.; and Xie, X

    Zhu, K.; Wang, J.; Zhou, J.; Wang, Z.; Chen, H.; Wang, Y.; Yang, L.; Ye, W.; Zhang, Y.; Gong, N. Z.; and Xie, X. 2023. PromptRobust : Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts . Version Number: 5

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.