REVIEW 4 major objections 4 minor 144 references
Agentic AI trustworthiness is one problem observed from four vantage points, not four separate ones.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 15:02 UTC pith:WYAV2NTH
load-bearing objection A structured, useful survey that makes a plausible cross-domain case; the 'one problem' claim is a well-argued research program, not an established result. the 4 major comments →
Engineering Trustworthy Agentic AI for Critical Systems
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Observed across four constraint-bound domains, agentic AI systems fail in the same way: gaps at the reasoning-to-action boundary, robustness evidence that stops at curated benchmarks, audit trails that cover fragments rather than the full workflow, and security surfaces that grow with every tool or coordination channel. The paper reads this as evidence that trustworthiness is a single problem with domain-specific vocabulary, and it outlines a path toward a cross-domain graded assurance framework. The central claim is that the same pipeline, perception through audit, appears under different physical constraints, and that the location of the risky boundary, not the model architecture, determin
What carries the argument
The analytic scaffold is a ten-stage agentic workflow (perception, reasoning, planning, tool use, communication, memory update, validation, action, monitoring, audit) with five cross-cutting trustworthiness dimensions (safety and constraint satisfaction, robustness and reliability, transparency and interpretability, accountability and auditability, privacy and security). This workflow-to-dimension mapping lets the authors ask where trust mechanisms sit and where they are absent, and it produces the cross-domain comparison that yields the 'one problem' claim.
Load-bearing premise
The claim rests on the assumption that the selected papers fairly represent each domain; because the review is curated rather than systematic, the recurring failure modes could be an artifact of the sample.
What would settle it
A systematic review using explicit inclusion criteria that locates a critical-infrastructure domain where the four failure modes do not co-occur, or where a domain-specific failure mode carries most of the risk, would undermine the 'single problem' thesis.
If this is right
- A standard trustworthiness scorecard could be reported alongside task-success metrics in every agentic AI evaluation.
- Autonomy boundaries could be defined before deployment as graded levels, so that advisory, supervised, and fully autonomous authority are comparable across domains.
- Adversarial and rare-event test suites could be shared across power, robotics, HPC, and networks instead of being recreated separately in each field.
- Provenance standards for tools, memory, and multi-agent coordination could be built proactively, drawing on existing automotive and aviation traceability practice, before an incident forces them.
- A reusable framework would let a validation result for a grid-dispatch agent and a result for a UAV planner share a common unit of evidence.
Where Pith is reading between the lines
- A direct test of the paper's thesis: apply the same ten-stage, five-dimension analysis to a domain it did not review, such as healthcare or finance; if the same four failure modes appear, the 'single problem' claim is strengthened.
- The framework implies that capability benchmarks, such as reasoning or tool-use scores, should be subordinate to workflow-level assurance evidence, which would shift how agentic AI progress is measured.
- Because the recurring failure modes all live at the interpretation-to-action boundary, the paper suggests that future architectures may converge on a 'demoted orchestrator' pattern where the LLM plans but a verified component decides and executes.
- The graded-certification analogy carries a testable implication: agents in different domains should be comparable on the same axes, such as consequence severity, response time, and audit maturity, which the paper sketches but does not yet formalize.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a survey that proposes a trustworthiness model for agentic AI in critical systems, organized around five dimensions (safety/constraint satisfaction, robustness/reliability, transparency/interpretability, accountability/auditability, privacy/security) and a ten-point agentic workflow from perception to audit. It reviews agent architectures, trust mechanisms, evaluation approaches, and quantitative metrics, then applies the model to four domains: power systems, autonomous vehicles/robotics/UAVs, high-performance computing, and communication networks. On the basis of recurring design patterns and shared failure modes in these domains, the paper argues that agentic AI trustworthiness is a single problem observed from four vantage points, and it proposes a reusable cross-domain graded-certification framework analogous to aviation and automotive safety assurance.
Significance. If the cross-domain synthesis holds, the paper makes a genuinely useful contribution: it shifts the evaluation of agentic AI from capability-centric metrics toward assurance-oriented ones, provides a structured vocabulary for locating trust mechanisms in the agentic loop, and consolidates a wide range of mechanisms and metrics into actionable tables (notably Tables III and IV). The recurring-pattern analysis across four engineering domains is a valuable organizing device, and the proposal of a common assurance infrastructure is timely. The strengths are the breadth of the surveyed literature, the explicit workflow-level framing, and the attempt to identify structurally identical failure modes behind domain-specific vocabulary. However, the central 'single problem' claim is stronger than the evidence presented: the domain selection is not systematic, the analysis is partly circular, and one of the paper's own tables contradicts a key supporting assertion. The survey is therefore a useful reference but needs either a substantially stronger evidentiary basis or a more carefully scoped claim.
major comments (4)
- [Section II-B and Section XIII-G] The central claim that agentic AI trustworthiness is 'not four separate problems but one problem observed from four vantage points' rests on the representativeness of the four domains selected in Section II-B. No inclusion/exclusion criteria, search protocol, or justification of domain coverage is given; the domains are selected because they are 'mature engineering fields with well-defined operational constraints,' which may select for domains that already share a safety-critical engineering culture. Healthcare, finance, and other critical domains with different regulatory and validation regimes are not examined, yet the abstract and Section XIII-G generalize to agentic AI as a class. This is load-bearing: if the four-domain sample is biased, the cross-domain assurance framework is not established. The authors should either add a systematic review methodology with explicit domain-selecti
- [Section X and Section XII-B] The HPC domain analysis is too thin to carry the cross-domain synthesis. Section X identifies only two failure modes (hardware failures and memory latency) and does not map HPC systems onto the five trustworthiness dimensions in the same way as the other domains. Yet Section XII-B includes HPC in all four shared failure modes, citing HPC references [122]–[125] as evidence of an 'expanding security surface' and 'partial auditability' — claims that those references do not visibly support. If the HPC analysis does not independently exhibit the same structure, the claim that the failure modes recur across all four domains is weakened. The authors need to either strengthen the HPC section with agentic-security and accountability evidence or mark HPC as a partial/emerging case rather than a full vantage point.
- [Section XII and Table V] The first paragraph of Section XII claims that 'with the exception of dedicated evaluation frameworks such as Claw-Eval [68], whose entire purpose is to probe all five dimensions at once [136], [137], no deployed or applied agentic system reviewed in this survey scores strongly on more than two or three trustworthiness dimensions simultaneously.' This is contradicted by Table V, where ToolEmu [136] and Agent Security Bench [137] are shown with checkmarks across all five dimensions. Additionally, the citation [136], [137] is attached to Claw-Eval, but Claw-Eval is reference [68], not [136]/[137]. Since this sentence is used as evidence for the shared-gap structure, the contradiction undermines a specific load-bearing point and must be corrected.
- [Section III and Section XII] The cross-domain conclusion is partly circular: the five trustworthiness dimensions defined in Section III are used as the analytical lens for every domain in Sections VIII–XI, and Section XII then concludes that the same categories recur. Finding 'the same five dimensions' in each domain is unsurprising if the framework was imposed a priori. The paper does not test whether alternative dimensions (e.g., regulatory approval, clinical validation, market manipulation) might be equally or more salient in other critical domains. The authors should acknowledge this limitation explicitly and provide a falsifiable test of the 'single problem' thesis — for example, by applying the framework to an out-of-sample domain and showing that the same four failure modes emerge rather than being assumed.
minor comments (4)
- [Section V-C1] The subsection 'Generalizing to Unforeseen Adversaries' contains an incomplete sentence: 'This is because' is followed by a line break and then a description of ImageNet-UA. This appears to be a formatting or drafting error and should be fixed.
- [Section IV-B] Typo: 'only partially observablas' should read 'only partially observable.'
- [Section VII-A1] The TUE formula includes 'adjusted tunable weights' without specifying normalization or how the component metrics (Tool Selection Accuracy, Tool Usage Efficiency, API Call Precision) are measured consistently. As a survey of quantitative metrics, the paper should at least note that these weights are application-specific and that no standardized units are claimed.
- [References] Several references are duplicated (e.g., [11] and [14] are both IEC 61508, with identical text) and some citations are to technical blogs or industry web pages (e.g., [34], [70]–[74]) that will not be stable archival sources. For a survey aiming to be a foundation for assurance, the authors should prefer peer-reviewed or arXiv-stable sources where possible.
Circularity Check
Survey synthesis with background self-citations only; no load-bearing circularity.
full rationale
The paper is a survey, not a derivation: it contains no fitted parameters, no quantitative predictions, and no equation whose output is equivalent to its input. Its central claim—"agentic AI trustworthiness is not four separate engineering problems but one problem observed from four vantage points" (Section I and Section XII-G)—is an inductive synthesis of the surveyed literature, not a result forced by construction. The five trustworthiness dimensions in Section III are explicitly adopted as an analytical lens ("we organize this view around five cross-cutting dimensions"), and the paper acknowledges the domain sample is bounded rather than exhaustive ("While these four domains do not encompass every application of agentic AI..."). Self-citations such as [69] (supporting Figure 5) and [85]-[88] (listed as background power-system analytics work) are used as supporting references, but the cross-domain unity thesis does not rest on them, and no uniqueness theorem or prior-work-derived ansatz is invoked to forbid alternatives. The skeptical concern about non-systematic domain selection is an evidentiary/selection-bias critique, not circularity, and the reviewing rules explicitly exclude "not standard consensus" and sampling concerns from the circularity classification. No circular step can be exhibited with a quote and reduction, so the appropriate finding is no significant circularity.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption The five trustworthiness dimensions (safety, robustness, transparency, accountability, privacy) are a sufficient decomposition for agentic AI trustworthiness.
- domain assumption The four engineering domains surveyed are representative of critical agentic AI applications.
- ad hoc to paper The ten-point workflow (perception through audit) is a valid abstraction of agentic systems.
read the original abstract
Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic consequences. This survey addresses a gap in current literature by treating trustworthiness, whether agentic behavior can be verified, audited, and trusted under the constraints that engineering practice actually requires, as a first-class engineering property, rather than evaluating agentic AI by task capability alone. The study adopts a trustworthiness model organized around five cross-cutting dimensions: safety and constraint satisfaction; robustness and reliability; transparency and interpretability; accountability and auditability; and privacy and security. This is mapped onto an agentic assurance workflow spanning perception through audit. Building on this foundation, agentic systems architectures, threats, concrete trust mechanisms, and quantitative metrics are surveyed for direct application in agentic systems development and evaluation. These principles are then examined across four constraint-bound engineering domains: power systems, autonomous vehicles/robotics/UAVs, high-performance computing, and communication networks, identifying recurring design patterns, shared failure modes, and domain-specific gaps. Synthesizing across those domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.
Figures
Reference graph
Works this paper leans on
-
[1]
The rise of agentic ai: A review of definitions, frameworks, architectures, applications, evaluation metrics, and challenges,
A. Bandi, B. Kongari, R. Naguru, S. Pasnoor, and S. V . Vilipala, “The rise of agentic ai: A review of definitions, frameworks, architectures, applications, evaluation metrics, and challenges,”Future Internet, vol. 17, no. 9, 2025. [Online]. Available: https://www.mdpi.com/1999-5903/17/9/404
2025
-
[2]
React: Synergizing reasoning and acting in language models,
S. Yaoet al., “React: Synergizing reasoning and acting in language models,” 2023. [Online]. Available: https://arxiv.org/abs/2210.03629
Pith/arXiv arXiv 2023
-
[3]
Reflexion: Language agents with verbal reinforcement learning,
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” 2023. [Online]. Available: https://arxiv.org/abs/2303.11366
Pith/arXiv arXiv 2023
-
[4]
Tree of thoughts: Deliberate problem solving with large language models,
S. Yaoet al., “Tree of thoughts: Deliberate problem solving with large language models,” 2023. [Online]. Available: https: //arxiv.org/abs/2305.10601
Pith/arXiv arXiv 2023
-
[6]
Ai agents are breaking bad and CISOs aren’t ready,
A. Nithrakashyap, “Ai agents are breaking bad and CISOs aren’t ready,” Fast Company, September 2025, impact Council. [Online]. Available: https://www.fastcompany.com/91404298/ai-agents-are-bre aking-bad-and-cisos-arent-ready
arXiv 2025
-
[7]
Are we sleepwalking into an agentic AI crisis?
S. Damle, “Are we sleepwalking into an agentic AI crisis?”ABA Banking Journal, December 2025, american Bankers Association. [Online]. Available: https://bankingjournal.aba.com/2025/12/are-we-s leepwalking-into-an-agentic-ai-crisis/
2025
-
[8]
Agentic ai: Autonomous in- telligence for complex goals—a comprehensive survey,
D. B. Acharya, K. Kuppan, and B. Divya, “Agentic ai: Autonomous in- telligence for complex goals—a comprehensive survey,”IEEE Access, vol. 13, pp. 18 912–18 936, 2025
2025
-
[9]
Design of a web-based platform to leverage matlab functionality for digital engineering,
O. Al-Refai, R. Sufian, K. Abo Rubeieh, N. A. M. Abu Rmaileh, O. M. F. Abu-Sharkh, and S. Krishnan, “Design of a web-based platform to leverage matlab functionality for digital engineering,” inSoft Computing and Its Engineering Applications, K. K. Patel, K. Santosh, G. Gomes de Oliveira, A. Patel, and A. Ghosh, Eds. Cham: Springer Nature Switzerland, 2026...
2026
-
[10]
Agentic ai systems: Opportunities, chal- lenges, and trustworthiness,
T. Raheem and G. Hossain, “Agentic ai systems: Opportunities, chal- lenges, and trustworthiness,” in2025 IEEE International Conference on Electro Information Technology (eIT), 2025, pp. 618–624
2025
-
[12]
C. Shinde. (2024) Navigating SOTIF (ISO 21448) and ensuring safety in autonomous driving. Automotive IQ. Accessed: 2025-02-16. [Online]. Available: https://www.automotive-iq.com/functional-safety/ articles/navigating-sotif-iso-21448-and-ensuring-safety-in-autonomou s-driving
2024
-
[13]
Constrained policy optimization,
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” inProceedings of the 34th International Conference on Machine Learning - Volume 70, ser. ICML’17. JMLR.org, 2017, p. 22–31
2017
-
[14]
Iec 61508 & functional safety,
International Electrotechnical Commission, “Iec 61508 & functional safety,” IEC, Tech. Rep., 2022, accessed: 2025-02-16. [Online]. Available: https://assets.iec.ch/public/acos/IEC%2061508%20&%20Fu nctional%20Safety-2022.pdf?2023040501
2022
-
[15]
Towards deep learning models resistant to adversarial attacks,
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in International Conference on Learning Representations, 2018. [Online]. Available: https://openreview.net/forum?id=rJzIBfZAb
2018
-
[16]
(n.d.) Failure modes effects analysis (fmea)
Montana Department of Environmental Quality. (n.d.) Failure modes effects analysis (fmea). Montana Department of Environmental Quality. PDF document. [Online]. Available: https://deq.mt.gov/files/L and/Hardrock/Documents/TintinaMines/App%20R%20Failure%20Mo des%20Effects%20Analysis/App%20R%20Failure%20Modes%20Effe cts%20Analysis.pdf
-
[17]
(2020) ASEMS toolkit: FMEA/FMECA
UK Ministry of Defence. (2020) ASEMS toolkit: FMEA/FMECA. Defence Equipment and Support (DE&S). Archived version from August 13, 2020. [Online]. Available: https://web.archive.org/web/20 200813203140/https://www.asems.mod.uk/toolkit/fmeafmeca
2020
-
[18]
Explainable ai (xai): Core ideas, techniques, and solutions,
R. Dwivediet al., “Explainable ai (xai): Core ideas, techniques, and solutions,”ACM Comput. Surv., vol. 55, no. 9, Jan. 2023. [Online]. Available: https://doi.org/10.1145/3561048
doi:10.1145/3561048 2023
-
[19]
Explainable brain tumor classification using transfer learning of deep convolutional neural networks,
O. Al-Refai and A. Alabed, “Explainable brain tumor classification using transfer learning of deep convolutional neural networks,” in2025 16th International Conference on Information and Communication Systems (ICICS), 2025, pp. 1–6
2025
-
[20]
AI Pact: Organisations’ commitments,
European Artificial Intelligence Office, “AI Pact: Organisations’ commitments,” European Commission, Tech. Rep., Sep. 2024, accessed: 2025-02-17. [Online]. Available: https://artificialintelligenc eact.eu/wp-content/uploads/2025/03/2024.09.25-AI-Pact-final.pdf
2024
-
[21]
I. D. Rajiet al., “Closing the ai accountability gap: defining an end-to-end framework for internal algorithmic auditing,” in Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, ser. FAT* ’20. New York, NY , USA: Association for Computing Machinery, 2020, p. 33–44. [Online]. Available: https://doi.org/10.1145/3351095.3372873
arXiv 2020
-
[22]
(2025, May) AI act compliance checker flowchart (v1.0)
Future of Life Institute. (2025, May) AI act compliance checker flowchart (v1.0). Future of Life Institute. Accessed: 2025-02-17. [Online]. Available: https://artificialintelligenceact.eu/wp-content/upl oads/2025/07/AI-Act-Compliance-Checker-Flowchart-v1.0 compres sed.pdf
2025
-
[23]
Deep learning with differential privacy,
M. Abadiet al., “Deep learning with differential privacy,” in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’16. New York, NY , USA: Association for Computing Machinery, 2016, p. 308–318. [Online]. Available: https://doi.org/10.1145/2976749.2978318
arXiv 2016
-
[24]
An overview of catastrophic ai risks,
D. Hendrycks, M. Mazeika, and T. Woodside, “An overview of catastrophic ai risks,” 2023. [Online]. Available: https://arxiv.org/abs/ 2306.12001
Pith/arXiv arXiv 2023
-
[25]
(n.d.) What is ISO 26262 functional safety standard? Synopsys, Inc
Synopsys. (n.d.) What is ISO 26262 functional safety standard? Synopsys, Inc. Accessed: 2025-02-17. [Online]. Available: https: //www.synopsys.com/glossary/what-is-iso-26262.html
2025
-
[26]
Tamper-resistant safeguards for open-weight llms,
R. Tamirisaet al., “Tamper-resistant safeguards for open-weight llms,”
-
[27]
Testing robustness against unforeseen adversaries,
M. Kaufmannet al., “Testing robustness against unforeseen adversaries,” 2023. [Online]. Available: https://arxiv.org/abs/1908.080 16
2023
-
[28]
Improving alignment and robustness with circuit breakers,
A. Zouet al., “Improving alignment and robustness with circuit breakers,” 2024. [Online]. Available: https://arxiv.org/abs/2406.04313
Pith/arXiv arXiv 2024
-
[29]
Representation engineering: A top-down approach to ai transparency,
——, “Representation engineering: A top-down approach to ai transparency,” 2025. [Online]. Available: https://arxiv.org/abs/2310.0 1405
2025
-
[30]
Y . Y . Chiuet al., “Morebench: Evaluating procedural and pluralistic moral reasoning in language models, more than outcomes,” 2025. [Online]. Available: https://arxiv.org/abs/2510.16380
Pith/arXiv arXiv 2025
-
[31]
Zero trust for ai systems: A reference architecture and assurance framework,
R. Campbell, “Zero trust for ai systems: A reference architecture and assurance framework,”Preprints, February 2026. [Online]. Available: https://doi.org/10.20944/preprints202602.0085.v1
arXiv 2026
-
[32]
Universal and transferable adversarial attacks on aligned language models,
A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” 2023. [Online]. Available: https://arxiv.org/abs/2307.15043
Pith/arXiv arXiv 2023
-
[33]
The wmdp benchmark: Measuring and reducing malicious use with unlearning,
N. Liet al., “The wmdp benchmark: Measuring and reducing malicious use with unlearning,” 2024. [Online]. Available: https: //arxiv.org/abs/2403.03218
Pith/arXiv arXiv 2024
-
[34]
One step away from a massive data breach: What we found inside MoltBot,
M. Siman Tov Bustan and N. Zadok, “One step away from a massive data breach: What we found inside MoltBot,” OX Security Blog, Jan. 2026, accessed: 2026-02-04. [Online]. Available: https://www.ox.security/blog/one-step-away-from-a-massive-data-bre ach-what-we-found-inside-moltbot/
2026
-
[35]
Rest meets react: Self-improvement for multi-step reasoning llm agent,
R. Aksitovet al., “Rest meets react: Self-improvement for multi-step reasoning llm agent,” 2023. [Online]. Available: https: //arxiv.org/abs/2312.10003
Pith/arXiv arXiv 2023
-
[36]
Q*: Improving multi-step reasoning for llms with deliberative planning,
C. Wanget al., “Q*: Improving multi-step reasoning for llms with deliberative planning,” 2024. [Online]. Available: https: //arxiv.org/abs/2406.14283
Pith/arXiv arXiv 2024
-
[37]
Deepseek: Revolutionizing ai with open-source reasoning models -advancing innovation, accessibility, and competition with openai and gemini 2.0,
A. Ramachandran, “Deepseek: Revolutionizing ai with open-source reasoning models -advancing innovation, accessibility, and competition with openai and gemini 2.0,” 01 2025
2025
-
[38]
(2026) Reasoning model practical guide: Enterprise comparison and deployment strategies for deepseek r1, openai o3, and gemini 3
Meta Intelligence. (2026) Reasoning model practical guide: Enterprise comparison and deployment strategies for deepseek r1, openai o3, and gemini 3. Meta Intelligence. Accessed: 2026-06-30. [Online]. Available: https://www.meta-intelligence.tech/en/insight-reasoning-m odels
2026
-
[39]
Camel: Communicative agents for
G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “Camel: Communicative agents for ”mind” exploration of large language model society,” 2023. [Online]. Available: https://arxiv.org/abs/2303.17760
Pith/arXiv arXiv 2023
-
[40]
A survey on llm- based multi-agent systems: workflow, infrastructure, and challenges,
X. Li, S. Wang, S. Zeng, Y . Wu, and Y . Yang, “A survey on llm- based multi-agent systems: workflow, infrastructure, and challenges,” Vicinagearth, vol. 1, no. 1, p. 9, Oct 2024. [Online]. Available: https://doi.org/10.1007/s44336-024-00009-2
-
[41]
Toolformer: Language models can teach themselves to use tools,
T. Schicket al., “Toolformer: Language models can teach themselves to use tools,” 2023. [Online]. Available: https://arxiv.org/abs/2302.04761
Pith/arXiv arXiv 2023
-
[42]
V oyager: An open-ended embodied agent with large language models,
G. Wanget al., “V oyager: An open-ended embodied agent with large language models,” 2023. [Online]. Available: https: //arxiv.org/abs/2305.16291
Pith/arXiv arXiv 2023
-
[43]
Toolgym: an open-world tool-using environment for scalable agent testing and data curation,
Z. Xiet al., “Toolgym: an open-world tool-using environment for scalable agent testing and data curation,” 2026. [Online]. Available: https://arxiv.org/abs/2601.06328
Pith/arXiv arXiv 2026
-
[44]
A comparative study of modern AI frame- works based on architecture, integration, and scalability,
S.-H. Cho and Y .-S. Lee, “A comparative study of modern AI frame- works based on architecture, integration, and scalability,”International Journal of Advanced Smart Convergence, vol. 14, no. 4, pp. 158–167, Dec. 2025
2025
-
[45]
V . Mavroudis, “Langchain v0.3,”Preprints, November 2024. [Online]. Available: https://doi.org/10.20944/preprints202411.0566.v1
arXiv 2024
-
[46]
(2026) Building managed agents on the Gemini Enterprise Agent Platform
Google Cloud. (2026) Building managed agents on the Gemini Enterprise Agent Platform. Google Cloud Documentation. Accessed: 2026-06-30. [Online]. Available: https://docs.cloud.google.com/gemini -enterprise-agent-platform/build/managed-agents
2026
-
[47]
(2026) Building managed agents with the Gemini API
Google AI for Developers. (2026) Building managed agents with the Gemini API. Google AI Documentation. Accessed: 2026-06-30. [Online]. Available: https://ai.google.dev/gemini-api/docs/custom-age nts
2026
-
[48]
(2026) Interacting with managed agents on the Gemini Enterprise Agent Platform
Google Cloud. (2026) Interacting with managed agents on the Gemini Enterprise Agent Platform. Google Cloud Documentation. Accessed: 2026-06-30. [Online]. Available: https://docs.cloud.google.com/gemini -enterprise-agent-platform/build/managed-agents/interact-with-agents
2026
-
[49]
G. C. Developers and P. Team. (2026, May) I/O ’26 news for agent developers on Google Cloud. Google Cloud Blog. Accessed: 2026- 06-30. [Online]. Available: https://cloud.google.com/blog/topics/devel opers-practitioners/io26-news-for-agent-developers-on-google-cloud
2026
-
[50]
Trustworthy agentic ai: A survey and taxonomy of secure coordination and hallucination mitigation in multi-agent large language model systems,
T. Vangalapat and S. Shaikh, “Trustworthy agentic ai: A survey and taxonomy of secure coordination and hallucination mitigation in multi-agent large language model systems,”International Journal of Innovative Science and Research Technology, p. 1660, 02 2026
2026
-
[51]
(2024) Inside the push to standardize communication between AI agents
HackerNoon. (2024) Inside the push to standardize communication between AI agents. HackerNoon. Accessed: 2026-04-15. [Online]. Available: https://hackernoon.com/inside-the-push-to-standardize-com munication-between-ai-agents This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this ver...
2024
-
[52]
Enabling agents to communicate entirely in latent space,
Z. Duet al., “Enabling agents to communicate entirely in latent space,” 2026. [Online]. Available: https://openreview.net/forum?id=rm YbgsehTd
2026
-
[53]
Thought communication in multiagent collaboration,
Y . Zhenget al., “Thought communication in multiagent collaboration,” inThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Online]. Available: https://openreview.net /forum?id=tq9lyV9Cml
2025
-
[54]
Byzantine-robust decentralized coordination of llm agents,
Y . Jo and C. Park, “Byzantine-robust decentralized coordination of llm agents,” 2025. [Online]. Available: https://arxiv.org/abs/2507.14928
Pith/arXiv arXiv 2025
-
[55]
Blockagents: Towards byzantine-robust llm-based multi-agent coordination via blockchain,
B. Chen, G. Li, X. Lin, Z. Wang, and J. Li, “Blockagents: Towards byzantine-robust llm-based multi-agent coordination via blockchain,” inProceedings of the ACM Turing Award Celebration Conference - China 2024, ser. ACM-TURC ’24. New York, NY , USA: Association for Computing Machinery, 2024, p. 187–192. [Online]. Available: https://doi.org/10.1145/3674399.3674445
arXiv 2024
-
[56]
Safety pretraining: Toward the next generation of safe ai,
P. Mainiet al., “Safety pretraining: Toward the next generation of safe ai,” 2025. [Online]. Available: https://arxiv.org/abs/2504.16980
arXiv 2025
-
[57]
N. Muet al., “Can llms follow simple rules?” 2024. [Online]. Available: https://arxiv.org/abs/2311.04235
Pith/arXiv arXiv 2024
-
[58]
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark,
A. Panet al., “Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark,”
-
[59]
Utility engineering: Analyzing and controlling emergent value systems in ais,
M. Mazeikaet al., “Utility engineering: Analyzing and controlling emergent value systems in ais,” 2025. [Online]. Available: https: //arxiv.org/abs/2502.08640
Pith/arXiv arXiv 2025
-
[60]
Agentic AI: a comprehensive survey of architectures, applications, and future directions,
M. Abou Ali, F. Dornaika, and J. Charafeddine, “Agentic AI: a comprehensive survey of architectures, applications, and future directions,”Artificial Intelligence Review, vol. 59, no. 1, p. 11, Nov
-
[61]
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,
A. Srivastavaet al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,” 2023. [Online]. Available: https://arxiv.org/abs/2206.04615
Pith/arXiv arXiv 2023
-
[62]
L. Phanet al., “Humanity’s last exam,” 2025. [Online]. Available: https://arxiv.org/abs/2501.14249
Pith/arXiv arXiv 2025
-
[63]
Agentharm: A benchmark for measuring harmfulness of llm agents,
M. Andriushchenkoet al., “Agentharm: A benchmark for measuring harmfulness of llm agents,” 2025. [Online]. Available: https: //arxiv.org/abs/2410.09024
Pith/arXiv arXiv 2025
-
[64]
K. Huanget al., “A novel zero-trust identity framework for agentic ai: Decentralized authentication and fine-grained access control,” 2025. [Online]. Available: https://arxiv.org/abs/2505.19301
Pith/arXiv arXiv 2025
-
[65]
Generative ai for enhanced cybersecurity: building a zero- trust architecture with agentic ai,
A. Gurram, “Generative ai for enhanced cybersecurity: building a zero- trust architecture with agentic ai,”World J. Adv. Eng. Technol. Sci, vol. 15, no. 1, pp. 2380–2396, 2025
2025
-
[66]
The mask benchmark: Disentangling honesty from accuracy in ai systems,
R. Renet al., “The mask benchmark: Disentangling honesty from accuracy in ai systems,” 2026. [Online]. Available: https: //arxiv.org/abs/2503.03750
arXiv 2026
-
[67]
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,
M. Mazeikaet al., “Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,” 2024. [Online]. Available: https://arxiv.org/abs/2402.04249
Pith/arXiv arXiv 2024
-
[68]
Claw-eval: Toward trustworthy evaluation of autonomous agents,
B. Yeet al., “Claw-eval: Toward trustworthy evaluation of autonomous agents,” 2026. [Online]. Available: https://arxiv.org/abs/2604.06132
Pith/arXiv arXiv 2026
-
[69]
Composable trust in agentic ai: Bridging architectural capability and system-level assurance,
O. Al-Refai, I. Shahbaz, and E. Hammad, “Composable trust in agentic ai: Bridging architectural capability and system-level assurance,” in 2026 56th Annual IEEE International Conference on Dependable Systems and Networks Workshops (DSN-W), 2026, pp. 17–20
2026
-
[70]
(2025) RagaAI AAEF (Agentic Application Evaluation Framework)
RagaAI. (2025) RagaAI AAEF (Agentic Application Evaluation Framework). RagaAI Catalyst. Accessed: 2026-02-18. [Online]. Available: https://docs.raga.ai/ragaai-aaef-agentic-application-evaluat ion-framework
2025
-
[71]
Whitepaper: Agentic application evaluation framework (AAEF),
——, “Whitepaper: Agentic application evaluation framework (AAEF),” RagaAI, Inc., Tech. Rep., Jun. 2024, accessed: 2025- 02-18. [Online]. Available: https://raga.ai/resources/patentsandpublicat ions/whitepaper-agentic-application-evaluation-framework
2024
-
[72]
Starkloff, S
A.-G. Starkloff, S. Kokaina, and S. Rahimi. (2026, Jan.) Evaluations for the agentic world. QuantumBlack, AI by McKinsey. Accessed: 2026-02-18. [Online]. Available: https://medium.com/quantumblack/ evaluations-for-the-agentic-world-c3c150f0dd5a
2026
-
[73]
Morales Aguilera
F. Morales Aguilera. (2025, Jul.) Building real-world agentic AI systems: A practical guide. AI Simplified in Plain English. Accessed: 2026-02-18. [Online]. Available: https://medium.com/ai-simplified-i n-plain-english/building-real-world-agentic-ai-systems-a-practical-g uide-9748d572b58b
2025
-
[74]
(2025, May) MDR vs
Wiz Experts Team. (2025, May) MDR vs. SOC: What’s the difference? Wiz, Inc. Accessed: 2026-02-18. [Online]. Available: https://www.wiz.io/academy/detection-and-response/mdr-vs-soc
2025
-
[75]
St-webagentbench: A benchmark for evaluating safety and trustworthiness in web agents,
I. Levy, B. Wiesel, S. Marreed, A. Oved, A. Yaeli, and S. Shlomov, “St-webagentbench: A benchmark for evaluating safety and trustworthiness in web agents,” 2025. [Online]. Available: https://arxiv.org/abs/2410.06703
Pith/arXiv arXiv 2025
-
[76]
successes
L. R. Sammeta. (2025, Dec.) The AI agent report card you’ve been ignoring: Why 30% of your agent’s “successes” are actually failures. Accessed: 2026-02-23. [Online]. Available: https://laxmikumars.medi um.com/the-ai-agent-report-card-youve-been-ignoring-why-30-of-y our-agent-s-successes-are-actually-498fbebf44f9
2025
-
[77]
Evaluating agentic ai systems: A balanced framework for performance, robustness, safety and beyond,
M. Shukla, “Evaluating agentic ai systems: A balanced framework for performance, robustness, safety and beyond,”Preprints, August 2025. [Online]. Available: https://doi.org/10.20944/preprints202508.1847.v1
arXiv 2025
-
[78]
Technical report: Evaluating goal drift in language model agents,
R. Arike, E. Donoway, H. Bartsch, and M. Hobbhahn, “Technical report: Evaluating goal drift in language model agents,” 2025. [Online]. Available: https://arxiv.org/abs/2505.02709
Pith/arXiv arXiv 2025
-
[79]
M. A. Ferrag, A. Lakas, and M. Debbah, “α 3-bench: A unified benchmark of safety, robustness, and efficiency for llm-based uav agents over 6g networks,” 2026. [Online]. Available: https: //arxiv.org/abs/2601.03281
arXiv 2026
-
[80]
Autoadvexbench: Benchmarking autonomous exploitation of adversarial example defenses,
N. Carlini, J. Rando, E. Debenedetti, M. Nasr, and F. Tram`er, “Autoadvexbench: Benchmarking autonomous exploitation of adversarial example defenses,” 2025. [Online]. Available: https://arxiv.org/abs/2503.01811
Pith/arXiv arXiv 2025
-
[81]
L. Rijo. (2026, Feb.) UC berkeley unveils framework as AI agents threaten to outrun oversight. PPC Land. Accessed: 2026-02-24. [Online]. Available: https://ppc.land/uc-berkeley-unveils-framework-a s-ai-agents-threaten-to-outrun-oversight/
2026
-
[82]
N. Yadav. (2025, Nov.) 10 essential steps for evaluating the reliability of AI agents. Maxim AI. Accessed: 2026-02-24. [Online]. Available: https://www.getmaxim.ai/articles/10-essential-steps-for-evaluating-the -reliability-of-ai-agents/
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.