Pith. sign in

REVIEW 4 major objections 4 minor 144 references

Agentic AI trustworthiness is one problem observed from four vantage points, not four separate ones.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 15:02 UTC pith:WYAV2NTH

load-bearing objection A structured, useful survey that makes a plausible cross-domain case; the 'one problem' claim is a well-argued research program, not an established result. the 4 major comments →

arxiv 2607.18548 v1 pith:WYAV2NTH submitted 2026-07-20 cs.AI cs.MAcs.SYeess.SY

Engineering Trustworthy Agentic AI for Critical Systems

classification cs.AI cs.MAcs.SYeess.SY
keywords agentic AItrustworthinesscritical infrastructuresafety assuranceAI safety evaluationpower systemsautonomous vehiclescommunication networks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that trustworthy agentic AI in critical infrastructure should be treated as an engineering property, not measured by task capability alone. It claims that the same four failure modes appear in power systems, autonomous vehicles, robotics and UAVs, high-performance computing, and communication networks: unsafe translation from reasoning to action, narrow robustness evidence, fragmentary auditability, and an expanding security surface. Because these recur across domains, the paper concludes that agentic AI trustworthiness is a single problem and that a reusable cross-domain assurance framework, analogous to the graded certification regimes of aviation and automotive safety, is feasible. What differs across domains, the paper argues, is only where interpretation becomes consequence, not the underlying trust challenge.

Core claim

Observed across four constraint-bound domains, agentic AI systems fail in the same way: gaps at the reasoning-to-action boundary, robustness evidence that stops at curated benchmarks, audit trails that cover fragments rather than the full workflow, and security surfaces that grow with every tool or coordination channel. The paper reads this as evidence that trustworthiness is a single problem with domain-specific vocabulary, and it outlines a path toward a cross-domain graded assurance framework. The central claim is that the same pipeline, perception through audit, appears under different physical constraints, and that the location of the risky boundary, not the model architecture, determin

What carries the argument

The analytic scaffold is a ten-stage agentic workflow (perception, reasoning, planning, tool use, communication, memory update, validation, action, monitoring, audit) with five cross-cutting trustworthiness dimensions (safety and constraint satisfaction, robustness and reliability, transparency and interpretability, accountability and auditability, privacy and security). This workflow-to-dimension mapping lets the authors ask where trust mechanisms sit and where they are absent, and it produces the cross-domain comparison that yields the 'one problem' claim.

Load-bearing premise

The claim rests on the assumption that the selected papers fairly represent each domain; because the review is curated rather than systematic, the recurring failure modes could be an artifact of the sample.

What would settle it

A systematic review using explicit inclusion criteria that locates a critical-infrastructure domain where the four failure modes do not co-occur, or where a domain-specific failure mode carries most of the risk, would undermine the 'single problem' thesis.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A standard trustworthiness scorecard could be reported alongside task-success metrics in every agentic AI evaluation.
  • Autonomy boundaries could be defined before deployment as graded levels, so that advisory, supervised, and fully autonomous authority are comparable across domains.
  • Adversarial and rare-event test suites could be shared across power, robotics, HPC, and networks instead of being recreated separately in each field.
  • Provenance standards for tools, memory, and multi-agent coordination could be built proactively, drawing on existing automotive and aviation traceability practice, before an incident forces them.
  • A reusable framework would let a validation result for a grid-dispatch agent and a result for a UAV planner share a common unit of evidence.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the paper's thesis: apply the same ten-stage, five-dimension analysis to a domain it did not review, such as healthcare or finance; if the same four failure modes appear, the 'single problem' claim is strengthened.
  • The framework implies that capability benchmarks, such as reasoning or tool-use scores, should be subordinate to workflow-level assurance evidence, which would shift how agentic AI progress is measured.
  • Because the recurring failure modes all live at the interpretation-to-action boundary, the paper suggests that future architectures may converge on a 'demoted orchestrator' pattern where the LLM plans but a verified component decides and executes.
  • The graded-certification analogy carries a testable implication: agents in different domains should be comparable on the same axes, such as consequence severity, response time, and audit maturity, which the paper sketches but does not yet formalize.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper is a survey that proposes a trustworthiness model for agentic AI in critical systems, organized around five dimensions (safety/constraint satisfaction, robustness/reliability, transparency/interpretability, accountability/auditability, privacy/security) and a ten-point agentic workflow from perception to audit. It reviews agent architectures, trust mechanisms, evaluation approaches, and quantitative metrics, then applies the model to four domains: power systems, autonomous vehicles/robotics/UAVs, high-performance computing, and communication networks. On the basis of recurring design patterns and shared failure modes in these domains, the paper argues that agentic AI trustworthiness is a single problem observed from four vantage points, and it proposes a reusable cross-domain graded-certification framework analogous to aviation and automotive safety assurance.

Significance. If the cross-domain synthesis holds, the paper makes a genuinely useful contribution: it shifts the evaluation of agentic AI from capability-centric metrics toward assurance-oriented ones, provides a structured vocabulary for locating trust mechanisms in the agentic loop, and consolidates a wide range of mechanisms and metrics into actionable tables (notably Tables III and IV). The recurring-pattern analysis across four engineering domains is a valuable organizing device, and the proposal of a common assurance infrastructure is timely. The strengths are the breadth of the surveyed literature, the explicit workflow-level framing, and the attempt to identify structurally identical failure modes behind domain-specific vocabulary. However, the central 'single problem' claim is stronger than the evidence presented: the domain selection is not systematic, the analysis is partly circular, and one of the paper's own tables contradicts a key supporting assertion. The survey is therefore a useful reference but needs either a substantially stronger evidentiary basis or a more carefully scoped claim.

major comments (4)
  1. [Section II-B and Section XIII-G] The central claim that agentic AI trustworthiness is 'not four separate problems but one problem observed from four vantage points' rests on the representativeness of the four domains selected in Section II-B. No inclusion/exclusion criteria, search protocol, or justification of domain coverage is given; the domains are selected because they are 'mature engineering fields with well-defined operational constraints,' which may select for domains that already share a safety-critical engineering culture. Healthcare, finance, and other critical domains with different regulatory and validation regimes are not examined, yet the abstract and Section XIII-G generalize to agentic AI as a class. This is load-bearing: if the four-domain sample is biased, the cross-domain assurance framework is not established. The authors should either add a systematic review methodology with explicit domain-selecti
  2. [Section X and Section XII-B] The HPC domain analysis is too thin to carry the cross-domain synthesis. Section X identifies only two failure modes (hardware failures and memory latency) and does not map HPC systems onto the five trustworthiness dimensions in the same way as the other domains. Yet Section XII-B includes HPC in all four shared failure modes, citing HPC references [122]–[125] as evidence of an 'expanding security surface' and 'partial auditability' — claims that those references do not visibly support. If the HPC analysis does not independently exhibit the same structure, the claim that the failure modes recur across all four domains is weakened. The authors need to either strengthen the HPC section with agentic-security and accountability evidence or mark HPC as a partial/emerging case rather than a full vantage point.
  3. [Section XII and Table V] The first paragraph of Section XII claims that 'with the exception of dedicated evaluation frameworks such as Claw-Eval [68], whose entire purpose is to probe all five dimensions at once [136], [137], no deployed or applied agentic system reviewed in this survey scores strongly on more than two or three trustworthiness dimensions simultaneously.' This is contradicted by Table V, where ToolEmu [136] and Agent Security Bench [137] are shown with checkmarks across all five dimensions. Additionally, the citation [136], [137] is attached to Claw-Eval, but Claw-Eval is reference [68], not [136]/[137]. Since this sentence is used as evidence for the shared-gap structure, the contradiction undermines a specific load-bearing point and must be corrected.
  4. [Section III and Section XII] The cross-domain conclusion is partly circular: the five trustworthiness dimensions defined in Section III are used as the analytical lens for every domain in Sections VIII–XI, and Section XII then concludes that the same categories recur. Finding 'the same five dimensions' in each domain is unsurprising if the framework was imposed a priori. The paper does not test whether alternative dimensions (e.g., regulatory approval, clinical validation, market manipulation) might be equally or more salient in other critical domains. The authors should acknowledge this limitation explicitly and provide a falsifiable test of the 'single problem' thesis — for example, by applying the framework to an out-of-sample domain and showing that the same four failure modes emerge rather than being assumed.
minor comments (4)
  1. [Section V-C1] The subsection 'Generalizing to Unforeseen Adversaries' contains an incomplete sentence: 'This is because' is followed by a line break and then a description of ImageNet-UA. This appears to be a formatting or drafting error and should be fixed.
  2. [Section IV-B] Typo: 'only partially observablas' should read 'only partially observable.'
  3. [Section VII-A1] The TUE formula includes 'adjusted tunable weights' without specifying normalization or how the component metrics (Tool Selection Accuracy, Tool Usage Efficiency, API Call Precision) are measured consistently. As a survey of quantitative metrics, the paper should at least note that these weights are application-specific and that no standardized units are claimed.
  4. [References] Several references are duplicated (e.g., [11] and [14] are both IEC 61508, with identical text) and some citations are to technical blogs or industry web pages (e.g., [34], [70]–[74]) that will not be stable archival sources. For a survey aiming to be a foundation for assurance, the authors should prefer peer-reviewed or arXiv-stable sources where possible.

Circularity Check

0 steps flagged

Survey synthesis with background self-citations only; no load-bearing circularity.

full rationale

The paper is a survey, not a derivation: it contains no fitted parameters, no quantitative predictions, and no equation whose output is equivalent to its input. Its central claim—"agentic AI trustworthiness is not four separate engineering problems but one problem observed from four vantage points" (Section I and Section XII-G)—is an inductive synthesis of the surveyed literature, not a result forced by construction. The five trustworthiness dimensions in Section III are explicitly adopted as an analytical lens ("we organize this view around five cross-cutting dimensions"), and the paper acknowledges the domain sample is bounded rather than exhaustive ("While these four domains do not encompass every application of agentic AI..."). Self-citations such as [69] (supporting Figure 5) and [85]-[88] (listed as background power-system analytics work) are used as supporting references, but the cross-domain unity thesis does not rest on them, and no uniqueness theorem or prior-work-derived ansatz is invoked to forbid alternatives. The skeptical concern about non-systematic domain selection is an evidentiary/selection-bias critique, not circularity, and the reviewing rules explicitly exclude "not standard consensus" and sampling concerns from the circularity classification. No circular step can be exhibited with a quote and reduction, so the appropriate finding is no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

The paper introduces no free parameters or invented entities. Its framework rests on the three stated assumptions about the sufficiency of the five dimensions, the representativeness of the four domains, and the adequacy of the ten-stage workflow. These are reasonable qualitative assumptions for a survey, but they are not empirically validated.

axioms (3)
  • domain assumption The five trustworthiness dimensions (safety, robustness, transparency, accountability, privacy) are a sufficient decomposition for agentic AI trustworthiness.
    Stated in Section III as the paper's organizing model; the paper does not justify exhaustiveness or independence of the dimensions.
  • domain assumption The four engineering domains surveyed are representative of critical agentic AI applications.
    Stated in Section II-B ('They are selected because they are mature engineering fields...'); other critical domains (healthcare, finance) are excluded without detailed justification.
  • ad hoc to paper The ten-point workflow (perception through audit) is a valid abstraction of agentic systems.
    Introduced in Section III as the paper's own analytical structure ('This decomposition is not meant to prescribe...'); it is not derived from empirical data or prior frameworks.

pith-pipeline@v1.3.0-alltime-deepseek · 45232 in / 7158 out tokens · 80112 ms · 2026-08-01T15:02:51.538610+00:00 · methodology

0 comments
read the original abstract

Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic consequences. This survey addresses a gap in current literature by treating trustworthiness, whether agentic behavior can be verified, audited, and trusted under the constraints that engineering practice actually requires, as a first-class engineering property, rather than evaluating agentic AI by task capability alone. The study adopts a trustworthiness model organized around five cross-cutting dimensions: safety and constraint satisfaction; robustness and reliability; transparency and interpretability; accountability and auditability; and privacy and security. This is mapped onto an agentic assurance workflow spanning perception through audit. Building on this foundation, agentic systems architectures, threats, concrete trust mechanisms, and quantitative metrics are surveyed for direct application in agentic systems development and evaluation. These principles are then examined across four constraint-bound engineering domains: power systems, autonomous vehicles/robotics/UAVs, high-performance computing, and communication networks, identifying recurring design patterns, shared failure modes, and domain-specific gaps. Synthesizing across those domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.

Figures

Figures reproduced from arXiv: 2607.18548 by Adam Ali Husseinat, Eman Hammad, Ibrahim Shahbaz, Jaewon Kim, Michael Mandulak, Omar Al-Refai.

Figure 1
Figure 1. Figure 1: Generic agentic AI observe–reason–act loop. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Generic multi-agent architecture showing interconnected loops, shared memory, and collective interaction in a common environment. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Holistic framework for trustworthy agentic AI in critical engineering systems. The figure organizes the agentic workflow around perception, reasoning, [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Representative agentic AI architecture classes and their trustworthiness implications. Dashed arrows denote reflection, communication, or information [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Evolution of Benchmarks from static AI safety evaluation to agentic system-level evaluation. [PITH_FULL_IMAGE:figures/full_fig_p014_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Power-systems adaptation of the multi-agent agentic AI architecture. [PITH_FULL_IMAGE:figures/full_fig_p019_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Robotics adaptation of the multi-agent agentic AI architecture. [PITH_FULL_IMAGE:figures/full_fig_p021_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: High-performance computing adaptation of the multi-agent agentic [PITH_FULL_IMAGE:figures/full_fig_p023_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Communication networks adaptation of the multi-agent agentic AI [PITH_FULL_IMAGE:figures/full_fig_p024_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Human–AI autonomy spectrum, illustrating more collaborative [PITH_FULL_IMAGE:figures/full_fig_p028_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

144 extracted references · 1 canonical work pages

  1. [1]

    The rise of agentic ai: A review of definitions, frameworks, architectures, applications, evaluation metrics, and challenges,

    A. Bandi, B. Kongari, R. Naguru, S. Pasnoor, and S. V . Vilipala, “The rise of agentic ai: A review of definitions, frameworks, architectures, applications, evaluation metrics, and challenges,”Future Internet, vol. 17, no. 9, 2025. [Online]. Available: https://www.mdpi.com/1999-5903/17/9/404

  2. [2]

    React: Synergizing reasoning and acting in language models,

    S. Yaoet al., “React: Synergizing reasoning and acting in language models,” 2023. [Online]. Available: https://arxiv.org/abs/2210.03629

  3. [3]

    Reflexion: Language agents with verbal reinforcement learning,

    N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” 2023. [Online]. Available: https://arxiv.org/abs/2303.11366

  4. [4]

    Tree of thoughts: Deliberate problem solving with large language models,

    S. Yaoet al., “Tree of thoughts: Deliberate problem solving with large language models,” 2023. [Online]. Available: https: //arxiv.org/abs/2305.10601

  5. [6]

    Ai agents are breaking bad and CISOs aren’t ready,

    A. Nithrakashyap, “Ai agents are breaking bad and CISOs aren’t ready,” Fast Company, September 2025, impact Council. [Online]. Available: https://www.fastcompany.com/91404298/ai-agents-are-bre aking-bad-and-cisos-arent-ready

  6. [7]

    Are we sleepwalking into an agentic AI crisis?

    S. Damle, “Are we sleepwalking into an agentic AI crisis?”ABA Banking Journal, December 2025, american Bankers Association. [Online]. Available: https://bankingjournal.aba.com/2025/12/are-we-s leepwalking-into-an-agentic-ai-crisis/

  7. [8]

    Agentic ai: Autonomous in- telligence for complex goals—a comprehensive survey,

    D. B. Acharya, K. Kuppan, and B. Divya, “Agentic ai: Autonomous in- telligence for complex goals—a comprehensive survey,”IEEE Access, vol. 13, pp. 18 912–18 936, 2025

  8. [9]

    Design of a web-based platform to leverage matlab functionality for digital engineering,

    O. Al-Refai, R. Sufian, K. Abo Rubeieh, N. A. M. Abu Rmaileh, O. M. F. Abu-Sharkh, and S. Krishnan, “Design of a web-based platform to leverage matlab functionality for digital engineering,” inSoft Computing and Its Engineering Applications, K. K. Patel, K. Santosh, G. Gomes de Oliveira, A. Patel, and A. Ghosh, Eds. Cham: Springer Nature Switzerland, 2026...

  9. [10]

    Agentic ai systems: Opportunities, chal- lenges, and trustworthiness,

    T. Raheem and G. Hossain, “Agentic ai systems: Opportunities, chal- lenges, and trustworthiness,” in2025 IEEE International Conference on Electro Information Technology (eIT), 2025, pp. 618–624

  10. [12]

    C. Shinde. (2024) Navigating SOTIF (ISO 21448) and ensuring safety in autonomous driving. Automotive IQ. Accessed: 2025-02-16. [Online]. Available: https://www.automotive-iq.com/functional-safety/ articles/navigating-sotif-iso-21448-and-ensuring-safety-in-autonomou s-driving

  11. [13]

    Constrained policy optimization,

    J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” inProceedings of the 34th International Conference on Machine Learning - Volume 70, ser. ICML’17. JMLR.org, 2017, p. 22–31

  12. [14]

    Iec 61508 & functional safety,

    International Electrotechnical Commission, “Iec 61508 & functional safety,” IEC, Tech. Rep., 2022, accessed: 2025-02-16. [Online]. Available: https://assets.iec.ch/public/acos/IEC%2061508%20&%20Fu nctional%20Safety-2022.pdf?2023040501

  13. [15]

    Towards deep learning models resistant to adversarial attacks,

    A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in International Conference on Learning Representations, 2018. [Online]. Available: https://openreview.net/forum?id=rJzIBfZAb

  14. [16]

    (n.d.) Failure modes effects analysis (fmea)

    Montana Department of Environmental Quality. (n.d.) Failure modes effects analysis (fmea). Montana Department of Environmental Quality. PDF document. [Online]. Available: https://deq.mt.gov/files/L and/Hardrock/Documents/TintinaMines/App%20R%20Failure%20Mo des%20Effects%20Analysis/App%20R%20Failure%20Modes%20Effe cts%20Analysis.pdf

  15. [17]

    (2020) ASEMS toolkit: FMEA/FMECA

    UK Ministry of Defence. (2020) ASEMS toolkit: FMEA/FMECA. Defence Equipment and Support (DE&S). Archived version from August 13, 2020. [Online]. Available: https://web.archive.org/web/20 200813203140/https://www.asems.mod.uk/toolkit/fmeafmeca

  16. [18]

    Explainable ai (xai): Core ideas, techniques, and solutions,

    R. Dwivediet al., “Explainable ai (xai): Core ideas, techniques, and solutions,”ACM Comput. Surv., vol. 55, no. 9, Jan. 2023. [Online]. Available: https://doi.org/10.1145/3561048

  17. [19]

    Explainable brain tumor classification using transfer learning of deep convolutional neural networks,

    O. Al-Refai and A. Alabed, “Explainable brain tumor classification using transfer learning of deep convolutional neural networks,” in2025 16th International Conference on Information and Communication Systems (ICICS), 2025, pp. 1–6

  18. [20]

    AI Pact: Organisations’ commitments,

    European Artificial Intelligence Office, “AI Pact: Organisations’ commitments,” European Commission, Tech. Rep., Sep. 2024, accessed: 2025-02-17. [Online]. Available: https://artificialintelligenc eact.eu/wp-content/uploads/2025/03/2024.09.25-AI-Pact-final.pdf

  19. [21]

    Closing the ai accountability gap: defining an end-to-end framework for internal algorithmic auditing,

    I. D. Rajiet al., “Closing the ai accountability gap: defining an end-to-end framework for internal algorithmic auditing,” in Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, ser. FAT* ’20. New York, NY , USA: Association for Computing Machinery, 2020, p. 33–44. [Online]. Available: https://doi.org/10.1145/3351095.3372873

  20. [22]

    (2025, May) AI act compliance checker flowchart (v1.0)

    Future of Life Institute. (2025, May) AI act compliance checker flowchart (v1.0). Future of Life Institute. Accessed: 2025-02-17. [Online]. Available: https://artificialintelligenceact.eu/wp-content/upl oads/2025/07/AI-Act-Compliance-Checker-Flowchart-v1.0 compres sed.pdf

  21. [23]

    Deep learning with differential privacy,

    M. Abadiet al., “Deep learning with differential privacy,” in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’16. New York, NY , USA: Association for Computing Machinery, 2016, p. 308–318. [Online]. Available: https://doi.org/10.1145/2976749.2978318

  22. [24]

    An overview of catastrophic ai risks,

    D. Hendrycks, M. Mazeika, and T. Woodside, “An overview of catastrophic ai risks,” 2023. [Online]. Available: https://arxiv.org/abs/ 2306.12001

  23. [25]

    (n.d.) What is ISO 26262 functional safety standard? Synopsys, Inc

    Synopsys. (n.d.) What is ISO 26262 functional safety standard? Synopsys, Inc. Accessed: 2025-02-17. [Online]. Available: https: //www.synopsys.com/glossary/what-is-iso-26262.html

  24. [26]

    Tamper-resistant safeguards for open-weight llms,

    R. Tamirisaet al., “Tamper-resistant safeguards for open-weight llms,”

  25. [27]

    Testing robustness against unforeseen adversaries,

    M. Kaufmannet al., “Testing robustness against unforeseen adversaries,” 2023. [Online]. Available: https://arxiv.org/abs/1908.080 16

  26. [28]

    Improving alignment and robustness with circuit breakers,

    A. Zouet al., “Improving alignment and robustness with circuit breakers,” 2024. [Online]. Available: https://arxiv.org/abs/2406.04313

  27. [29]

    Representation engineering: A top-down approach to ai transparency,

    ——, “Representation engineering: A top-down approach to ai transparency,” 2025. [Online]. Available: https://arxiv.org/abs/2310.0 1405

  28. [30]

    Morebench: Evaluating procedural and pluralistic moral reasoning in language models, more than outcomes,

    Y . Y . Chiuet al., “Morebench: Evaluating procedural and pluralistic moral reasoning in language models, more than outcomes,” 2025. [Online]. Available: https://arxiv.org/abs/2510.16380

  29. [31]

    Zero trust for ai systems: A reference architecture and assurance framework,

    R. Campbell, “Zero trust for ai systems: A reference architecture and assurance framework,”Preprints, February 2026. [Online]. Available: https://doi.org/10.20944/preprints202602.0085.v1

  30. [32]

    Universal and transferable adversarial attacks on aligned language models,

    A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” 2023. [Online]. Available: https://arxiv.org/abs/2307.15043

  31. [33]

    The wmdp benchmark: Measuring and reducing malicious use with unlearning,

    N. Liet al., “The wmdp benchmark: Measuring and reducing malicious use with unlearning,” 2024. [Online]. Available: https: //arxiv.org/abs/2403.03218

  32. [34]

    One step away from a massive data breach: What we found inside MoltBot,

    M. Siman Tov Bustan and N. Zadok, “One step away from a massive data breach: What we found inside MoltBot,” OX Security Blog, Jan. 2026, accessed: 2026-02-04. [Online]. Available: https://www.ox.security/blog/one-step-away-from-a-massive-data-bre ach-what-we-found-inside-moltbot/

  33. [35]

    Rest meets react: Self-improvement for multi-step reasoning llm agent,

    R. Aksitovet al., “Rest meets react: Self-improvement for multi-step reasoning llm agent,” 2023. [Online]. Available: https: //arxiv.org/abs/2312.10003

  34. [36]

    Q*: Improving multi-step reasoning for llms with deliberative planning,

    C. Wanget al., “Q*: Improving multi-step reasoning for llms with deliberative planning,” 2024. [Online]. Available: https: //arxiv.org/abs/2406.14283

  35. [37]

    Deepseek: Revolutionizing ai with open-source reasoning models -advancing innovation, accessibility, and competition with openai and gemini 2.0,

    A. Ramachandran, “Deepseek: Revolutionizing ai with open-source reasoning models -advancing innovation, accessibility, and competition with openai and gemini 2.0,” 01 2025

  36. [38]

    (2026) Reasoning model practical guide: Enterprise comparison and deployment strategies for deepseek r1, openai o3, and gemini 3

    Meta Intelligence. (2026) Reasoning model practical guide: Enterprise comparison and deployment strategies for deepseek r1, openai o3, and gemini 3. Meta Intelligence. Accessed: 2026-06-30. [Online]. Available: https://www.meta-intelligence.tech/en/insight-reasoning-m odels

  37. [39]

    Camel: Communicative agents for

    G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “Camel: Communicative agents for ”mind” exploration of large language model society,” 2023. [Online]. Available: https://arxiv.org/abs/2303.17760

  38. [40]

    A survey on llm- based multi-agent systems: workflow, infrastructure, and challenges,

    X. Li, S. Wang, S. Zeng, Y . Wu, and Y . Yang, “A survey on llm- based multi-agent systems: workflow, infrastructure, and challenges,” Vicinagearth, vol. 1, no. 1, p. 9, Oct 2024. [Online]. Available: https://doi.org/10.1007/s44336-024-00009-2

  39. [41]

    Toolformer: Language models can teach themselves to use tools,

    T. Schicket al., “Toolformer: Language models can teach themselves to use tools,” 2023. [Online]. Available: https://arxiv.org/abs/2302.04761

  40. [42]

    V oyager: An open-ended embodied agent with large language models,

    G. Wanget al., “V oyager: An open-ended embodied agent with large language models,” 2023. [Online]. Available: https: //arxiv.org/abs/2305.16291

  41. [43]

    Toolgym: an open-world tool-using environment for scalable agent testing and data curation,

    Z. Xiet al., “Toolgym: an open-world tool-using environment for scalable agent testing and data curation,” 2026. [Online]. Available: https://arxiv.org/abs/2601.06328

  42. [44]

    A comparative study of modern AI frame- works based on architecture, integration, and scalability,

    S.-H. Cho and Y .-S. Lee, “A comparative study of modern AI frame- works based on architecture, integration, and scalability,”International Journal of Advanced Smart Convergence, vol. 14, no. 4, pp. 158–167, Dec. 2025

  43. [45]

    Langchain v0.3,

    V . Mavroudis, “Langchain v0.3,”Preprints, November 2024. [Online]. Available: https://doi.org/10.20944/preprints202411.0566.v1

  44. [46]

    (2026) Building managed agents on the Gemini Enterprise Agent Platform

    Google Cloud. (2026) Building managed agents on the Gemini Enterprise Agent Platform. Google Cloud Documentation. Accessed: 2026-06-30. [Online]. Available: https://docs.cloud.google.com/gemini -enterprise-agent-platform/build/managed-agents

  45. [47]

    (2026) Building managed agents with the Gemini API

    Google AI for Developers. (2026) Building managed agents with the Gemini API. Google AI Documentation. Accessed: 2026-06-30. [Online]. Available: https://ai.google.dev/gemini-api/docs/custom-age nts

  46. [48]

    (2026) Interacting with managed agents on the Gemini Enterprise Agent Platform

    Google Cloud. (2026) Interacting with managed agents on the Gemini Enterprise Agent Platform. Google Cloud Documentation. Accessed: 2026-06-30. [Online]. Available: https://docs.cloud.google.com/gemini -enterprise-agent-platform/build/managed-agents/interact-with-agents

  47. [49]

    G. C. Developers and P. Team. (2026, May) I/O ’26 news for agent developers on Google Cloud. Google Cloud Blog. Accessed: 2026- 06-30. [Online]. Available: https://cloud.google.com/blog/topics/devel opers-practitioners/io26-news-for-agent-developers-on-google-cloud

  48. [50]

    Trustworthy agentic ai: A survey and taxonomy of secure coordination and hallucination mitigation in multi-agent large language model systems,

    T. Vangalapat and S. Shaikh, “Trustworthy agentic ai: A survey and taxonomy of secure coordination and hallucination mitigation in multi-agent large language model systems,”International Journal of Innovative Science and Research Technology, p. 1660, 02 2026

  49. [51]

    (2024) Inside the push to standardize communication between AI agents

    HackerNoon. (2024) Inside the push to standardize communication between AI agents. HackerNoon. Accessed: 2026-04-15. [Online]. Available: https://hackernoon.com/inside-the-push-to-standardize-com munication-between-ai-agents This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this ver...

  50. [52]

    Enabling agents to communicate entirely in latent space,

    Z. Duet al., “Enabling agents to communicate entirely in latent space,” 2026. [Online]. Available: https://openreview.net/forum?id=rm YbgsehTd

  51. [53]

    Thought communication in multiagent collaboration,

    Y . Zhenget al., “Thought communication in multiagent collaboration,” inThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Online]. Available: https://openreview.net /forum?id=tq9lyV9Cml

  52. [54]

    Byzantine-robust decentralized coordination of llm agents,

    Y . Jo and C. Park, “Byzantine-robust decentralized coordination of llm agents,” 2025. [Online]. Available: https://arxiv.org/abs/2507.14928

  53. [55]

    Blockagents: Towards byzantine-robust llm-based multi-agent coordination via blockchain,

    B. Chen, G. Li, X. Lin, Z. Wang, and J. Li, “Blockagents: Towards byzantine-robust llm-based multi-agent coordination via blockchain,” inProceedings of the ACM Turing Award Celebration Conference - China 2024, ser. ACM-TURC ’24. New York, NY , USA: Association for Computing Machinery, 2024, p. 187–192. [Online]. Available: https://doi.org/10.1145/3674399.3674445

  54. [56]

    Safety pretraining: Toward the next generation of safe ai,

    P. Mainiet al., “Safety pretraining: Toward the next generation of safe ai,” 2025. [Online]. Available: https://arxiv.org/abs/2504.16980

  55. [57]

    Can llms follow simple rules?

    N. Muet al., “Can llms follow simple rules?” 2024. [Online]. Available: https://arxiv.org/abs/2311.04235

  56. [58]

    Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark,

    A. Panet al., “Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark,”

  57. [59]

    Utility engineering: Analyzing and controlling emergent value systems in ais,

    M. Mazeikaet al., “Utility engineering: Analyzing and controlling emergent value systems in ais,” 2025. [Online]. Available: https: //arxiv.org/abs/2502.08640

  58. [60]

    Agentic AI: a comprehensive survey of architectures, applications, and future directions,

    M. Abou Ali, F. Dornaika, and J. Charafeddine, “Agentic AI: a comprehensive survey of architectures, applications, and future directions,”Artificial Intelligence Review, vol. 59, no. 1, p. 11, Nov

  59. [61]

    Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,

    A. Srivastavaet al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,” 2023. [Online]. Available: https://arxiv.org/abs/2206.04615

  60. [62]

    Humanity’s last exam,

    L. Phanet al., “Humanity’s last exam,” 2025. [Online]. Available: https://arxiv.org/abs/2501.14249

  61. [63]

    Agentharm: A benchmark for measuring harmfulness of llm agents,

    M. Andriushchenkoet al., “Agentharm: A benchmark for measuring harmfulness of llm agents,” 2025. [Online]. Available: https: //arxiv.org/abs/2410.09024

  62. [64]

    A novel zero-trust identity framework for agentic ai: Decentralized authentication and fine-grained access control,

    K. Huanget al., “A novel zero-trust identity framework for agentic ai: Decentralized authentication and fine-grained access control,” 2025. [Online]. Available: https://arxiv.org/abs/2505.19301

  63. [65]

    Generative ai for enhanced cybersecurity: building a zero- trust architecture with agentic ai,

    A. Gurram, “Generative ai for enhanced cybersecurity: building a zero- trust architecture with agentic ai,”World J. Adv. Eng. Technol. Sci, vol. 15, no. 1, pp. 2380–2396, 2025

  64. [66]

    The mask benchmark: Disentangling honesty from accuracy in ai systems,

    R. Renet al., “The mask benchmark: Disentangling honesty from accuracy in ai systems,” 2026. [Online]. Available: https: //arxiv.org/abs/2503.03750

  65. [67]

    Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,

    M. Mazeikaet al., “Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,” 2024. [Online]. Available: https://arxiv.org/abs/2402.04249

  66. [68]

    Claw-eval: Toward trustworthy evaluation of autonomous agents,

    B. Yeet al., “Claw-eval: Toward trustworthy evaluation of autonomous agents,” 2026. [Online]. Available: https://arxiv.org/abs/2604.06132

  67. [69]

    Composable trust in agentic ai: Bridging architectural capability and system-level assurance,

    O. Al-Refai, I. Shahbaz, and E. Hammad, “Composable trust in agentic ai: Bridging architectural capability and system-level assurance,” in 2026 56th Annual IEEE International Conference on Dependable Systems and Networks Workshops (DSN-W), 2026, pp. 17–20

  68. [70]

    (2025) RagaAI AAEF (Agentic Application Evaluation Framework)

    RagaAI. (2025) RagaAI AAEF (Agentic Application Evaluation Framework). RagaAI Catalyst. Accessed: 2026-02-18. [Online]. Available: https://docs.raga.ai/ragaai-aaef-agentic-application-evaluat ion-framework

  69. [71]

    Whitepaper: Agentic application evaluation framework (AAEF),

    ——, “Whitepaper: Agentic application evaluation framework (AAEF),” RagaAI, Inc., Tech. Rep., Jun. 2024, accessed: 2025- 02-18. [Online]. Available: https://raga.ai/resources/patentsandpublicat ions/whitepaper-agentic-application-evaluation-framework

  70. [72]

    Starkloff, S

    A.-G. Starkloff, S. Kokaina, and S. Rahimi. (2026, Jan.) Evaluations for the agentic world. QuantumBlack, AI by McKinsey. Accessed: 2026-02-18. [Online]. Available: https://medium.com/quantumblack/ evaluations-for-the-agentic-world-c3c150f0dd5a

  71. [73]

    Morales Aguilera

    F. Morales Aguilera. (2025, Jul.) Building real-world agentic AI systems: A practical guide. AI Simplified in Plain English. Accessed: 2026-02-18. [Online]. Available: https://medium.com/ai-simplified-i n-plain-english/building-real-world-agentic-ai-systems-a-practical-g uide-9748d572b58b

  72. [74]

    (2025, May) MDR vs

    Wiz Experts Team. (2025, May) MDR vs. SOC: What’s the difference? Wiz, Inc. Accessed: 2026-02-18. [Online]. Available: https://www.wiz.io/academy/detection-and-response/mdr-vs-soc

  73. [75]

    St-webagentbench: A benchmark for evaluating safety and trustworthiness in web agents,

    I. Levy, B. Wiesel, S. Marreed, A. Oved, A. Yaeli, and S. Shlomov, “St-webagentbench: A benchmark for evaluating safety and trustworthiness in web agents,” 2025. [Online]. Available: https://arxiv.org/abs/2410.06703

  74. [76]

    successes

    L. R. Sammeta. (2025, Dec.) The AI agent report card you’ve been ignoring: Why 30% of your agent’s “successes” are actually failures. Accessed: 2026-02-23. [Online]. Available: https://laxmikumars.medi um.com/the-ai-agent-report-card-youve-been-ignoring-why-30-of-y our-agent-s-successes-are-actually-498fbebf44f9

  75. [77]

    Evaluating agentic ai systems: A balanced framework for performance, robustness, safety and beyond,

    M. Shukla, “Evaluating agentic ai systems: A balanced framework for performance, robustness, safety and beyond,”Preprints, August 2025. [Online]. Available: https://doi.org/10.20944/preprints202508.1847.v1

  76. [78]

    Technical report: Evaluating goal drift in language model agents,

    R. Arike, E. Donoway, H. Bartsch, and M. Hobbhahn, “Technical report: Evaluating goal drift in language model agents,” 2025. [Online]. Available: https://arxiv.org/abs/2505.02709

  77. [79]

    α 3-bench: A unified benchmark of safety, robustness, and efficiency for llm-based uav agents over 6g networks,

    M. A. Ferrag, A. Lakas, and M. Debbah, “α 3-bench: A unified benchmark of safety, robustness, and efficiency for llm-based uav agents over 6g networks,” 2026. [Online]. Available: https: //arxiv.org/abs/2601.03281

  78. [80]

    Autoadvexbench: Benchmarking autonomous exploitation of adversarial example defenses,

    N. Carlini, J. Rando, E. Debenedetti, M. Nasr, and F. Tram`er, “Autoadvexbench: Benchmarking autonomous exploitation of adversarial example defenses,” 2025. [Online]. Available: https://arxiv.org/abs/2503.01811

  79. [81]

    L. Rijo. (2026, Feb.) UC berkeley unveils framework as AI agents threaten to outrun oversight. PPC Land. Accessed: 2026-02-24. [Online]. Available: https://ppc.land/uc-berkeley-unveils-framework-a s-ai-agents-threaten-to-outrun-oversight/

  80. [82]

    N. Yadav. (2025, Nov.) 10 essential steps for evaluating the reliability of AI agents. Maxim AI. Accessed: 2026-02-24. [Online]. Available: https://www.getmaxim.ai/articles/10-essential-steps-for-evaluating-the -reliability-of-ai-agents/

Showing first 80 references.