Pith. sign in

REVIEW 3 major objections 5 minor 74 references

SoK: How Frontier AI Reshapes System-Level Security Risk Dynamics in Critical Infrastructure

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Frontier AI reshapes critical-infrastructure security through five interacting risk dynamics, not a single threat class.

desk verdict A thoughtful, well-scoped SoK that gives CI security a usable five-dimensional risk-dynamics lens; the central framework is partly circular and RD3 is thin, but the paper earns referee time. read the letter →

arxiv 2608.04033 v1 pith:NH2FHNN5 submitted 2026-08-03 cs.CR

classification cs.CR
keywords frontierAIcriticalinfrastructuresecurityriskdynamicssystem-levelassurancesupplychainagenticresearch-practicemismatchSoK
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This Systematization-of-Knowledge paper argues that frontier AI — large language models, multimodal foundation models, and agentic systems — does not just add new vulnerabilities to critical infrastructure; it reconfigures how risk emerges, spreads, and is controlled. The authors propose five interacting risk-dynamics dimensions — capability emergence, infiltration pathways, cross-system propagation, control authority, and response capacity — as the correct unit of analysis, replacing taxonomies organized by attack type, lifecycle stage, or asset class. Why it would matter if true: the lens determines what gets secured. If the paper is right, assurance for AI in railways, grids, hospitals, and water systems must be judged at the system level — by time from detection to containment, provenance visibility, and whether human override survives — rather than by model robustness scores. The paper draws that conclusion itself, deriving deployment-oriented assurance criteria and a research agenda grounded in continuous operation, partial observability, and safety-critical timing.

What carries the argument

The load-bearing machinery is the five-dimensional risk-dynamics framework, a lifecycle-aligned set of lenses whose unit of analysis is the recurring mechanism through which FAI changes CI security outcomes. Each of the five dimensions is operationalized by ten sub-question probes — fifty probes in all — which serve as a coding scaffold applied through interpretive, non-exclusive coding of 226 retained records from a screened corpus of 250, with consolidation into twenty-one analytical categories. A second-pass synthesis then compares AI-security research assumptions against recurring CI operational constraints — continuous operation, limited testing windows, partial observability, vendor opacity, safety-critical timing, and low false-positive tolerance — yielding four translation gaps and five deployment-relevant assurance criteria: detection-to-containment latency, explainability and interpretability, adversarial robustness, supply-chain risk mitigation, and incident response.

What would settle it

A concrete disconfirmation would be a documented or testbed-reproduced CI security event whose driving mechanism maps onto none of the five dimensions — for example, a damaging AI-mediated failure carried purely by a contractual, regulatory, or financial coupling, with no data, model, supply-chain, control, or response-capacity component. Short of that, the framework's completeness claim stands or falls with cross-system propagation: if shared-model and software-layer couplings never produce the cross-utility cascades the paper highlights, its high-consequence case weakens, and a study logging all AI-related CI incidents over a defined period to check whether each maps to at least one of RD1–RD5 would settle the matter.

Watch

Extended reading notes

Core claim

The paper's central claim is that frontier AI reshapes critical-infrastructure security through interacting dynamics rather than as a single threat class or a model-level robustness problem. Specifically: (RD1) new AI-enabled attack and defense capabilities, (RD2) infiltration through data, models, and AI supply chains, (RD3) cross-system propagation through shared models, data flows, and automation loops, (RD4) degradation of effective human and technical control authority, and (RD5) strain on institutional response capacity. The dimensions are analytically distinct but operationally coupled, and the proper unit of analysis is the recurring risk-dynamics mechanism, not the individual attack, asset, or lifecycle phase. From this, the authors conclude that CI security must shift from model-centric evaluation to lifecycle-structured, system-level assurance, with bounded deployment as the default posture for high-risk, safety-critical, or irreversible actions.

Load-bearing premise

The framework stands on the premise that five predefined risk-dynamics dimensions and fifty coding probes, applied to a 226-record corpus, capture the mechanisms that actually determine how frontier AI changes critical-infrastructure security — a premise the paper itself flags as exposed to framework-induced bias, with evidence for cross-system propagation notably sparse.

Editorial extensions

If this is right

  • Assurance for AI in critical infrastructure must shift from model robustness benchmarks to system-level criteria: detection-to-containment latency, fallback quality, auditability, and cascade containment.
  • Bounded deployment — AI authority that is explicit, limited, observable, and reversible — should be the default posture for high-risk, safety-critical, or irreversible CI actions rather than unconstrained autonomy.
  • Supply-chain provenance becomes a central control: because compromise enters upstream through data, models, retrieval corpora, and shadow AI, inventories such as AI bills of materials, artifact signing, and update governance are primary defenses.
  • The binding constraint on CI security is structural, not technical: a recurring research–practice mismatch between offline, bounded-adversary evaluation and continuous, black-box, safety-critical operations means deployable assurance cannot be imported from model-centric research as-is.
  • Policy should prioritize implementation-oriented governance — defined assurance evidence, provenance requirements, incident-reporting triggers, procurement standards, and cross-sector coordination — over high-level principles.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable corollary the paper leaves implicit: once RD2–RD5 dominate, model-level robustness scores should be near-orthogonal to system-level CI safety, so two deployments with identical model defenses but different provenance and response capacity should show measurably different outage and containment outcomes.
  • The framework could be operationalized as an audit instrument — scoring a deployment on each dimension, such as dependency visibility, control-handoff latency, and response-drill frequency — giving regulators a concrete alternative to model-centric evaluation; the paper does not build that instrument.
  • Because evidence for cross-system propagation is thin (23 coded sources), the framework's most consequential claim is also its least tested; building cascade testbeds that couple AI dispatch tools with protection logic across two simulated utilities would directly probe the predicted pathways.
  • A plausible generalization is that the same five lenses apply to narrower, non-frontier AI in CI (for example, conventional predictive-maintenance models), suggesting the framework captures a general property of AI-in-infrastructure rather than a frontier-specific one; the paper does not make or test this claim.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a five-dimensional 'risk-dynamics' framework (RD1–RD5: Capability Emergence, Infiltration Pathways, Cross-System Propagation, Control Authority, Response Capacity) for organizing how frontier AI (FAI) reshapes security risk in critical infrastructure (CI). It develops a coding methodology with 50 sub-question probes, applies it to a corpus of 226 screened sources, synthesizes the coded evidence into 21 analytical categories, identifies a research–practice mismatch, derives five deployment-oriented assurance criteria and six research directions, and concludes that FAI should be addressed at the system level rather than as a model-level robustness problem alone.

Significance. If the framework holds, it provides a useful organizing structure for a fragmented literature, and the paper is careful and transparent about its limitations (Section 3.7). The explicit definitions, the detailed probe-based coding scaffold with convergence mappings in Appendix Tables 7–11, and the translation of CI operational constraints into concrete assurance criteria (Section 5.2) are genuine contributions that go beyond simple threat enumeration. The paper does not ship machine-checked proofs or reproducible code, but its qualitative methodology is described in enough detail to be audited and extended, and the acknowledgment of framework-induced bias is a notable strength. However, the empirical pipeline cannot independently validate the sufficiency of the five dimensions, and the paper would need either stronger evidence for RD3 or a more modest conclusion to fully support its central claim.

major comments (3)
  1. [§3.4–3.5, §6] The empirical synthesis is a closed loop: the fifty probes (§3.4) are derived from the five RD dimensions, sources are coded against those probes, and the consolidation (§3.5) returns twenty-one categories that are then presented in Section 4 as the risk landscape supporting the conclusion in Section 6 that FAI 'reshapes CI security through interacting dynamics' of exactly those five dimensions. Because the framework is both the measurement instrument and the output, the coding exercise cannot independently validate the claim that RD1–RD5 are the right or sufficient organizing structure. The acknowledged 'framework-induced bias' (§3.7) is a real concession, but it is not operationalized; the paper would be strengthened by a negative-case analysis (deliberately searching for and reporting any identified mechanisms that do not map to any probe) or by applying the coding instrument to a held-out set of sources. As written, the conclusion overstates the epistemic weight of the empirical pipeline.
  2. [§4.3, Table 3] RD3 (Cross-System Propagation) is the dimension most directly tied to the paper's headline claim that FAI risk must be addressed at the 'system level' rather than the 'model level,' yet it rests on only 23 coded sources (of 226) and §3.7 concedes the evidence is sparse. The mechanisms in §4.3 (e.g., 'shared model dependency propagation' and 'protective system paradoxes') are described as plausible but are supported by a very thin evidence base, and the text itself says 'empirical evidence remains sparse.' The current corpus therefore does not provide strong support for the central claim that cross-system propagation is a defining, recurring system-level dynamic of FAI in CI. The authors should either substantially expand the RD3 evidence base (e.g., systematic inclusion of grey-literature incident reports and adjacent safety-critical domains) or explicitly reframe RD3 as a hypothesis requiring prospective validation.
  3. [§2.2, §4, Table 4] The manuscript repeatedly asserts that the dimensions are 'operationally coupled' and that the framework captures 'cross-dimensional interactions' (§1.1, §2.2, §4), and the conclusion in §6 is explicitly about 'interacting dynamics.' However, the only cross-dimensional artifact in the paper is the set of three bridge concepts in Appendix Table 4; there is no systematic analysis of how the dimensions interact (e.g., which RD1 mechanisms co-occur with RD4 failures in the coded corpus, or how RD5 capacity constraints alter RD3 cascade outcomes). The 'interacting dynamics' claim is asserted rather than demonstrated. The paper should either restrict its conclusion to the claim that the five dimensions are a useful decomposition, or it should present a cross-dimension analysis (e.g., a co-occurrence matrix from the non-exclusive coding) to substantiate the interaction claim.
minor comments (5)
  1. [Figure 2] The figure contains a typo: 'infrastrcuture' should be 'infrastructure.'
  2. [References] Several references have inconsistent spacing, such as 'V . Y' (Refs. [30], [17]) and 'P. Kelley' with stray spaces; a consistent reference style is needed.
  3. [§3.3, Table 2] The screening numbers are internally consistent (250 screened, 24 excluded with reasons, 226 retained), but the wording 'Initial screening (title, abstract, conclusion); 250 records screened' would be clearer if it stated that this count is after de-duplication.
  4. [§1.2] Contribution 3 says 'This motivates the deployment-oriented research directions,' but the research directions are actually derived in Section 5.3; consider phrasing this as 'we derive' to match the methodology.
  5. [Appendix Tables 7–11] The long merged cells in the appendix tables are difficult to read; consider reformatting the convergence rationale entries as separate bullet points or paragraphs.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the five-dimensional framework is a transparently introduced analytical lens, and the paper's own Section 3.7 discloses the framework-induced bias that a skeptic would cite.

full rationale

The derivation chain is a qualitative SoK synthesis, not a formal derivation. Section 1.1 introduces RD1–RD5 as the paper's proposed lens; Section 3.4 explicitly states that the fifty probes are operationalizations of these pre-defined dimensions, and Section 3.5 consolidates probe evidence into categories. This means the risk landscape presented in Section 4 is scaffolded by the framework, but the paper never claims the dimensions were discovered from the corpus or that the framework is uniquely forced by the literature. The conclusion in Section 6 is an interpretive synthesis, not a quantity computed from fitted inputs. There are no equations, no fitted parameters renamed as predictions, and no load-bearing self-citations (the author list does not appear in the references). The sparse RD3 evidence (23 of 226 sources, disclosed in Section 3.7) shows the coding had discriminatory power: sources were not forced into all dimensions. The acknowledged limitation in Section 3.7—'Reliance on predefined dimensions and probes may introduce framework-induced bias'—is a transparency statement, not a hidden circular step. The research-practice mismatch and assurance criteria are anchored in external CI operational constraints (continuous operation, partial observability, safety-critical timing, etc.), not merely in the RD definitions. Under the hard rules, framework-induced bias in a SoK does not reduce the central claim to its inputs by construction; the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No physical entities are introduced. The central claim rests on the adequacy of the five-dimensional decomposition and the representativeness of the corpus; these are domain assumptions, not fitted parameters. The framework is a literature synthesis, so there are no free parameters in the machine-learning sense.

assumptions (4)
  • domain assumption The five predefined risk-dynamics dimensions and fifty probes form a sufficient decomposition of FAI-related CI security risk.
    Introduced as a 'pragmatic design choice' in Section 3.4; the authors acknowledge in Section 3.7 that this may introduce framework-induced bias.
  • domain assumption The screened corpus of 250 records, with 226 retained, is representative of the broader FAI-CI security risk phenomenon.
    Section 3.3 describes corpus construction; Section 3.7 concedes the corpus is 'extensive but not comprehensive' and evidence is uneven across sectors.
  • domain assumption Grey literature and practitioner reports can serve as valid evidence for operational constraints and emerging practice.
    Section 3.3 and Section 3.7 state that grey literature is used for deployment constraints, not as stand-alone validation of technical mechanisms.
  • domain assumption Interpretive, non-exclusive probe-based coding produces stable analytical categories.
    Section 3.5 describes the coding process, but no inter-coder reliability measure or independent audit is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SoK: How Frontier AI Reshapes System-Level Security Risk Dynamics in Critical Infrastructure." pith.science (2026). https://pith.science/paper/NH2FHNN5

@misc{pith2026260804033,
  author       = {Pith},
  title        = {Pith review of: SoK: How Frontier AI Reshapes System-Level Security Risk Dynamics in Critical Infrastructure},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NH2FHNN5}},
  note         = {Machine review of arXiv:2608.04033}
}
read the original abstract

Frontier artificial intelligence (FAI), encompassing large-scale, general-purpose AI systems, including large language models, multimodal foundation models, and agentic systems, is increasingly integrated into critical infrastructure (CI). This challenges long-standing security assumptions of bounded behavior, segmented networks, component transparency, and human-paced decision-making. Existing AI-security literature typically organizes risks by attack type, lifecycle stage, or asset class, but fails to capture the system-level dynamics, such as how risk emerges, spreads, and is controlled, through which FAI reshapes CI security outcomes. This Systematization of Knowledge (SoK) introduces a five-dimensional risk-dynamics framework that characterizes how FAI reconfigures CI security across the lifecycle: (i) Capability Emergence through new FAI-enabled attack and defense capabilities, (ii) Infiltration Pathways through data, models and AI supply chains, (iii) Cross-System Propagation across interconnected infrastructures and dependencies, (iv) degradation of effective technical and human Control Authority, and (v) strain on institutional Response Capacity under operational pressure. Rather than enumerating threats, the framework identifies recurring mechanisms that jointly determine system-level risk. We further identify a structural mismatch between academic AI-security research and CI operational constraints, and derive a deployment-oriented research agenda grounded in system-level assurance criteria. Collectively, this work shifts attention from model-centric robustness to lifecycle-structured, system-level assurance in interconnected CI environments.

Figures

Figures reproduced from arXiv: 2608.04033 by the authors.

Figure 1
Figure 1. Overview of the five security risk-dynamics dimensions. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Conceptual framework linking research questions to risk-dynamic dimensions, deployment gaps, assurance criteria, [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

74 extracted references · 66 canonical work pages

  1. [1]

    Strengthening AI Critical Infrastructure Security with the MIT AI Risk Repository and MITRE ATLAS Frameworks,

    J. Carroll, “Strengthening AI Critical Infrastructure Security with the MIT AI Risk Repository and MITRE ATLAS Frameworks,” in Proceedings of the 24th European Conference on Cyber Warfare and Security. ECCWS, 2025. [Online]. Available: https://papers.a cademic-conferences.org/index.php/eccws/article/view/3713

  2. [2]

    AI-Based Framework for Predictive Vulnerability As- sessment and Risk Prioritization in Critical Infrastructure Systems,

    M. Okoebor, “AI-Based Framework for Predictive Vulnerability As- sessment and Risk Prioritization in Critical Infrastructure Systems,” International Journal of Scientific Research and Modern Technology, vol. 4, no. 10, pp. 81–85, 2025

  3. [3]

    AI-Powered Incident Response Automation in Critical Infrastructure Protection,

    E. Obuse, E. D. Etim, I. A. Essien, and et al., “AI-Powered Incident Response Automation in Critical Infrastructure Protection,”Int. j. adv. multidisc. res. stud., vol. 3, no. 1, pp. 1156–1171, 2023

  4. [4]

    eGridGPT: Trustworthy AI in the Control Room,

    S. L. Choi, R. Jain, P. Emami, K. Wadsack, F. Ding, H. Sun, K. Gruchalla, J. Hong, H. Zhang, X. Zhu, and B. Kroposki, “eGridGPT: Trustworthy AI in the Control Room,”Technical Report,

  5. [5]

    Watchdogs and oracles: Runtime verification meets large language models for autonomous systems,

    A. Ferrando, “Watchdogs and oracles: Runtime verification meets large language models for autonomous systems,”Electronic Proceedings in Theoretical Computer Science, vol. 436, p. 80–87, Nov. 2025. [Online]. Available: http://dx.doi.org/10.4204/EPTCS.4 36.8

  6. [6]

    Ai for critical infras- tructure security: Concepts, challenges, and future directions,

    M. Al-Hawawreh, Z. Baig, and S. Zeadally, “Ai for critical infras- tructure security: Concepts, challenges, and future directions,”IEEE Internet of Things Magazine, vol. 7, no. 4, pp. 136–142, 2024

  7. [7]

    Building cyber resilience in critical infrastructure,

    pwc, “Building cyber resilience in critical infrastructure,”pwc report,

  8. [8]

    Digital twin-driven intrusion detection for industrial scada: A cyber-physical case study,

    A. Sayghe, “Digital twin-driven intrusion detection for industrial scada: A cyber-physical case study,”Sensors, vol. 25, no. 16, 2025. [Online]. Available: https://www.mdpi.com/1424-8220/25/16/4963

Show all 74 references
  1. [9]

    Ai agents vs. agentic ai: A conceptual taxonomy, applications and challenges,

    R. Sapkota, K. I. Roumeliotis, and M. Karkee, “Ai agents vs. agentic ai: A conceptual taxonomy, applications and challenges,”Information Fusion, vol. 126, p. 103599, 2026

  2. [10]

    Alignment, agency and autonomy in frontier ai: A systems engineering perspective,

    K. Tallam, “Alignment, agency and autonomy in frontier ai: A systems engineering perspective,” 2025. [Online]. Available: https://arxiv.org/abs/2503.05748

  3. [11]

    Autonomous ai-based cybersecurity framework for critical infrastructure: Real-time threat mitigation,

    J. Paulraj, B. Raghuraman, N. Gopalakrishnan, and Y . Otoum, “Autonomous ai-based cybersecurity framework for critical infrastructure: Real-time threat mitigation,”2025 IEEE/ACIS 29th International Conference on Software Engineering, Artificial Intelligence, Networking and Par...

  4. [12]

    Large language models and their applications in roadway safety and mobility enhancement: A comprehensive review,

    M. M. Karim, Y . Shi, S. Zhang, B. Wang, M. Nasri, and Y . Wang, “Large language models and their applications in roadway safety and mobility enhancement: A comprehensive review,”Artificial Intelli- gence for Transportation, vol. 1, 2025

  5. [13]

    AI Security Institute Frontier AI Trends Report,

    AI Security Institute (AISI), “AI Security Institute Frontier AI Trends Report,”Report, 2025. [Online]. Available: https: //www.aisi.gov.uk/frontier-ai-trends-report

  6. [14]

    From transformers to large language models: A systematic review of ai applications in the energy sector towards agentic digital twins,

    G. Antonesi, T. Cioara, I. Anghel, V . Michalakopoulos, E. Sarmas, and L. Toderean, “From transformers to large language models: A systematic review of ai applications in the energy sector towards agentic digital twins,” 2025. [Online]. Available: https: //arxiv.org/abs/2506.06359

  7. [15]

    TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems,

    S. Raza, R. Sapkota, M. Karkee, and C. Emmanouilidis, “TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems,” 2025. [Online]. Available: https://arxiv.org/abs/2506.04133

  8. [16]

    Agentic ai for autonomous anomaly management in complex systems,

    R. V . Barenji and S. Khoshgoftar, “Agentic ai for autonomous anomaly management in complex systems,” 2025. [Online]. Available: https://arxiv.org/abs/2507.15676

  9. [17]

    Trustworthy agentic AI systems: a cross-layer review of architectures, threat models, and governance strategies for real-world deployment,

    I. ADABARA, B. O. Sadiq, A. N. Shuaibu, Y . I. Danjuma, and V . Maninti, “Trustworthy agentic AI systems: a cross-layer review of architectures, threat models, and governance strategies for real-world deployment,”f1000research, vol. 14, 2025. [Online]. Available: https://europ...

  10. [18]

    Standardized threat taxonomy for ai security, governance, and regulatory compliance,

    H. Huwyler, “Standardized threat taxonomy for ai security, governance, and regulatory compliance,” 2025. [Online]. Available: https://arxiv.org/abs/2511.21901

  11. [19]

    ETSI EN304223: Securing Artificial Intelligence (SAI); Baseline Cyber Security Requirements for AI Models and Systems,

    ETSI, “ETSI EN304223: Securing Artificial Intelligence (SAI); Baseline Cyber Security Requirements for AI Models and Systems,”Report European Standard, 2025. [Online]. Available: https://www.etsi.org/deliver/etsi TS/104200 104299/104223/01.01.0 1 60/ts 104223v010101p.pdf

  12. [20]

    Risk Taxonomy and Thresholds for Frontier AI Frameworks,

    Frontier Model Forum (FMF), “Risk Taxonomy and Thresholds for Frontier AI Frameworks,”Technical Report, 2025. [Online]. Available: https://www.frontiermodelforum.org/technical-reports/ris k-taxonomy-and-thresholds/

  13. [21]

    A FAIR Taxonomy for Cyber Risk Scenarios: An Analyst’s Guide for Defining Risk Scenarios for Continuous Risk Management,

    FAIR Institute, “A FAIR Taxonomy for Cyber Risk Scenarios: An Analyst’s Guide for Defining Risk Scenarios for Continuous Risk Management,”Report, 2025. [Online]. Available: https: //www.fairinstitute.org/

  14. [22]

    Securing Critical Infrastructure in the Age of AI,

    Centre for Security and Emerging Technology (CEST), “Securing Critical Infrastructure in the Age of AI,”Workshop Report, 2024. [Online]. Available: https://cset.georgetown.edu/publication/securing -critical-infrastructure-in-the-age-of-ai/

  15. [23]

    Shadow AI: The hidden risk inside your enterprise (and how to manage it),

    Z. Banach, “Shadow AI: The hidden risk inside your enterprise (and how to manage it),” 2025. [Online]. Available: https://www.invicti. com/blog/web-security/shadow-ai-risks-challenges-solutions-for

  16. [24]

    Industry temperature check: barriers and enablers to AI assurance,

    Gov UK, “Industry temperature check: barriers and enablers to AI assurance,” 2022. [Online]. Available: https://www.gov.uk/governm ent/publications/industry-temperature-check-barriers-and-enablers-t o-ai-assurance

  17. [25]

    Supply Chain Security and AI Risk Governance Model for Critical Infrastructure under NIS2, CER, and CRA,

    N. Parlov, G. Akrap, and J. Esterhajer, “Supply Chain Security and AI Risk Governance Model for Critical Infrastructure under NIS2, CER, and CRA,”Applied Cybersecurity and Internet Governance (ACIG), vol. 4, 2025

  18. [26]

    Exploiting Trust in Open-Source AI: The Hidden Supply Chain Risk No One Is Watching,

    Ashish Verma and Deep Patel, “Exploiting Trust in Open-Source AI: The Hidden Supply Chain Risk No One Is Watching,” 2025. [Online]. Available: https://www.trendaisecurity.com/en-us/resource s-insights/research/exploiting-trust-in-open-source-ai-the-hidden-s upply-chain-risk-no...

  19. [27]

    AI-Driven Governance of Shadow AI Systems for Enhanced Enterprise Cybersecurity,

    E. Whitford, S. Marcel, and O. Oladele, “AI-Driven Governance of Shadow AI Systems for Enhanced Enterprise Cybersecurity,” Report, 10 2025. [Online]. Available: https://www.researchgate.net/p ublication/397014702 AI-Driven Governance of Shadow AI Syste ms for Enhanced Enterpri...

  20. [28]

    Military AI Cyber Agents (MAICAs) Constitute a Global Threat to Critical Infrastructure,

    T. Dubber and S. Lazar, “Military AI Cyber Agents (MAICAs) Constitute a Global Threat to Critical Infrastructure,” 2025. [Online]. Available: https://arxiv.org/abs/2506.12094

  21. [29]

    A Framework for Evaluating Emerging Cyberattack Capabilities of AI,

    M. Rodriguez, R. A. Popa, F. Flynn, L. Liang, A. Dafoe, and A. Wang, “A Framework for Evaluating Emerging Cyberattack Capabilities of AI,” 2025. [Online]. Available: https://arxiv.org/abs/2503.11917

  22. [30]

    Frontier ai’s impact on the cybersecurity landscape,

    Y . Potter, W. Guo, Z. Wang, T. Shi, H. Li, A. Zhang, P. G. Kelley, K. Thomas, and D. Song, “Frontier ai’s impact on the cybersecurity landscape,” 2025. [Online]. Available: https: //arxiv.org/abs/2504.05408

  23. [31]

    Dual-use of large language models (llms) and generative ai (genai) in cybersecurity: Risks, defenses, and governance strategies,

    K. Ahi, V . Agrawal, and S. Valizadeh, “Dual-use of large language models (llms) and generative ai (genai) in cybersecurity: Risks, defenses, and governance strategies,”TechRxiv, vol. 2025, no. 0826,

  24. [32]

    How does ai transform cyber risk management?

    S. Zeijlemaker, Y . K. Lemiesa, S. L. Schr ¨oer, A. Abhishta, and M. Siegel, “How does ai transform cyber risk management?”Systems, vol. 13, no. 10, 2025

  25. [33]

    Principles for the Secure Integration of Artificial Intelligence in Operational Technology,

    Various Gov., “Principles for the Secure Integration of Artificial Intelligence in Operational Technology,” 2025. [Online]. Available: https://www.cisa.gov/resources-tools/resources/principles-secure-int egration-artificial-intelligence-operational-technology

  26. [34]

    Future Risks of Frontier AI: Which capabilities and risks could emerge at the cutting edge of AI in the future?

    Gov. Office for Science UK, “Future Risks of Frontier AI: Which capabilities and risks could emerge at the cutting edge of AI in the future?” 2023. [Online]. Available: https://assets.publishing.service. gov.uk/media/653bc393d10f3500139a6ac5/future-risks-of-frontier-a i-annex-a.pdf

  27. [35]

    A survey on the applications of frontier ai, foundation models, and large language models to intelligent transportation systems,

    M. R. Shoaib, H. M. Emara, and J. Zhao, “A survey on the applications of frontier ai, foundation models, and large language models to intelligent transportation systems,” inInternational Conference on Computer and Applications (ICCA). IEEE, 2023. [Online]. Available: https://i...

  28. [36]

    Critical Infrastructure Sectors,

    CISA, “Critical Infrastructure Sectors,” 2026. [Online]. Available: https://www.cisa.gov/topics/critical-infrastructure-security-and-resil ience/critical-infrastructure-sectors

  29. [37]

    Critical infrastructure resilience at EU-level,

    E. Union, “Critical infrastructure resilience at EU-level,” 2026. [Online]. Available: https://home-affairs.ec.europa.eu/policies/intern al-security/counter-terrorism-and-radicalisation/protection/critical-i nfrastructure-resilience-eu-level en

  30. [38]

    Challenges in the vulnerability and risk analysis of critical infrastructures,

    E. Zio, “Challenges in the vulnerability and risk analysis of critical infrastructures,”Reliability Engineering & System Safety, vol. 152, pp. 137–150, 2016

  31. [39]

    Kufeoglu and A

    S. Kufeoglu and A. T. Akgun,Cyber Resilience in Critical Infras- tructure, 1st ed. CRC Press Taylor & Francis Group, 2024

  32. [40]

    Critical infrastructure, panarchies and the vulnerability paths of cascading disasters,

    G. Pescaroli and D. Alexander, “Critical infrastructure, panarchies and the vulnerability paths of cascading disasters,”Nat Hazards, vol. 82, pp. 175–192, 2016

  33. [41]

    Shadow AI: Gover- nance, Risk, and Organisational Resilience,

    J. A. J. Ross, L. Hibbert, and E. J. Moss, “Shadow AI: Gover- nance, Risk, and Organisational Resilience,” in2025 International Conference on Artificial Intelligence, Computer, Data Sciences and Applications (ACDSA), 2025, pp. 1–9

  34. [42]

    Roles and Responsibilities Framework for Artificial Intelligence in Critical Infrastructure,

    US Department of Homeland Security, “Roles and Responsibilities Framework for Artificial Intelligence in Critical Infrastructure,” 2024. [Online]. Available: https://www.dhs.gov/publication/roles-and-respo nsibilities-framework-artificial-intelligence-critical-infrastructure

  35. [43]

    Framework for Improving Critical Infrastructure Cybersecurity,

    NIST, “Framework for Improving Critical Infrastructure Cybersecurity,” 2018. [Online]. Available: https://nvlpubs.nist.g ov/nistpubs/cswp/nist.cswp.04162018.pdf

  36. [44]

    Artificial Intelligence Risk Management Framework (AI RMF 1.0),

    ——, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” 2023. [Online]. Available: https://nvlpubs.nist.gov/nistp ubs/ai/nist.ai.100-1.pdf

  37. [45]

    LLMs for Cybersecurity in the Big Data Era: A Comprehensive Review of Applications, Challenges, and Future Directions,

    A. Karras, L. Theodorakopoulos, C. Karras, A. Theodoropoulou, I. Kalliampakou, and G. Kalogeratos, “LLMs for Cybersecurity in the Big Data Era: A Comprehensive Review of Applications, Challenges, and Future Directions,”Information, vol. 16, no. 11, 2025

  38. [46]

    Generative AI in cybersecurity: A comprehensive review of LLM applications and vulnerabilities,

    M. A. Ferrag, F. Alwahedi, A. Battah, B. Cherif, A. Mechri, N. Ti- hanyi, T. Bisztray, and M. Debbah, “Generative AI in cybersecurity: A comprehensive review of LLM applications and vulnerabilities,” Internet of Things and Cyber-Physical Systems, vol. 5, pp. 1–46, 2025

  39. [47]

    Managing extreme AI risks amid rapid progress,

    Y . Bengio, G. Hinton, and et al., “Managing extreme AI risks amid rapid progress,”Science AAAS, vol. 384(6698), p. 842–845, 2024

  40. [48]

    Agentic AI and the Cyber Arms Race,

    S. Oesch, J. Hutchins, P. Austria, and A. Chaulagain, “Agentic AI and the Cyber Arms Race,”Computer, vol. 58, no. 5, pp. 82–85, 2025

  41. [49]

    Dual use concerns of generative AI and large language models,

    A. Grinbaum and L. Adomaitis, “Dual use concerns of generative AI and large language models,”Journal of Responsible Innovation, vol. 11, no. 1, p. 2304381, 2024

  42. [50]

    Risk thresholds for frontier AI,

    L. Koessler, J. Schuett, and M. Anderljung, “Risk thresholds for frontier AI,” 2024. [Online]. Available: https://arxiv.org/abs/2406.1 4713

  43. [51]

    A Master Attack Methodology for an AI-Based Automated Attack Planner for Smart Cities,

    G. Falco, A. Viswanathan, C. Caldera, and H. Shrobe, “A Master Attack Methodology for an AI-Based Automated Attack Planner for Smart Cities,”IEEE Access, vol. 6, pp. 48 360–48 373, 2018

  44. [52]

    Poisoning and Backdooring Contrastive Learning,

    N. Carlini and A. Terzis, “Poisoning and Backdooring Contrastive Learning,” inInternational Conference on Learning Representations,

  45. [53]

    Extracting Training Data from Large Language Models,

    N. Carlini, F. Tram `er, E. Wallace, M. Jagielski, A. Herbert-V oss, K. Lee, A. Roberts, T. B. Brown, D. X. Song, ´U. Erlingsson, A. Oprea, and C. Raffel, “Extracting Training Data from Large Language Models,” inUSENIX Security Symposium, 2020. [Online]. Available: https://api...

  46. [54]

    What We Know about AIBOMs: Results from a Multivocal Litera- ture Review on Artificial Intelligence Bill of Materials,

    S. Nocera, M. D. Penta, F. Ahmed, S. Romano, and G. Scanniello, “What We Know about AIBOMs: Results from a Multivocal Litera- ture Review on Artificial Intelligence Bill of Materials,”ACM Trans. Softw. Eng. Methodol., Dec. 2026 (online 2025), just Accepted

  47. [55]

    Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI,

    A. Barredo Arrieta, N. D ´ıaz-Rodr´ıguez, J. Del Ser, A. Bennetot, S. Tabik, A. Barbado, S. Garcia, S. Gil-Lopez, D. Molina, R. Ben- jamins, R. Chatila, and F. Herrera, “Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward respon...

  48. [56]

    Large Language Models Integration in Smart Grids,

    S. Madani, A. Tavasoli, Z. Khoshtarash Astaneh, and P.-O. Pineau, “Large Language Models Integration in Smart Grids,”Energy Re- ports, vol. 14, pp. 1562–1577, 2025

  49. [57]

    On Large Language Models in Mission-Critical IT Governance: Are We Ready Yet? ,

    M. Esposito, F. Palagiano, V . Lenarduzzi, and D. Taibi, “On Large Language Models in Mission-Critical IT Governance: Are We Ready Yet? ,” in2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). Los Alamitos, CA, USA...

  50. [58]

    Identifying, un- derstanding and analyzing critical infrastructure interdependencies,

    S. M. Rinaldi, J. P. Peerenboom, and T. K. Kelly, “Identifying, un- derstanding and analyzing critical infrastructure interdependencies,” IEEE Control Systems Magazine, pp. 11–25, 2001. [Online]. Avail- able: https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=969131

  51. [59]

    Hybrid learning-based fault prediction and cascading failure mitigation in multi-network energy systems,

    X. Wu, Y . Cao, H. Wu, and et al., “Hybrid learning-based fault prediction and cascading failure mitigation in multi-network energy systems,”Scientific reports, 2025

  52. [60]

    Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents,

    V . S. Narajala and O. Narayan, “Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents,” inInternational conference on algorithms, computing and artificial intelligence (ACAI). IEEE, 2025. [Online]. Available: https://ieeexplore.ie...

  53. [61]

    Agentic Large Language Models, a Survey,

    A. Plaat, M. Van Duijn, N. Van Stein, M. Preuss, P. Van der Putten, and K. J. Batenburg, “Agentic Large Language Models, a Survey,” Journal of Artificial Intelligence Research, vol. 84, Dec. 2025

  54. [62]

    Agentic AI Frameworks: Architectures, Protocols, and Design Challenges,

    H. Derouiche, Z. Brahmi, and H. Mazeni, “Agentic AI Frameworks: Architectures, Protocols, and Design Challenges,” 2025. [Online]. Available: https://arxiv.org/abs/2508.10146

  55. [63]

    AI-Enhanced Cyber Threat Detection and Response Advancing National Security in Critical Infrastructure,

    M. A. Goffer, M. S. Uddin, J. kaur, S. N. Hasan, C. R. Barikdar, J. Hassan, N. Das, P. Chakraborty, and R. Hasan, “AI-Enhanced Cyber Threat Detection and Response Advancing National Security in Critical Infrastructure,”Journal of Posthumanism, vol. 5, no. 3, p. 1667–1689, Apr. 2025

  56. [64]

    Frontier Model Forum: Advancing frontier AI safety and security

    F. M. Forum, “Frontier Model Forum: Advancing frontier AI safety and security.” [Online]. Available: https://www.frontiermodelforum.o rg/

  57. [65]

    Securing Critical Infrastructure in the Age of AI,

    K. Crichton, P. Eke, J. Ji, and et al., “Securing Critical Infrastructure in the Age of AI,”Workshop Report, 2024. [Online]. Available: https://cset.georgetown.edu/publication/securing-critical-infrastructur e-in-the-age-of-ai/

  58. [66]

    Independent review of the security of critical infrastructure act 2018,

    Jill Slay AM, “Independent review of the security of critical infrastructure act 2018,” 2018. [Online]. Available: https://www.ho meaffairs.gov.au/cyber-security-subsite/files/independent-review-soc i-act-final-report.pdf

  59. [67]

    Data Poisoning Vulnerabilities Across Health Care Artificial Intelligence Architec- tures: Analytical Security Framework and Defense Strategies,

    F. Abtahi, F. Seoane, I. Pau, and M. Vega-Barbas, “Data Poisoning Vulnerabilities Across Health Care Artificial Intelligence Architec- tures: Analytical Security Framework and Defense Strategies,”Jour- nal of medical Internet research, vol. 28, 2026

  60. [68]

    Guidelines for Next-Generation Grid Communications Architecture,

    U.S. Department of Energy, Office of Electricity, “Guidelines for Next-Generation Grid Communications Architecture,” 2024. [Online]. Available: https://www.energy.gov/sites/default/files/202 4-10/Guidelines%20for%20Next-Generation%20Grid%20Communi cations%20Architecture.pdf

  61. [69]

    ChatGPT and Other Large Language Models for Cybersecurity of Smart Grid Applications,

    A. Zaboli, S. L. Choi, T.-J. Song, and J. Hong, “ChatGPT and Other Large Language Models for Cybersecurity of Smart Grid Applications,” 2024. [Online]. Available: https://ieeexplore.ieee.org/ stamp/stamp.jsp?arnumber=10688863

  62. [70]

    A Formal Verification Frame- work for Ensuring Safety in Autonomous Cyber-Physical Systems Operating in Unstructured Environments,

    F. Mehmood, U. Riaz, and N. Khalid, “A Formal Verification Frame- work for Ensuring Safety in Autonomous Cyber-Physical Systems Operating in Unstructured Environments,”Journal of Computational Intelligence for Hybrid Cloud and Edge Computing Networks, vol. 9, 2025. Appendix TA...

  63. [2021]

    Available: https://www.pwc.com.au/cyber-security-d igital-trust/critical-infrastructure/critical-infrastructure-resilience.ht ml

    [Online]. Available: https://www.pwc.com.au/cyber-security-d igital-trust/critical-infrastructure/critical-infrastructure-resilience.ht ml

  64. [2022]

    Available: https://openreview.net/forum?id=iC4UHb Q01Mp

    [Online]. Available: https://openreview.net/forum?id=iC4UHb Q01Mp

  65. [2024]

    Available: https://docs.nlr.gov/docs/fy24osti/87740.p df

    [Online]. Available: https://docs.nlr.gov/docs/fy24osti/87740.p df

  66. [2025]

    Available: https://www.techrxiv.org/doi/abs/10.3622 7/techrxiv.175616948.85236631/v1

    [Online]. Available: https://www.techrxiv.org/doi/abs/10.3622 7/techrxiv.175616948.85236631/v1

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.