Pith. sign in

REVIEW 4 major objections 5 minor 97 references

Secure human oversight of AI: Threat modeling in a socio-technical context

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read Human oversight of AI creates a new attack surface: attackers can undermine oversight by targeting the AI system, the communication channels, or the personnel themselves, so oversight must be designed and hardened as a security-relevant com

desk verdict A useful security framing for human oversight—mapping known cyber attacks onto oversight requirements—but the mapping is heuristic and the 'systematic threat modeling' claim oversells the body; still worth a serious referee. read the letter →

arxiv 2509.12290 v3 pith:JHED3N6F submitted 2025-09-15 cs.CR cs.CYcs.HC

classification cs.CRcs.CYcs.HC
keywords humanoversightattacksurfacethreatmodelingAIsecurityvectorshardeningstrategiesgovernancesocio-technical
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that human oversight of AI, promoted as a safeguard and increasingly required in high-stakes regulation, is itself a target. Malicious actors can undermine oversight by attacking the AI system, the channels through which oversight personnel see and control it, or the personnel directly. The paper models human oversight as an IT application and extends the cybersecurity concept of an attack vector to any pathway that degrades effective oversight. It maps eleven such vectors to the four requirements of effective oversight—knowing what the AI is doing, being able to intervene, staying in control of one's own actions, and holding the right intentions—and pairs each with hardening strategies. A sympathetic reader should take away that secure oversight is a design problem in its own right, not a side effect of effective oversight.

What carries the argument

The carrying mechanism is the modeling of human oversight as an IT application and the corresponding widening of 'attack vector' to mean any pathway that undermines the effectiveness of human oversight. This yields the human oversight attack surface: the union of the AI system's technical infrastructure, the communication channels between system and personnel, and the personnel themselves. The organizing device is a four-requirement model of effective oversight; each of the eleven listed attack vectors is classified by which requirement it attacks, and each of the seven listed hardening strategies is mapped to the vectors it counters.

What would settle it

A finding that would settle the claim: an oversight setup that satisfies all four requirements—personnel well informed, able to intervene, unpressured, and well-intentioned—but is still successfully disabled by an attack would contradict the paper's mapping. Alternatively, an empirical study showing that the listed hardening strategies, such as personnel training, fail to reduce the success of coercion or bribery in oversight roles would undercut the practical claim.

Watch

Extended reading notes

Core claim

The central claim is that the human oversight architecture creates a new attack surface inside the safety, security, and accountability structure of AI operations. Attacking oversight is a way to attack AI operations: an actor can degrade the four requirements of effective oversight—epistemic access, causal power, self-control, and fitting intentions—using vectors such as poisoning, adversarial inputs, manipulated explanations, denial of service, man-in-the-middle interference, malware, social engineering, coercion, bribery, insider threats, and vulnerability exploitation. Because oversight is becoming a regulatory and operational cornerstone in high-risk domains, the paper argues that faili

Load-bearing premise

The load-bearing premise is that the human and organizational sides of oversight can be analyzed with the same attack-vector and hardening vocabulary used for digital systems; if psychological and social processes such as fear, fatigue, or bribery do not behave like technical components with corresponding fixes, the central mapping loses its foundation.

Editorial extensions

If this is right

  • Designers can use the vector-to-requirement mapping as a threat-modeling checklist when building oversight for high-risk AI.
  • Securing oversight requires protecting all four requirements; an attack that disables even one—for example, blocking the ability to intervene—can make oversight ineffective.
  • Security measures must cover the full loop of model, interface, network, and personnel, not just the AI model itself.
  • The paper's non-exhaustive list can serve as a starting point for governance and audit frameworks to require oversight-specific security testing.
  • If oversight is left unhardened, human oversight may increase rather than decrease the riskiness of AI operations, since it becomes a single high-value point of attack.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy invites a concrete next step: build red-team exercises for oversight pipelines in fields like medicine or public administration, then measure which of the four requirements fails first; the paper does not run such tests.
  • The same framework can classify attacks by future AI agents that model and try to evade oversight; the paper mentions this as a future scenario, but the mapping already supplies a vocabulary for it.
  • General-purpose hardening measures such as personnel training may be the weakest link: the paper lists training as countering social engineering, coercion, and bribery, but evidence that training changes behavior under real coercion or bribery remains thin.
  • If regulators adopt this view, oversight would shift from a procedural checkbox to a security-critical component with its own assurance and auditing requirements.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper argues that human oversight of AI, promoted as a safeguard against AI risks, itself creates an attack surface that malicious actors can exploit to undermine the four requirements of effective oversight from Sterz et al. (epistemic access, causal power, self-control, fitting intentions). It identifies eleven attack vectors and seven hardening strategies, organized in Tables 1 and 2, and discusses each in Sections 3 and 4. The paper's stated contributions are introducing a security perspective on human oversight and providing an overview of attack vectors and hardening strategies.

Significance. If the proposed perspective is taken up, it fills a genuine gap: existing human-oversight research focuses on effectiveness and largely ignores adversarial exploits of the oversight process. The paper provides a useful compilation and clearly links to established cybersecurity taxonomies. Its explicit use of a well-known oversight-requirements framework is a strength, as is the candid acknowledgment that the lists are not exhaustive and that future work is needed. However, the current manuscript is better described as a structured position paper/checklist than as a systematic threat model; the central contribution depends on mapping tables whose entries are asserted rather than derived. With revision to supply a method and support for the mappings, the paper could be a useful starting point for the community.

major comments (4)
  1. [Abstract and §3, §4] The abstract promises 'systematic threat modeling,' but the body presents an explicitly non-exhaustive list of vectors 'inspired by' the literature with no threat-modeling methodology, no adversary model, and no selection criteria. For example, §3 states the list 'is not exhaustive but illustrates the breadth.' This mismatches the stated contribution. The authors should either (a) supply a systematic method (e.g., define actors, assets, trust boundaries, and derive vectors from the four requirements) or (b) revise the claims to 'overview of attack vectors' as the internal abstract already does.
  2. [Table 1, §3.1] The mapping of poisoning attacks to 'self-control' is unsupported. The text posits that triggered outputs cause 'excessive micro-notifications or confirmation requests,' but that is a generic UI/notification effect independent of data poisoning; no mechanism connects poisoning to the oversight interface. Similar unstated assumptions appear in other table entries. Because Table 1 is the systematic output that supports the 'secure human oversight' conclusion, each mapping should be justified via an explicit causal chain or at least flagged as a hypothesis.
  3. [Table 2, §4.4] Transparency is mapped to 'all attack vectors,' including coercion and bribery. Section 4.4 gives no mechanism for how transparency prevents physical threats or bribery, and it is not clear that it can. This overclaim weakens the credibility of the hardening table. The mapping should be restricted to vectors for which a plausible mechanism exists, and other strategies should be developed for human-targeted vectors.
  4. [§2.2] The definitional move — extending 'attack vector' to 'methods or pathways that undermine the effectiveness of human oversight' — makes the central claim that oversight is attackable trivially true. A more persuasive approach would start from an adversary model and show how an adversary can use specific capabilities to compromise each of the four requirements. As written, the paper risks relabeling known socio-technical failure modes as 'attacks' without demonstrating that they are part of an attack surface in the cybersecurity sense.
minor comments (5)
  1. [Abstract vs. full text] The arXiv abstract says 'for the purpose of systematic threat modeling' while the full-text abstract says 'we analyze attack vectors' and 'overview'; align the two versions.
  2. [§3.5] 'Snarfing' appears to be a typo for 'sniffing' in the context of WLAN-based attacks.
  3. [§4.7] 'Training of human oversight personal' should be 'personnel.'
  4. [References] Some references have formatting issues, e.g., entry [45] appears as 'Y amagishi.' A final proofreading pass would help.
  5. [Figure 1] The figure is cited in the introduction but not reproduced in the text; ensure it is legible and referenced at the appropriate point.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a conceptual threat-modeling exercise that applies external frameworks and makes explicit, substantive mapping claims rather than deriving predictions from fitted inputs or self-citation chains.

full rationale

The paper does not fit parameters, make quantitative predictions, or derive an empirical result from its own inputs. Its central move is to extend the definition of 'attack vector' to include 'methods or pathways that undermine the effectiveness of human oversight' (Section 2.2). This extension is explicit and stipulative; the paper does not hide a derivation behind it. The subsequent mapping in Table 1 is not a consequence of that definition alone: each entry is argued with a concrete reasoning chain (e.g., poisoning attacks affect self-control only via the additional assumption that triggered outputs generate excessive micro-notifications). Those steps are substantive, if sometimes under-specified, and they are not forced by the definition. The four requirements of effective oversight are imported from Sterz et al. [1]. Although two present authors co-authored that paper, [1] is a published, independently accessible FAccT paper, and the current work applies it rather than rehearsing or assuming its proof; no load-bearing conclusion depends on an unverified self-citation chain. Other self-citations ([5], [25], [74]) provide background or prior definitions and are not used to close an argument. The hardening strategies in Table 2 are also proposed applications of known cybersecurity concepts, not predictions validated by the paper's own method. The paper explicitly acknowledges its lists are 'not exhaustive' and 'inspired by' the broader literature, and it flags open questions such as counterproductive behavior as future work. Overbreadth or lack of mechanistic detail in some mappings is a rigor or correctness concern, not circularity. No equation, fitted parameter, or defined term is surreptitiously reused to produce the conclusion.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper rests on imported building blocks: the four oversight requirements from Sterz et al. [1] and the transferability of cybersecurity taxonomies to the human oversight context. Both are taken without validation. No numerical fits or new entities are introduced.

assumptions (4)
  • domain assumption The four requirements from Sterz et al. [1] (epistemic access, causal power, self-control, fitting intentions) constitute the relevant goals that effective human oversight must satisfy.
    Invoked in Section 2.1 and used to structure Figure 1 and Table 1; no independent justification or completeness argument is given in this paper.
  • domain assumption Human oversight can be modeled as an IT application whose components (system, communication, personnel) are attackable using the same conceptual tools as digital systems.
    Stated in the Abstract and operationalized in Section 2.2, where the definition of attack vector is extended to pathways undermining oversight effectiveness.
  • domain assumption Cybersecurity attack vector and hardening strategy classifications from the cited literature transfer to human oversight without fundamental modification.
    The paper contextualizes existing strategies such as encryption, intrusion detection, and red teaming without a formal argument that their effectiveness and semantics are preserved in the socio-technical oversight setting.
  • domain assumption Adversaries include cybercriminals, foreign governments, and AI agents that are sufficiently capable to exploit the listed vectors.
    Introduced in Section 3 without a formal adversary model or capability assessment; the threat model is qualitative.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Secure human oversight of AI: Threat modeling in a socio-technical context." pith.science (2026). https://pith.science/paper/JHED3N6F

@misc{pith2026250912290,
  author       = {Pith},
  title        = {Pith review of: Secure human oversight of AI: Threat modeling in a socio-technical context},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JHED3N6F}},
  note         = {Machine review of arXiv:2509.12290}
}
read the original abstract

Human oversight of AI is promoted as a safeguard against risks such as inaccurate outputs, system malfunctions, or violations of fundamental rights, and is mandated in regulation like the European AI Act. Yet debates on human oversight have largely focused on its effectiveness, while overlooking a critical dimension: the security of human oversight. We argue that human oversight creates a new attack surface within the safety, security, and accountability architecture of AI operations. Drawing on cybersecurity perspectives, we model human oversight as an IT application for the purpose of systematic threat modeling of the human oversight process. Threat modeling allows us to identify security risks within human oversight and points towards possible mitigation strategies. Our contributions are: (1) introducing a security perspective on human oversight, (2) offering researchers and practitioners guidance on how to approach their human oversight applications from a security point of view, and (3) providing a systematic overview of attack vectors and hardening strategies to enable secure human oversight of AI.

Figures

Figures reproduced from arXiv: 2509.12290 by the authors.

Figure 1
Figure 1. The human oversight attack surface encompasses the AI system’s technical infrastructure, the human oversight [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

97 extracted references · 5 linked inside Pith

  1. [1]

    On the quest for effectiveness in human oversight: Interdisciplinary perspectives

    Sarah Sterz, Kevin Baum, Sebastian Biewer, Holger Hermanns, Anne Lauber-Rönsberg, Philip Meinel, and Markus Langer. On the quest for effectiveness in human oversight: Interdisciplinary perspectives. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 2495–2507, 2024

  2. [2]

    The flaws of policies requiring human oversight of government algorithms.Computer Law & Security Review, 45:105681, 2022

    Ben Green. The flaws of policies requiring human oversight of government algorithms.Computer Law & Security Review, 45:105681, 2022

  3. [3]

    Johann Laux. Institutionalised distrust and human oversight of artificial intelligence: towards a democratic design of AI governance under the European Union AI Act.AI & SOCIETY, 39(6):2853–2866, December 2024

  4. [4]

    The dangers of faulty, biased, or malicious algorithms requires independent oversight

    Ben Shneiderman. The dangers of faulty, biased, or malicious algorithms requires independent oversight. Proceedings of the National Academy of Sciences, 113(48):13538–13540, November 2016

  5. [5]

    Effective human oversight of AI-Based systems: A signal detection perspective on the detection of inaccurate and unfair outputs.Minds and Machines, 2024

    Markus Langer, Kevin Baum, and Nadine Schlicker. Effective human oversight of AI-Based systems: A signal detection perspective on the detection of inaccurate and unfair outputs.Minds and Machines, 2024

  6. [6]

    ‘Human oversight’ in the EU artificial intelligence act: what, when and by whom?Law, Innovation and Technology, 15(2):508–535, July 2023

    Lena Enqvist. ‘Human oversight’ in the EU artificial intelligence act: what, when and by whom?Law, Innovation and Technology, 15(2):508–535, July 2023

  7. [7]

    The global landscape of ai ethics guidelines.Nature machine intelligence, 1(9):389–399, 2019

    Anna Jobin, Marcello Ienca, and Effy Vayena. The global landscape of ai ethics guidelines.Nature machine intelligence, 1(9):389–399, 2019

  8. [8]

    Vera Liao, and Chenhao Tan

    Vivian Lai, Chacha Chen, Alison Smith-Renner, Q. Vera Liao, and Chenhao Tan. Towards a science of human-AI decision making: An overview of design space in empirical human-subject studies. In2023 ACM Conference on Fairness, Accountability, and Transparency, pages 1369–1385, Chicago IL USA, June 2023. ACM

Show all 97 references
  1. [9]

    Does the whole exceed its parts? The effect of AI explanations on complementary team performance

    Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. Does the whole exceed its parts? The effect of AI explanations on complementary team performance. InProceedings of the 2021 CHI Conference on Human Factors in ...

  2. [10]

    Pont, and Alessandro Bozzon

    Mireia Yurrita, Himanshu Verma, Agathe Balayn, Ujwal Gadiraju, Sylvia C. Pont, and Alessandro Bozzon. Towards effective human intervention in algorithmic decision-making: Understanding the effect of decision- makers’ configuration on decision-subjects’ fairness perceptions. In...

  3. [11]

    Vera Liao, Elizabeth Anne Watkins, Carina Manger, Hal Daumé III, Andreas Riener, and Mark O Riedl

    Upol Ehsan, Philipp Wintersberger, Q. Vera Liao, Elizabeth Anne Watkins, Carina Manger, Hal Daumé III, Andreas Riener, and Mark O Riedl. Human-centered explainable ai (hcxai): Beyond opening the black-box of ai. InExtended Abstracts of the 2022 CHI Conference on Human Factors ...

  4. [12]

    Vera Liao, Daniel Gruen, and Sarah Miller

    Q. Vera Liao, Daniel Gruen, and Sarah Miller. Questioning the ai: Informing design practices for explainable ai user experiences. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems, CHI ’20, page 1–15, New York, NY , USA, 2020. Association for Compu...

  5. [13]

    Vera Liao, Ricardo Baeza-Yates, Lora Aroyo, Jess Holbrook, Ewa Luger, Michael Madaio, Ilana Golbin Blumenfeld, Maria De-Arteaga, Jessica Vitak, and Alexandra Olteanu

    Mohammad Tahaei, Marios Constantinides, Daniele Quercia, Sean Kennedy, Michael Muller, Simone Stumpf, Q. Vera Liao, Ricardo Baeza-Yates, Lora Aroyo, Jess Holbrook, Ewa Luger, Michael Madaio, Ilana Golbin Blumenfeld, Maria De-Arteaga, Jessica Vitak, and Alexandra Olteanu. Human...

  6. [14]

    Human-centered explainable ai (xai): From algorithms to user experiences

    Q Vera Liao and Kush R Varshney. Human-centered explainable ai (xai): From algorithms to user experiences. arXiv preprint arXiv:2110.10790, 2021

  7. [15]

    Optimizing decision-maker’s intrinsic motivation for effective human-ai decision-making

    Zana Buçinca. Optimizing decision-maker’s intrinsic motivation for effective human-ai decision-making. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA ’24, New York, NY , USA, 2024. Association for Computing Machinery

  8. [16]

    McBride, Wendy A

    Sara E. McBride, Wendy A. Rogers, and Arthur D. Fisk. Understanding the effect of workload on automation use for younger and older adults.Human Factors: The Journal of the Human Factors and Ergonomics Society, 53(6):672–686, December 2011

  9. [17]

    Work with ai and work for ai: Autonomous vehicle safety drivers’ lived experiences

    Mengdi Chu, Keyu Zong, Xin Shu, Jiangtao Gong, Zhicong Lu, Kaimin Guo, Xinyi Dai, and Guyue Zhou. Work with ai and work for ai: Autonomous vehicle safety drivers’ lived experiences. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New Yo...

  10. [18]

    this chatbot would never

    Joel Wester, Henning Pohl, Simo Hosio, and Niels van Berkel. "this chatbot would never...": Perceived moral agency of mental health chatbots.Proc. ACM Hum.-Comput. Interact., 8(CSCW1), April 2024

  11. [19]

    Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy.International Journal of Human–Computer Interaction, 36(6):495–504, April 2020

    Ben Shneiderman. Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy.International Journal of Human–Computer Interaction, 36(6):495–504, April 2020

  12. [20]

    Why does automation adoption in organizations remain a fallacy?: Scrutinizing practitioners’ imaginaries in an international airport

    Garoa Gomez-Beldarrain, Himanshu Verma, Euiyoung Kim, and Alessandro Bozzon. Why does automation adoption in organizations remain a fallacy?: Scrutinizing practitioners’ imaginaries in an international airport. In Proceedings of the 2025 CHI Conference on Human Factors in Comp...

  13. [21]

    The more you know: Trust dynamics and calibration in highly automated driving and the effects of take-overs, system malfunction, and system transparency

    Johannes Kraus, David Scholz, Dina Stiegemeier, and Martin Baumann. The more you know: Trust dynamics and calibration in highly automated driving and the effects of take-overs, system malfunction, and system transparency. Human Factors: The Journal of the Human Factors and Erg...

  14. [22]

    (beyond) reasonable doubt: Challenges that public defenders face in scrutinizing ai in court

    Angela Jin and Niloufar Salehi. (beyond) reasonable doubt: Challenges that public defenders face in scrutinizing ai in court. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY , USA, 2024. Association for Computing Machinery

  15. [23]

    Exploring what people need to know to be ai literate: Tailoring for a diversity of ai roles and responsibilities

    Shixian Xie, John Zimmerman, and Motahhare Eslami. Exploring what people need to know to be ai literate: Tailoring for a diversity of ai roles and responsibilities. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, New York, NY , USA, 202...

  16. [24]

    Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz

    Saleema Amershi, Dan Weld, Mihaela V orvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. Guidelines for human-AI interaction. InProceedings of the 2019 CHI Conference on ...

  17. [25]

    Markus Langer, Daniel Oster, Timo Speith, Holger Hermanns, Lena Kästner, Eva Schmidt, Andreas Sesing, and Kevin Baum. What do we want from Explainable artificial intelligence (XAI)? A stakeholder perspective on XAI and a conceptual model guiding interdisciplinary XAI research....

  18. [26]

    Explanation in artificial intelligence: Insights from the social sciences.Artificial Intelligence, 267:1–38, February 2019

    Tim Miller. Explanation in artificial intelligence: Insights from the social sciences.Artificial Intelligence, 267:1–38, February 2019

  19. [27]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making.Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1):1–21, April 2021

  20. [28]

    Chen, Afshin Nikzad, and Niloufar Salehi

    Bhada Yun, Dana Feng, Ace S. Chen, Afshin Nikzad, and Niloufar Salehi. Generative ai in knowledge work: Design implications for data navigation and decision-making. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, New York, NY , USA, 202...

  21. [29]

    Parker and Gudela Grote

    Sharon K. Parker and Gudela Grote. Automation, algorithms, and beyond: Why work design matters more than ever in a digital world.Applied Psychology, 71(4):1171–1204, February 2020

  22. [30]

    Attack surface definitions: A systematic literature review.Information and Software Technology, 104:94–103, 2018

    Christopher Theisen, Nuthan Munaiah, Mahran Al-Zyoud, Jeffrey C Carver, Andrew Meneely, and Laurie Williams. Attack surface definitions: A systematic literature review.Information and Software Technology, 104:94–103, 2018

  23. [31]

    Autonomous driving security: State of the art and challenges.IEEE Internet of Things Journal, 9(10):7572–7595, 2021

    Cong Gao, Geng Wang, Weisong Shi, Zhongmin Wang, and Yanping Chen. Autonomous driving security: State of the art and challenges.IEEE Internet of Things Journal, 9(10):7572–7595, 2021

  24. [32]

    Federated learning attack surface: taxonomy, cyber defences, challenges, and future directions.Artificial Intelligence Review, 55(5):3569–3606, 2022

    Attia Qammar, Jianguo Ding, and Huansheng Ning. Federated learning attack surface: taxonomy, cyber defences, challenges, and future directions.Artificial Intelligence Review, 55(5):3569–3606, 2022

  25. [33]

    Poison frogs! targeted clean-label poisoning attacks on neural networks.Advances in neural information processing systems, 31, 2018

    Ali Shafahi, W Ronny Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Dumitras, and Tom Goldstein. Poison frogs! targeted clean-label poisoning attacks on neural networks.Advances in neural information processing systems, 31, 2018

  26. [34]

    Badnets: Evaluating backdooring attacks on deep neural networks.Ieee Access, 7:47230–47244, 2019

    Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. Badnets: Evaluating backdooring attacks on deep neural networks.Ieee Access, 7:47230–47244, 2019

  27. [35]

    Subpopulation data poisoning attacks

    Matthew Jagielski, Giorgio Severi, Niklas Pousette Harger, and Alina Oprea. Subpopulation data poisoning attacks. InProceedings of the 2021 ACM SIGSAC conference on computer and communications security, pages 3104–3122, 2021

  28. [36]

    Adversarial attacks on neural network policies.International conference on learning representations, 2017

    Sandy Huang, Nicolas Papernot, Ian Goodfellow, Yan Duan, and Pieter Abbeel. Adversarial attacks on neural network policies.International conference on learning representations, 2017. 13 Secure Human Oversight of AI: Exploring the Attack Surface of Human Oversight

  29. [37]

    Adversarial attacks on neural networks for graph data

    Daniel Zügner, Amir Akbarnejad, and Stephan Günnemann. Adversarial attacks on neural networks for graph data. InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pages 2847–2856, 2018

  30. [38]

    Threat of adversarial attacks on deep learning in computer vision: A survey

    Naveed Akhtar and Ajmal Mian. Threat of adversarial attacks on deep learning in computer vision: A survey. Ieee Access, 6:14410–14430, 2018

  31. [39]

    Fooling lime and shap: Adversarial attacks on post hoc explanation methods

    Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. Fooling lime and shap: Adversarial attacks on post hoc explanation methods. InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pages 180–186, 2020

  32. [40]

    Fairwashing: the risk of rationalization

    Ulrich Aïvodji, Hiromi Arai, Olivier Fortineau, Sébastien Gambs, Satoshi Hara, and Alain Tapp. Fairwashing: the risk of rationalization. InInternational Conference on Machine Learning, pages 161–170. PMLR, 2019

  33. [41]

    Counterfactual explanations can be manipulated.Advances in neural information processing systems, 34:62–75, 2021

    Dylan Slack, Anna Hilgard, Himabindu Lakkaraju, and Sameer Singh. Counterfactual explanations can be manipulated.Advances in neural information processing systems, 34:62–75, 2021

  34. [42]

    Distributed denial of service attacks

    Felix Lau, Stuart H Rubin, Michael H Smith, and Ljiljana Trajkovic. Distributed denial of service attacks. InSmc 2000 conference proceedings. 2000 ieee international conference on systems, man and cybernetics. ’cybernetics evolving to systems, humans, organizations, and their ...

  35. [43]

    An analysis of using reflectors for distributed denial-of-service attacks.ACM SIGCOMM Computer Communication Review, 31(3):38–47, 2001

    Vern Paxson. An analysis of using reflectors for distributed denial-of-service attacks.ACM SIGCOMM Computer Communication Review, 31(3):38–47, 2001

  36. [44]

    A survey of man in the middle attacks.IEEE communications surveys & tutorials, 18(3):2027–2051, 2016

    Mauro Conti, Nicola Dragoni, and Viktor Lesyk. A survey of man in the middle attacks.IEEE communications surveys & tutorials, 18(3):2027–2051, 2016

  37. [45]

    The world of malware: An overview

    Anitta Patience Namanya, Andrea Cullen, Irfan U Awan, and Jules Pagna Disso. The world of malware: An overview. In2018 IEEE 6th international conference on future Internet of Things and cloud (FiCloud), pages 420–427. IEEE, 2018

  38. [46]

    Collaborative work in malware analysis: Understanding the roles and challenges of malware analysts

    Rei Yamagishi, Shota Fujii, Shingo Yasuda, Takayuki Sato, and Ayako A Hasegawa. Collaborative work in malware analysis: Understanding the roles and challenges of malware analysts. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pages 1–15, 2025

  39. [47]

    delaying

    Fatemeh Alizadeh, Gunnar Stevens, Timo Jakobi, and Jana Krüger. Catch me if you can:" delaying" as a social engineering technique in the post-attack phase.Proceedings of the ACM on Human-Computer Interaction, 7(CSCW1):1–25, 2023

  40. [48]

    The influence of context on response to spear-phishing attacks: an in-situ deception study

    Verena Distler. The influence of context on response to spear-phishing attacks: an in-situ deception study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–18, 2023

  41. [49]

    Corruption as a cybersecurity threat in the new world order.Connections: The Quarterly Journal, 20(2):75–87, 2021

    Bohdan M Holovkin, Oleksii V Tavolzhanskyi, and Oleksandr V Lysodyed. Corruption as a cybersecurity threat in the new world order.Connections: The Quarterly Journal, 20(2):75–87, 2021

  42. [50]

    Coercion in cybersecurity: What public health models reveal.Journal of Cybersecurity, 3(3):173– 183, 2017

    Steven Weber. Coercion in cybersecurity: What public health models reveal.Journal of Cybersecurity, 3(3):173– 183, 2017

  43. [51]

    The dark triad and insider threats in cyber security.Communications of the ACM, 63(12):64–80, 2020

    Michele Maasberg, Craig Van Slyke, Selwyn Ellis, and Nicole Beebe. The dark triad and insider threats in cyber security.Communications of the ACM, 63(12):64–80, 2020

  44. [52]

    Developing cyber resilient systems: a systems security engineering approach

    Ron Ross, Victoria Pillitteri, Richard Graubart, Deborah Bodeau, and Rosalie McQuaid. Developing cyber resilient systems: a systems security engineering approach. Technical report, National Institute of Standards and Technology, 2019

  45. [53]

    A taxonomy of cyber defence strategies against false data attacks in smart grids.ACM Computing Surveys, 55(14s):1–37, 2023

    Haftu Tasew Reda, Adnan Anwar, Abdun Naser Mahmood, and Zahir Tari. A taxonomy of cyber defence strategies against false data attacks in smart grids.ACM Computing Surveys, 55(14s):1–37, 2023

  46. [54]

    Automated adversary-in-the-loop cyber-physical defense planning.ACM Transactions on Cyber-Physical Systems, 7(3):1–25, 2023

    Sandeep Banik, Thiagarajan Ramachandran, Arnab Bhattacharya, and Shaunak D Bopardikar. Automated adversary-in-the-loop cyber-physical defense planning.ACM Transactions on Cyber-Physical Systems, 7(3):1–25, 2023

  47. [55]

    Prentice-Hall, Inc., 1995

    William Stallings.Network and internetwork security: principles and practice. Prentice-Hall, Inc., 1995

  48. [56]

    Reconsidering network management interfaces for communities

    Ndinelao Iitumba, Siddhant Shinde, Deysi Ortega, Naveen Bagalkot, Nervo Verdezoto, Ganief Manuel, TB Dinesh, and Melissa Densmore. Reconsidering network management interfaces for communities. InProceedings of the 4th African Human Computer Interaction Conference, pages 162–169, 2023

  49. [57]

    CRC press, 2018

    Alfred J Menezes, Paul C Van Oorschot, and Scott A Vanstone.Handbook of applied cryptography. CRC press, 2018. 14 Secure Human Oversight of AI: Exploring the Attack Surface of Human Oversight

  50. [58]

    The impact of risk appeal approaches on users’ sharing confidential information

    Elham Al Qahtani, Peter Story, and Mohamed Shehab. The impact of risk appeal approaches on users’ sharing confidential information. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems, pages 1–21, 2024

  51. [59]

    Survey of intrusion detection systems: techniques, datasets and challenges.Cybersecurity, 2(1):1–22, 2019

    Ansam Khraisat, Iqbal Gondal, Peter Vamplew, and Joarder Kamruzzaman. Survey of intrusion detection systems: techniques, datasets and challenges.Cybersecurity, 2(1):1–22, 2019

  52. [60]

    Transformers and large language models for efficient intrusion detection systems: A comprehen- sive survey.Information Fusion, page 103347, 2025

    Hamza Kheddar. Transformers and large language models for efficient intrusion detection systems: A comprehen- sive survey.Information Fusion, page 103347, 2025

  53. [61]

    Apress, 2019

    Jacob G Oakley.Professional red teaming: conducting successful cybersecurity engagements. Apress, 2019

  54. [62]

    The human factor in ai red teaming: Perspectives from social and collaborative computing

    Alice Qian Zhang, Ryland Shaw, Jacy Reese Anthis, Ashlee Milton, Emily Tseng, Jina Suh, Lama Ahmad, Ram Shankar Siva Kumar, Julian Posada, Benjamin Shestakofsky, et al. The human factor in ai red teaming: Perspectives from social and collaborative computing. InCompanion Public...

  55. [63]

    John Wiley & Sons, 2023

    Chris Hughes and Tony Turner.Software Transparency: supply chain security in an era of a software-driven society. John Wiley & Sons, 2023

  56. [64]

    security by obscurity

    Peter Hall, Olivia Mundahl, and Sunoo Park. The pitfalls of “security by obscurity” and what they mean for transparent ai. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 28042–28051, 2025

  57. [65]

    A comprehensive cybersecurity audit model to improve cybersecurity assurance: The cybersecurity audit model (csam)

    Regner Sabillon, Jordi Serra-Ruiz, Victor Cavaller, and Jeimy Cano. A comprehensive cybersecurity audit model to improve cybersecurity assurance: The cybersecurity audit model (csam). In2017 International Conference on Information Systems and Computer Science (INCISCOS), pages...

  58. [66]

    Governing ai safety through independent audits.Nature Machine Intelligence, 3(7):566–571, 2021

    Gregory Falco, Ben Shneiderman, Julia Badger, Ryan Carrier, Anton Dahbura, David Danks, Martin Eling, Alwyn Goodloe, Jerry Gupta, Christopher Hart, et al. Governing ai safety through independent audits.Nature Machine Intelligence, 3(7):566–571, 2021

  59. [67]

    Train as you fight: Evaluating authentic cybersecurity training in cyber ranges

    Magdalena Glas, Manfred Vielberth, and Guenther Pernul. Train as you fight: Evaluating authentic cybersecurity training in cyber ranges. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–19, 2023

  60. [68]

    Understanding and improving user adoption and security awareness in password checkup services

    Sanghak Oh, Heewon Baek, Jun Ho Huh, Taeyoung Kim, Woojin Jeon, Ian Oakley, and Hyoungshick Kim. Understanding and improving user adoption and security awareness in password checkup services. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pages...

  61. [69]

    Concrete problems in ai safety.arXiv preprint arXiv:1606.06565, 2016

    Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. Concrete problems in ai safety.arXiv preprint arXiv:1606.06565, 2016

  62. [70]

    Frontier models are capable of in-context scheming.arXiv preprint arXiv:2412.04984, 2024

    Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. Frontier models are capable of in-context scheming.arXiv preprint arXiv:2412.04984, 2024

  63. [71]

    Balancing transparency and risk: An overview of the security and privacy risks of open-source machine learning models

    Dominik Hintersdorf, Lukas Struppek, and Kristian Kersting. Balancing transparency and risk: An overview of the security and privacy risks of open-source machine learning models. InInternational Conference on Bridging the Gap between AI and Reality, pages 269–283. Springer Nat...

  64. [72]

    Generative ai models: Opportunities and risks for industry and authorities.arXiv preprint arXiv:2406.04734, 2024

    Tobias Alt, Andrea Ibisch, Clemens Meiser, Anna Wilhelm, Raphael Zimmer, Jonas Ditz, Dominique Dresen, Christoph Droste, Jens Karschau, Friederike Laus, et al. Generative ai models: Opportunities and risks for industry and authorities.arXiv preprint arXiv:2406.04734, 2024

  65. [73]

    Invisible backdoor attack with sample-specific triggers

    Yuezun Li, Yiming Li, Baoyuan Wu, Longkang Li, Ran He, and Siwei Lyu. Invisible backdoor attack with sample-specific triggers. InProceedings of the IEEE/CVF international conference on computer vision, pages 16463–16472, 2021

  66. [74]

    Explainable artificial intelligence in an adversarial context

    Jonas Ditz and Matthias Heck. Explainable artificial intelligence in an adversarial context. Technical report, Federal Office for Information Security, 2025

  67. [75]

    Bridging the transparency gap: What can explainable ai learn from the ai act? InECAI 2023, pages 964–971

    Balint Gyevnar, Nick Ferguson, and Burkhard Schafer. Bridging the transparency gap: What can explainable ai learn from the ai act? InECAI 2023, pages 964–971. IOS Press, 2023

  68. [76]

    Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai.Information Fusion, 58:82–115, June 2020

    Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador Garcia, Sergio Gil-Lopez, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera. Explainable artificial intelligence (xai): Concepts, taxo...

  69. [77]

    Sides: Separating idealization from deceptive’explanations’ in xai

    Emily Sullivan. Sides: Separating idealization from deceptive’explanations’ in xai. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 1714–1724, 2024. 15 Secure Human Oversight of AI: Exploring the Attack Surface of Human Oversight

  70. [78]

    Upol Ehsan and Mark O. Riedl. Explainability pitfalls: Beyond dark patterns in explainable ai.Patterns, 5(6):100971, June 2024

  71. [79]

    Ddos attacks in cloud computing: Issues, taxonomy, and future directions.Computer Communications, 107:30–48, July 2017

    Gaurav Somani, Manoj Singh Gaur, Dheeraj Sanghi, Mauro Conti, and Rajkumar Buyya. Ddos attacks in cloud computing: Issues, taxonomy, and future directions.Computer Communications, 107:30–48, July 2017

  72. [80]

    Dos- and ddos attacks, 2022

    Bundesamt für Sicherheit in der Informationstechnik. Dos- and ddos attacks, 2022. [Online; accessed 8-July-2025]

  73. [81]

    Before we knew it: an empirical study of zero-day attacks in the real world

    Leyla Bilge and Tudor Dumitra¸ s. Before we knew it: an empirical study of zero-day attacks in the real world. In Proceedings of the 2012 ACM Conference on Computer and Communications Security, CCS ’12, page 833–844, New York, NY , USA, 2012. Association for Computing Machinery

  74. [82]

    What is malware?, 2024

    Bundesamt für Sicherheit in der Informationstechnik. What is malware?, 2024. [Online; accessed 16-July-2025]

  75. [83]

    Position is power: System prompts as a mechanism of bias in large language models (llms)

    Anna Neumann, Elisabeth Kirsten, Muhammad Bilal Zafar, and Jatinder Singh. Position is power: System prompts as a mechanism of bias in large language models (llms). InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, pages 573–598, 2025

  76. [84]

    Advanced social engineering attacks

    Katharina Krombholz, Heidelinde Hobel, Markus Huber, and Edgar Weippl. Advanced social engineering attacks. Journal of Information Security and Applications, 22:113–122, June 2015

  77. [85]

    Social engineering - the human factor, 2022

    Bundesamt für Sicherheit in der Informationstechnik. Social engineering - the human factor, 2022. [Online; accessed 11-July-2025]

  78. [86]

    Gheyas and Ali E

    Iffat A. Gheyas and Ali E. Abdallah. Detection and prediction of insider threats to cyber security: a systematic literature review and meta-analysis.Big Data Analytics, 1(1), August 2016

  79. [87]

    Distributed denial of service (ddos) resilience in cloud: Review and conceptual cloud ddos mitigation framework.Journal of Network and Computer Applications, 67:147–165, 2016

    Opeyemi Osanaiye, Kim-Kwang Raymond Choo, and Mqhele Dlodlo. Distributed denial of service (ddos) resilience in cloud: Review and conceptual cloud ddos mitigation framework.Journal of Network and Computer Applications, 67:147–165, 2016

  80. [88]

    Ddos attack detection and mitigation using sdn: methods, practices, and solutions.Arabian Journal for Science and Engineering, 42(2):425–441, 2017

    Narmeen Zakaria Bawany, Jawwad A Shamsi, and Khaled Salah. Ddos attack detection and mitigation using sdn: methods, practices, and solutions.Arabian Journal for Science and Engineering, 42(2):425–441, 2017

  81. [89]

    Intrusion detection systems, 2001

    Rebecca Gurley Bace, Peter Mell, et al. Intrusion detection systems, 2001

  82. [90]

    Cryptool 2.0: Open-source kryptologie für jedermann.Datenschutz und Datensicherheit-DuD, 38(10):701–708, 2014

    Nils Kopal, Olga Kieselmann, Arno Wacker, and Bernhard Esslinger. Cryptool 2.0: Open-source kryptologie für jedermann.Datenschutz und Datensicherheit-DuD, 38(10):701–708, 2014

  83. [91]

    Cryptographic mechanisms: Recommendations and key lengths (bsi tr-02102-1), 2025

    Bundesamt für Sicherheit in der Informationstechnik. Cryptographic mechanisms: Recommendations and key lengths (bsi tr-02102-1), 2025

  84. [92]

    Information security, cybersecurity and privacy protection — information security management systems — requirements, 2022

    International Organization for Standardization. Information security, cybersecurity and privacy protection — information security management systems — requirements, 2022

  85. [93]

    The best form of defence–the benefits of red teaming.Computer Fraud & Security, 2018(10):8–12, 2018

    Steve Mansfield-Devine. The best form of defence–the benefits of red teaming.Computer Fraud & Security, 2018(10):8–12, 2018

  86. [94]

    Social engineering defence mechanisms and counteracting training strategies.Information and Computer Security, 25(2):206–222, June 2017

    Peter Schaab, Kristian Beckers, and Sebastian Pape. Social engineering defence mechanisms and counteracting training strategies.Information and Computer Security, 25(2):206–222, June 2017

  87. [95]

    Automation bias in the ai act: On the legal implications of attempting to de-bias human oversight of ai.European Journal of Risk Regulation, page 1–16, July 2025

    Johann Laux and Hannah Ruschemeier. Automation bias in the ai act: On the legal implications of attempting to de-bias human oversight of ai.European Journal of Risk Regulation, page 1–16, July 2025

  88. [96]

    An approach to technical agi safety and security.arXiv preprint arXiv:2504.01849, 2025

    Rohin Shah, Alex Irpan, Alexander Matt Turner, Anna Wang, Arthur Conmy, David Lindner, Jonah Brown-Cohen, Lewis Ho, Neel Nanda, Raluca Ada Popa, et al. An approach to technical agi safety and security.arXiv preprint arXiv:2504.01849, 2025

  89. [97]

    Anita Taylor, Stephanie E

    Bernd Marcus, O. Anita Taylor, Stephanie E. Hastings, Alexandra Sturm, and Oliver Weigelt. The structure of counterproductive work behavior: A review, a structural meta-analysis, and a primary study.Journal of Management, 42(1):203–233, September 2013. 16

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.