REVIEW 3 major objections 6 minor 1 cited by
Safety Co-Option and Compromised National Security: The Self-Fulfilling Prophecy of Weakened AI Risk Thresholds
T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper argues that AI technologists are engaging in 'safety revisionism' — redefining safety-engineering terms so that foundation models can enter military use without meeting established risk thresholds — and that this will weaken US…
desk verdict A pointed, well-sourced policy argument about 'safety revisionism' in AI defense work, though it overreaches by asserting without proof that traditional risk thresholds apply to foundation models. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the risk threshold — a quantified level of risk exposure above which action must be taken — together with the linguistic drift that detaches it from assurance practice. On one side stands the historical framework in which thresholds are fixed by societal deliberation and then substantiated by safety cases: structured arguments, backed by evidence, that a system is acceptably safe for a defined application in a defined environment. On the other side stand the contemporary 'frontier safety frameworks' that convert safety cases into capability targets and substitute red-teaming for boundary testing, with no threshold to test against. The comparison between those two sides is the gear that turns missing AI governance into lowered real-world safety for defense and, through precedent, for civilian critical infrastructure.
What would settle it
The central claim would fail if a military AI system that passed a current capabilities-evaluation framework was then assessed under a standard military system-safety analysis and met the same quantified risk thresholds required of the conventional system it replaces, such as a probability of a fatal hazardous event per deployment at or below the legacy-system threshold. It would also be weakened by a documented case in which an AI 'red-teaming' exercise uncovered a previously unknown hazard and measurably reduced estimated risk, showing that the redefined practice can carry real assurance weight.
Extended reading notes
Core claim
The core discovery is a mechanism with a name: 'safety revisionism.' When societally accepted risk tolerances for AI do not exist, technologists become the de facto arbiters of how much harm is acceptable, and they exercise that power by quietly replacing the vocabulary of safety engineering. Risk thresholds that were once set through democratic deliberation — such as the quantified fatal-risk criteria derived for nuclear power or the Safety Integrity Levels used in defense system safety — are superseded by 'frontier safety frameworks' that state unverifiable goals about model intent, such as not causing 'catastrophe,' without any numeric threshold a system can be tested against. The paper shows that the methodologies meant to substantiate safety, particularly safety cases, have been redefined into 'affirmative cases' and red-teaming exercises that cannot demonstrate risk reduction. The conclusion is that this trajectory is a self-fulfilling prophecy: the weakened thresholds adopted to preserve a supposed AI advantage are exactly what will compromise the safety and security of the military systems that depend on them.
Load-bearing premise
The argument assumes that the risk thresholds developed for earlier safety-critical technologies — nuclear power plants, military hardware — apply to AI-based military systems and that foundation models are 'no exception'; if AI failure modes are so different that those thresholds cannot meaningfully be transferred, the charge that AI firms are 'revising' safety loses its force.
Editorial extensions
If this is right
- Military evaluation frameworks built on 'capabilities evaluations' cannot demonstrate that a foundation model is acceptably safe, because they never commit to a quantified risk threshold.
- Foundation models in targeting and decision-support roles will carry documented failure modes — poor performance on non-English data, supply-chain attack surfaces, brittle safeguards — into settings where those failures can kill civilians and violate international humanitarian law.
- If defense adoption sets the precedent, civilian safety-critical uses of AI will be judged against the same lowered bar, since military and civilian assurance standards influence each other.
- Democratic bodies, not technologists, need to set explicit AI risk tolerances, including the number of societally accepted fatalities a given deployment assumes.
- The national-security justification for accelerated adoption is self-defeating: a brittle AI system is an asset adversaries can exploit, so weakening thresholds to win an AI race undermines the very advantage the race is meant to secure.
Reading between the lines
- Going beyond the paper: its logic yields a testable prediction — among militaries that field AI-enabled targeting and decision support, those with weaker assurance thresholds will show more civilian casualties and friendly-fire incidents, then loosen thresholds further to keep pace.
- A practical audit rule follows: any AI safety framework that cannot state a numeric threshold (for example, a maximum probability of a fatal hazardous event per deployment) is not a safety framework in the engineering sense, whatever its name; applying that test to current policy documents would give governance bodies a concrete checklist.
- The same capabilities-evaluation logic is already visible outside defense, in proposals to run government administration with foundation models, so the missing-threshold problem is broader than the military case the paper examines.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is an argumentative policy analysis arguing that, because no societally deliberated risk thresholds have been set for AI, AI technologists—mainly industry labs and self-styled 'AI safety' organizations—have been able to redefine safety terminology and substitute 'capabilities' and 'alignment' evaluations for traditional safety-engineering assurance. The authors reconstruct the history of risk thresholds from Chauncey Starr's nuclear-era work through MIL-STD-882 and Safety Integrity Levels, then argue that current AI evaluation frameworks (Google DeepMind's Frontier Safety Framework, Anthropic's Responsible Scaling Policy, Task Force Lima, the UK AI Security Institute) hollow out the meaning of 'safety', 'safety case', and 'red-teaming'. They use the Gaza deployments and supply-chain vulnerabilities to argue that the resulting absence of enforceable thresholds enables deployment of foundation models in military contexts at lowered safety and security thresholds, ultimately undermining US national security and international humanitarian law. The paper concludes by calling for democratic deliberation to set AI risk thresholds and for preserving traditional TEVV standards in military AI evaluation.
Significance. If the argument holds, the paper makes a valuable contribution by connecting AI safety discourse to the history and vocabulary of safety engineering and by identifying a concrete mechanism—terminological revisionism—through which existing assurance standards can be bypassed. Its strengths are the historically grounded account of risk thresholds, the concrete documentation of military AI deployment failures, and a clear normative proposal: preserve established assurance frameworks and set AI risk thresholds through democratic deliberation rather than industry self-assessment. The paper is not an empirical study and does not ship machine-checked proofs or reproducible code; its contribution is conceptual and evidentiary. The central claim is defensible but requires an important qualification: the transferability of quantitative risk thresholds to foundation models is asserted rather than demonstrated, and the national-security conclusion is hedged in some places but stated categorically in others.
major comments (3)
- [Section 2 / Table 1] The paper's central charge of 'lowered' risk thresholds presupposes that established quantitative thresholds (Starr's 10^-4 deaths per person per year, MIL-STD-882E categories, SIL levels) can be populated for foundation models. The paper asserts this ('their risk categories fall firmly within the scope of traditional risk frameworks') but never shows how a probability of a hazardous event can be assigned to a context-dependent, nondeterministic model at assurance time. Without such a demonstration, the 'lowered threshold' claim is not established; the situation may instead be that no applicable threshold has been operationalized. This distinction is load-bearing for the 'revisionism' charge, because mislabeling an absence of a standard as a weakening of a standard changes the normative force of the argument. Please either provide a worked example for a concrete use case (e.g., target nomination or intelligence analysis) or reframe the argument as 'no applicable risk threshold has been operationalized'.
- [Section 4, final paragraph] The claim that AI systems are 'deployed at levels far below the risk thresholds that would be deemed acceptable through standardized safety processes' is not supported by a quantitative comparison. The cited evidence—Arabic mistranslation, cell-tower-based civilian casualty estimates, and evasion failures—demonstrates qualitative failures, but the paper provides no probability-of-failure estimates, exposure analysis, or comparison against any specific threshold. As stated, the claim is not falsifiable. It should either be quantified for a specific deployment scenario or explicitly presented as a qualitative judgment about the absence of demonstrated compliance.
- [Sections 1 and 5] The title's 'self-fulfilling prophecy' and the national-security conclusion are causal claims that the evidence does not fully support. The paper shows that some military AI uses have failed and that some frameworks use nonstandard terminology, but it does not systematically consider alternative explanations for the observed outcomes, such as bureaucratic incentives, technical immaturity, or good-faith disagreements about how to adapt assurance methods to AI. The hedging in Section 1 ('may be precisely what disadvantages') is honest, but the categorical conclusion in Section 5 ('will imperil' and 'will result in a significant civilian death toll') goes beyond the evidence presented. Please separate the well-supported claim that current frameworks do not satisfy traditional assurance standards from the more speculative claim that this trajectory will compromise US national security.
minor comments (6)
- [Section 2] The paper defines risk tolerance and risk threshold as distinct concepts but later uses them interchangeably (e.g., Section 3: 'risk thresholds are derived from risk tolerances'). Please maintain the distinction consistently.
- [Section 3] The statement that a safety case 'is not intended to produce any 'targets' or thresholds' is too categorical; some safety-case methodologies incorporate quantitative safety requirements as part of the argument. The substantive criticism of Google DeepMind's framework does not depend on this categorical claim, so it can be softened without harming the argument.
- [Section 3.1] The characterization of Task Force Lima as taking a 'capabilities evaluation' approach relies on the task force's executive summary [55]. Please verify the citation and, if possible, quote the relevant language, since the executive summary may not be publicly accessible.
- [Section 4] The claim that foundation models have 'poor accuracy with non-English languages, especially for Arabic' is supported by general LLM cultural-bias studies [61, 62]; the connection to the specific IDF systems described in [9] should be made more explicit, since those systems may include non-LLM components.
- [Section 5] The phrase 'existential risks that are very real and present' sits awkwardly with the paper's earlier rejection of 'speculative existential risks' in Sections 1 and 2. Please clarify whether this refers to concrete military harms rather than the speculative AI-existential-risk literature.
- [Throughout] Please standardize spelling and formatting inconsistencies, including 'DOD' vs 'DoD' and 'LAWs' vs 'LAWS', and check the column headings in Table 1, which appear truncated.
Circularity Check
No significant circularity: the paper is a standards-based policy critique whose argument is externally anchored, with no fitted parameters or self-referential derivation.
full rationale
The paper does not present a formal derivation, a fitted model, or a predictive claim that reduces to its own inputs by construction. Its central charge—that AI technologists have engaged in 'safety revisionism' by substituting ill-defined safety terminology for established frameworks—is argued by comparison to external standards (MIL-STD-882E, Safety Integrity Levels, Starr's risk analysis, TEVV) and to documented deployment failures (e.g., IDF systems, red-teaming brittleness). The authors' self-citations, such as [50], [51], [52], and [53], are used as supporting background references on AI risk and military proliferation rather than as load-bearing uniqueness theorems or as definitions that presuppose the conclusion. The closest thing to a potentially circular premise is the assertion that foundation models' failure modes 'fall firmly within the scope of traditional risk frameworks' (Section 2), but this is an unproven assumption about applicability, not a self-referential derivation; it could be challenged as a correctness or evidence concern, not as circularity. No equation, threshold, or evaluation result is shown to be equivalent to its own input, and no fitted parameter is renamed as a prediction. The paper is self-contained as a policy argument and scores 0 on circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Societal risk thresholds for technology should be determined through democratic deliberation rather than by technologists.
- domain assumption Established safety-engineering frameworks (e.g., MIL-STD-882, SIL, safety cases) are applicable to AI-based military systems and should be preserved.
- domain assumption The cited examples of AI military failure (e.g., IDF targeting in Gaza) are representative and accurately reported.
Cite this review
Pith. "Pith review of Safety Co-Option and Compromised National Security: The Self-Fulfilling Prophecy of Weakened AI Risk Thresholds." pith.science (2026). https://pith.science/paper/UJWTMH75
@misc{pith2026250415088,
author = {Pith},
title = {Pith review of: Safety Co-Option and Compromised National Security: The Self-Fulfilling Prophecy of Weakened AI Risk Thresholds},
year = {2026},
howpublished = {\url{https://pith.science/paper/UJWTMH75}},
note = {Machine review of arXiv:2504.15088}
}
read the original abstract
Risk thresholds provide a measure of the level of risk exposure that a society or individual is willing to withstand, ultimately shaping how we determine the safety of technological systems. Against the backdrop of the Cold War, the first risk analyses, such as those devised for nuclear systems, cemented societally accepted risk thresholds against which safety-critical and defense systems are now evaluated. But today, the appropriate risk tolerances for AI systems have yet to be agreed on by global governing efforts, despite the need for democratic deliberation regarding the acceptable levels of harm to human life. Absent such AI risk thresholds, AI technologists-primarily industry labs, as well as "AI safety" focused organizations-have instead advocated for risk tolerances skewed by a purported AI arms race and speculative "existential" risks, taking over the arbitration of risk determinations with life-or-death consequences, subverting democratic processes. In this paper, we demonstrate how such approaches have allowed AI technologists to engage in "safety revisionism," substituting traditional safety methods and terminology with ill-defined alternatives that vie for the accelerated adoption of military AI uses at the cost of lowered safety and security thresholds. We explore how the current trajectory for AI risk determination and evaluation for foundation model use within national security is poised for a race to the bottom, to the detriment of the US's national security interests. Safety-critical and defense systems must comply with assurance frameworks that are aligned with established risk thresholds, and foundation models are no exception. As such, development of evaluation frameworks for AI-based military systems must preserve the safety and security of US critical and defense infrastructure, and remain in alignment with international humanitarian law.
Forward citations
Cited by 1 Pith paper
-
Harmonizing AI Safety Thresholds
The authors propose harmonized AI capability floors: non-zero full-chain TLO cyber completion triggers safeguards, and AI progress at 5× trend for 3 months triggers safeguards, with biorisk left as a diagnostic.
Reference graph
Works this paper leans on
-
[1]
‘Lavender’: The AI machine directing Israel’s bombing spree in Gaza, April 2024
Yuval Abraham. ‘Lavender’: The AI machine directing Israel’s bombing spree in Gaza, April 2024. URL https://www.972mag.com/ lavender-ai-israeli-army-gaza/
2024
-
[2]
Canada launches first AI strategy for federal public service - Global Government Forum, April 2025
Jack Aldane. Canada launches first AI strategy for federal public service - Global Government Forum, April 2025. URL https: //www.globalgovernmentforum.com/canada-launches-first-ai-strategy-for-federal-public-service/
2025
-
[3]
Trump Can Keep America’s AI Advantage
Dario Amodei and Matt Pottinger. Trump Can Keep America’s AI Advantage. Wall Street Journal , January 2025. URL https: //www.wsj.com/opinion/trump-can-keep-americas-ai-advantage-china-chips-data-eccdce91
2025
-
[4]
Frontier AI Regulation: Managing Emerging Risks to Public Safety
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, and Kevin Wolf. Fr...
-
[5]
Anduril Partners with OpenAI to Advance U.S
Anduril. Anduril Partners with OpenAI to Advance U.S. Artificial Intelligence Leadership and Protect U.S. and Allied Forces, December
-
[6]
Anthropic’s Responsible Scaling Policy, September 2023
Anthropic. Anthropic’s Responsible Scaling Policy, September 2023. URL https://www.anthropic.com/news/anthropics-responsible- scaling-policy
2023
-
[7]
Security-informed safety, January 2023
National Protective Security Authority. Security-informed safety, January 2023. URL https://www.npsa.gov.uk/security-informed- safety
2023
-
[8]
Why policy makers should beware claims of new ‘arms races’
Haydn Belfield and Christian Ruhl. Why policy makers should beware claims of new ‘arms races’. Bulletin of the Atomic Scientists , July
Show all 106 references
-
[9]
As Israel uses US-made AI models in war, concerns arise about tech’s role in who lives and who dies
Michael Biesecker, Sam Mednick, and Garance Burke. As Israel uses US-made AI models in war, concerns arise about tech’s role in who lives and who dies. AP News , February 2025. URL https://apnews.com/article/israel-palestinians-ai-technology- 737bc17af7b03e98c29cec4e15d0f108
2025
- [10]
- [11]
-
[12]
How a billionaire-backed network of AI advisers took over Washington
Brendan Bordelon. How a billionaire-backed network of AI advisers took over Washington. POLITICO, October 2023. URL https://www.politico.com/news/2023/10/13/open-philanthropy-funding-ai-policy-00121362
2023
-
[13]
Key Congress staffers in AI debate are funded by tech giants like Google and Microsoft
Brendan Bordelon. Key Congress staffers in AI debate are funded by tech giants like Google and Microsoft. POLITICO, March 2023. URL https://www.politico.com/news/2023/12/03/congress-ai-fellows-tech-companies-00129701
2023
-
[14]
The Risks of Integrating Generative AI into Weapon Systems
Vincent Boulanin. The Risks of Integrating Generative AI into Weapon Systems. Technical report, GC REAIM Expert Policy Note Series, April 2025
2025
-
[15]
Michael Brenes and William D. Hartung. Private Finance and the Quest to Remake Modern Warfare, June 2024. URL https: //quincyinst.org/research/private-finance-and-the-quest-to-remake-modern-warfare/
2024
-
[16]
Britain dances to JD Vance’s tune as it renames AI institute
Tom Bristow. Britain dances to JD Vance’s tune as it renames AI institute. POLITICO, February 2025. URL https://www.politico.eu/ article/jd-vance-britain-ai-safety-institute-aisi-security/
2025
- [17]
-
[18]
Extracting training data from diffusion models
Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, and Eric Wallace. Extracting training data from diffusion models. In Proceedings of the 32nd USENIX Conference on Security Symposium (SEC ’23). USENIX Ass...
2023
-
[19]
Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr
Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr. Poisoning Web-Scale Training Datasets is Practical. In 2024 IEEE Symposium on Security and Privacy (SP), pages 407–4...
2024
-
[20]
Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Itay Yona, Eric Wallace, David Rolnick, and Florian Tramèr
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Itay Yona, Eric Wallace, David Rolnick, and Florian Tramèr. Stealing Part of a Production Language Model. ...
2024 arXiv
-
[21]
ÓhÉigeartaigh
Stephen Cave and Seán S. ÓhÉigeartaigh. An AI Race for Strategic Advantage: Rhetoric and Risks. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society , pages 36–40, New Orleans LA USA, December 2018. ACM. ISBN 9781450360128. doi: 10.1145/3278721.3278780. UR...
2018
-
[22]
Brian J. Chen. Dispelling Myths of AI and Efficiency. Data & Society, 2025. URL https://datasociety.net/library/dispelling-myths-of-ai- and-efficiency/?trk=feed_main-feed-card_feed-article-content
2025
- [23]
-
[24]
Nuclear Regulatory Commission
U.S. Nuclear Regulatory Commission. Safety Goals for Nuclear Power Plant Operation. U.S. Nuclear Regulatory Commission, May
-
[25]
Marine Corps Warfighting Laboratory: Experiment Division
The United States Marine Corps. Marine Corps Warfighting Laboratory: Experiment Division. URL https://www.mcwl.marines.mil/ Divisions/Experiment/
-
[26]
The AI that Wasn’t There: Global Order and the (Mis)Perception of Powerful AI
Mary (Missy) Cummings. The AI that Wasn’t There: Global Order and the (Mis)Perception of Powerful AI. Policy Roundtable: Artificial Intelligence and International Security, Texas National Security Review , June 2020. URL https://tnsr.org/roundtable/policy-roundtable- artificia...
2020
-
[27]
Cummings and Ben Bauchwitz
M.L. Cummings and Ben Bauchwitz. Identifying Research Gaps through Self-Driving Car Data Analysis.IEEE Transactions on Intelligent Vehicles, pages 1–10, 2024. ISSN 2379-8904. doi: 10.1109/TIV.2024.3506936. URL https://ieeexplore.ieee.org/document/10778107/ ?arnumber=10778107
2024
-
[28]
Introducing the Frontier Safety Framework, April 2025
Google DeepMind. Introducing the Frontier Safety Framework, April 2025. URL https://deepmind.google/discover/blog/introducing- the-frontier-safety-framework/
2025
-
[29]
Behind China’s Plans to Build AI for the World
Bill Drexel and Hannah Kelley. Behind China’s Plans to Build AI for the World. POLITICO, November 2023. URL https://www.politico. com/news/magazine/2023/11/30/china-global-ai-plans-00129160
2023
-
[30]
Israel built an ‘AI factory’ for war
Elizabeth Dwoskin. Israel built an ‘AI factory’ for war. It unleashed it in Gaza. The Washington Post , December 2024. URL https://www.washingtonpost.com/technology/2024/12/29/ai-israel-war-gaza-idf/
2024
-
[31]
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfi...
2022 arXiv
-
[32]
Paul M. Grant. Chauncey Starr (1912–2007). Nature, 447(7146):789–789, June 2007. ISSN 1476-4687. doi: 10.1038/447789a. URL https://www.nature.com/articles/447789a
1912 doi
-
[33]
AI Safety: Navigating the Expanding Landscape of Potential Harms
Ibrahim Habli and John Alexander McDermid. AI Safety: Navigating the Expanding Landscape of Potential Harms. Safety-Critical Systems Club Newsletter, May 2024. URL https://eprints.whiterose.ac.uk/id/eprint/213035/
2024
- [34]
-
[35]
The Artificial Intelligence Arms Race: Trends and World Leaders in Autonomous Weapons Development
Justin Haner and Denise Garcia. The Artificial Intelligence Arms Race: Trends and World Leaders in Autonomous Weapons Development. Global Policy, 10(3):331–337, September 2019. ISSN 1758-5880, 1758-5899. doi: 10.1111/1758-5899.12713. URL https://onlinelibrary. wiley.com/doi/10...
2019
-
[36]
The Nuclear-Level Risk of Superintelligent AI
Dan Hendrycks and Eric Schmidt. The Nuclear-Level Risk of Superintelligent AI. TIME, March 2025. URL https://time.com/7265056/ nuclear-level-risk-of-superintelligent-ai/
2025
-
[37]
Read: JD Vance’s full speech on AI and the EU.The Spectator, February 2025
Coffee House. Read: JD Vance’s full speech on AI and the EU.The Spectator, February 2025. URL https://www.spectator.co.uk/article/read- jd-vances-full-speech-on-ai-and-the-eu/
2025
-
[38]
Tackling AI Security Risks to Unleash Growth and Deliver Plan for Change
AI Security Institute. Tackling AI Security Risks to Unleash Growth and Deliver Plan for Change. Department for Science, Innovation and Technology, February 2025. URL https://www.gov.uk/government/news/tackling-ai-security-risks-to-unleash-growth-and-deliver- plan-for-change
2025
-
[39]
Safety Cases at AISI
Geoffrey Irving. Safety Cases at AISI. AI Security Institute, August 2024. URL https://www.aisi.gov.uk/work/safety-cases-at-aisi
2024
-
[40]
The Fifth Branch
Sheila Jasanoff. The Fifth Branch. Harvard University Press, 1994. URL https://www.hup.harvard.edu/books/9780674300620
1994
-
[41]
Containing the Atom: Sociotechnical Imaginaries and Nuclear Power in the United States and South Korea
Sheila Jasanoff and Sang-Hyun Kim. Containing the Atom: Sociotechnical Imaginaries and Nuclear Power in the United States and South Korea. Minerva, 47(2):119–146, June 2009. ISSN 0026-4695, 1573-1871. doi: 10.1007/s11024-009-9124-4. URL http: //link.springer.com/10.1007/s11024...
2009 doi
-
[42]
Under the radar? examining the evaluation of foundation models
Elliot Jones, Mahi Hardalupas, and William Agnew. Under the radar? examining the evaluation of foundation models. Ada Lovelace Institute, July 2024. URL https://www.adalovelaceinstitute.org/wp-content/uploads/2024/09/Ada-Lovelace-Institute-Under-the-radar- 230924.pdf
2024
-
[43]
2023 Landscape: Confronting Tech Power
Amba Kak and Sarah Myers West. 2023 Landscape: Confronting Tech Power. AI Now Institute, April 2023. URL https://ainowinstitute. org/wp-content/uploads/2023/04/AI-Now-2023-Landscape-Report-FINAL.pdf
2023
-
[44]
Nidhi Kalra and Susan M. Paddock. Driving to Safety: How Many Miles of Driving Would It Take to Demonstrate Autonomous Vehicle Reliability? Technical report, RAND Corporation, April 2016. URL https://www.rand.org/pubs/research_reports/RR1478.html
2016
-
[45]
Measurement challenges in AI catastrophic risk governance and safety frameworks
Atoosa Kasirzadeh. Measurement challenges in AI catastrophic risk governance and safety frameworks. Tech Policy Press, Sep 2024. URL https://www.techpolicy.press/measurement-challenges-in-ai-catastrophic-risk-governance-and-safety-frameworks/
2024
-
[46]
Dependability analysis of safety critical systems: Issues and challenges
Raj kamal Kaur, Babita Pandey, and Lalit Kumar Singh. Dependability analysis of safety critical systems: Issues and challenges. Annals of Nuclear Energy, 120:127–154, October 2018. ISSN 0306-4549. doi: 10.1016/j.anucene.2018.05.027. URL https://www.sciencedirect. com/science/a...
2018 doi
-
[47]
Doomed to Repeat History? Lessons from the Crypto Wars of the 1990s
Danielle Kehl, Andi Wilson, and Kevin Bankston. Doomed to Repeat History? Lessons from the Crypto Wars of the 1990s. URL http: //newamerica.org/cybersecurity-initiative/policy-papers/doomed-to-repeat-history-lessons-from-the-crypto-wars-of-the-1990s/
-
[48]
Elon Musk Ally Tells Staff ’AI-First’ Is the Future of Key Government Agency
Makena Kelly. Elon Musk Ally Tells Staff ’AI-First’ Is the Future of Key Government Agency. Wired, 2025. ISSN 1059-1028. URL https://www.wired.com/story/elon-musk-lieutenant-gsa-ai-agency/
2025
-
[49]
A Systematic Approach to Safety Case Management
Tim Kelly. A Systematic Approach to Safety Case Management. SAE Transactions, 113:257–266, 2004. ISSN 0096-736X. URL https://www.jstor.org/stable/44699541
2004
-
[50]
How AI Can Be Regulated Like Nuclear Energy
Heidy Khlaaf. How AI Can Be Regulated Like Nuclear Energy. TIME, October 2023. URL https://time.com/6327635/ai-needs-to-be- regulated-like-nuclear-weapons/
2023
-
[51]
Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems
Heidy Khlaaf. Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems. Trail of Bits , March 2023. URL https://www.trailofbits.com/documents/Toward_comprehensive_risk_assessments.pdf
2023
-
[52]
A Hazard Analysis Framework for Code Synthesis Large Language Models
Heidy Khlaaf, Pamela Mishkin, Joshua Achiam, Gretchen Krueger, and Miles Brundage. A Hazard Analysis Framework for Code Synthesis Large Language Models. arXiv, July 2022. URL https://arxiv.org/abs/2207.14157v1
2022 arXiv
-
[53]
Mind the Gap: Foundation Models and the Covert Proliferation of Military Intelligence, Surveillance, and Targeting
Heidy Khlaaf, Sarah Myers West, and Meredith Whittaker. Mind the Gap: Foundation Models and the Covert Proliferation of Military Intelligence, Surveillance, and Targeting. arXiv, October 2024. doi: 10.48550/arXiv.2410.14831. URL http://arxiv.org/abs/2410.14831
-
[54]
Nato use of civil standards, October 2018
Steven Lapsley. Nato use of civil standards, October 2018. URL https://www.dsp.dla.mil/Portals/26/Documents/Publications/ Conferences/2018/2018%20International%20Standardization%20Workshop/20181031-Item4-UK_NATO_UseOfCivilStandards- IntlStdznWorkshop_Lapsely.pdf?ver=2018-11-06...
2018
-
[55]
Task Force Lima Executive Summary
Task Force Lima. Task Force Lima Executive Summary. Technical report, US Department of Defense, December 2024
2024
-
[56]
MacKenzie
Donald A. MacKenzie. Inventing Accuracy: A Historical Sociology of Nuclear Missile Guidance . MIT Press, 1990. ISBN 9780262132589
1990
-
[57]
Pasquale
Gianclaudio Malgieri and Frank A. Pasquale. From transparency to justification: Toward ex ante accountability for ai. SSRN Electronic Journal, 2022. doi: 10.2139/ssrn.4099657
2022 doi
-
[58]
AI should replace some work of civil servants, Starmer to announce
Rowena Mason and Rowena Mason Whitehall editor. AI should replace some work of civil servants, Starmer to announce. The Guardian, March 2025. ISSN 0261-3077. URL https://www.theguardian.com/technology/2025/mar/12/ai-should-replace-some-work- of-civil-servants-under-new-rules-k...
2025
- [59]
-
[60]
Anthropic’s Dario Amodei: Democracies must maintain the lead in AI
Madhumita Murgia. Anthropic’s Dario Amodei: Democracies must maintain the lead in AI. Financial Times, December 2024. URL https://www.ft.com/content/e75e3388-4700-413d-ab67-778410c2d977
2024
-
[61]
On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
Tarek Naous and Wei Xu. On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena. arXiv, January 2025. doi: 10.48550/arXiv.2501.04662. URL http://arxiv.org/abs/2501.04662
-
[62]
Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
Tarek Naous, Michael J Ryan, Alan Ritter, and Wei Xu. Having Beer after Prayer? Measuring Cultural Bias in Large Language Models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (...
2024 doi
-
[63]
AI as Normal Technology
Arvind Narayanan and Sayash Kapoor. AI as Normal Technology. Columbia University, The Knight First Amendment Institute , 2025. URL https://www.aisnakeoil.com/p/ai-as-normal-technology
2025
-
[64]
Feder Cooper, Daphne Ippolito, Christopher A
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. Scalable Extraction of Training Data from (Production) Language Models, 2023. URL 18 • Khlaaf and...
2023 arXiv
-
[65]
NIST. U.S. AI Safety Institute Establishes New U.S. Government Taskforce to Collaborate on Research and Testing of AI Models to Manage National Security Capabilities & Risks, November 2024. URL https://www.nist.gov/news-events/news/2024/11/us-ai-safety- institute-establishes-n...
2024
-
[66]
Implementation of improved Air System Safety Case Regulation - RA 1205, 2019
Ministry of Defence. Implementation of improved Air System Safety Case Regulation - RA 1205, 2019. URL https://www.gov.uk/ government/news/implementation-of-improved-air-system-safety-case-regulation-ra-1205
2019
-
[67]
Department of Defense
The U.S. Department of Defense. Standard Practice: System Safety. Technical Report MIL-STD-882E, The U.S. Department of Defense, May 2012
2012
-
[68]
DoD Modeling and Simulation (M&S) Verification, Validation, and Accreditation (VV&A),
U.S. Department of Defense. DoD Instruction 5000.61, “DoD Modeling and Simulation (M&S) Verification, Validation, and Accreditation (VV&A), ”. U.S. Department of Defense, September 2024. URL https://www.esd.whs.mil/Portals/54/Documents/DD/issuances/dodi/ 500061p.pdf
2024
-
[69]
Artificial Intelligence Safety and Security Board, 2025
Department of Homeland Security. Artificial Intelligence Safety and Security Board, 2025. URL https://www.dhs.gov/artificial- intelligence-safety-and-security-board
2025
-
[70]
2024 ICRC IHL Challenges Report | ICRC, September 2024
International Committee of the Red Cross. 2024 ICRC IHL Challenges Report | ICRC, September 2024. URL https://www.icrc.org/en/ report/2024-icrc-report-ihl-challenges
2024
-
[71]
The Safety Goals of the U.S
David Okrent. The Safety Goals of the U.S. Nuclear Regulatory Commission. Science, 236(4799):296–300, April 1987. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.3563510. URL https://www.science.org/doi/10.1126/science.3563510
1987 doi
-
[72]
Nathan Matias
Amy Orben and J. Nathan Matias. Fixing the science of digital technology harms. Science, 388(6743):152–155, April 2025. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.adt6807. URL https://www.science.org/doi/10.1126/science.adt6807
2025 doi
-
[73]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022 arXiv
-
[74]
Palantir IR, 2024
Palantir. Palantir IR, 2024. URL https://investors.palantir.com/news-details/2024/Anthropic-and-Palantir-Partner-to-Bring-Claude-AI- Models-to-AWS-for-U.S.-Government-Intelligence-and-Defense-Operations/
2024
-
[75]
Red Teaming Language Models with Language Models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red Teaming Language Models with Language Models. arXiv, February 2022. doi: 10.48550/arXiv.2202.03286. URL http: //arxiv.org/abs/2202.03286
-
[76]
Chauncey Starr: A personal memoir, June 2007
POWER. Chauncey Starr: A personal memoir, June 2007. URL https://www.powermag.com/chauncey-starr-a-personal-memoir/
2007
-
[77]
Introducing Primer Delta: Our next-gen AI-native platform to transform information overload into decision advantage, April
Primer. Introducing Primer Delta: Our next-gen AI-native platform to transform information overload into decision advantage, April
-
[78]
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! arXiv, October 2023
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! arXiv, October 2023. URL https://arxiv.org/abs/2310.03693v1
2023 arXiv
- [79]
-
[80]
Jane O. Rathbun. Don guidance on the use of generative artificial intelligence and large language models. Department of Navy Chief Information Officer, 2023. URL https://www.doncio.navy.mil/ContentView.aspx?id=16442
2023
-
[81]
arms race
Heather M. Roff. The frame problem: The AI “arms race” isn’t one. Bulletin of the Atomic Scientists , 75(3):95–98, May 2019. ISSN 0096-3402, 1938-3282. doi: 10.1080/00963402.2019.1604836. URL https://www.tandfonline.com/doi/full/10.1080/00963402.2019.1604836
2019
-
[82]
Fintech Founder Charged with Fraud after “AI” Shopping App Found to be Powered by Humans in the Philippines
Charles Rollet. Fintech Founder Charged with Fraud after “AI” Shopping App Found to be Powered by Humans in the Philippines. TechCrunch, Apr 2025. URL https://techcrunch.com/2025/04/10/fintech-founder-charged-with-fraud-after-ai-shopping-app-found-to- be-powered-by-humans-in-t...
2025
-
[83]
Anatomy of an AI Coup
Eryk Salvaggio. Anatomy of an AI Coup. Tech Policy Press, February 2025. URL https://techpolicy.press/anatomy-of-an-ai-coup
2025
-
[84]
Silicon Valley Goes to War: Artificial Intelligence, Weapons Systems, and Moral Agency
Elke Schwarz and DePaul University. Silicon Valley Goes to War: Artificial Intelligence, Weapons Systems, and Moral Agency. Philosophy Today, 65(3):549–569, 2021. ISSN 0031-8256. doi: 10.5840/philtoday2021519407. URL http://www.pdcnet.org/oom/service? url_ver=Z39.88-2004&rft_v...
2021 doi
-
[85]
AI Red Teaming: Applying Software TEVV for AI Evaluations | CISA, November 2024
Jonathan Spring and Divjot Singh Bawa. AI Red Teaming: Applying Software TEVV for AI Evaluations | CISA, November 2024. URL https://www.cisa.gov/news-events/news/ai-red-teaming-applying-software-tevv-ai-evaluations. Safety Co-Option and Compromised National Security: The Self-...
2024
-
[86]
Risk Criteria for Nuclear Power Plants: A Pragmatic Proposal
Chauncey Starr. Risk Criteria for Nuclear Power Plants: A Pragmatic Proposal. Risk Analysis, 1(2):113–120, June 1981. ISSN 0272-4332, 1539-6924. doi: 10.1111/j.1539-6924.1981.tb01406.x. URL https://onlinelibrary.wiley.com/doi/10.1111/j.1539-6924.1981.tb01406.x
1981
-
[87]
Risks of Risk Decisions
Chauncey Starr and Chris Whipple. Risks of Risk Decisions. In Risk In The Technological Society . Routledge, 1982. ISBN 9780429304873
1982
-
[88]
Philosophical Basis for Risk Analysis
Chauncey Starr, Richard Rudman, and Chris Whipple. Philosophical Basis for Risk Analysis. Annual Review of Energy, 1(1):629–662, November 1976. ISSN 0362-1626. doi: 10.1146/annurev.eg.01.110176.003213. URL https://www.annualreviews.org/doi/10.1146/annurev. eg.01.110176.003213
1976
-
[89]
How can we know a self-driving car is safe? Ethics and Information Technology, 23(4):635–647, December 2021
Jack Stilgoe. How can we know a self-driving car is safe? Ethics and Information Technology, 23(4):635–647, December 2021. ISSN 1572-8439. doi: 10.1007/s10676-021-09602-1. URL https://doi.org/10.1007/s10676-021-09602-1
2021 doi
-
[90]
The Illusion of China’s AI Prowess
Helen Toner, Jenny Xiao, and Jeffrey Ding. The Illusion of China’s AI Prowess. Foreign Affairs, June 2023. URL https://www. foreignaffairs.com/china/illusion-chinas-ai-prowess-regulation-helen-toner
2023
-
[91]
Thunderforge Project: Integrating Commercial AI-Powered Decision-Making, 2025
Defense Innovation Unit. Thunderforge Project: Integrating Commercial AI-Powered Decision-Making, 2025. URL https://www.diu. mil/latest/dius-thunderforge-project-to-integrate-commercial-ai-powered-decision-making
2025
-
[92]
‘Breeder’ Reactor Plan Facing Delays
McElheny Victor K. ‘Breeder’ Reactor Plan Facing Delays. The New York Times, December 1973
1973
-
[93]
Inside Task Force Lima’s exploration of 180-plus generative AI use cases for DOD
Brandi Vincent. Inside Task Force Lima’s exploration of 180-plus generative AI use cases for DOD. DefenseScoop, November 2023. URL https://defensescoop.com/2023/11/06/inside-task-force-limas-exploration-of-180-plus-generative-ai-use-cases-for-dod/
2023
-
[94]
The risks and inefficacies of AI systems in military targeting support
Jimena Sofía Viveros Álvarez. The risks and inefficacies of AI systems in military targeting support. ICRC Humanitarian Law & Policy Blog, September 2024. URL https://blogs.icrc.org/law-and-policy/2024/09/04/the-risks-and-inefficacies-of-ai-systems-in-military- targeting-support/
2024
-
[95]
Gaza: Israeli military’s digital tools risk civilian harm, 2024
Human Rights Watch. Gaza: Israeli military’s digital tools risk civilian harm, 2024. URL https://www.hrw.org/news/2024/09/10/gaza- israeli-militarys-digital-tools-risk-civilian-harm
2024
-
[96]
Ethical and social risks of harm from Language Models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks...
-
[97]
The China Trap: U.S
Jessica Chen Weiss. The China Trap: U.S. Foreign Policy and the Perilous Logic of Zero-Sum Competition. Foreign Affairs, August
-
[98]
DOGE’s Plans to Replace Humans With AI Are Already Under Way, March 2025
Matteo Wong. DOGE’s Plans to Replace Humans With AI Are Already Under Way, March 2025. URL https://www.theatlantic.com/ technology/archive/2025/03/gsa-chat-doge-ai/681987/
2025
-
[99]
Work and Greg Grant
Robert O. Work and Greg Grant. Beating the Americans at their Own Game. An Offset Strategy with Chinese Characteristics. Center for a New American Security , 2019. ISSN 2510-2648, 2510-263X. doi: 10.1515/sirius-2019-4022
2019 doi
-
[100]
Persistent Pre-Training Poisoning of LLMs
Yiming Zhang, Javier Rando, Ivan Evtimov, Jianfeng Chi, Eric Michael Smith, Nicholas Carlini, Florian Tramèr, and Daphne Ippolito. Persistent Pre-Training Poisoning of LLMs. arXiv, October 2024. doi: 10.48550/arXiv.2410.13722. URL http://arxiv.org/abs/2410.13722
-
[101]
Zico Kolter, and Matt Fredrikson
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv, December 2023. doi: 10.48550/arXiv.2307.15043. URL http://arxiv.org/abs/2307.15043
-
[102]
URL https://www.foreignaffairs.com/china/china-trap-us-foreign-policy-zero-sum-competition
-
[1983]
URL https://www.nrc.gov/docs/ML0717/ML071770230.pdf
-
[2022]
URL https://thebulletin.org/2022/07/why-policy-makers-should-beware-claims-of-new-arms-races/
2022
-
[2023]
URL https://primer.ai/business-solutions/introducing-primer-delta-our-next-gen-ai-native-platform-to-transform-information- overload-into-decision-advantage/
-
[2024]
URL https://www.anduril.com/anduril-partners-with-openai-to-advance-u-s-artificial-intelligence-leadership-and-protect-u-s/
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.