REVIEW 3 major objections 5 minor 87 references
Safety Features for a Centralised AGI Project
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The report argues that a US government-led AGI project can reduce catastrophic risk by adopting seven institutional safety features, including tripwire reporting, pause protocols, and board approval for training runs.
desk verdict A candid, well-referenced policy synthesis for a centralized US AGI project; the risk-reduction claim is weaker than the analysis admits, but the paper is worth engaging. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the 'safety case' checkpoint system. Before pre-training, after pre-training, before internal deployment, and before each expansion of a model's permissions, technical teams must present an affirmative, evidence-backed argument that the system is safe enough, and a board of technical experts must approve proceeding by supermajority. Around this core sit the tripwire capabilities and intolerable risk thresholds that trigger reporting and emergency pauses, plus a designated point of contact who can execute a pause instruction from any staff member. This machinery is what turns the report's priorities into operational constraints.
What would settle it
A concrete falsifier would be a demonstration that a frontier model can pass every pre-deployment safety-case evaluation, including limit evals, while concealing a capability that later causes catastrophic harm—for example, a model that strategically underperforms on all dangerous-capability tests but exhibits situational awareness and deceptive goal-seeking once deployed.
Extended reading notes
Core claim
The paper's central claim is that a government-led AGI project can be made safer through institutional choices rather than only through technical alignment research. Its seven features are: information escalation with explicit reporting thresholds and protected dissent; pause protocols with bottom-up and top-down emergency authority; a technical board with binding authority over training and deployment decisions; separate internal audit and risk-monitoring bodies; an intelligence and scenario-planning division; a designated verification project for international agreements; and a plan for automated research. The report argues that each feature compensates for a specific failure mode of high-stakes technology development, and that a centralized project has a key advantage over private labs: it can standardize one set of pause thresholds without competitive pressure to lower them.
Load-bearing premise
The entire architecture presupposes that dangerous capabilities in advanced AI can be detected and evaluated reliably before they become catastrophic; the report itself concedes that the science of model evaluation may not mature quickly and that models can hide their abilities through sandbagging or deceptive behavior.
Editorial extensions
If this is right
- A centralized project would halt at each defined checkpoint unless an affirmative safety case passes a board supermajority, so development would slow or stop whenever safety evidence is insufficient.
- Reporting requirements triggered by regular cadence, compute scaling, and capability thresholds would surface risks that private-sector frameworks currently leave undefined.
- Bottom-up pause authority, modelled on stop-work authority in other high-risk industries, would let technical staff halt training runs before senior leadership recognizes a threat.
- A designated verification project, including hardware-enabled mechanisms on AI chips, would make international agreements to limit AI development potentially verifiable.
- If evaluation science does not mature in time, adherence to these protocols would delay AGI development—a delay the report says the project should accept and plan for.
Reading between the lines
- Beyond the report, the same safety-case and tripwire logic could be tested retroactively against existing frontier labs by asking how often their training runs would have triggered pauses under these thresholds.
- The report's institutional features are likely transferable to a multilateral or private-consortium setting, not only a US government project, since the underlying failure modes—competitive pressure, opaque development, and unclear thresholds—are not unique to government.
- A hidden implication is that the value of all seven features is epistemic, not mechanical: they buy time and force evidence, but they cannot compensate for a fundamental inability to evaluate a deceptive model.
- One testable extension would be to run a sandboxed simulation of the pause-and-checkpoint protocol on an existing large model training run to measure how much latency and cost the safety cases would add.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper analyzes a hypothetical scenario in which the US government centralizes AGI development into a single project under its direct control, and proposes four high-level priorities (assessing alignment difficulty, maintaining optionality, envisioning the end-state, and a fourth implicit priority related to information flows) and seven safety features: information escalation, pause protocols, board oversight, audit and risk monitoring, a geopolitics division, a designated verification project, and a plan for automated research. The report draws heavily on analogies from nuclear regulation, aviation safety, manufacturing quality control, and existing frontier-AI safety frameworks, and it repeatedly acknowledges the scientific immaturity of model evaluation and safety cases. Its central claim is that implementing some or all of these features could reduce catastrophic risk from loss-of-control or egregious misuse, while explicitly accepting that adherence to the protocols may delay AGI development until scientific breakthroughs occur.
Significance. If the central claim is accepted, the paper provides a concrete institutional blueprint for a high-stakes policy decision, with actionable recommendations on reporting thresholds, emergency pauses, binding board authority, audit separation, verification technologies, and geopolitical contingency planning. Its strengths are the specificity of the proposed mechanisms, the use of documented precedents (NRC, CRITIC, DNFSB, RAND HEMs), and its unusual candor about the limits of current evaluation science and safety-case methodology. The paper does not present new empirical evidence or formal models, but it is a serious and well-grounded policy analysis. Its main weakness is that the loss-of-control risk-reduction claim relies on the reliability of capability detection, which the paper itself repeatedly concedes is not currently attainable; the paper needs to either articulate a non-detection-based mechanism for reducing loss-of-control risk or narrow its claim.
major comments (3)
- [§5.2.3, §6] The paper's central claim in §6 that implementing the proposed features 'could reduce catastrophic risk stemming from loss-of-control or egregious misuse' presupposes that dangerous capabilities can be detected before they become catastrophic. Yet §5.2.1 concedes that 'a mature science of model evaluation may not be possible in a short timeframe,' §5.2.3 states that 'building full safety cases for models significantly more advanced than today's is not yet possible,' and §5.4.2 acknowledges the lack of well-established methods for forecasting AI-caused catastrophes. In the regime the paper itself treats as live, the pause checkpoints and tripwires may only produce delay, and the paper does not explain how delay alone reduces loss-of-control risk when detection is unreliable. Please either specify a fallback mechanism that does not depend on capability evaluation (for example, compute-based or hardware-enforced hard ceilings) or restrict the conclusion to misuse risks and to risk reduction through slowing development.
- [§5.2.1] The report recommends that 'clear evidence' that evaluation science will not mature 'should itself trigger a pause in advancing model capabilities.' This tripwire is underspecified and potentially circular: recognizing that a mature evaluation science is impossible is itself an evaluative judgment subject to the same unreliability and expert disagreement the report documents elsewhere. Please define what evidence would qualify, who would make the determination, and what safeguards would prevent this determination from being overridden by the race dynamics described in the same section.
- [§5.3.2] The Board's binding authority is a load-bearing element of the design, but the report does not address the failure mode in which the Executive Branch, which controls the project, disregards or overrides a Board veto—for example, by ordering resumption of training after the Board votes against it. The NRC analogy is imperfect because the NRC regulates a private industry rather than an executive-controlled project; please discuss the legal and practical enforceability of Board decisions against the President or the project's head, or explain why the analogy remains apt despite this difference.
minor comments (5)
- [Throughout] There are repeated typographical errors in the name 'Anthropic,' spelled 'Antrophic' in several places (e.g., §5.2.1, §5.2.3, §5.3.1); these should be corrected.
- [§5.1] The sentence 'This will need to be efficiently escalated up to project leadership, and to the risk monitoring and internal compliance teams (see Figure 1)' lacks Figure 2; the organizational chart is Figure 2, while Figure 1 shows the 'How hard is AI safety?' graphic. The cross-reference should be fixed.
- [§1.2, §3, §5] The executive summary's numbering (1.2 High-level priorities, 1.3 Safety Features) does not match the body's numbering (Section 3 for priorities, Section 5 for safety features); the numbering and cross-references should be harmonized.
- [§2] The paper states it is 'informed by semi-structured interviews with experts in AI safety and governance,' but it provides no information on the number of interviewees, their selection, the interview questions, or the analysis method, so the reader cannot assess this evidence base.
- [§5.2.3] The phrase 'in each of in a list of predetermined thresholds' contains a grammatical error; it should read 'at each checkpoint in a list of predetermined thresholds.'
Circularity Check
No significant circularity: the report's recommendations are anchored to external precedents and its central conclusion is a conditional deduction from openly stated premises.
full rationale
This paper is a qualitative policy report with no equations, fitted parameters, or formal derivations. Its seven safety features are justified by external institutional precedents (NRC, CRITIC, RAND HEMs, frontier-lab safety frameworks, DNFSB, SIGAR) and by semi-structured expert interviews, rather than by importing conclusions from the author's own prior work. The central conclusion that pauses are likely to be necessary is a straightforward conditional: if the project adopts an affirmative safety standard, and if watertight safety cases are not yet possible for advanced models, then progression to later checkpoints will be blocked. The report explicitly acknowledges the fragility of its own premise regarding evaluation science, noting in Section 5.2.1 that 'a mature science of model evaluation may not be possible in a short timeframe' and in Section 5.2.3 that 'building full safety cases for models significantly more advanced than today's is not yet possible.' Those admissions weaken the report's practical risk-reduction claim, but they do not make the reasoning circular; they identify an unverified assumption on which the recommendations depend. No self-citations are load-bearing, no known result is renamed as a new framework, and no fitted input is relabeled as a prediction. The appropriate finding is therefore no significant circularity.
Assumptions & free parameters
assumptions (6)
- domain assumption Uncontrolled AGI would pose extreme, potentially existential global risks.
- domain assumption A US government-led centralized AGI project is plausible enough to justify detailed institutional design.
- domain assumption Dangerous capabilities can be detected early enough for reporting and pause triggers to matter.
- domain assumption An affirmative safety case standard is an appropriate and enforceable gateway for proceeding.
- domain assumption The Board and oversight bodies will retain real authority under political and competitive pressure.
- domain assumption Hardware-enabled verification mechanisms can be developed and deployed before or during a treaty window.
invented entities (4)
-
A binding technical Safety Board
-
A designated point of contact for pause requests
-
A risk monitoring team separate from risk management
-
A designated verification project with hardware-enabled mechanisms
Cite this review
Pith. "Pith review of Safety Features for a Centralised AGI Project." pith.science (2026). https://pith.science/paper/4CAFGYGI
@misc{pith2026250721082,
author = {Pith},
title = {Pith review of: Safety Features for a Centralised AGI Project},
year = {2026},
howpublished = {\url{https://pith.science/paper/4CAFGYGI}},
note = {Machine review of arXiv:2507.21082}
}
read the original abstract
Recent AI progress has outpaced expectations, with some experts now predicting AI that matches or exceeds human capabilities in all cognitive areas (AGI) could emerge this decade, potentially posing grave national and global security threats. AI development is currently occurring primarily in the private sector with minimal oversight. This report analyzes a scenario where the US government centralizes AGI development under its direct control, and identifies four high-level priorities and seven safety features to reduce risks.
Figures
Reference graph
Works this paper leans on
-
[1]
OpenAI o3 and o3-mini—12 Days of OpenAI: Day 12,
OpenAI, “OpenAI o3 and o3-mini—12 Days of OpenAI: Day 12,” Dec. 2024. [Online]. Available: https://www.youtube.com/watch?v=SKBG1sqdyIU
2024
-
[2]
Can AI Scaling Continue Through 2030?
J. Sevilla, “Can AI Scaling Continue Through 2030?” Aug. 2024. [Online]. Available: https: //epoch.ai/blog/can-ai-scaling-continue-through-2030
2024
-
[3]
Statement on AI Risk | CAIS,
Center for AI Safety, “Statement on AI Risk | CAIS,” May 2023. [Online]. Available: https: //safe.ai/work/statement-on-ai-risk
2023
-
[4]
Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence,
J. R. Biden, “Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence,” Oct. 2023, executive Order 14110. [Online]. Available: https://www.federalregister.gov/documents/2023/11/01/2023-24283/ safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence
2023
-
[5]
Removing Barriers to American Leadership in Artificial Intelligence,
D. J. Trump, “Removing Barriers to American Leadership in Artificial Intelligence,” Jan. 2025, executive Order signed January 23, 2025. [Online]. Available: https://www.whitehouse.gov/presidential-actions/2025/01/ removing-barriers-to-american-leadership-in-artificial-intelligence/
2025
-
[6]
Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models,
S. Nevo, D. Lahav, A. Karpur, Y . Bar-On, H. A. Bradley, and J. Alstott, “Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models,” Tech. Rep., May 2024. [Online]. Available: https://www.rand.org/pubs/research_reports/RRA2849-1.html
2024
-
[7]
2024 REPORT TO CONGRESS of the U.S.-CHINA ECONOMIC AND SECURITY REVIEW COMMISSION,
U.S.-China Economic and Security Review Commission, “2024 REPORT TO CONGRESS of the U.S.-CHINA ECONOMIC AND SECURITY REVIEW COMMISSION,” Nov. 2024, publisher: U.S.-China Economic and Security Review Commission. [Online]. Available: https://www.uscc.gov/sites/default/files/2024-11/2024_ Annual_Report_to_Congress.pdf
2024
-
[8]
Soft Nationalization: How the US Government Will Control AI Labs | Convergence Analysis,
D. Cheng and C. Katzke, “Soft Nationalization: How the US Government Will Control AI Labs | Convergence Analysis,” Aug. 2024. [Online]. Available: https://www.convergenceanalysis.org/publications/ soft-nationalization-how-the-us-government-will-control-ai-labs
work page 2024
Show all 87 references
-
[9]
Core Views on AI Safety: When, Why, What, and How,
Anthropic, “Core Views on AI Safety: When, Why, What, and How,” Mar. 2023. [Online]. Available: https://www.anthropic.com/news/core-views-on-ai-safety
2023
-
[10]
How hard is AI safety?
C. Olah, “How hard is AI safety?” Jul. 2023. [Online]. Available: https://x.com/ch402/status/ 1666482929772666880?lang=en
2023
-
[11]
Defining and Characterizing Reward Hacking,
J. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger, “Defining and Characterizing Reward Hacking,” Mar. 2025, arXiv:2209.13085. [Online]. Available: http://arxiv.org/abs/2209.13085
2025 arXiv
-
[12]
AI Deception: A Survey of Examples, Risks, and Potential Solutions,
P. S. Park, S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks, “AI Deception: A Survey of Examples, Risks, and Potential Solutions,” Aug. 2023. [Online]. Available: https://arxiv.org/abs/2308.14752v1
2023 arXiv
-
[13]
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models,
W. Zhou, X. Wang, L. Xiong, H. Xia, Y . Gu, M. Chai, F. Zhu, C. Huang, S. Dou, Z. Xi, R. Zheng, S. Gao, Y . Zou, H. Yan, Y . Le, R. Wang, L. Li, J. Shao, T. Gui, Q. Zhang, and X. Huang, “EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models,” Mar. 2024, arX...
2024 arXiv
-
[14]
Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence Strategy,
D. Hendrycks, E. Schmidt, and A. Wang, “Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence Strategy,” 2024. [Online]. Available: https://www.nationalsecurity.ai/chapter/ deterrence-with-mutual-assured-ai-malfunction-maim
2024
-
[15]
AI Behind Closed Doors: a Primer on The Governance of Internal Deployment,
C. Stix, M. Pistillo, G. Sastry, M. Hobbhahn, A. Ortega, M. Balesni, A. Hallensleben, N. Goldowsky-Dill, and K. Sharkey, “AI Behind Closed Doors: a Primer on The Governance of Internal Deployment,” Apr. 2025, arXiv:2504.12170. [Online]. Available: http://arxiv.org/abs/2504.12170
2025 arXiv
-
[16]
Evaluating frontier AI R&D capabilities of language model agents against human experts,
METR, “Evaluating frontier AI R&D capabilities of language model agents against human experts,”METR Blog, Nov. 2024. [Online]. Available: https://metr.org/blog/2024-11-22-evaluating-r-d-capabilities-of-llms/
2024
-
[17]
Measuring the Persuasiveness of Language Models,
E. Durmus, L. Lovitt, A. Tamkin, S. Ritchie, J. Clark, and D. Ganguli, “Measuring the Persuasiveness of Language Models,” Apr. 2024, publisher: Anthropic. [Online]. Available: https://www.anthropic.com/research/ measuring-model-persuasiveness
2024
-
[18]
Emergent Abilities of Large Language Models,
J. Wei, Y . Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent Abilities of Large Language Models,” Oct. 2022, arXiv:2206.07682. [Online]. Available: ht...
2022 arXiv
-
[19]
Training Compute Thresholds — Features and Functions in AI Regulation,
L. Heim, “Training Compute Thresholds — Features and Functions in AI Regulation,” Apr. 2024. [Online]. Available: https://blog.heim.xyz/training-compute-thresholds/ 23 SAFETY FEATURES FOR A CENTRALIZEDAGIPROJECT
2024
-
[20]
Will there be a discontinuity in AI capabilities?
AI Safety Info, “Will there be a discontinuity in AI capabilities?” 2024. [Online]. Available: https://aisafety.info/questions/7729/Will-there-be-a-discontinuity-in-AI-capabilities
2024
-
[21]
Response to BIS AI Reporting Requirements RFC — MIRI Technical Governance Team,
MIRI Technical Governance Team, “Response to BIS AI Reporting Requirements RFC — MIRI Technical Governance Team,” 2025. [Online]. Available: https://techgov.intelligence.org/research/ response-to-bis-ai-reporting-requirements-rfc
2025
-
[22]
AI Risk Management Framework,
National Institute of Standards and Technology, “AI Risk Management Framework,”NIST, Jan. 2023. [Online]. Available: https://www.nist.gov/itl/ai-risk-management-framework
2023
-
[23]
Common Elements of Frontier AI Safety Policies,
METR, “Common Elements of Frontier AI Safety Policies,”METR Blog, Mar. 2025. [Online]. Available: https://metr.org/blog/2025-03-26-common-elements-of-frontier-ai-safety-policies/
2025
-
[24]
Announcing our updated Responsible Scaling Policy,
Anthropic, “Announcing our updated Responsible Scaling Policy,” Oct. 2024. [Online]. Available: https://www.anthropic.com/news/announcing-our-updated-responsible-scaling-policy
2024
-
[25]
xAI Risk Management Framework (Draft),
xAI, “xAI Risk Management Framework (Draft),” Feb. 2025, publisher: xAI. [Online]. Available: https://x.ai/documents/2025.02.20-RMF-Draft.pdf
2025
-
[26]
Intolerable Risk Threshold Recommendations for Artificial Intelligence,
D. Raman, N. Madkour, E. Murphy, K. Jackson, and J. Newman, “Intolerable Risk Threshold Recommendations for Artificial Intelligence,” Jan. 2025. [Online]. Available: https://cltc.berkeley.edu/publication/ intolerable-ai-risk-thresholds/
2025
-
[27]
5 FAH-2 H-430 HANDLING SYMBOLS,
U.S. Department of State, “5 FAH-2 H-430 HANDLING SYMBOLS,” 2024. [Online]. Available: https://fam.state.gov/fam/05fah02/05fah020430.html
2024
-
[28]
Notes on the Critic System,
W. Tidewell, “Notes on the Critic System,” 1980. [Online]. Available: https://www.cia.gov/resources/csi/static/ Notes-on-Critic-System.pdf
1980
-
[29]
Handling of Critical (CRITIC) Information,
Central Intelligence Agency, “Handling of Critical (CRITIC) Information,” Jul. 1979, publisher: Central Intelli- gence Agency. [Online]. Available: https://www.cia.gov/readingroom/docs/CIA-RDP83-00156R000200040001-1. pdf
1979
-
[30]
2 FAM 070 DISSENT CHANNEL,
U.S. Department of State, “2 FAM 070 DISSENT CHANNEL,” 2024, publisher: US Department of State. [Online]. Available: https://fam.state.gov/fam/02fam/02fam0070.html
2024
-
[31]
NRC DIFFERING PROFESSIONAL OPINION PROGRAM,
Nuclear Regulatory Commission, “NRC DIFFERING PROFESSIONAL OPINION PROGRAM,” Aug. 2015, publisher: Nuclear Regulatory Commission. [Online]. Available: https://www.nrc.gov/docs/ml1513/ml15132a664. pdf
2015
-
[32]
DOE Differing Professional Opinions,
U.S. Department of Energy, “DOE Differing Professional Opinions,” 2024. [Online]. Available: https://www.energy.gov/ehss/doe-differing-professional-opinions
2024
-
[33]
Stifling Dissent,
D. Van Schooten and N. Schwellenbach, “Stifling Dissent,” 2021. [Online]. Available: https: //www.pogo.org/reports/stifling-dissent
2021
-
[34]
Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts,
S. Field, “Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts,” Jan. 2025, arXiv:2502.14870. [Online]. Available: http://arxiv.org/abs/2502.14870
2025 arXiv
-
[35]
AI Whistleblowers,
H. Wu, “AI Whistleblowers,” Mar. 2025. [Online]. Available: https://papers.ssrn.com/sol3/papers.cfm?abstract_ id=4790511
2025
-
[36]
Our updated Preparedness Framework,
OpenAI, “Our updated Preparedness Framework,” Dec. 2024, publisher: OpenAI. [Online]. Available: https://openai.com/index/updating-our-preparedness-framework/
2024
-
[37]
A Sketch of Potential Tripwire Capabili- ties for AI,
Carnegie Endowment for International Peace, “A Sketch of Potential Tripwire Capabili- ties for AI,” Dec. 2024. [Online]. Available: https://carnegieendowment.org/research/2024/12/ a-sketch-of-potential-tripwire-capabilities-for-ai?lang=en
2024
-
[38]
Risk Thresholds for Frontier AI | GovAI,
GovAI, “Risk Thresholds for Frontier AI | GovAI,” 2024. [Online]. Available: https://www.governance.ai/ research-paper/risk-thresholds-for-frontier-ai
2024
-
[39]
AI models can be dangerous before public deployment,
METR, “AI models can be dangerous before public deployment,”METR Blog, Jan. 2025. [Online]. Available: https://metr.org/blog/2025-01-17-ai-models-dangerous-before-public-deployment/
2025
-
[40]
The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models,
A. Pan, K. Bhatia, and J. Steinhardt, “The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models,” Feb. 2022, arXiv:2201.03544. [Online]. Available: http://arxiv.org/abs/2201.03544
2022 arXiv
-
[41]
Towards Understanding Sycophancy in Language Models,
M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, N. Cheng, E. Durmus, Z. Hatfield-Dodds, S. R. Johnston, S. Kravec, T. Maxwell, S. McCandlish, K. Ndousse, O. Rausch, N. Schiefer, D. Yan, M. Zhang, and E. Perez, “Towards Understanding Sycophancy in Language ...
2025 arXiv
-
[42]
Andon - Toyota Production System guide,
Toyota UK Magazine, “Andon - Toyota Production System guide,” May 2016, publisher: Toyota UK Magazine. [Online]. Available: https://mag.toyota.co.uk/andon-toyota-production-system/
2016
-
[43]
Stop Work Authority: How It Works,
TRADESAFE, “Stop Work Authority: How It Works,” Sep. 2024. [Online]. Available: https: //trdsf.com/blogs/news/stop-work-authority
2024
-
[44]
Whirlpool Corp. v. Marshall, 445 U.S. 1 (1980),
U.S. Supreme Court, “Whirlpool Corp. v. Marshall, 445 U.S. 1 (1980),” 1980. [Online]. Available: https://supreme.justia.com/cases/federal/us/445/1/
1980
-
[45]
OSH Act of 1970,
U.S. Congress, “OSH Act of 1970,” Dec. 1970, publisher: Occupational Safety and Health Administration
1970
-
[46]
Racing to the precipice: a model of artificial intelligence development,
S. Armstrong, N. Bostrom, and C. Shulman, “Racing to the precipice: a model of artificial intelligence development,”AI & SOCIETY, vol. 31, no. 2, pp. 201–206, May 2016. [Online]. Available: https://doi.org/10.1007/s00146-015-0590-y
2016 doi
-
[47]
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs,
R. Laine, B. Chughtai, J. Betley, K. Hariharan, J. Scheurer, M. Balesni, M. Hobbhahn, A. Meinke, and O. Evans, “Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs,” Jul. 2024, arXiv:2407.04694. [Online]. Available: http://arxiv.org/abs/2407.04694
2024 arXiv
-
[48]
AI Sandbagging: Language Models can Strategically Underperform on Evaluations,
T. v. d. Weij, F. Hofstätter, O. Jaffe, S. F. Brown, and F. R. Ward, “AI Sandbagging: Language Models can Strategically Underperform on Evaluations,” Feb. 2025, arXiv:2406.07358. [Online]. Available: http://arxiv.org/abs/2406.07358
2025 arXiv
-
[49]
Alignment faking in large language models,
R. Greenblatt, C. Denison, B. Wright, F. Roger, M. MacDiarmid, S. Marks, J. Treutlein, T. Belonax, J. Chen, D. Duvenaud, A. Khan, J. Michael, S. Mindermann, E. Perez, L. Petrini, J. Uesato, J. Kaplan, B. Shlegeris, S. R. Bowman, and E. Hubinger, “Alignment faking in large lang...
2024
-
[50]
Safety Goals for Nuclear Power Plant Operation,
Nuclear Regulatory Commission, “Safety Goals for Nuclear Power Plant Operation,” 1986, publisher: Nuclear Regulatory Commission. [Online]. Available: https://www.nrc.gov/docs/ML0717/ML071770230.pdf
1986
-
[51]
Affirmative safety: An approach to risk management for high-risk AI,
A. R. Wasil, J. Clymer, D. Krueger, E. Dardaman, S. Campos, and E. R. Murphy, “Affirmative safety: An approach to risk management for high-risk AI,” Apr. 2024. [Online]. Available: https://arxiv.org/abs/2406.15371v1
2024 arXiv
-
[52]
Towards A Rigorous Science of Interpretable Machine Learning,
F. Doshi-Velez and B. Kim, “Towards A Rigorous Science of Interpretable Machine Learning,” Feb. 2017. [Online]. Available: https://arxiv.org/abs/1702.08608v2
2017 arXiv
-
[53]
Provably safe systems: the only path to controllable AGI,
M. Tegmark and S. Omohundro, “Provably safe systems: the only path to controllable AGI,” Sep. 2023, arXiv:2309.01933. [Online]. Available: http://arxiv.org/abs/2309.01933
2023 arXiv
-
[54]
Backgrounder on Emergency Preparedness at Nuclear Power Plants,
U.S. Nuclear Regulatory Commission, “Backgrounder on Emergency Preparedness at Nuclear Power Plants,” 2024. [Online]. Available: https://www.nrc.gov/reading-rm/doc-collections/fact-sheets/emerg-plan-prep-nuc-power. html
2024
-
[55]
Safety cases at AISI | AISI Work,
UK AI Safety Institute, “Safety cases at AISI | AISI Work,” 2024. [Online]. Available: https: //www.aisi.gov.uk/work/safety-cases-at-aisi
2024
-
[56]
National Defense Authorization Act for Fiscal Year 1989,
U.S. Congress, “National Defense Authorization Act for Fiscal Year 1989,” Mar. 1988. [Online]. Available: https://www.congress.gov/bill/100th-congress/house-bill/4264/summary/17
1989
-
[57]
42 U.S. Code § 5841 - Establishment and transfers,
——, “42 U.S. Code § 5841 - Establishment and transfers,” 1974. [Online]. Available: https: //www.law.cornell.edu/uscode/text/42/5841
1974
-
[58]
Survey of 2,778 AI authors: six parts in pictures,
K. Grace, “Survey of 2,778 AI authors: six parts in pictures,” 2024. [Online]. Available: https: //blog.aiimpacts.org/p/2023-ai-survey-of-2778-six-things
2024
-
[59]
"Existential risk from AI
R. Bensinger, “"Existential risk from AI" survey results,” Jun. 2021. [Online]. Available: https: //www.alignmentforum.org/posts/QvwSr5LsxyDeaPK5s/existential-risk-from-ai-survey-results
2021
-
[60]
Anthropic’s Recommendations to OSTP for the U.S. AI Action Plan,
Anthropic, “Anthropic’s Recommendations to OSTP for the U.S. AI Action Plan,” Jan. 2025. [Online]. Available: https://www.anthropic.com/news/anthropic-s-recommendations-ostp-u-s-ai-action-plan
2025
-
[61]
Tweet by @nabla_theta,
L. Gao, “Tweet by @nabla_theta,” Dec. 2024. [Online]. Available: https://x.com/nabla_theta/status/ 1869144832595431553
2024
-
[62]
AI-Enabled Coups: How a Small Group Could Use AI to Seize Power,
T. Davidson, L. Finnveden, and R. Hadshar, “AI-Enabled Coups: How a Small Group Could Use AI to Seize Power,” Apr. 2025. [Online]. Available: https://www.forethought.org/research/ ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power
2025
-
[63]
New Survey: Broad Expert Consensus for Many AGI Safety and Governance Practices | GovAI,
J. Schuett, N. Dreksler, M. Anderljung, D. McCaffary, L. Heim, E. Bluemke, and B. Garfinkel, “New Survey: Broad Expert Consensus for Many AGI Safety and Governance Practices | GovAI,” Jun. 2023. [Online]. Available: https: //www.governance.ai/analysis/broad-expert-consensus-fo...
2023
-
[64]
Inspector General Act of 1978,
U.S. Congress, “Inspector General Act of 1978,” 1978. [Online]. Available: https://www.congress.gov/bill/ 95th-congress/house-bill/8588
1978
-
[65]
Operation Warp Speed: Accelerated COVID-19 Vaccine Development Status and Efforts to Address Manufacturing Challenges | U.S. GAO,
U.S. Government Accountability Office, “Operation Warp Speed: Accelerated COVID-19 Vaccine Development Status and Efforts to Address Manufacturing Challenges | U.S. GAO,” Feb. 2021. [Online]. Available: https://www.gao.gov/products/gao-21-319
2021
-
[66]
Update on ARC’s recent eval efforts,
Alignment Research Center, “Update on ARC’s recent eval efforts,”METR Blog, Mar. 2023. [Online]. Available: https://metr.org/blog/2023-03-18-update-on-recent-evals/
2023
-
[67]
Introducing Alignment Stress-Testing at Anthropic,
evhub, “Introducing Alignment Stress-Testing at Anthropic,” Jan. 2024. [Online]. Available: https: //www.alignmentforum.org/posts/EPDSdXr8YbsDkgsDG/introducing-alignment-stress-testing-at-anthropic
2024
-
[68]
Distribution Shifts and The Importance of AI Safety,
L. Lang, “Distribution Shifts and The Importance of AI Safety,” Sep. 2022. [Online]. Available: https: //www.alignmentforum.org/posts/TRKF9g65nhPBQoxJu/distribution-shifts-and-the-importance-of-ai-safety
2022
-
[69]
Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy,
B. Shlegeris and R. Greenblatt, “Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy,” Jul. 2023. [Online]. Available: https://www.alignmentforum.org/posts/ MbWWKbyD5gLhJgfwn/meta-level-adversarial-evaluation-of-oversight-...
2023
-
[70]
Organizational Arrangements for Risk Assessment,
National Research Council (US) Committee on the Institutional Means for Assessment of Risks to Public Health, “Organizational Arrangements for Risk Assessment,” inRisk Assessment in the Federal Government: Managing the Process. National Academies Press (US), 1983. [Online]. Av...
1983
-
[71]
Working together to face humanity’s greatest threats: Introduction to The Future of Research on Catastrophic and Existential Risk
A. M. Currie and S. Ó hÉigeartaigh, “Working together to face humanity’s greatest threats: Introduction to The Future of Research on Catastrophic and Existential Risk.” Sep. 2018. [Online]. Available: https://www.repository.cam.ac.uk/handle/1810/280193
2018
-
[72]
Existential Risk Prevention as a Global Priority,
N. Bostrom, “Existential Risk Prevention as a Global Priority,” Feb. 2013. [Online]. Available: https://existential-risk.com/concept.pdf
2013
-
[73]
Roots of Disagreement on AI Risk: Exploring the Potential and Pitfalls of Adversarial Collaboration,
J. Rosenberg, E. Karger, A. Morris, M. Hickman, R. Hadshar, Z. Jacobs, and P. Tetlock, “Roots of Disagreement on AI Risk: Exploring the Potential and Pitfalls of Adversarial Collaboration,” 2022, publisher: Forecasting Research Institute. [Online]. Available: https://static1.s...
2022
-
[74]
Superhuman Automated Forecasting | CAIS,
Center for AI Safety, “Superhuman Automated Forecasting | CAIS,” Jan. 2025. [Online]. Available: https://safe.ai/blog/forecasting
2025
-
[75]
Technical Options for Flexible Hardware-Enabled Guarantees,
J. Petrie and O. Aarne, “Technical Options for Flexible Hardware-Enabled Guarantees,” Jun. 2025. [Online]. Available: https://arxiv.org/abs/2506.03409v1
2025 arXiv
-
[76]
Verification methods for international AI agreements,
A. R. Wasil, T. Reed, J. W. Miller, and P. Barnett, “Verification methods for international AI agreements,” Aug
-
[77]
Global Security Remote Sensing and Verification,
Sandia National Laboratories, “Global Security Remote Sensing and Verification,” 2024. [Online]. Available: https://www.sandia.gov/missions/global-security-remote-sensing-and-verification/
2024
-
[78]
Verification and other safeguards activities,
International Atomic Energy Agency, “Verification and other safeguards activities,” Jun. 2016. [Online]. Available: https://www.iaea.org/topics/verification-and-other-safeguards-activities
2016
-
[79]
Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090,
G. Kulp, D. Gonzales, E. Smith, L. Heim, P. Puri, M. J. D. Vermeer, and Z. Winkelman, “Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090,” Tech. Rep., Jan. 202...
2024
-
[80]
Why policy makers should beware claims of new ’arms races’,
H. Belfield and C. Ruhl, “Why policy makers should beware claims of new ’arms races’,” Jul. 2022, publisher: July 2022. [Online]. Available: https://thebulletin.org/2022/07/ why-policy-makers-should-beware-claims-of-new-arms-races/
2022
-
[81]
Who is behind DeepSeek and how did it achieve its AI ’Sputnik moment’?
A. Hawkins, “Who is behind DeepSeek and how did it achieve its AI ’Sputnik moment’?” The Guardian, Jan. 2025. [Online]. Available: https://www.theguardian.com/technology/2025/jan/28/ who-is-behind-deepseek-and-how-did-it-achieve-its-ai-sputnik-moment
2025
-
[82]
On DeepSeek and Export Controls,
D. Amodei, “On DeepSeek and Export Controls,” Jan. 2025. [Online]. Available: https://www.darioamodei.com/ post/on-deepseek-and-export-controls
2025
-
[83]
What Is DeepSeek? New Chinese Artificial Intelligence Rivals Chat- GPT, OpenAI,
M. W. Roeloffs, “What Is DeepSeek? New Chinese Artificial Intelligence Rivals Chat- GPT, OpenAI,” Jan. 2025. [Online]. Available: https://www.forbes.com/sites/maryroeloffs/2025/01/27/ what-is-deepseek-new-chinese-ai-startup-rivals-openai-and-claims-its-far-cheaper/ 26 SAFETY F...
2025
-
[84]
Trends in U.S. Intention-to-Stay Rates of International Ph.D. Gradu- ates Across Nationality and STEM Fields,
R. Zwetsloot, J. Feldgoise, and J. Dunham, “Trends in U.S. Intention-to-Stay Rates of International Ph.D. Gradu- ates Across Nationality and STEM Fields,” Sep. 2019. [Online]. Available: https://cset.georgetown.edu/publication/ trends-in-u-s-intention-to-stay-rates-of-internat...
2019
-
[85]
AI Pioneer Geoffrey Hinton Talks About AI Gaining Con- trol,
A. Morris, “AI Pioneer Geoffrey Hinton Talks About AI Gaining Con- trol,” May 2023. [Online]. Available: https://www.forbes.com/sites/andreamorris/2023/05/03/ ai-pioneer-geoffrey-hinton-talks-at-mit-about-ai-gaining-control/
2023
-
[86]
Introducing Superalignment,
OpenAI, “Introducing Superalignment,” Jul. 2023. [Online]. Available: https://openai.com/index/ introducing-superalignment/ 27
2023
-
[2024]
Available: https://arxiv.org/abs/2408.16074v2
[Online]. Available: https://arxiv.org/abs/2408.16074v2
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.