Pith. sign in

REVIEW 3 major objections 5 minor 87 references

Safety Features for a Centralised AGI Project

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The report argues that a US government-led AGI project can reduce catastrophic risk by adopting seven institutional safety features, including tripwire reporting, pause protocols, and board approval for training runs.

desk verdict A candid, well-referenced policy synthesis for a centralized US AGI project; the risk-reduction claim is weaker than the analysis admits, but the paper is worth engaging. read the letter →

arxiv 2507.21082 v1 pith:4CAFGYGI submitted 2025-06-17 cs.CY

classification cs.CY
keywords artificialgeneralintelligenceAIsafetygovernancepolicynationalsecuritycasespauseprotocolsverification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The report argues that if the US government ever centralized AGI development in a single project under its direct control, the risk of catastrophic outcomes—loss of control over systems more intelligent than humans, or egregious misuse of those systems—could be substantially reduced by deliberate institutional design. It identifies four high-level priorities and seven concrete safety features, from reporting requirements and dissent channels to board oversight and hardware-based verification. The strongest claim, stated in the conclusion, is that implementing some or all of these features could reduce catastrophic risk. A sympathetic reader should care because the report translates an abstract worry about AI risk into specific, implementable mechanisms that could be put in place before such a project exists.

What carries the argument

The load-bearing mechanism is the 'safety case' checkpoint system. Before pre-training, after pre-training, before internal deployment, and before each expansion of a model's permissions, technical teams must present an affirmative, evidence-backed argument that the system is safe enough, and a board of technical experts must approve proceeding by supermajority. Around this core sit the tripwire capabilities and intolerable risk thresholds that trigger reporting and emergency pauses, plus a designated point of contact who can execute a pause instruction from any staff member. This machinery is what turns the report's priorities into operational constraints.

What would settle it

A concrete falsifier would be a demonstration that a frontier model can pass every pre-deployment safety-case evaluation, including limit evals, while concealing a capability that later causes catastrophic harm—for example, a model that strategically underperforms on all dangerous-capability tests but exhibits situational awareness and deceptive goal-seeking once deployed.

Watch

Extended reading notes

Core claim

The paper's central claim is that a government-led AGI project can be made safer through institutional choices rather than only through technical alignment research. Its seven features are: information escalation with explicit reporting thresholds and protected dissent; pause protocols with bottom-up and top-down emergency authority; a technical board with binding authority over training and deployment decisions; separate internal audit and risk-monitoring bodies; an intelligence and scenario-planning division; a designated verification project for international agreements; and a plan for automated research. The report argues that each feature compensates for a specific failure mode of high-stakes technology development, and that a centralized project has a key advantage over private labs: it can standardize one set of pause thresholds without competitive pressure to lower them.

Load-bearing premise

The entire architecture presupposes that dangerous capabilities in advanced AI can be detected and evaluated reliably before they become catastrophic; the report itself concedes that the science of model evaluation may not mature quickly and that models can hide their abilities through sandbagging or deceptive behavior.

Editorial extensions

If this is right

  • A centralized project would halt at each defined checkpoint unless an affirmative safety case passes a board supermajority, so development would slow or stop whenever safety evidence is insufficient.
  • Reporting requirements triggered by regular cadence, compute scaling, and capability thresholds would surface risks that private-sector frameworks currently leave undefined.
  • Bottom-up pause authority, modelled on stop-work authority in other high-risk industries, would let technical staff halt training runs before senior leadership recognizes a threat.
  • A designated verification project, including hardware-enabled mechanisms on AI chips, would make international agreements to limit AI development potentially verifiable.
  • If evaluation science does not mature in time, adherence to these protocols would delay AGI development—a delay the report says the project should accept and plan for.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the report, the same safety-case and tripwire logic could be tested retroactively against existing frontier labs by asking how often their training runs would have triggered pauses under these thresholds.
  • The report's institutional features are likely transferable to a multilateral or private-consortium setting, not only a US government project, since the underlying failure modes—competitive pressure, opaque development, and unclear thresholds—are not unique to government.
  • A hidden implication is that the value of all seven features is epistemic, not mechanical: they buy time and force evidence, but they cannot compensate for a fundamental inability to evaluate a deceptive model.
  • One testable extension would be to run a sandboxed simulation of the pause-and-checkpoint protocol on an existing large model training run to measure how much latency and cost the safety cases would add.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper analyzes a hypothetical scenario in which the US government centralizes AGI development into a single project under its direct control, and proposes four high-level priorities (assessing alignment difficulty, maintaining optionality, envisioning the end-state, and a fourth implicit priority related to information flows) and seven safety features: information escalation, pause protocols, board oversight, audit and risk monitoring, a geopolitics division, a designated verification project, and a plan for automated research. The report draws heavily on analogies from nuclear regulation, aviation safety, manufacturing quality control, and existing frontier-AI safety frameworks, and it repeatedly acknowledges the scientific immaturity of model evaluation and safety cases. Its central claim is that implementing some or all of these features could reduce catastrophic risk from loss-of-control or egregious misuse, while explicitly accepting that adherence to the protocols may delay AGI development until scientific breakthroughs occur.

Significance. If the central claim is accepted, the paper provides a concrete institutional blueprint for a high-stakes policy decision, with actionable recommendations on reporting thresholds, emergency pauses, binding board authority, audit separation, verification technologies, and geopolitical contingency planning. Its strengths are the specificity of the proposed mechanisms, the use of documented precedents (NRC, CRITIC, DNFSB, RAND HEMs), and its unusual candor about the limits of current evaluation science and safety-case methodology. The paper does not present new empirical evidence or formal models, but it is a serious and well-grounded policy analysis. Its main weakness is that the loss-of-control risk-reduction claim relies on the reliability of capability detection, which the paper itself repeatedly concedes is not currently attainable; the paper needs to either articulate a non-detection-based mechanism for reducing loss-of-control risk or narrow its claim.

major comments (3)
  1. [§5.2.3, §6] The paper's central claim in §6 that implementing the proposed features 'could reduce catastrophic risk stemming from loss-of-control or egregious misuse' presupposes that dangerous capabilities can be detected before they become catastrophic. Yet §5.2.1 concedes that 'a mature science of model evaluation may not be possible in a short timeframe,' §5.2.3 states that 'building full safety cases for models significantly more advanced than today's is not yet possible,' and §5.4.2 acknowledges the lack of well-established methods for forecasting AI-caused catastrophes. In the regime the paper itself treats as live, the pause checkpoints and tripwires may only produce delay, and the paper does not explain how delay alone reduces loss-of-control risk when detection is unreliable. Please either specify a fallback mechanism that does not depend on capability evaluation (for example, compute-based or hardware-enforced hard ceilings) or restrict the conclusion to misuse risks and to risk reduction through slowing development.
  2. [§5.2.1] The report recommends that 'clear evidence' that evaluation science will not mature 'should itself trigger a pause in advancing model capabilities.' This tripwire is underspecified and potentially circular: recognizing that a mature evaluation science is impossible is itself an evaluative judgment subject to the same unreliability and expert disagreement the report documents elsewhere. Please define what evidence would qualify, who would make the determination, and what safeguards would prevent this determination from being overridden by the race dynamics described in the same section.
  3. [§5.3.2] The Board's binding authority is a load-bearing element of the design, but the report does not address the failure mode in which the Executive Branch, which controls the project, disregards or overrides a Board veto—for example, by ordering resumption of training after the Board votes against it. The NRC analogy is imperfect because the NRC regulates a private industry rather than an executive-controlled project; please discuss the legal and practical enforceability of Board decisions against the President or the project's head, or explain why the analogy remains apt despite this difference.
minor comments (5)
  1. [Throughout] There are repeated typographical errors in the name 'Anthropic,' spelled 'Antrophic' in several places (e.g., §5.2.1, §5.2.3, §5.3.1); these should be corrected.
  2. [§5.1] The sentence 'This will need to be efficiently escalated up to project leadership, and to the risk monitoring and internal compliance teams (see Figure 1)' lacks Figure 2; the organizational chart is Figure 2, while Figure 1 shows the 'How hard is AI safety?' graphic. The cross-reference should be fixed.
  3. [§1.2, §3, §5] The executive summary's numbering (1.2 High-level priorities, 1.3 Safety Features) does not match the body's numbering (Section 3 for priorities, Section 5 for safety features); the numbering and cross-references should be harmonized.
  4. [§2] The paper states it is 'informed by semi-structured interviews with experts in AI safety and governance,' but it provides no information on the number of interviewees, their selection, the interview questions, or the analysis method, so the reader cannot assess this evidence base.
  5. [§5.2.3] The phrase 'in each of in a list of predetermined thresholds' contains a grammatical error; it should read 'at each checkpoint in a list of predetermined thresholds.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the report's recommendations are anchored to external precedents and its central conclusion is a conditional deduction from openly stated premises.

full rationale

This paper is a qualitative policy report with no equations, fitted parameters, or formal derivations. Its seven safety features are justified by external institutional precedents (NRC, CRITIC, RAND HEMs, frontier-lab safety frameworks, DNFSB, SIGAR) and by semi-structured expert interviews, rather than by importing conclusions from the author's own prior work. The central conclusion that pauses are likely to be necessary is a straightforward conditional: if the project adopts an affirmative safety standard, and if watertight safety cases are not yet possible for advanced models, then progression to later checkpoints will be blocked. The report explicitly acknowledges the fragility of its own premise regarding evaluation science, noting in Section 5.2.1 that 'a mature science of model evaluation may not be possible in a short timeframe' and in Section 5.2.3 that 'building full safety cases for models significantly more advanced than today's is not yet possible.' Those admissions weaken the report's practical risk-reduction claim, but they do not make the reasoning circular; they identify an unverified assumption on which the recommendations depend. No self-citations are load-bearing, no known result is renamed as a new framework, and no fitted input is relabeled as a prediction. The appropriate finding is therefore no significant circularity.

Assumptions & free parameters 0 free parameters · 6 assumptions · 4 invented entities

The report has no mathematical or empirical derivation, so the ledger captures the assumptions its recommendations depend on. Free parameters: none, because there are no fitted numbers. Axioms: the central axioms are that AGI risk is catastrophic, that centralization is plausible enough to plan for, that reliable capability evaluation can be developed or that safety cases can be made, that staff pause authority will not be overturned by insiders or political pressure, and that verification technology can be deployed in time. Invented entities: the proposed institutions such as the Board, pause point of contact, risk monitoring team, and verification project are treated as new organizational constructs introduced by the paper; none has independent evidence of effectiveness.

assumptions (6)
  • domain assumption Uncontrolled AGI would pose extreme, potentially existential global risks.
    Motivates the whole report; asserted in Section 2 with citations to expert statements, not demonstrated.
  • domain assumption A US government-led centralized AGI project is plausible enough to justify detailed institutional design.
    Section 2 footnote discusses Defense Production Act authorities but acknowledges feasibility is contested.
  • domain assumption Dangerous capabilities can be detected early enough for reporting and pause triggers to matter.
    Required for features 1 and 2; Section 5.2.3 notes evaluations are unreliable and models may sandbag or deceive.
  • domain assumption An affirmative safety case standard is an appropriate and enforceable gateway for proceeding.
    Section 5.2.3 requires affirmative safety cases before checkpoints, while admitting such cases are not currently possible for advanced models.
  • domain assumption The Board and oversight bodies will retain real authority under political and competitive pressure.
    Section 5.3.2 proposes congressional authorization, but the report acknowledges a clean separation between technical safety and strategic use is not always feasible.
  • domain assumption Hardware-enabled verification mechanisms can be developed and deployed before or during a treaty window.
    Section 5.5.2 cites RAND that such mechanisms are possible but may take years; no firm timeline is given.
invented entities (4)
  • A binding technical Safety Board
    purpose: To approve or disapprove training runs and internal deployment based on safety cases.
    Proposed in Section 5.3; no existing equivalent with an AGI mandate.
  • A designated point of contact for pause requests
    purpose: Sole intermediary authorized to execute a halt instruction from technical staff.
    Proposed in Section 5.2.2 to streamline bottom-up pauses.
  • A risk monitoring team separate from risk management
    purpose: To quantify catastrophic and existential risk at checkpoints.
    Proposed in Section 5.4.2; relies on forecasting methods that are explicitly immature.
  • A designated verification project with hardware-enabled mechanisms
    purpose: To verify international agreements limiting AI development.
    Proposed in Section 5.5; hardware-enabled mechanisms are at concept stage.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Safety Features for a Centralised AGI Project." pith.science (2026). https://pith.science/paper/4CAFGYGI

@misc{pith2026250721082,
  author       = {Pith},
  title        = {Pith review of: Safety Features for a Centralised AGI Project},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4CAFGYGI}},
  note         = {Machine review of arXiv:2507.21082}
}
read the original abstract

Recent AI progress has outpaced expectations, with some experts now predicting AI that matches or exceeds human capabilities in all cognitive areas (AGI) could emerge this decade, potentially posing grave national and global security threats. AI development is currently occurring primarily in the private sector with minimal oversight. This report analyzes a scenario where the US government centralizes AGI development under its direct control, and identifies four high-level priorities and seven safety features to reduce risks.

Figures

Figures reproduced from arXiv: 2507.21082 by the authors.

Figure 1
Figure 1. How hard is AI safety? [10] 3 High-level priorities This section proposes four guiding principles that should inform the design and execution of a centralized AGI project. They rest on the assumption that AGI is likely to pose extreme risks that threaten global security if not developed cautiously. 3.1 Assess the difficulty of alignment A central priority of the project should be assessing the difficulty of aligning… view at source ↗
Figure 2
Figure 2. Organizational chart 4 Organizational structure and information flows [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

87 extracted references · 65 canonical work pages

  1. [1]

    OpenAI o3 and o3-mini—12 Days of OpenAI: Day 12,

    OpenAI, “OpenAI o3 and o3-mini—12 Days of OpenAI: Day 12,” Dec. 2024. [Online]. Available: https://www.youtube.com/watch?v=SKBG1sqdyIU

  2. [2]

    Can AI Scaling Continue Through 2030?

    J. Sevilla, “Can AI Scaling Continue Through 2030?” Aug. 2024. [Online]. Available: https: //epoch.ai/blog/can-ai-scaling-continue-through-2030

  3. [3]

    Statement on AI Risk | CAIS,

    Center for AI Safety, “Statement on AI Risk | CAIS,” May 2023. [Online]. Available: https: //safe.ai/work/statement-on-ai-risk

  4. [4]

    Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence,

    J. R. Biden, “Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence,” Oct. 2023, executive Order 14110. [Online]. Available: https://www.federalregister.gov/documents/2023/11/01/2023-24283/ safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence

  5. [5]

    Removing Barriers to American Leadership in Artificial Intelligence,

    D. J. Trump, “Removing Barriers to American Leadership in Artificial Intelligence,” Jan. 2025, executive Order signed January 23, 2025. [Online]. Available: https://www.whitehouse.gov/presidential-actions/2025/01/ removing-barriers-to-american-leadership-in-artificial-intelligence/

  6. [6]

    Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models,

    S. Nevo, D. Lahav, A. Karpur, Y . Bar-On, H. A. Bradley, and J. Alstott, “Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models,” Tech. Rep., May 2024. [Online]. Available: https://www.rand.org/pubs/research_reports/RRA2849-1.html

  7. [7]

    2024 REPORT TO CONGRESS of the U.S.-CHINA ECONOMIC AND SECURITY REVIEW COMMISSION,

    U.S.-China Economic and Security Review Commission, “2024 REPORT TO CONGRESS of the U.S.-CHINA ECONOMIC AND SECURITY REVIEW COMMISSION,” Nov. 2024, publisher: U.S.-China Economic and Security Review Commission. [Online]. Available: https://www.uscc.gov/sites/default/files/2024-11/2024_ Annual_Report_to_Congress.pdf

  8. [8]

    Soft Nationalization: How the US Government Will Control AI Labs | Convergence Analysis,

    D. Cheng and C. Katzke, “Soft Nationalization: How the US Government Will Control AI Labs | Convergence Analysis,” Aug. 2024. [Online]. Available: https://www.convergenceanalysis.org/publications/ soft-nationalization-how-the-us-government-will-control-ai-labs

Show all 87 references
  1. [9]

    Core Views on AI Safety: When, Why, What, and How,

    Anthropic, “Core Views on AI Safety: When, Why, What, and How,” Mar. 2023. [Online]. Available: https://www.anthropic.com/news/core-views-on-ai-safety

  2. [10]

    How hard is AI safety?

    C. Olah, “How hard is AI safety?” Jul. 2023. [Online]. Available: https://x.com/ch402/status/ 1666482929772666880?lang=en

  3. [11]

    Defining and Characterizing Reward Hacking,

    J. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger, “Defining and Characterizing Reward Hacking,” Mar. 2025, arXiv:2209.13085. [Online]. Available: http://arxiv.org/abs/2209.13085

  4. [12]

    AI Deception: A Survey of Examples, Risks, and Potential Solutions,

    P. S. Park, S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks, “AI Deception: A Survey of Examples, Risks, and Potential Solutions,” Aug. 2023. [Online]. Available: https://arxiv.org/abs/2308.14752v1

  5. [13]

    EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models,

    W. Zhou, X. Wang, L. Xiong, H. Xia, Y . Gu, M. Chai, F. Zhu, C. Huang, S. Dou, Z. Xi, R. Zheng, S. Gao, Y . Zou, H. Yan, Y . Le, R. Wang, L. Li, J. Shao, T. Gui, Q. Zhang, and X. Huang, “EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models,” Mar. 2024, arX...

  6. [14]

    Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence Strategy,

    D. Hendrycks, E. Schmidt, and A. Wang, “Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence Strategy,” 2024. [Online]. Available: https://www.nationalsecurity.ai/chapter/ deterrence-with-mutual-assured-ai-malfunction-maim

  7. [15]

    AI Behind Closed Doors: a Primer on The Governance of Internal Deployment,

    C. Stix, M. Pistillo, G. Sastry, M. Hobbhahn, A. Ortega, M. Balesni, A. Hallensleben, N. Goldowsky-Dill, and K. Sharkey, “AI Behind Closed Doors: a Primer on The Governance of Internal Deployment,” Apr. 2025, arXiv:2504.12170. [Online]. Available: http://arxiv.org/abs/2504.12170

  8. [16]

    Evaluating frontier AI R&D capabilities of language model agents against human experts,

    METR, “Evaluating frontier AI R&D capabilities of language model agents against human experts,”METR Blog, Nov. 2024. [Online]. Available: https://metr.org/blog/2024-11-22-evaluating-r-d-capabilities-of-llms/

  9. [17]

    Measuring the Persuasiveness of Language Models,

    E. Durmus, L. Lovitt, A. Tamkin, S. Ritchie, J. Clark, and D. Ganguli, “Measuring the Persuasiveness of Language Models,” Apr. 2024, publisher: Anthropic. [Online]. Available: https://www.anthropic.com/research/ measuring-model-persuasiveness

  10. [18]

    Emergent Abilities of Large Language Models,

    J. Wei, Y . Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent Abilities of Large Language Models,” Oct. 2022, arXiv:2206.07682. [Online]. Available: ht...

  11. [19]

    Training Compute Thresholds — Features and Functions in AI Regulation,

    L. Heim, “Training Compute Thresholds — Features and Functions in AI Regulation,” Apr. 2024. [Online]. Available: https://blog.heim.xyz/training-compute-thresholds/ 23 SAFETY FEATURES FOR A CENTRALIZEDAGIPROJECT

  12. [20]

    Will there be a discontinuity in AI capabilities?

    AI Safety Info, “Will there be a discontinuity in AI capabilities?” 2024. [Online]. Available: https://aisafety.info/questions/7729/Will-there-be-a-discontinuity-in-AI-capabilities

  13. [21]

    Response to BIS AI Reporting Requirements RFC — MIRI Technical Governance Team,

    MIRI Technical Governance Team, “Response to BIS AI Reporting Requirements RFC — MIRI Technical Governance Team,” 2025. [Online]. Available: https://techgov.intelligence.org/research/ response-to-bis-ai-reporting-requirements-rfc

  14. [22]

    AI Risk Management Framework,

    National Institute of Standards and Technology, “AI Risk Management Framework,”NIST, Jan. 2023. [Online]. Available: https://www.nist.gov/itl/ai-risk-management-framework

  15. [23]

    Common Elements of Frontier AI Safety Policies,

    METR, “Common Elements of Frontier AI Safety Policies,”METR Blog, Mar. 2025. [Online]. Available: https://metr.org/blog/2025-03-26-common-elements-of-frontier-ai-safety-policies/

  16. [24]

    Announcing our updated Responsible Scaling Policy,

    Anthropic, “Announcing our updated Responsible Scaling Policy,” Oct. 2024. [Online]. Available: https://www.anthropic.com/news/announcing-our-updated-responsible-scaling-policy

  17. [25]

    xAI Risk Management Framework (Draft),

    xAI, “xAI Risk Management Framework (Draft),” Feb. 2025, publisher: xAI. [Online]. Available: https://x.ai/documents/2025.02.20-RMF-Draft.pdf

  18. [26]

    Intolerable Risk Threshold Recommendations for Artificial Intelligence,

    D. Raman, N. Madkour, E. Murphy, K. Jackson, and J. Newman, “Intolerable Risk Threshold Recommendations for Artificial Intelligence,” Jan. 2025. [Online]. Available: https://cltc.berkeley.edu/publication/ intolerable-ai-risk-thresholds/

  19. [27]

    5 FAH-2 H-430 HANDLING SYMBOLS,

    U.S. Department of State, “5 FAH-2 H-430 HANDLING SYMBOLS,” 2024. [Online]. Available: https://fam.state.gov/fam/05fah02/05fah020430.html

  20. [28]

    Notes on the Critic System,

    W. Tidewell, “Notes on the Critic System,” 1980. [Online]. Available: https://www.cia.gov/resources/csi/static/ Notes-on-Critic-System.pdf

  21. [29]

    Handling of Critical (CRITIC) Information,

    Central Intelligence Agency, “Handling of Critical (CRITIC) Information,” Jul. 1979, publisher: Central Intelli- gence Agency. [Online]. Available: https://www.cia.gov/readingroom/docs/CIA-RDP83-00156R000200040001-1. pdf

  22. [30]

    2 FAM 070 DISSENT CHANNEL,

    U.S. Department of State, “2 FAM 070 DISSENT CHANNEL,” 2024, publisher: US Department of State. [Online]. Available: https://fam.state.gov/fam/02fam/02fam0070.html

  23. [31]

    NRC DIFFERING PROFESSIONAL OPINION PROGRAM,

    Nuclear Regulatory Commission, “NRC DIFFERING PROFESSIONAL OPINION PROGRAM,” Aug. 2015, publisher: Nuclear Regulatory Commission. [Online]. Available: https://www.nrc.gov/docs/ml1513/ml15132a664. pdf

  24. [32]

    DOE Differing Professional Opinions,

    U.S. Department of Energy, “DOE Differing Professional Opinions,” 2024. [Online]. Available: https://www.energy.gov/ehss/doe-differing-professional-opinions

  25. [33]

    Stifling Dissent,

    D. Van Schooten and N. Schwellenbach, “Stifling Dissent,” 2021. [Online]. Available: https: //www.pogo.org/reports/stifling-dissent

  26. [34]

    Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts,

    S. Field, “Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts,” Jan. 2025, arXiv:2502.14870. [Online]. Available: http://arxiv.org/abs/2502.14870

  27. [35]

    AI Whistleblowers,

    H. Wu, “AI Whistleblowers,” Mar. 2025. [Online]. Available: https://papers.ssrn.com/sol3/papers.cfm?abstract_ id=4790511

  28. [36]

    Our updated Preparedness Framework,

    OpenAI, “Our updated Preparedness Framework,” Dec. 2024, publisher: OpenAI. [Online]. Available: https://openai.com/index/updating-our-preparedness-framework/

  29. [37]

    A Sketch of Potential Tripwire Capabili- ties for AI,

    Carnegie Endowment for International Peace, “A Sketch of Potential Tripwire Capabili- ties for AI,” Dec. 2024. [Online]. Available: https://carnegieendowment.org/research/2024/12/ a-sketch-of-potential-tripwire-capabilities-for-ai?lang=en

  30. [38]

    Risk Thresholds for Frontier AI | GovAI,

    GovAI, “Risk Thresholds for Frontier AI | GovAI,” 2024. [Online]. Available: https://www.governance.ai/ research-paper/risk-thresholds-for-frontier-ai

  31. [39]

    AI models can be dangerous before public deployment,

    METR, “AI models can be dangerous before public deployment,”METR Blog, Jan. 2025. [Online]. Available: https://metr.org/blog/2025-01-17-ai-models-dangerous-before-public-deployment/

  32. [40]

    The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models,

    A. Pan, K. Bhatia, and J. Steinhardt, “The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models,” Feb. 2022, arXiv:2201.03544. [Online]. Available: http://arxiv.org/abs/2201.03544

  33. [41]

    Towards Understanding Sycophancy in Language Models,

    M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, N. Cheng, E. Durmus, Z. Hatfield-Dodds, S. R. Johnston, S. Kravec, T. Maxwell, S. McCandlish, K. Ndousse, O. Rausch, N. Schiefer, D. Yan, M. Zhang, and E. Perez, “Towards Understanding Sycophancy in Language ...

  34. [42]

    Andon - Toyota Production System guide,

    Toyota UK Magazine, “Andon - Toyota Production System guide,” May 2016, publisher: Toyota UK Magazine. [Online]. Available: https://mag.toyota.co.uk/andon-toyota-production-system/

  35. [43]

    Stop Work Authority: How It Works,

    TRADESAFE, “Stop Work Authority: How It Works,” Sep. 2024. [Online]. Available: https: //trdsf.com/blogs/news/stop-work-authority

  36. [44]

    Whirlpool Corp. v. Marshall, 445 U.S. 1 (1980),

    U.S. Supreme Court, “Whirlpool Corp. v. Marshall, 445 U.S. 1 (1980),” 1980. [Online]. Available: https://supreme.justia.com/cases/federal/us/445/1/

  37. [45]

    OSH Act of 1970,

    U.S. Congress, “OSH Act of 1970,” Dec. 1970, publisher: Occupational Safety and Health Administration

  38. [46]

    Racing to the precipice: a model of artificial intelligence development,

    S. Armstrong, N. Bostrom, and C. Shulman, “Racing to the precipice: a model of artificial intelligence development,”AI & SOCIETY, vol. 31, no. 2, pp. 201–206, May 2016. [Online]. Available: https://doi.org/10.1007/s00146-015-0590-y

  39. [47]

    Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs,

    R. Laine, B. Chughtai, J. Betley, K. Hariharan, J. Scheurer, M. Balesni, M. Hobbhahn, A. Meinke, and O. Evans, “Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs,” Jul. 2024, arXiv:2407.04694. [Online]. Available: http://arxiv.org/abs/2407.04694

  40. [48]

    AI Sandbagging: Language Models can Strategically Underperform on Evaluations,

    T. v. d. Weij, F. Hofstätter, O. Jaffe, S. F. Brown, and F. R. Ward, “AI Sandbagging: Language Models can Strategically Underperform on Evaluations,” Feb. 2025, arXiv:2406.07358. [Online]. Available: http://arxiv.org/abs/2406.07358

  41. [49]

    Alignment faking in large language models,

    R. Greenblatt, C. Denison, B. Wright, F. Roger, M. MacDiarmid, S. Marks, J. Treutlein, T. Belonax, J. Chen, D. Duvenaud, A. Khan, J. Michael, S. Mindermann, E. Perez, L. Petrini, J. Uesato, J. Kaplan, B. Shlegeris, S. R. Bowman, and E. Hubinger, “Alignment faking in large lang...

  42. [50]

    Safety Goals for Nuclear Power Plant Operation,

    Nuclear Regulatory Commission, “Safety Goals for Nuclear Power Plant Operation,” 1986, publisher: Nuclear Regulatory Commission. [Online]. Available: https://www.nrc.gov/docs/ML0717/ML071770230.pdf

  43. [51]

    Affirmative safety: An approach to risk management for high-risk AI,

    A. R. Wasil, J. Clymer, D. Krueger, E. Dardaman, S. Campos, and E. R. Murphy, “Affirmative safety: An approach to risk management for high-risk AI,” Apr. 2024. [Online]. Available: https://arxiv.org/abs/2406.15371v1

  44. [52]

    Towards A Rigorous Science of Interpretable Machine Learning,

    F. Doshi-Velez and B. Kim, “Towards A Rigorous Science of Interpretable Machine Learning,” Feb. 2017. [Online]. Available: https://arxiv.org/abs/1702.08608v2

  45. [53]

    Provably safe systems: the only path to controllable AGI,

    M. Tegmark and S. Omohundro, “Provably safe systems: the only path to controllable AGI,” Sep. 2023, arXiv:2309.01933. [Online]. Available: http://arxiv.org/abs/2309.01933

  46. [54]

    Backgrounder on Emergency Preparedness at Nuclear Power Plants,

    U.S. Nuclear Regulatory Commission, “Backgrounder on Emergency Preparedness at Nuclear Power Plants,” 2024. [Online]. Available: https://www.nrc.gov/reading-rm/doc-collections/fact-sheets/emerg-plan-prep-nuc-power. html

  47. [55]

    Safety cases at AISI | AISI Work,

    UK AI Safety Institute, “Safety cases at AISI | AISI Work,” 2024. [Online]. Available: https: //www.aisi.gov.uk/work/safety-cases-at-aisi

  48. [56]

    National Defense Authorization Act for Fiscal Year 1989,

    U.S. Congress, “National Defense Authorization Act for Fiscal Year 1989,” Mar. 1988. [Online]. Available: https://www.congress.gov/bill/100th-congress/house-bill/4264/summary/17

  49. [57]

    42 U.S. Code § 5841 - Establishment and transfers,

    ——, “42 U.S. Code § 5841 - Establishment and transfers,” 1974. [Online]. Available: https: //www.law.cornell.edu/uscode/text/42/5841

  50. [58]

    Survey of 2,778 AI authors: six parts in pictures,

    K. Grace, “Survey of 2,778 AI authors: six parts in pictures,” 2024. [Online]. Available: https: //blog.aiimpacts.org/p/2023-ai-survey-of-2778-six-things

  51. [59]

    "Existential risk from AI

    R. Bensinger, “"Existential risk from AI" survey results,” Jun. 2021. [Online]. Available: https: //www.alignmentforum.org/posts/QvwSr5LsxyDeaPK5s/existential-risk-from-ai-survey-results

  52. [60]

    Anthropic’s Recommendations to OSTP for the U.S. AI Action Plan,

    Anthropic, “Anthropic’s Recommendations to OSTP for the U.S. AI Action Plan,” Jan. 2025. [Online]. Available: https://www.anthropic.com/news/anthropic-s-recommendations-ostp-u-s-ai-action-plan

  53. [61]

    Tweet by @nabla_theta,

    L. Gao, “Tweet by @nabla_theta,” Dec. 2024. [Online]. Available: https://x.com/nabla_theta/status/ 1869144832595431553

  54. [62]

    AI-Enabled Coups: How a Small Group Could Use AI to Seize Power,

    T. Davidson, L. Finnveden, and R. Hadshar, “AI-Enabled Coups: How a Small Group Could Use AI to Seize Power,” Apr. 2025. [Online]. Available: https://www.forethought.org/research/ ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power

  55. [63]

    New Survey: Broad Expert Consensus for Many AGI Safety and Governance Practices | GovAI,

    J. Schuett, N. Dreksler, M. Anderljung, D. McCaffary, L. Heim, E. Bluemke, and B. Garfinkel, “New Survey: Broad Expert Consensus for Many AGI Safety and Governance Practices | GovAI,” Jun. 2023. [Online]. Available: https: //www.governance.ai/analysis/broad-expert-consensus-fo...

  56. [64]

    Inspector General Act of 1978,

    U.S. Congress, “Inspector General Act of 1978,” 1978. [Online]. Available: https://www.congress.gov/bill/ 95th-congress/house-bill/8588

  57. [65]

    Operation Warp Speed: Accelerated COVID-19 Vaccine Development Status and Efforts to Address Manufacturing Challenges | U.S. GAO,

    U.S. Government Accountability Office, “Operation Warp Speed: Accelerated COVID-19 Vaccine Development Status and Efforts to Address Manufacturing Challenges | U.S. GAO,” Feb. 2021. [Online]. Available: https://www.gao.gov/products/gao-21-319

  58. [66]

    Update on ARC’s recent eval efforts,

    Alignment Research Center, “Update on ARC’s recent eval efforts,”METR Blog, Mar. 2023. [Online]. Available: https://metr.org/blog/2023-03-18-update-on-recent-evals/

  59. [67]

    Introducing Alignment Stress-Testing at Anthropic,

    evhub, “Introducing Alignment Stress-Testing at Anthropic,” Jan. 2024. [Online]. Available: https: //www.alignmentforum.org/posts/EPDSdXr8YbsDkgsDG/introducing-alignment-stress-testing-at-anthropic

  60. [68]

    Distribution Shifts and The Importance of AI Safety,

    L. Lang, “Distribution Shifts and The Importance of AI Safety,” Sep. 2022. [Online]. Available: https: //www.alignmentforum.org/posts/TRKF9g65nhPBQoxJu/distribution-shifts-and-the-importance-of-ai-safety

  61. [69]

    Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy,

    B. Shlegeris and R. Greenblatt, “Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy,” Jul. 2023. [Online]. Available: https://www.alignmentforum.org/posts/ MbWWKbyD5gLhJgfwn/meta-level-adversarial-evaluation-of-oversight-...

  62. [70]

    Organizational Arrangements for Risk Assessment,

    National Research Council (US) Committee on the Institutional Means for Assessment of Risks to Public Health, “Organizational Arrangements for Risk Assessment,” inRisk Assessment in the Federal Government: Managing the Process. National Academies Press (US), 1983. [Online]. Av...

  63. [71]

    Working together to face humanity’s greatest threats: Introduction to The Future of Research on Catastrophic and Existential Risk

    A. M. Currie and S. Ó hÉigeartaigh, “Working together to face humanity’s greatest threats: Introduction to The Future of Research on Catastrophic and Existential Risk.” Sep. 2018. [Online]. Available: https://www.repository.cam.ac.uk/handle/1810/280193

  64. [72]

    Existential Risk Prevention as a Global Priority,

    N. Bostrom, “Existential Risk Prevention as a Global Priority,” Feb. 2013. [Online]. Available: https://existential-risk.com/concept.pdf

  65. [73]

    Roots of Disagreement on AI Risk: Exploring the Potential and Pitfalls of Adversarial Collaboration,

    J. Rosenberg, E. Karger, A. Morris, M. Hickman, R. Hadshar, Z. Jacobs, and P. Tetlock, “Roots of Disagreement on AI Risk: Exploring the Potential and Pitfalls of Adversarial Collaboration,” 2022, publisher: Forecasting Research Institute. [Online]. Available: https://static1.s...

  66. [74]

    Superhuman Automated Forecasting | CAIS,

    Center for AI Safety, “Superhuman Automated Forecasting | CAIS,” Jan. 2025. [Online]. Available: https://safe.ai/blog/forecasting

  67. [75]

    Technical Options for Flexible Hardware-Enabled Guarantees,

    J. Petrie and O. Aarne, “Technical Options for Flexible Hardware-Enabled Guarantees,” Jun. 2025. [Online]. Available: https://arxiv.org/abs/2506.03409v1

  68. [76]

    Verification methods for international AI agreements,

    A. R. Wasil, T. Reed, J. W. Miller, and P. Barnett, “Verification methods for international AI agreements,” Aug

  69. [77]

    Global Security Remote Sensing and Verification,

    Sandia National Laboratories, “Global Security Remote Sensing and Verification,” 2024. [Online]. Available: https://www.sandia.gov/missions/global-security-remote-sensing-and-verification/

  70. [78]

    Verification and other safeguards activities,

    International Atomic Energy Agency, “Verification and other safeguards activities,” Jun. 2016. [Online]. Available: https://www.iaea.org/topics/verification-and-other-safeguards-activities

  71. [79]

    Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090,

    G. Kulp, D. Gonzales, E. Smith, L. Heim, P. Puri, M. J. D. Vermeer, and Z. Winkelman, “Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090,” Tech. Rep., Jan. 202...

  72. [80]

    Why policy makers should beware claims of new ’arms races’,

    H. Belfield and C. Ruhl, “Why policy makers should beware claims of new ’arms races’,” Jul. 2022, publisher: July 2022. [Online]. Available: https://thebulletin.org/2022/07/ why-policy-makers-should-beware-claims-of-new-arms-races/

  73. [81]

    Who is behind DeepSeek and how did it achieve its AI ’Sputnik moment’?

    A. Hawkins, “Who is behind DeepSeek and how did it achieve its AI ’Sputnik moment’?” The Guardian, Jan. 2025. [Online]. Available: https://www.theguardian.com/technology/2025/jan/28/ who-is-behind-deepseek-and-how-did-it-achieve-its-ai-sputnik-moment

  74. [82]

    On DeepSeek and Export Controls,

    D. Amodei, “On DeepSeek and Export Controls,” Jan. 2025. [Online]. Available: https://www.darioamodei.com/ post/on-deepseek-and-export-controls

  75. [83]

    What Is DeepSeek? New Chinese Artificial Intelligence Rivals Chat- GPT, OpenAI,

    M. W. Roeloffs, “What Is DeepSeek? New Chinese Artificial Intelligence Rivals Chat- GPT, OpenAI,” Jan. 2025. [Online]. Available: https://www.forbes.com/sites/maryroeloffs/2025/01/27/ what-is-deepseek-new-chinese-ai-startup-rivals-openai-and-claims-its-far-cheaper/ 26 SAFETY F...

  76. [84]

    Trends in U.S. Intention-to-Stay Rates of International Ph.D. Gradu- ates Across Nationality and STEM Fields,

    R. Zwetsloot, J. Feldgoise, and J. Dunham, “Trends in U.S. Intention-to-Stay Rates of International Ph.D. Gradu- ates Across Nationality and STEM Fields,” Sep. 2019. [Online]. Available: https://cset.georgetown.edu/publication/ trends-in-u-s-intention-to-stay-rates-of-internat...

  77. [85]

    AI Pioneer Geoffrey Hinton Talks About AI Gaining Con- trol,

    A. Morris, “AI Pioneer Geoffrey Hinton Talks About AI Gaining Con- trol,” May 2023. [Online]. Available: https://www.forbes.com/sites/andreamorris/2023/05/03/ ai-pioneer-geoffrey-hinton-talks-at-mit-about-ai-gaining-control/

  78. [86]

    Introducing Superalignment,

    OpenAI, “Introducing Superalignment,” Jul. 2023. [Online]. Available: https://openai.com/index/ introducing-superalignment/ 27

  79. [2024]

    Available: https://arxiv.org/abs/2408.16074v2

    [Online]. Available: https://arxiv.org/abs/2408.16074v2

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.