Pith. sign in

REVIEW 1 major objections 5 minor 47 references

Bare Minimum Mitigations for Autonomous AI Development

T0 review · 1 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Frontier AI labs should put four minimum safeguards in place before AI agents automate most internal R&D or gain rapid self-improvement to catastrophic capability.

desk verdict A concrete, operational threshold package for autonomous AI R&D that is worth engaging seriously, even though the 'act now' urgency rests on a two-year lead-time claim the paper doesn't support. read the letter →

arxiv 2504.15416 v2 pith:4AKGTRHY submitted 2025-04-21 cs.CY

classification cs.CY
keywords autonomousAIR&Dfrontiersafetyagentscomputemisusedetectioncatastrophicriskdisclosuremodelweightsecuritygovernancethresholdsresponsiblescaling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that frontier AI developers should implement four minimum safeguards before two capability thresholds are crossed: before AI agents automate most internal research and engineering, and before AI agents can rapidly improve to catastrophic capabilities with little compute and human assistance. The safeguards are understanding safety-critical training and assurance details, detecting internal AI agents that egregiously misuse compute, rapidly disclosing catastrophic risks to home governments, and preventing theft of critical AI software. The thresholds are framed as communication tools rather than triggers, with the immediate takeaway that most safety practices require over two years of preparation, so the time to act is now. A sympathetic reader would care because the paper tries to turn broad warnings about self-improving AI into concrete, implementable obligations tied to observable capability milestones.

What carries the argument

The central machinery is a pair of capability thresholds paired with threat models. Threshold One defines the automation of most internal research and engineering in terms of a productivity-loss comparison against laying off half of a lab's engineers, and it triggers recommendations to maintain understanding of safety-critical details and to detect egregious compute misuse. Threshold Two defines rapid improvement to catastrophic capabilities in terms of speed (about a year), compute (under 100x frontier training compute), and human assistance (a handful of people with generic skills), and it triggers disclosure to home governments and protection of critical AI software from theft. The threat models that carry the argument are safety sabotage, unauthorized internal deployment, adaptation lag, and capability proliferation, each explaining how a specific failure becomes possible once a threshold is crossed.

What would settle it

A falsifying observation would be a sustained plateau: if, over the next several years, the time horizon of tasks AI agents can complete autonomously stops growing and the share of lab R&D productivity attributable to agents remains far below the lay-off-half-of-engineers threshold, with no agent demonstrating a low-compute software-only path to catastrophic capability, the paper's central timeline would fail.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that the risks of autonomous AI R&D fall into two categories—harms from automating AI development itself and harms from rapid autonomous improvement—and that each category can be managed by a specific minimum safeguard if adopted before a defined threshold. Threshold One is crossed when the productivity loss from not using AI agents in internal software R&D exceeds the loss from laying off half of the lab's software engineers and researchers. Threshold Two is crossed when AI agents can improve to catastrophic capabilities within about a year, using less than 100x the compute of frontier training and no more than a handful of people with generic technical skills as assistance. The paper argues that these thresholds define natural deadlines for preparing safety, security, and governance measures, and that most such measures take years to build.

Load-bearing premise

The argument rests on the premise that AI agents will meaningfully automate most lab R&D or reach rapid catastrophic improvement soon enough that preparations begun now are the ones that matter; if agent capability growth stalls or long-horizon benchmarks overestimate progress, the urgency and the deadlines both recede.

Editorial extensions

If this is right

  • If Threshold One is near, labs should already be building monitoring of training data, experimental code, and internal agent behavior so they can detect sabotage and unauthorized compute use before agents control most R&D.
  • If Threshold Two is crossed, governments need visibility into internally deployed models, not just pre-deployment evaluations, to avoid adaptation lag.
  • Measures to secure AI software against well-resourced cyberattacks and insider threats take years to build, so they should be planned before rapid self-improvement becomes possible.
  • The thresholds, used as communication tools, give researchers and policymakers a shared language for when safety practices are no longer optional.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the productivity-loss definition of Threshold One is measurable in principle—a lab could track what fraction of its R&D output would vanish if agents were removed—which makes the threshold a testable governance target rather than a rhetorical one.
  • Editorial inference: because the paper treats rapid improvement as software-driven, progress on algorithmic efficiency rather than raw compute is the leading indicator to watch; if efficiency gains plateau, Threshold Two recedes.
  • Editorial inference: the adaptation-lag argument suggests a natural extension to a near-miss disclosure regime, since waiting until catastrophic capabilities are demonstrated may already be too late for coordinated responses.
  • Editorial inference: a natural next test is whether international coordination can form around the two thresholds, since the paper relies on home-government disclosure but leaves open how governments would share and act on those disclosures across borders.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. This policy paper argues that autonomous AI research and development (R&D) is a near-term possibility and identifies two broad risk categories: risks from automating AI R&D itself (safety sabotage, unauthorized internal deployment) and risks from rapid autonomous improvement (adaptation lag, capability proliferation). To address these, it defines two thresholds—Threshold One, when AI agents automate most internal research and engineering, and Threshold Two, when AI agents can rapidly improve to catastrophic capabilities with little compute and human assistance—and proposes four minimum safeguards for frontier AI developers to implement before these thresholds are reached: (1) thoroughly understand safety-critical training and assurance details, (2) detect internal AI agents egregiously misusing compute, (3) rapidly disclose catastrophic risks to home governments, and (4) prevent theft or exfiltration of critical AI software. The thresholds are explicitly framed as communication tools rather than triggers, and the paper concludes that because most safety practices require over two years of preparation, action should begin now.

Significance. The paper's main contribution is translating a broad 'red line' concern about autonomous AI R&D into concrete, named recommendations with definitions, threat models, and implementation indicators. The authors are appropriately transparent about uncertainty: Section 3.2 labels recursive improvement scenarios speculative, footnote 2 acknowledges that benchmark extrapolation may over- or underestimate progress, and Section 3.3 notes that proliferation can have benefits. The two-threshold structure is a useful coordination device, and Recommendation Four is grounded in an external detailed cost estimate (Nevo et al., 2024). If one accepts the threat model, the recommendations are coherent and largely follow from prior cited analyses; the paper does not derive them in a circular way. The central weakness is the unsupported lead-time claim underlying the 'act now' conclusion, which is load-bearing because the thresholds are explicitly not triggers for action.

major comments (1)
  1. [Section 1.4 and Section 3.3] The sentence 'Most safety practices require over two years of preparation, so the time to act is now' is the only explicit action-forcing conclusion in the paper, since the thresholds are described as 'communication tools, not triggers for action' (Section 1.4). However, the cited timeline evidence concerns only Recommendation Four, where Nevo et al. (2024) is invoked for 'years of concerted effort' (Section 3.3). No timeline evidence is offered for Recommendations One, Two, or Three; these might be implementable in months or could also require years. For example, Section 2.4 itself recommends that control and security measures be 'incrementally enhanced' as automation increases, which points away from a hard two-year lead time for the full package. The authors should either provide evidence for the multi-year preparation claim across the recommended measures or temper the 'time to act now' conclusion to match the support actually provided.
minor comments (5)
  1. [Section 1.1, Figure 2] The text states that projecting trends in Figure 2 indicates that by early 2027 AI agents might complete week-long software engineering tasks, but no quantitative projection method, data points, or uncertainty intervals are given; a brief description of the extrapolation would help readers assess the claim's robustness.
  2. [Section 2.2] The phrase 'AI agents internally deployed in plausibly pose several risks' appears to contain a stray 'in'; it should read 'AI agents internally deployed plausibly pose several risks.'
  3. [Section 2.4] The suggested indicators for Threshold One (volume of autonomously generated code, task-completion timelines, qualitative staff evaluations) are reasonable, but no guidance is given for what values would signal that the threshold is approaching; even for a communication tool, illustrative calibration points would be useful.
  4. [Section 3.1] Defining 'little compute' as less than 100 times the compute used for training by frontier developers is surprising, since 100x is not obviously 'little' in absolute terms; a sentence justifying this relative definition would reduce potential misunderstanding.
  5. [Section 3.4] The paper deliberately declines to propose a specific metric for Threshold Two; given the definition's importance, it would be helpful to list at least candidate directional indicators (e.g., self-improvement benchmark trends, compute efficiency gains) and explain why none is decisive.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's recommendations are normative proposals grounded in workshop consensus and external evidence, and no derivation reduces to its own inputs.

full rationale

The paper does not claim a formal derivation from first principles; it reports a two-day expert workshop whose participants agreed on risk pathways, two thresholds, and four policy recommendations. The thresholds are explicitly 'communication tools, not triggers for action,' so they are not fitted parameters or predicted quantities. The supporting evidence for near-term autonomous R&D possibility comes from external benchmark extrapolations (Kwa et al., 2025) and expert surveys (Grace et al., 2024), not from the paper's own conclusions. The self-citations (Clymer et al., 2024a; Clymer et al., 2024b; Greenblatt et al., 2024b) are used to point to existing safety-case, rogue-replication, and AI-control analyses; these are supporting references for proposed mitigations, not the justification for the central recommendation package itself. Footnote 2 candidly notes that benchmarks might overestimate or underestimate capabilities, which is an epistemic limitation rather than a circular step. The claim that 'most safety practices require over two years of preparation' is asserted without direct evidence, but an unsupported premise is a correctness or evidence gap, not a circularity, because the paper does not define the recommendations in terms of that claim nor fit the claim to the recommendations. No equation, definitional identity, or self-citation chain forces the outputs to equal the inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces conceptual thresholds rather than physical entities; no new objects are postulated. The threat model depends on the near-term autonomy projection and on the normative premise that lab and government intervention is appropriate; both are domain assumptions, not derived results.

assumptions (3)
  • domain assumption AI agents may plausibly automate most internal AI research and engineering in the near term, based on trend extrapolation from Kwa et al. (2025) and Grace et al. (2024).
    Load-bearing for Threshold One; if this projection is wrong, the deadlines lose urgency. The paper itself notes in Section 1.1 that benchmarks may over- or underestimate capability trajectories.
  • domain assumption Autonomous AI R&D could lead to catastrophic risks through safety sabotage, unauthorized deployment, adaptation lag, or capability proliferation.
    Threat models in Sections 2.2 and 3.2; the paper cites Benton et al. and Meinke et al. for sabotage and scheming, but the catastrophic concretization remains speculative.
  • domain assumption Frontier AI developers have a responsibility to implement the proposed safeguards, and governments have a legitimate role in receiving rapid disclosures and responding.
    Normative premise underlying all four recommendations; not defended beyond general safety and governance considerations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bare Minimum Mitigations for Autonomous AI Development." pith.science (2026). https://pith.science/paper/4AKGTRHY

@misc{pith2026250415416,
  author       = {Pith},
  title        = {Pith review of: Bare Minimum Mitigations for Autonomous AI Development},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4AKGTRHY}},
  note         = {Machine review of arXiv:2504.15416}
}
read the original abstract

Artificial intelligence (AI) is advancing rapidly, with the potential for significantly automating AI research and development itself in the near future. In 2024, international scientists, including Turing Award recipients, warned of risks from autonomous AI research and development (R&D), suggesting a red line such that no AI system should be able to improve itself or other AI systems without explicit human approval and assistance. However, the criteria for meaningful human approval remain unclear, and there is limited analysis on the specific risks of autonomous AI R&D, how they arise, and how to mitigate them. In this brief paper, we outline how these risks may emerge and propose four minimum safeguard recommendations applicable when AI agents significantly automate or accelerate AI development.

Figures

Figures reproduced from arXiv: 2504.15416 by the authors.

Figure 1
Figure 1. Summary of risk models and policy recommendations. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. AI agents are becoming increasingly autonomous (Kwa et al., 2025). [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 23 canonical work pages

  1. [1]

    Ziegler, Elizabeth Barnes, and Lawrence Chan

    Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Neev Parikh, David Rein, Lucas Jun Koba Sato, Hjalmar Wijk, Daniel M. Ziegler, Elizabeth Barnes, and Lawrence ...

  2. [2]

    Thousands of AI authors on the future of AI , 2024

    Katja Grace, Harlan Stewart, Julia Fabienne Sandkühler, Stephen Thomas, Ben Weinstein-Raun, and Jan Brauner. Thousands of AI authors on the future of AI , 2024. URL https://arxiv.org/abs/2401.02843

  3. [3]

    How AI can automate AI research and development

    Gaurav Sett. How AI can automate AI research and development. https://www.rand.org/pubs/commentary/2024/10/how-ai-can-automate-ai-research-and-development.html, October 2024. URL https://www.rand.org/pubs/commentary/2024/10/how-ai-can-automate-ai-research-and-development.html. RAND Corporation Commentary

  4. [4]

    On deepseek and export controls

    Dario Amodei. On deepseek and export controls. https://www.darioamodei.com/post/on-deepseek-and-export-controls, January 2025. URL https://www.darioamodei.com/post/on-deepseek-and-export-controls. Personal blog post

  5. [5]

    Preparedness framework version 2

    OpenAI . Preparedness framework version 2. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf, April 2025. URL https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf. White Paper

  6. [6]

    Frontier safety framework version 2.0

    Google DeepMind . Frontier safety framework version 2.0. https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/updating-the-frontier-safety-framework/Frontier URL https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/updating-the-frontier-safety-framework/Frontier White Paper

  7. [7]

    Responsible scaling policy version 2.1

    Anthropic . Responsible scaling policy version 2.1. https://www-cdn.anthropic.com/17310f6d70ae5627f55313ed067afc1a762a4068.pdf, March 2025. URL https://www-cdn.anthropic.com/17310f6d70ae5627f55313ed067afc1a762a4068.pdf. White Paper

  8. [8]

    Frontier governance framework

    Microsoft Corporation . Frontier governance framework. https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/microsoft/msc/documents/presentations/CSR/Frontier-Governance-Framework.pdf, February 2025. URL https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/microsoft/msc/documents/presentations/CSR/Frontier-Governance-Framework.pdf. Version 1

Show all 47 references
  1. [9]

    Amazon’s frontier model safety framework, 2025

    Amazon. Amazon’s frontier model safety framework, 2025. URL https://www.amazon.science/publications/amazons-frontier-model-safety-framework

  2. [10]

    AI Safety Institute and U.K

    U.S. AI Safety Institute and U.K. AI Safety Institute . Joint pre-deployment evaluation of openai's o1 model. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6763fac97cd22a9484ac3c37_o1_uk_us_december_publication_final.pdf, December 2024. URL https://cdn.prod.websi...

  3. [11]

    Idais-beijing statement, March 2024

    International Dialogues on AI Safety . Idais-beijing statement, March 2024. URL https://idais.ai/dialogue/idais-beijing/. IDAIS-Beijing, March 10--11, 2024

  4. [12]

    Ai behind closed doors: a primer on the governance of internal deployment, 2025

    Charlotte Stix, Matteo Pistillo, Girish Sastry, Marius Hobbhahn, Alejandro Ortega, Mikita Balesni, Annika Hallensleben, Nix Goldowsky-Dill, and Lee Sharkey. Ai behind closed doors: a primer on the governance of internal deployment, 2025. URL https://arxiv.org/abs/2504.12170

  5. [13]

    Bowman, and David Duvenaud

    Joe Benton, Misha Wagner, Eric Christiansen, Cem Anil, Ethan Perez, Jai Srivastav, Esin Durmus, Deep Ganguli, Shauna Kravec, Buck Shlegeris, Jared Kaplan, Holden Karnofsky, Evan Hubinger, Roger Grosse, Samuel R. Bowman, and David Duvenaud. Sabotage evaluations for frontier mod...

  6. [14]

    Bowman, and Evan Hubinger

    Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Sam...

  7. [15]

    Frontier models are capable of in-context scheming, 2025

    Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. Frontier models are capable of in-context scheming, 2025. URL https://arxiv.org/abs/2412.04984

  8. [16]

    Will AI R&D automation cause a software intelligence explosion?, 2025

    Daniel Eth and Tom Davidson. Will AI R&D automation cause a software intelligence explosion?, 2025. URL https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion. Accessed: 2025-04-23

  9. [17]

    Assessing potential future artificial intelligence risks, benefits and policy imperatives

    Organisation for Economic Co-operation and Development . Assessing potential future artificial intelligence risks, benefits and policy imperatives. Technical Report 27, OECD Publishing, November 2024. URL https://doi.org/10.1787/3f4e3dfb-en. OECD Artificial Intelligence Papers, No. 27

  10. [18]

    AI control: Improving safety despite intentional subversion, 2024 b

    Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. AI control: Improving safety despite intentional subversion, 2024 b . URL https://arxiv.org/abs/2312.06942

  11. [19]

    Towards evaluations-based safety cases for AI scheming, 2024

    Mikita Balesni, Marius Hobbhahn, David Lindner, Alexander Meinke, Tomek Korbak, Joshua Clymer, Buck Shlegeris, Jérémy Scheurer, Charlotte Stix, Rusheb Shah, Nicholas Goldowsky-Dill, Dan Braun, Bilal Chughtai, Owain Evans, Daniel Kokotajlo, and Lucius Bushnaq. Towards evaluatio...

  12. [20]

    Adversarial machine learning: A taxonomy and terminology of attacks and mitigations

    Apostol Vassilev, Alina Oprea, Alie Fordyce, and Hyrum Anderson. Adversarial machine learning: A taxonomy and terminology of attacks and mitigations. Technical Report NIST AI 100-2e2023, National Institute of Standards and Technology, January 2024. URL https://doi.org/10.6028/...

  13. [21]

    Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan P...

  14. [22]

    AI catastrophes and rogue deployments, June 2024 a

    Buck Shlegeris. AI catastrophes and rogue deployments, June 2024 a . URL https://redwoodresearch.substack.com/p/ai-catastrophes-and-rogue-deployments. Redwood Research Blog

  15. [23]

    Frontier AI systems have surpassed the self-replicating red line, 2024

    Xudong Pan, Jiarun Dai, Yihe Fan, and Min Yang. Frontier AI systems have surpassed the self-replicating red line, 2024. URL https://arxiv.org/abs/2412.12140

  16. [24]

    Knapp and Joel Thomas Langill

    Eric D. Knapp and Joel Thomas Langill. Chapter 3 - industrial cyber security history and trends. In Eric D. Knapp and Joel Thomas Langill, editors, Industrial Network Security (Second Edition), pages 41--57. Syngress, Boston, second edition edition, 2015. ISBN 978-0-12-420114-...

  17. [25]

    Safety cases: How to justify the safety of advanced AI systems, 2024 a

    Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen. Safety cases: How to justify the safety of advanced AI systems, 2024 a . URL https://arxiv.org/abs/2403.10462

  18. [26]

    Safe beyond sale: post-deployment monitoring of AI , June 2024

    Merlin Stein and Connor Dunlop. Safe beyond sale: post-deployment monitoring of AI , June 2024. URL https://www.adalovelaceinstitute.org/blog/post-deployment-monitoring-of-ai/. Ada Lovelace Institute Blog

  19. [27]

    Deployment corrections: An incident response framework for frontier AI models, 2023

    Joe O'Brien, Shaun Ee, and Zoe Williams. Deployment corrections: An incident response framework for frontier AI models, 2023. URL https://arxiv.org/abs/2310.00328

  20. [28]

    Re-bench: Evaluating frontier AI R&D capabilities of language model agents against human experts, 2024

    Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, Michael Chen, Josh Clymer, Jai Dhyani, Elena Ericheva, Katharyn Garcia, Brian Goodrich, Nikola Jurkovic, Megan Kinniment, Aron Lajko, Seraphina Nix, Lucas Sato, William Saunders, Maks...

  21. [29]

    Interviewing AI researchers on automation of AI R&D , 2024

    David Owen. Interviewing AI researchers on automation of AI R&D , 2024. URL https://epoch.ai/blog/interviewing-ai-researchers-on-automation-of-ai-rnd. Accessed: 2025-04-16

  22. [30]

    What a compute-centric framework says about takeoff speeds

    Tom Davidson. What a compute-centric framework says about takeoff speeds. Technical report, Open Philanthropy, June 2023. URL https://www.openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds/. Research Report

  23. [31]

    Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek- AI , Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei...

  24. [32]

    AI achieves silver-medal standard solving international mathematical olympiad problems, July 2024

    AlphaProof and AlphaGeometry teams . AI achieves silver-medal standard solving international mathematical olympiad problems, July 2024. URL https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/. DeepMind Blog

  25. [33]

    Increased compute efficiency and the diffusion of AI capabilities, 2024

    Konstantin Pilz, Lennart Heim, and Nicholas Brown. Increased compute efficiency and the diffusion of AI capabilities, 2024. URL https://arxiv.org/abs/2311.15377

  26. [34]

    Racing to the precipice: A model of artificial intelligence development

    Stuart Armstrong, Nick Bostrom, and Carl Shulman. Racing to the precipice: A model of artificial intelligence development. Technical Report 2013-1, Future of Humanity Institute, University of Oxford, 2013

  27. [35]

    The rogue replication threat model

    Josh Clymer, Hjalmar Wijk, and Beth Barnes. The rogue replication threat model. https://metr.org/blog/2024-11-12-rogue-replication-threat-model/, 11 2024 b

  28. [36]

    Announcing our updated responsible scaling policy, October 2024

    Anthropic . Announcing our updated responsible scaling policy, October 2024. URL https://www.anthropic.com/news/announcing-our-updated-responsible-scaling-policy. Anthropic Blog

  29. [37]

    Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schroeder, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Aaron Purewal, Csaba Botos, Fabro Steibel, Fazel Keshtkar, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph I...

  30. [38]

    AI now 2023 landscape: Confronting tech power

    Amba Kak and Sarah Myers West. AI now 2023 landscape: Confronting tech power. Technical report, AI Now Institute, April 2023. URL https://ainowinstitute.org/2023-landscape. Technical Report

  31. [39]

    Elizabeth Seger, Noemi Dreksler, Richard Moulange, Emily Dardaman, Jonas Schuett, K. Wei, Christoph Winter, Mackenzie Arnold, Seán Ó hÉigeartaigh, Anton Korinek, Markus Anderljung, Ben Bucknall, Alan Chan, Eoghan Stafford, Leonie Koessler, Aviv Ovadya, Ben Garfinkel, Emma Blue...

  32. [40]

    Considerations influencing offense-defense dynamics from artificial intelligence, 2024

    Giulio Corsi, Kyle Kilian, and Richard Mallah. Considerations influencing offense-defense dynamics from artificial intelligence, 2024. URL https://arxiv.org/abs/2412.04029

  33. [41]

    Societal adaptation to advanced AI , 2025

    Jamie Bernardi, Gabriel Mukobi, Hilary Greaves, Lennart Heim, and Markus Anderljung. Societal adaptation to advanced AI , 2025. URL https://arxiv.org/abs/2405.10295

  34. [42]

    Understanding frontier AI capabilities and risks through semi-structured interviews

    Akash Wasil, Lukas Berglund, Tom Reed, Miro Plueckebaum, and Everett Smith. Understanding frontier AI capabilities and risks through semi-structured interviews. Technical report, SSRN, July 2024 a . URL https://ssrn.com/abstract=4881729. SSRN Working Paper

  35. [43]

    AI emergency preparedness: Examining the federal government's ability to detect and respond to ai-related national security threats, 2024 b

    Akash Wasil, Everett Smith, Corin Katzke, and Justin Bullock. AI emergency preparedness: Examining the federal government's ability to detect and respond to ai-related national security threats, 2024 b . URL https://arxiv.org/abs/2407.17347

  36. [44]

    Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models

    Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, and Jeff Alstott. Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models . RAND Corporation, Santa Monica, CA, 2024. doi:10.7249/RRA2849-1

  37. [45]

    Access to powerful AI might make computer security radically easier, June 2024 b

    Buck Shlegeris. Access to powerful AI might make computer security radically easier, June 2024 b . URL https://redwoodresearch.substack.com/p/access-to-powerful-ai-might-make. Redwood Research Blog

  38. [46]

    AI capabilities can be significantly improved without expensive retraining, 2023

    Tom Davidson, Jean-Stanislas Denain, Pablo Villalobos, and Guillem Bas. AI capabilities can be significantly improved without expensive retraining, 2023. URL https://arxiv.org/abs/2312.07413

  39. [47]

    AI benchmarking hub, April 2025

    Epoch AI . AI benchmarking hub, April 2025. URL https://epoch.ai/data/ai-benchmarking-dashboard. Last updated April 16, 2025

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.