REVIEW 1 major objections 5 minor 47 references
Bare Minimum Mitigations for Autonomous AI Development
T0 review · 1 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Frontier AI labs should put four minimum safeguards in place before AI agents automate most internal R&D or gain rapid self-improvement to catastrophic capability.
desk verdict A concrete, operational threshold package for autonomous AI R&D that is worth engaging seriously, even though the 'act now' urgency rests on a two-year lead-time claim the paper doesn't support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a pair of capability thresholds paired with threat models. Threshold One defines the automation of most internal research and engineering in terms of a productivity-loss comparison against laying off half of a lab's engineers, and it triggers recommendations to maintain understanding of safety-critical details and to detect egregious compute misuse. Threshold Two defines rapid improvement to catastrophic capabilities in terms of speed (about a year), compute (under 100x frontier training compute), and human assistance (a handful of people with generic skills), and it triggers disclosure to home governments and protection of critical AI software from theft. The threat models that carry the argument are safety sabotage, unauthorized internal deployment, adaptation lag, and capability proliferation, each explaining how a specific failure becomes possible once a threshold is crossed.
What would settle it
A falsifying observation would be a sustained plateau: if, over the next several years, the time horizon of tasks AI agents can complete autonomously stops growing and the share of lab R&D productivity attributable to agents remains far below the lay-off-half-of-engineers threshold, with no agent demonstrating a low-compute software-only path to catastrophic capability, the paper's central timeline would fail.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that the risks of autonomous AI R&D fall into two categories—harms from automating AI development itself and harms from rapid autonomous improvement—and that each category can be managed by a specific minimum safeguard if adopted before a defined threshold. Threshold One is crossed when the productivity loss from not using AI agents in internal software R&D exceeds the loss from laying off half of the lab's software engineers and researchers. Threshold Two is crossed when AI agents can improve to catastrophic capabilities within about a year, using less than 100x the compute of frontier training and no more than a handful of people with generic technical skills as assistance. The paper argues that these thresholds define natural deadlines for preparing safety, security, and governance measures, and that most such measures take years to build.
Load-bearing premise
The argument rests on the premise that AI agents will meaningfully automate most lab R&D or reach rapid catastrophic improvement soon enough that preparations begun now are the ones that matter; if agent capability growth stalls or long-horizon benchmarks overestimate progress, the urgency and the deadlines both recede.
Editorial extensions
If this is right
- If Threshold One is near, labs should already be building monitoring of training data, experimental code, and internal agent behavior so they can detect sabotage and unauthorized compute use before agents control most R&D.
- If Threshold Two is crossed, governments need visibility into internally deployed models, not just pre-deployment evaluations, to avoid adaptation lag.
- Measures to secure AI software against well-resourced cyberattacks and insider threats take years to build, so they should be planned before rapid self-improvement becomes possible.
- The thresholds, used as communication tools, give researchers and policymakers a shared language for when safety practices are no longer optional.
Reading between the lines
- Editorial inference: the productivity-loss definition of Threshold One is measurable in principle—a lab could track what fraction of its R&D output would vanish if agents were removed—which makes the threshold a testable governance target rather than a rhetorical one.
- Editorial inference: because the paper treats rapid improvement as software-driven, progress on algorithmic efficiency rather than raw compute is the leading indicator to watch; if efficiency gains plateau, Threshold Two recedes.
- Editorial inference: the adaptation-lag argument suggests a natural extension to a near-miss disclosure regime, since waiting until catastrophic capabilities are demonstrated may already be too late for coordinated responses.
- Editorial inference: a natural next test is whether international coordination can form around the two thresholds, since the paper relies on home-government disclosure but leaves open how governments would share and act on those disclosures across borders.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This policy paper argues that autonomous AI research and development (R&D) is a near-term possibility and identifies two broad risk categories: risks from automating AI R&D itself (safety sabotage, unauthorized internal deployment) and risks from rapid autonomous improvement (adaptation lag, capability proliferation). To address these, it defines two thresholds—Threshold One, when AI agents automate most internal research and engineering, and Threshold Two, when AI agents can rapidly improve to catastrophic capabilities with little compute and human assistance—and proposes four minimum safeguards for frontier AI developers to implement before these thresholds are reached: (1) thoroughly understand safety-critical training and assurance details, (2) detect internal AI agents egregiously misusing compute, (3) rapidly disclose catastrophic risks to home governments, and (4) prevent theft or exfiltration of critical AI software. The thresholds are explicitly framed as communication tools rather than triggers, and the paper concludes that because most safety practices require over two years of preparation, action should begin now.
Significance. The paper's main contribution is translating a broad 'red line' concern about autonomous AI R&D into concrete, named recommendations with definitions, threat models, and implementation indicators. The authors are appropriately transparent about uncertainty: Section 3.2 labels recursive improvement scenarios speculative, footnote 2 acknowledges that benchmark extrapolation may over- or underestimate progress, and Section 3.3 notes that proliferation can have benefits. The two-threshold structure is a useful coordination device, and Recommendation Four is grounded in an external detailed cost estimate (Nevo et al., 2024). If one accepts the threat model, the recommendations are coherent and largely follow from prior cited analyses; the paper does not derive them in a circular way. The central weakness is the unsupported lead-time claim underlying the 'act now' conclusion, which is load-bearing because the thresholds are explicitly not triggers for action.
major comments (1)
- [Section 1.4 and Section 3.3] The sentence 'Most safety practices require over two years of preparation, so the time to act is now' is the only explicit action-forcing conclusion in the paper, since the thresholds are described as 'communication tools, not triggers for action' (Section 1.4). However, the cited timeline evidence concerns only Recommendation Four, where Nevo et al. (2024) is invoked for 'years of concerted effort' (Section 3.3). No timeline evidence is offered for Recommendations One, Two, or Three; these might be implementable in months or could also require years. For example, Section 2.4 itself recommends that control and security measures be 'incrementally enhanced' as automation increases, which points away from a hard two-year lead time for the full package. The authors should either provide evidence for the multi-year preparation claim across the recommended measures or temper the 'time to act now' conclusion to match the support actually provided.
minor comments (5)
- [Section 1.1, Figure 2] The text states that projecting trends in Figure 2 indicates that by early 2027 AI agents might complete week-long software engineering tasks, but no quantitative projection method, data points, or uncertainty intervals are given; a brief description of the extrapolation would help readers assess the claim's robustness.
- [Section 2.2] The phrase 'AI agents internally deployed in plausibly pose several risks' appears to contain a stray 'in'; it should read 'AI agents internally deployed plausibly pose several risks.'
- [Section 2.4] The suggested indicators for Threshold One (volume of autonomously generated code, task-completion timelines, qualitative staff evaluations) are reasonable, but no guidance is given for what values would signal that the threshold is approaching; even for a communication tool, illustrative calibration points would be useful.
- [Section 3.1] Defining 'little compute' as less than 100 times the compute used for training by frontier developers is surprising, since 100x is not obviously 'little' in absolute terms; a sentence justifying this relative definition would reduce potential misunderstanding.
- [Section 3.4] The paper deliberately declines to propose a specific metric for Threshold Two; given the definition's importance, it would be helpful to list at least candidate directional indicators (e.g., self-improvement benchmark trends, compute efficiency gains) and explain why none is decisive.
Circularity Check
No circularity: the paper's recommendations are normative proposals grounded in workshop consensus and external evidence, and no derivation reduces to its own inputs.
full rationale
The paper does not claim a formal derivation from first principles; it reports a two-day expert workshop whose participants agreed on risk pathways, two thresholds, and four policy recommendations. The thresholds are explicitly 'communication tools, not triggers for action,' so they are not fitted parameters or predicted quantities. The supporting evidence for near-term autonomous R&D possibility comes from external benchmark extrapolations (Kwa et al., 2025) and expert surveys (Grace et al., 2024), not from the paper's own conclusions. The self-citations (Clymer et al., 2024a; Clymer et al., 2024b; Greenblatt et al., 2024b) are used to point to existing safety-case, rogue-replication, and AI-control analyses; these are supporting references for proposed mitigations, not the justification for the central recommendation package itself. Footnote 2 candidly notes that benchmarks might overestimate or underestimate capabilities, which is an epistemic limitation rather than a circular step. The claim that 'most safety practices require over two years of preparation' is asserted without direct evidence, but an unsupported premise is a correctness or evidence gap, not a circularity, because the paper does not define the recommendations in terms of that claim nor fit the claim to the recommendations. No equation, definitional identity, or self-citation chain forces the outputs to equal the inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption AI agents may plausibly automate most internal AI research and engineering in the near term, based on trend extrapolation from Kwa et al. (2025) and Grace et al. (2024).
- domain assumption Autonomous AI R&D could lead to catastrophic risks through safety sabotage, unauthorized deployment, adaptation lag, or capability proliferation.
- domain assumption Frontier AI developers have a responsibility to implement the proposed safeguards, and governments have a legitimate role in receiving rapid disclosures and responding.
Cite this review
Pith. "Pith review of Bare Minimum Mitigations for Autonomous AI Development." pith.science (2026). https://pith.science/paper/4AKGTRHY
@misc{pith2026250415416,
author = {Pith},
title = {Pith review of: Bare Minimum Mitigations for Autonomous AI Development},
year = {2026},
howpublished = {\url{https://pith.science/paper/4AKGTRHY}},
note = {Machine review of arXiv:2504.15416}
}
read the original abstract
Artificial intelligence (AI) is advancing rapidly, with the potential for significantly automating AI research and development itself in the near future. In 2024, international scientists, including Turing Award recipients, warned of risks from autonomous AI research and development (R&D), suggesting a red line such that no AI system should be able to improve itself or other AI systems without explicit human approval and assistance. However, the criteria for meaningful human approval remain unclear, and there is limited analysis on the specific risks of autonomous AI R&D, how they arise, and how to mitigate them. In this brief paper, we outline how these risks may emerge and propose four minimum safeguard recommendations applicable when AI agents significantly automate or accelerate AI development.
Figures
Reference graph
Works this paper leans on
-
[1]
Ziegler, Elizabeth Barnes, and Lawrence Chan
Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Neev Parikh, David Rein, Lucas Jun Koba Sato, Hjalmar Wijk, Daniel M. Ziegler, Elizabeth Barnes, and Lawrence ...
arXiv 2025
-
[2]
Thousands of AI authors on the future of AI , 2024
Katja Grace, Harlan Stewart, Julia Fabienne Sandkühler, Stephen Thomas, Ben Weinstein-Raun, and Jan Brauner. Thousands of AI authors on the future of AI , 2024. URL https://arxiv.org/abs/2401.02843
arXiv 2024
-
[3]
How AI can automate AI research and development
Gaurav Sett. How AI can automate AI research and development. https://www.rand.org/pubs/commentary/2024/10/how-ai-can-automate-ai-research-and-development.html, October 2024. URL https://www.rand.org/pubs/commentary/2024/10/how-ai-can-automate-ai-research-and-development.html. RAND Corporation Commentary
work page 2024
-
[4]
On deepseek and export controls
Dario Amodei. On deepseek and export controls. https://www.darioamodei.com/post/on-deepseek-and-export-controls, January 2025. URL https://www.darioamodei.com/post/on-deepseek-and-export-controls. Personal blog post
work page 2025
-
[5]
Preparedness framework version 2
OpenAI . Preparedness framework version 2. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf, April 2025. URL https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf. White Paper
work page 2025
-
[6]
Frontier safety framework version 2.0
Google DeepMind . Frontier safety framework version 2.0. https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/updating-the-frontier-safety-framework/Frontier URL https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/updating-the-frontier-safety-framework/Frontier White Paper
-
[7]
Responsible scaling policy version 2.1
Anthropic . Responsible scaling policy version 2.1. https://www-cdn.anthropic.com/17310f6d70ae5627f55313ed067afc1a762a4068.pdf, March 2025. URL https://www-cdn.anthropic.com/17310f6d70ae5627f55313ed067afc1a762a4068.pdf. White Paper
work page 2025
-
[8]
Microsoft Corporation . Frontier governance framework. https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/microsoft/msc/documents/presentations/CSR/Frontier-Governance-Framework.pdf, February 2025. URL https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/microsoft/msc/documents/presentations/CSR/Frontier-Governance-Framework.pdf. Version 1
work page 2025
Show all 47 references
-
[9]
Amazon’s frontier model safety framework, 2025
Amazon. Amazon’s frontier model safety framework, 2025. URL https://www.amazon.science/publications/amazons-frontier-model-safety-framework
2025
-
[10]
AI Safety Institute and U.K
U.S. AI Safety Institute and U.K. AI Safety Institute . Joint pre-deployment evaluation of openai's o1 model. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6763fac97cd22a9484ac3c37_o1_uk_us_december_publication_final.pdf, December 2024. URL https://cdn.prod.websi...
2024
-
[11]
Idais-beijing statement, March 2024
International Dialogues on AI Safety . Idais-beijing statement, March 2024. URL https://idais.ai/dialogue/idais-beijing/. IDAIS-Beijing, March 10--11, 2024
2024
-
[12]
Ai behind closed doors: a primer on the governance of internal deployment, 2025
Charlotte Stix, Matteo Pistillo, Girish Sastry, Marius Hobbhahn, Alejandro Ortega, Mikita Balesni, Annika Hallensleben, Nix Goldowsky-Dill, and Lee Sharkey. Ai behind closed doors: a primer on the governance of internal deployment, 2025. URL https://arxiv.org/abs/2504.12170
2025 arXiv
-
[13]
Bowman, and David Duvenaud
Joe Benton, Misha Wagner, Eric Christiansen, Cem Anil, Ethan Perez, Jai Srivastav, Esin Durmus, Deep Ganguli, Shauna Kravec, Buck Shlegeris, Jared Kaplan, Holden Karnofsky, Evan Hubinger, Roger Grosse, Samuel R. Bowman, and David Duvenaud. Sabotage evaluations for frontier mod...
2024 arXiv
-
[14]
Bowman, and Evan Hubinger
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Sam...
2024 arXiv
-
[15]
Frontier models are capable of in-context scheming, 2025
Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. Frontier models are capable of in-context scheming, 2025. URL https://arxiv.org/abs/2412.04984
2025 arXiv
-
[16]
Will AI R&D automation cause a software intelligence explosion?, 2025
Daniel Eth and Tom Davidson. Will AI R&D automation cause a software intelligence explosion?, 2025. URL https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion. Accessed: 2025-04-23
2025
-
[17]
Assessing potential future artificial intelligence risks, benefits and policy imperatives
Organisation for Economic Co-operation and Development . Assessing potential future artificial intelligence risks, benefits and policy imperatives. Technical Report 27, OECD Publishing, November 2024. URL https://doi.org/10.1787/3f4e3dfb-en. OECD Artificial Intelligence Papers, No. 27
2024 doi
-
[18]
AI control: Improving safety despite intentional subversion, 2024 b
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. AI control: Improving safety despite intentional subversion, 2024 b . URL https://arxiv.org/abs/2312.06942
2024 arXiv
-
[19]
Towards evaluations-based safety cases for AI scheming, 2024
Mikita Balesni, Marius Hobbhahn, David Lindner, Alexander Meinke, Tomek Korbak, Joshua Clymer, Buck Shlegeris, Jérémy Scheurer, Charlotte Stix, Rusheb Shah, Nicholas Goldowsky-Dill, Dan Braun, Bilal Chughtai, Owain Evans, Daniel Kokotajlo, and Lucius Bushnaq. Towards evaluatio...
2024 arXiv
-
[20]
Adversarial machine learning: A taxonomy and terminology of attacks and mitigations
Apostol Vassilev, Alina Oprea, Alie Fordyce, and Hyrum Anderson. Adversarial machine learning: A taxonomy and terminology of attacks and mitigations. Technical Report NIST AI 100-2e2023, National Institute of Standards and Technology, January 2024. URL https://doi.org/10.6028/...
2024 doi
-
[21]
Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan P...
2022 arXiv
-
[22]
AI catastrophes and rogue deployments, June 2024 a
Buck Shlegeris. AI catastrophes and rogue deployments, June 2024 a . URL https://redwoodresearch.substack.com/p/ai-catastrophes-and-rogue-deployments. Redwood Research Blog
2024
-
[23]
Frontier AI systems have surpassed the self-replicating red line, 2024
Xudong Pan, Jiarun Dai, Yihe Fan, and Min Yang. Frontier AI systems have surpassed the self-replicating red line, 2024. URL https://arxiv.org/abs/2412.12140
2024 arXiv
-
[24]
Knapp and Joel Thomas Langill
Eric D. Knapp and Joel Thomas Langill. Chapter 3 - industrial cyber security history and trends. In Eric D. Knapp and Joel Thomas Langill, editors, Industrial Network Security (Second Edition), pages 41--57. Syngress, Boston, second edition edition, 2015. ISBN 978-0-12-420114-...
2015 doi
-
[25]
Safety cases: How to justify the safety of advanced AI systems, 2024 a
Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen. Safety cases: How to justify the safety of advanced AI systems, 2024 a . URL https://arxiv.org/abs/2403.10462
2024 arXiv
-
[26]
Safe beyond sale: post-deployment monitoring of AI , June 2024
Merlin Stein and Connor Dunlop. Safe beyond sale: post-deployment monitoring of AI , June 2024. URL https://www.adalovelaceinstitute.org/blog/post-deployment-monitoring-of-ai/. Ada Lovelace Institute Blog
2024
-
[27]
Deployment corrections: An incident response framework for frontier AI models, 2023
Joe O'Brien, Shaun Ee, and Zoe Williams. Deployment corrections: An incident response framework for frontier AI models, 2023. URL https://arxiv.org/abs/2310.00328
2023 arXiv
-
[28]
Re-bench: Evaluating frontier AI R&D capabilities of language model agents against human experts, 2024
Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, Michael Chen, Josh Clymer, Jai Dhyani, Elena Ericheva, Katharyn Garcia, Brian Goodrich, Nikola Jurkovic, Megan Kinniment, Aron Lajko, Seraphina Nix, Lucas Sato, William Saunders, Maks...
2024 arXiv
-
[29]
Interviewing AI researchers on automation of AI R&D , 2024
David Owen. Interviewing AI researchers on automation of AI R&D , 2024. URL https://epoch.ai/blog/interviewing-ai-researchers-on-automation-of-ai-rnd. Accessed: 2025-04-16
2024
-
[30]
What a compute-centric framework says about takeoff speeds
Tom Davidson. What a compute-centric framework says about takeoff speeds. Technical report, Open Philanthropy, June 2023. URL https://www.openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds/. Research Report
2023
-
[31]
Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
DeepSeek- AI , Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei...
2025 arXiv
-
[32]
AI achieves silver-medal standard solving international mathematical olympiad problems, July 2024
AlphaProof and AlphaGeometry teams . AI achieves silver-medal standard solving international mathematical olympiad problems, July 2024. URL https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/. DeepMind Blog
2024
-
[33]
Increased compute efficiency and the diffusion of AI capabilities, 2024
Konstantin Pilz, Lennart Heim, and Nicholas Brown. Increased compute efficiency and the diffusion of AI capabilities, 2024. URL https://arxiv.org/abs/2311.15377
2024 arXiv
-
[34]
Racing to the precipice: A model of artificial intelligence development
Stuart Armstrong, Nick Bostrom, and Carl Shulman. Racing to the precipice: A model of artificial intelligence development. Technical Report 2013-1, Future of Humanity Institute, University of Oxford, 2013
2013
-
[35]
The rogue replication threat model
Josh Clymer, Hjalmar Wijk, and Beth Barnes. The rogue replication threat model. https://metr.org/blog/2024-11-12-rogue-replication-threat-model/, 11 2024 b
2024
-
[36]
Announcing our updated responsible scaling policy, October 2024
Anthropic . Announcing our updated responsible scaling policy, October 2024. URL https://www.anthropic.com/news/announcing-our-updated-responsible-scaling-policy. Anthropic Blog
2024
-
[37]
Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schroeder, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Aaron Purewal, Csaba Botos, Fabro Steibel, Fazel Keshtkar, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph I...
2024 arXiv
-
[38]
AI now 2023 landscape: Confronting tech power
Amba Kak and Sarah Myers West. AI now 2023 landscape: Confronting tech power. Technical report, AI Now Institute, April 2023. URL https://ainowinstitute.org/2023-landscape. Technical Report
2023
-
[39]
Elizabeth Seger, Noemi Dreksler, Richard Moulange, Emily Dardaman, Jonas Schuett, K. Wei, Christoph Winter, Mackenzie Arnold, Seán Ó hÉigeartaigh, Anton Korinek, Markus Anderljung, Ben Bucknall, Alan Chan, Eoghan Stafford, Leonie Koessler, Aviv Ovadya, Ben Garfinkel, Emma Blue...
2023 arXiv
-
[40]
Considerations influencing offense-defense dynamics from artificial intelligence, 2024
Giulio Corsi, Kyle Kilian, and Richard Mallah. Considerations influencing offense-defense dynamics from artificial intelligence, 2024. URL https://arxiv.org/abs/2412.04029
2024 arXiv
-
[41]
Societal adaptation to advanced AI , 2025
Jamie Bernardi, Gabriel Mukobi, Hilary Greaves, Lennart Heim, and Markus Anderljung. Societal adaptation to advanced AI , 2025. URL https://arxiv.org/abs/2405.10295
2025 arXiv
-
[42]
Understanding frontier AI capabilities and risks through semi-structured interviews
Akash Wasil, Lukas Berglund, Tom Reed, Miro Plueckebaum, and Everett Smith. Understanding frontier AI capabilities and risks through semi-structured interviews. Technical report, SSRN, July 2024 a . URL https://ssrn.com/abstract=4881729. SSRN Working Paper
2024
-
[43]
AI emergency preparedness: Examining the federal government's ability to detect and respond to ai-related national security threats, 2024 b
Akash Wasil, Everett Smith, Corin Katzke, and Justin Bullock. AI emergency preparedness: Examining the federal government's ability to detect and respond to ai-related national security threats, 2024 b . URL https://arxiv.org/abs/2407.17347
2024 arXiv
-
[44]
Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models
Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, and Jeff Alstott. Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models . RAND Corporation, Santa Monica, CA, 2024. doi:10.7249/RRA2849-1
2024 doi
-
[45]
Access to powerful AI might make computer security radically easier, June 2024 b
Buck Shlegeris. Access to powerful AI might make computer security radically easier, June 2024 b . URL https://redwoodresearch.substack.com/p/access-to-powerful-ai-might-make. Redwood Research Blog
2024
-
[46]
AI capabilities can be significantly improved without expensive retraining, 2023
Tom Davidson, Jean-Stanislas Denain, Pablo Villalobos, and Guillem Bas. AI capabilities can be significantly improved without expensive retraining, 2023. URL https://arxiv.org/abs/2312.07413
2023 arXiv
-
[47]
AI benchmarking hub, April 2025
Epoch AI . AI benchmarking hub, April 2025. URL https://epoch.ai/data/ai-benchmarking-dashboard. Last updated April 16, 2025
2025
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.