Pith. sign in

REVIEW 2 major objections 5 minor 109 references

In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?

T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Verification mechanisms and shared protocols are the least risky areas for AI safety cooperation between geopolitical rivals; shared infrastructure and joint evaluations carry higher risks of capability transfer, sensitive-information…

desk verdict A genuinely useful risk-mapping for US–China AI safety cooperation, but the headline ranking leans on a qualitative table whose Protocol row is coded against the paper's own evidence. read the letter →

arxiv 2504.12914 v1 pith:YQILN66F submitted 2025-04-17 cs.CY

classification cs.CY
keywords technicalAIsafetyinternationalcooperationgeopoliticalrivalryverificationmechanismsprotocolsandstandardsevaluationinfrastructureUS-China
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks which areas of technical AI safety research are safe enough for geopolitical rivals—chiefly the United States and China—to cooperate on, and it argues that riskiness is area-specific rather than uniform. It builds a four-part risk typology—advancing the global capabilities frontier, differentially advancing a rival's capabilities, exposing sensitive information, and creating opportunities for harmful action—and applies it to four proposed cooperation areas: verification mechanisms, shared protocols and best practices, shared infrastructure, and evaluation methodologies. Its central finding is that research on AI verification mechanisms and on shared protocols is less challenging for rival cooperation than infrastructure and evaluations, because these areas mostly attest to or codify existing knowledge instead of extending it. The paper is best read as a risk-screening device: it identifies where the downside of cooperation is smallest, while explicitly setting aside whether the benefits in each area justify the effort.

What carries the argument

The analytical core is a two-way risk matrix. Four risk categories—advancing the global frontier of AI capabilities, differentially advancing a rival's capabilities, exposing nationally strategic information, and enabling harmful action by motivated actors—are crossed with four candidate cooperation areas: verification mechanisms, codified protocols and best practices, shared infrastructure, and evaluation methodologies. Each area is unpacked into subareas such as formal verification, verifiable audits, compute attestation, watermarking, safety frameworks, incident-reporting standards, secure-weight standards, evaluation infrastructure, shared compute clusters, and capability evaluations. The matrix does the argumentative work: the paper's verdict that verification and protocols are 'less-challenging areas for international cooperation' is a summary of the risk columns of that matrix, with historical precedents such as the Open Skies Treaty and Permissive Action Links supplying the reason to believe verification cooperation can stay low-risk.

What would settle it

A concrete test would be to run the same four-risk analysis on documented cooperation projects—for example, a bilateral compute-attestation prototype or the joint US–UK evaluation exercise—and find that a verification or protocol project exposed sensitive information, advanced a rival's capabilities, or enabled a covert backdoor; a structured expert elicitation that ranks verification or protocols riskier than infrastructure or evaluations would likewise falsify the paper's ordering.

Watch

Extended reading notes

Core claim

The paper's central claim is that the risks of international cooperation on technical AI safety vary systematically by area, and that verification mechanisms and shared protocols are the two areas where such risks are lowest. Verification is comparatively low-risk, the authors argue, because it predominantly certifies claims about systems rather than improving them, and the properties verified can be restricted in advance to those all parties already know, as with the certified sensors of the Open Skies Treaty. Protocols are low-risk because they codify existing knowledge into agreed procedures rather than pushing the research frontier, and they need not involve direct access to live AI systems or sensitive infrastructure. Infrastructure and evaluations, by contrast, carry higher risk of misuse, backdoor insertion, leakage of sensitive domain knowledge, and differential capability gain. The paper does not say the riskier areas should be avoided, nor that the safer areas are the most valuable; it limits its conclusion to a risk-based assessment of which areas are less challenging, and explicitly leaves benefit analysis to future work.

Load-bearing premise

The paper equates suitability for cooperation with low exposure to its four risks and explicitly leaves the benefits of cooperation unexamined; if benefits differ by area, a low-risk area could still be a poor place to cooperate, and the ranking could change.

Editorial extensions

If this is right

  • Governments can treat AI verification mechanisms—compute attestation, watermarks, verifiable audits—as the first candidates for structured cooperation between rival AI safety institutes and in track-II dialogues.
  • Protocol and standards development (safety frameworks, incident-reporting definitions, secure-weight standards) can be pursued without requiring parties to disclose sensitive details about their own systems.
  • Cooperation on shared infrastructure and joint evaluations of dangerous capabilities should be paired with explicit mitigations, such as open-source development, bug bounties, restricted access, and secure evaluation enclaves.
  • Vetting processes for international research collaborations can be supplemented with the paper's four technology-specific risk categories, rather than relying only on due diligence and sanctions screening.
  • The finding gives intergovernmental cooperation on AI safety a concrete starting agenda: verification and protocol work can build trust before harder questions of shared compute or capability evaluation are attempted.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the paper assesses only risks, its 'suitable' verdict should be read as a risk ranking rather than a priority list; a low-risk area could still be low-value, and verifying the benefits of verification and protocol work is the natural next step.
  • The historical analogies suggest a sharper test than the paper states: verification cooperation is easiest when the verified object is standardized, mutually observable, and low-complexity, so early compute attestation or watermarking is a more realistic first target than formal verification of frontier models.
  • The protocols category's low technical risk may be offset by political risk, since standards bodies have repeatedly been used to advance national advantage; the paper's verdict for protocols therefore depends on institutional design that resists capture.
  • A finer-grained risk map—scoring the paper's named subareas across the four risks through expert elicitation—would turn its four-area ranking into a portfolio tool for choosing which specific projects to propose to rivals.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper asks where US-China rivals can safely cooperate on technical AI safety. It defines four risk dimensions relevant to cooperation—advancing the global capabilities frontier, differentially advancing a rival's capabilities, exposing sensitive information, and providing opportunities for harmful action—and applies them to four candidate areas: verification mechanisms, codified protocols and best practices, shared infrastructure, and evaluation methodologies. Using historical analogies, survey evidence on US-China AI collaboration, and a qualitative assessment table, the paper concludes that verification and protocols are less-challenging areas for cooperation than infrastructure and evaluations and 'may be well-suited' for international cooperation. The paper explicitly disclaims any analysis of the benefits of cooperation and notes that its risk categories are not comprehensive, so the conclusion is best read as a risk-based suitability judgment rather than a full cost-benefit analysis.

Significance. If the ranking holds, the paper gives policymakers a usable starting point for directing rival cooperation toward verification and protocols while treating infrastructure and evaluations as riskier. The paper's strengths are its clear definitions, its structured four-dimensional typology, its use of concrete historical analogues (Open Skies, PALs, the Joint Verification Experiment), and the explicit hedged language ('may,' 'less-challenging'). Table 1 makes the underlying qualitative judgments checkable, and the paper is transparent that the list of areas and risks is non-exhaustive. The central limitation is that suitability is inferred from low risk alone, without weighing benefits or political/process risks; this does not invalidate the paper but does bound the strength of the policy conclusion.

major comments (2)
  1. [Introduction, §4, Table 1] The paper's central conclusion—that verification and protocols 'may be well-suited' for cooperation—is derived from a risk-only assessment, but the Introduction explicitly states that the paper does 'not aim to definitively identify the most suitable areas for cooperation, nor to investigate specific benefits.' If benefits differ across areas (e.g., protocols may be easy to agree on but weak in changing behavior, while verification may require deep system access to be valuable), a higher-risk, higher-benefit area could be more suitable than a lower-risk, lower-benefit one. I recommend either reframing the conclusion as 'less risky areas for cooperation' or adding a qualitative benefit discussion for each area to support the suitability claim.
  2. [§4.2.1, Table 1] The Protocols row codes 'Provides opportunity for harmful action' as 'Minimal,' yet the same subsection states that 'both states and industry actors tend to use the international standardisation process to advance their own interests, potentially to the detriment of other actors' and that protocols 'could be more politicised... leading to a degradation in scientific rigour.' The table's own note concedes that 'standardisation has sometimes been used to advance unilateral interests.' At minimum, the coding and the text are in tension; if the harmful-action category is meant to include harm through process capture, a 'Minimal/moderate' or 'Moderate' coding would be more consistent, and this would materially weaken the paper's ranking of Protocols as uniformly lowest-risk.
minor comments (5)
  1. [Abstract] 'We begin by why nations historically cooperate' should read 'We begin by examining why nations historically cooperate.'
  2. [Figure 1] The caption does not define the denominator for the percentage (e.g., percentage of US-authored AI-safety papers with at least one co-author from the indicated country), and the phrase 'co-authorship instances' is ambiguous; please clarify.
  3. [§2.3, footnote 9 vs. §5] Footnote 9 states 'we do not take a position on whether existing measures are sufficient,' but the Conclusion says the four risk sources are 'under-addressed in current risk mitigation strategies'; please reconcile these statements.
  4. [§4.2.1] The TBT/ITU example would be clearer if it distinguished 'international standards developed in the ITU' from the legal requirement in the WTO TBT Agreement, since the current phrasing could be read as claiming the TBT itself endorses ITU standards.
  5. [Appendix A, Table 1] The table is hard to parse because row labels and ratings are interleaved in the provided rendering; consider presenting a standard matrix with explicit row and column headers.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper offers a transparent qualitative risk assessment, not a derivation that reduces to its own inputs.

full rationale

The paper's central claim that verification mechanisms and protocols are less-challenging areas for international cooperation follows directly from the authors' explicit risk ratings in Table 1 and the accompanying prose in Section 4. These ratings are the paper's analytical inputs, not quantities derived from elsewhere, and the conclusion is presented as an assessment rather than as a prediction extracted from a fit. There is no fitted parameter later relabeled as a finding, no definition that smuggles the target conclusion into the premises, and no uniqueness theorem imported from the authors' prior work to force a particular choice. The paper does contain self-citations, including Reuel et al. (Open Problems in Technical AI Governance, [83]) and the International AI Safety Report ([10]), but these are used for background definitions, taxonomies, and context, and the verdict about cooperation suitability does not depend on accepting those citations as load-bearing evidence. A reader may disagree with the risk codings, and the Protocols row of Table 1 may sit uneasily with the text's own discussion of standards capture, but that is a consistency or correctness concern, not circularity.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The paper is a qualitative policy analysis with no equations, no fitted parameters, and no new technical entities. Its conclusions rest on domain assumptions about which risks matter and about the nature of verification and protocol work; these are clearly stated but not empirically demonstrated.

assumptions (6)
  • domain assumption The four risk categories (advancing global frontier capabilities, differential advancement of a rival, exposure of sensitive information, opportunities for harmful action) are sufficient to assess cooperation risks.
    Section 3 defines the typology and the paper explicitly states it makes no claim that these categories are comprehensive, yet the conclusion depends on them being the decisive risks.
  • domain assumption Safety research can have 'capability externalities' that improve model performance, for example RLHF.
    Section 3 cites RLHF as an example; this assumption justifies the first risk category of advancing the global capabilities frontier.
  • domain assumption Verification mechanisms that attest to properties of a system, rather than demonstrate existence of properties, are unlikely to advance AI capabilities.
    Section 4.1.1 uses this distinction to assign 'minimal' capability risk to verification; if attestation methods require capability-relevant innovation, the rating could change.
  • domain assumption Developing protocols codifies existing knowledge rather than extending the knowledge frontier.
    Table 1 and Section 4.2.1 rely on this to assign 'minimal' capability risk to protocols.
  • domain assumption The paper's adopted definitions of coordination, collaboration, and cooperation are valid despite acknowledged lack of academic consensus.
    Section 1.1.1 states the definitions and notes limited consensus in the literature.
  • domain assumption Game-theoretic accounts of international cooperation provide a sound basis for why rivals cooperate and what risks matter.
    Section 2.1 grounds the motivation and risk framework in works such as Jervis (1978) and Coe and Vaynman (2020).

how reviews work

0 comments
Cite this review

Pith. "Pith review of In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?." pith.science (2026). https://pith.science/paper/YQILN66F

@misc{pith2026250412914,
  author       = {Pith},
  title        = {Pith review of: In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YQILN66F}},
  note         = {Machine review of arXiv:2504.12914}
}
read the original abstract

International cooperation is common in AI research, including between geopolitical rivals. While many experts advocate for greater international cooperation on AI safety to address shared global risks, some view cooperation on AI with suspicion, arguing that it can pose unacceptable risks to national security. However, the extent to which cooperation on AI safety poses such risks, as well as provides benefits, depends on the specific area of cooperation. In this paper, we consider technical factors that impact the risks of international cooperation on AI safety research, focusing on the degree to which such cooperation can advance dangerous capabilities, result in the sharing of sensitive information, or provide opportunities for harm. We begin by why nations historically cooperate on strategic technologies and analyse current US-China cooperation in AI as a case study. We further argue that existing frameworks for managing associated risks can be supplemented with consideration of key risks specific to cooperation on technical AI safety research. Through our analysis, we find that research into AI verification mechanisms and shared protocols may be suitable areas for such cooperation. Through this analysis we aim to help researchers and governments identify and mitigate the risks of international cooperation on AI safety research, so that the benefits of cooperation can be fully realised.

Figures

Figures reproduced from arXiv: 2504.12914 by the authors.

Figure 1
Figure 1. AI safety co-authorship instances with American researchers (%). Incomplete data from 2023 and 2024 [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

109 extracted references · 42 canonical work pages

  1. [1]

    ISO/IEC 19790:2012

    2012. ISO/IEC 19790:2012. https://www.iso.org/standard/52906.html

  2. [2]

    Dealing with Risks in International Research Cooperation: Recommendations from the Deutsche Forschungsgemeinschaft

    2023. Dealing with Risks in International Research Cooperation: Recommendations from the Deutsche Forschungsgemeinschaft. https://www.dfg.de/resource/blob/289704/585cb3b48bb8e9f5b6e57e0e0a0d700e/risiken- int-kooperationen-en-data.pdf

  3. [3]

    US AISI and UK AISI Joint Pre-Deployment Test: OpenAI o1

    2024. US AISI and UK AISI Joint Pre-Deployment Test: OpenAI o1 . Technical Report. https://www.nist.gov/system/ files/documents/2024/12/18/US_UK_AI%20Safety%20Institute_%20December_Publication-OpenAIo1.pdf

  4. [4]

    Onni Aarne, Tim Fist, and Caleb Withers. 2024. Secure, Governable Chips: Using On-Chip Mechanisms to Manage National Security Risks from AI & Advanced Computing . Technical Report. Center for a New American Security. https://www.cnas.org/publications/reports/secure-governable-chips

  5. [5]

    Adan, Robert Trager, Kayla Blomquist, Claire Dennis, Gemma Edom, Lucia Velasco, Cecil Abungu, Ben Garfinkel, Julian Jacobs, Chinasa T

    Sumaya N. Adan, Robert Trager, Kayla Blomquist, Claire Dennis, Gemma Edom, Lucia Velasco, Cecil Abungu, Ben Garfinkel, Julian Jacobs, Chinasa T. Okolo, Boxi Wu, and Jai Vipra. 2024. Voice and Access in AI: Global AI Majority Participation in Artificial Intelligence Development and Governance . Technical Report. Oxford Martin AI Governance Initiative, Oxfo...

  6. [6]

    Jide Alaga, Jonas Schuett, and Markus Anderljung. 2024. A Grading Rubric for AI Safety Frameworks. doi:10.48550/ arXiv.2409.08751 arXiv:2409.08751 [cs]

  7. [7]

    Anthropic. 2024. Responsible Scaling Policy. https://assets.anthropic.com/m/24a47b00f10301cd/original/Anthropic- Responsible-Scaling-Policy-2024-10-15.pdf

  8. [8]

    Arms Control Association. 2021. The Open Skies Treaty at a Glance. https://www.armscontrol.org/factsheets/ openskies

Show all 109 references
  1. [9]

    Christel Baier and Joost-Pieter Katoen. 2008. Principles of Model Checking . The MIT Press. https://mitpress.mit.edu/ 9780262026499/principles-of-model-checking/

  2. [10]

    Okolo, Deborah Raji, Girish Sastry, Elizabeth Seger, Theodora Skeadas, Tobin South, Emma Strubell, Florian Tramèr, Lucia Velasco, and Nicole Wheeler

    Yoshua Bengio, Sören Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Philip Fox, Ben Garfinkel, Danielle Goldfarb, Hoda Heidari, Anson Ho, Sayash Kapoor, Leila Khalatbari, Shayne Longpre, Sam Manning, Vasilios Mavroudis, Mantas Mazei...

  3. [11]

    Okolo, Deborah Raji, Theodora Skeadas, and Florian Tramèr

    Yoshua Bengio, Sören Mindermann, Daniel Privitera, Rishi Bommasani, Stephen Casper, Yejin Choi, Danielle Goldfarb, Hoda Heidari, Leila Khalatbari, Shayne Longpre, Vasilios Mavroudis, Mantas Mazeika, Kwan Yee Ng, Chinasa T. Okolo, Deborah Raji, Theodora Skeadas, and Florian Tra...

  4. [12]

    Emily Benson. 2023. Updated October 7 Semiconductor Export Controls. https://www.csis.org/analysis/updated- october-7-semiconductor-export-controls

  5. [13]

    Tamay Besiroglu, Sage Andrus Bergerson, Amelia Michael, Lennart Heim, Xueyun Luo, and Neil Thompson. 2024. The Compute Divide in Machine Learning: A Threat to Academic Contribution and Scrutiny? doi:10.48550/arXiv. 2401.02452 arXiv:2401.02452 [cs]

  6. [14]

    Fateh Boudardara, Abderraouf Boussif, Pierre-Jean Meyer, and Mohamed Ghazel. 2024. A Review of Abstraction Methods Toward Verifying Neural Networks. ACM Trans. Embed. Comput. Syst. 23, 4 (June 2024), 58:1–58:19. doi:10.1145/3617508

  7. [15]

    Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, Tegan Maharaj, Pang Wei Koh, Sara Hooker, Jade Leung, Andrew Trask, Emma Bluemke, Jonathan Lebensold, Cullen O’Keefe, Mark Koren,...

  8. [16]

    John Burden. 2024. Evaluating AI Evaluation: Perils and Prospects. doi:10.48550/arXiv.2407.09221 arXiv:2407.09221 [cs] version: 1

  9. [17]

    Artificial Intelligence Security Commitment

    CAICT. 2024. Protecting AI security and building a model of industry self-discipline - the first batch of 17 companies signed the "Artificial Intelligence Security Commitment". https://mp.weixin.qq.com/s/s-XFKQCWhu0uye4opgb3Ng In Which Areas of Technical AI Safety Could Geopol...

  10. [18]

    Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell

    Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Charbel-Raphaël Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stew...

  11. [19]

    Xavier Castañer and Nuno Oliveira. 2020. Collaboration, Coordination, and Cooperation Among Organizations: Establishing the Distinctive Meanings of These Terms Through a Systematic Literature Review.Journal of Management 46, 6 (July 2020), 965–1001. doi:10.1177/014920632090156...

  12. [20]

    Seth Center and Emma Bates. 2019. Tech-Politik: Historical Perspectives on Innovation, Technology, and Strategic Competition. Technical Report. Center for Strategic & International Studies. https://www.csis.org/analysis/tech- politik-historical-perspectives-innovation-technolo...

  13. [21]

    Yu, Qiang Yang, and Xing Xie

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2024. A Survey on Evaluation of Large Language Models. ACM Trans. Intell. Syst. Technol...

  14. [22]

    Dami Choi, Yonadav Shavit, and David Duvenaud. 2023. Tools for Verifying Neural Models’ Training Data. doi:10. 48550/arXiv.2307.00682 arXiv:2307.00682 [cs]

  15. [23]

    Cathleen D Cimino-Isaacs and Karen M Sutter. 2024. Committee on Foreign Investment in the United States (CFIUS). https://crsreports.congress.gov/product/pdf/IF/IF10177

  16. [24]

    Daniel Clery. 2024. Giant fusion project is in big trouble: ITER operations delayed to 2034, with energy-producing reactions expected 5 years later. Science 385, 6704 (July 2024), 10–11. https://www.science.org/content/article/giant- international-fusion-project-big-trouble

  17. [25]

    Coe and Jane Vaynman

    Andrew J. Coe and Jane Vaynman. 2020. Why Arms Control Is So Rare. American Political Science Review 114, 2 (May 2020), 342–355. doi:10.1017/S000305541900073X

  18. [26]

    Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. 2023. Safe RLHF: Safe Reinforcement Learning from Human Feedback. doi:10.48550/arXiv.2310.12773 arXiv:2310.12773 [cs]

  19. [27]

    David "davidad" Dalrymple. 2024. Safeguarded AI: Constructing guaranteed safety . Technical Report. Advanced Research + Invention Agency. https://www.aria.org.uk/media/3nhijno4/aria-safeguarded-ai-programme-thesis- v1.pdf

  20. [28]

    David "davidad" Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, Alessandro Abate, Joe Halpern, Clark Barrett, Ding Zhao, Tan Zhi-Xuan, Jeannette Wing, and Joshua Tenenbaum. 2024....

  21. [29]

    Sumanth Dathathri, Abigail See, Sumedh Ghaisas, Po-Sen Huang, Rob McAdam, Johannes Welbl, Vandana Bachani, Alex Kaskasoli, Robert Stanforth, Tatiana Matejovicova, Jamie Hayes, Nidhi Vyas, Majd Al Merey, Jonah Brown-Cohen, Rudy Bunel, Borja Balle, Taylan Cemgil, Zahra Ahmed, Ki...

  22. [30]

    Google DeepMind. 2024. Frontier Safety Framework v1.0

  23. [31]

    Jeffrey Ding. 2024. Keep your enemies safer: technical cooperation and transferring nuclear safety and security technologies. European Journal of International Relations 30, 4 (Dec. 2024), 918–945. doi:10.1177/13540661241246622 Publisher: SAGE Publications Ltd

  24. [32]

    Diego Dorn, Alexandre Variengien, Charbel-Raphaël Segerie, and Vincent Corruble. 2024. BELLS: A Framework To- wards Future Proof Benchmarks for the Evaluation of LLM Safeguards. doi:10.48550/arXiv.2406.01364 arXiv:2406.01364 [cs]

  25. [33]

    Lipton, and Hoda Heidari

    Michael Feffer, Anusha Sinha, Wesley Hanwen Deng, Zachary C. Lipton, and Hoda Heidari. 2024. Red-Teaming for Generative AI: Silver Bullet or Security Theater? doi:10.48550/arXiv.2401.15897 arXiv:2401.15897 [cs]

  26. [34]

    Jaden Fiotto-Kaufman, Alexander R. Loftus, Eric Todd, Jannik Brinkmann, Koyena Pal, Dmitrii Troitskii, Michael Ripa, Adam Belfki, Can Rager, Caden Juang, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Nikhil Prakash, Carla Brodley, Arjun Guha, Jonathan Bel...

  27. [35]

    FLI. 2024. FLI AI Safety Index 2024: Independent experts evaluate safety practices of leading AI companies across critical domains. Technical Report. Future of Life Institute. http://futureoflife.org/index

  28. [36]

    Center for Arms Control and Non-Proliferation. 2017. Fact Sheet: The Threshold Test Ban Treaty (TTBT). https: //armscontrolcenter.org/fact-sheet-threshold-test-ban-treaty-ttbt/ 18 Bucknall, Siddiqui et al

  29. [37]

    Department for Science Innovation and Technology. 2023. The Bletchley Declaration by Countries Attending the AI Safety Summit, 1-2 November 2023. https://www.gov.uk/government/publications/ai-safety-summit-2023- the-bletchley-declaration/the-bletchley-declaration-by-countries-...

  30. [38]

    Department for Science Innovation and Technology. 2024. Frontier AI Safety Commitments, AI Seoul Summit

  31. [39]

    Department for Science Innovation and Technology. 2024. Seoul Ministerial Statement for advancing AI safety, innovation and inclusivity: AI Seoul Summit 2024. https://www.gov.uk/government/publications/seoul- ministerial-statement-for-advancing-ai-safety-innovation-and-inclusi...

  32. [40]

    Gallagher

    Nancy W. Gallagher. 1997. The politics of verification: Why ‘how much?’ Is not enough. Contemporary Security Policy 18, 2 (Aug. 1997), 138–170. doi:10.1080/13523269708404165 Publisher: Routledge

  33. [41]

    Sanjam Garg, Aarushi Goel, Somesh Jha, Saeed Mahloujifar, Mohammad Mahmoody, Guru-Vamsi Policharla, and Mingyuan Wang. 2023. Experimenting with Zero-Knowledge Proofs of Training. https://eprint.iacr.org/2023/1345 Publication info: Published elsewhere. Major revision. ACM CCS 2023

  34. [42]

    Soumya Suvra Ghosal, Souradip Chakraborty, Jonas Geiping, Furong Huang, Dinesh Manocha, and Amrit Bedi. 2023. A Survey on the Possibilities & Impossibilities of AI-generated Text Detection. Transactions on Machine Learning Research (Oct. 2023). https://openreview.net/forum?id=...

  35. [43]

    Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. 2024. AI Control: Improving Safety Despite Intentional Subversion. doi:10.48550/arXiv.2312.06942 arXiv:2312.06942 [cs]

  36. [44]

    Charlie Griffin, Louis Thomson, Buck Shlegeris, and Alessandro Abate. 2024. Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols. doi:10.48550/arXiv.2409.07985 arXiv:2409.07985 [cs]

  37. [45]

    Oliver Guest, Michael Aird, and Seán Ó hÉigeartaigh. 2023. Safeguarding the Safeguards: How best to promote AI alignment in the public interest . Technical Report. Institute for AI Policy and Strategy. https://www.iaps.ai/research/ safeguarding-the-safeguards

  38. [46]

    Oliver Guest and Zoe Williams. 2024. Topics for track IIs: What can be discussed in dialogues about ad- vanced AI risks without leaking sensitive information? Technical Report. Institute for AI Policy and Strat- egy. https://static1.squarespace.com/static/64edf8e7f2b10d716b5ba...

  39. [47]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. doi:10.48550/arXiv.1512.03385 arXiv:1512.03385 [cs]

  40. [48]

    Dan Hendrycks. 2024. Introduction to AI Safety, Ethics and Society . Taylor & Francis. https://www.aisafetybook.com/

  41. [49]

    Dan Hendrycks, Mantas Mazeika, and Thomas Woodside. 2023. An Overview of Catastrophic AI Risks. doi:10.48550/ arXiv.2306.12001 arXiv:2306.12001 [cs]

  42. [50]

    Hirschhorn, Brian J

    Eric L. Hirschhorn, Brian J. Egan, Edward J. Krauland, Eric L. Hirschhorn, Brian J. Egan, and Edward J. Krauland. 2022. U.S. Export Controls and Economic Sanctions (fourth edition, fourth edition ed.). Oxford University Press, Oxford, New York

  43. [51]

    C. A. R. Hoare. 1969. An axiomatic basis for computer programming. Commun. ACM 12, 10 (Oct. 1969), 576–580. doi:10.1145/363235.363259

  44. [52]

    The White House. 2022. Blueprint for an AI Bill of Rights: Making automated systems work for the American people. https://www.whitehouse.gov/ostp/ai-bill-of-rights/

  45. [53]

    Kardon, and Matt Sheehan

    Yukon Huang, Isaac B. Kardon, and Matt Sheehan. 2023. Three Takeaways From the Biden-Xi Meeting. https: //carnegieendowment.org/posts/2023/11/three-takeaways-from-the-biden-xi-meeting?lang=en

  46. [54]

    IAEA. 2014. IAEA Safeguards Overview. https://www.iaea.org/publications/factsheets/iaea-safeguards-overview Publisher: IAEA

  47. [55]

    UK AI Safety Institute. 2024. Conference on frontier AI safety frameworks. https://www.aisi.gov.uk/work/conference- on-frontier-ai-safety-frameworks

  48. [56]

    UK AI Safety Institute. 2024. Early lessons from evaluating frontier AI systems. https://www.aisi.gov.uk/work/early- lessons-from-evaluating-frontier-ai-systems

  49. [57]

    Robert Jervis. 1978. Cooperation Under the Security Dilemma.World Politics 30, 2 (1978), 167–214. doi:10.2307/2009958 Publisher: [Trustees of Princeton University, The Johns Hopkins University Press]

  50. [58]

    Wang Jingjing. 2016. The Whampoa Academy of China’s Internet. https://weibo.com/p/1001643998598932131471 English commentary and translation by Jeffrey Ding available at https://chinai.substack.com/p/chinai-37-happy-20th- anniversary

  51. [59]

    Daniel Kang, Tatsunori Hashimoto, Ion Stoica, and Yi Sun. 2022. Scaling up Trustless DNN Inference with Zero- Knowledge Proofs. doi:10.48550/arXiv.2210.08674 arXiv:2210.08674 [cs]. In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate? 19

  52. [60]

    Holden Karnofsky. 2024. If-Then Commitments for AI Risk Reduction. https://carnegieendowment.org/research/ 2024/09/if-then-commitments-for-ai-risk-reduction?lang=en

  53. [61]

    Leonie Koessler, Jonas Schuett, and Markus Anderljung. 2024. Risk thresholds for frontier AI. doi:10.48550/arXiv. 2406.14713 arXiv:2406.14713 [cs]

  54. [62]

    Stephen D. Krasner. 1991. Global Communications and National Power: Life on the Pareto Frontier. World Politics 43, 3 (April 1991), 336–366. doi:10.2307/2010398

  55. [63]

    Gabriel Kulp, Daniel Gonzales, Everett Smith, Lennart Heim, Prateek Puri, Michael J. D. Vermeer, and Zev Winkelman

  56. [64]

    Leveson and John P

    Nancy G. Leveson and John P. Thomas. 2023. Certification of Safety-Critical Systems. Commun. ACM 66, 10 (Sept. 2023), 22–26. doi:10.1145/3615860

  57. [65]

    Technical Report

    Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090 . Technical Report. RAND Corporation. https: //www.rand.org/pubs/working_papers/WRA3056-1.html

  58. [66]

    Jean-Christophe Mauduit. 2017. Collaboration around the International Space Station: Science for diplomacy and its implication for U.S.-Russia and China relations. In Proceedings of 7th Annual SAIS Asia Conference (SAIS 2018) , Vol. 462. Washington, DC, 412–413. doi:10.1038/462412a

  59. [67]

    Michael Martina and Trevor Hunnicutt. 2024. US, China meet in Geneva to discuss AI risks. Reuters (May 2024). https://www.reuters.com/technology/us-china-meet-geneva-discuss-ai-risks-2024-05-13/

  60. [68]

    Mouton, Caleb Lucas, and Ella Guest

    Christopher A. Mouton, Caleb Lucas, and Ella Guest. 2023.The Operational Risks of AI in Large-Scale Biological Attacks: A Red-Team Approach. Technical Report. RAND Corporation. https://www.rand.org/pubs/research_reports/RRA2977- 1.html

  61. [69]

    Chris Miller. 2023. Chip War: The Fight for the World’s Most Critical Technology . Simon & Schuster. https://www. simonandschuster.co.uk/books/Chip-War/Chris-Miller/9781398504127

  62. [70]

    Peter Naur. 1966. Proof of algorithms by general snapshots. BIT Numerical Mathematics 6, 4 (July 1966), 310–316. doi:10.1007/BF01966091

  63. [71]

    National Institute of Standards and Technology (US). 2024. Artificial Intelligence Risk Management Framework: Gener- ative Artificial Intelligence Profile. Technical Report NIST AI 600-1. National Institute of Standards and Technology (U.S.), Gaithersburg, MD. error: 600–1 pag...

  64. [72]

    Emerging Technology Observatory. 2024. Country Activity Tracker (CAT): Artificial Intelligence. https://cat.eto.tech/

  65. [73]

    Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, and Jeff Alstott. 2024. Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models . Technical Report. RAND Corporation. https: //www.rand.org/pubs/research_reports/RRA2849-1.html

  66. [74]

    Department of Commerce

    U.S. Department of Commerce. 2024. U.S. Secretary of Commerce Raimondo and U.S. Secretary of State Blinken Announce Inaugural Convening of International Network of AI Safety Institutes in San Francisco | U.S. Department of Commerce. https://www.commerce.gov/news/press-releases...

  67. [75]

    International Network of AI Safety Institutes. 2024. Improving International Testing of Foundation Models: A Pilot Testing Exercise from the International Network of AI Safety Institutes . Technical Report. San Fran- cisco. https://www.nist.gov/system/files/documents/2024/11/2...

  68. [76]

    World Health Organization. 2021. Ethics and Governance of Artificial Intelligence for Health: WHO guidance. https://www.who.int/publications/i/item/9789240029200

  69. [77]

    OpenAI. 2023. Preparedness Framework (Beta). https://cdn.openai.com/openai-preparedness-framework-beta.pdf

  70. [78]

    Sean O’Connor. 2019. How Chinese Companies Facilitate Technology Transfer from the United States . Staff Research Report. U.S.-China Economic and Security Review Commission, Washington, DC. https://www.uscc.gov/sites/ default/files/Research/How%20Chinese%20Companies%20Facilita...

  71. [79]

    2024.Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI

    Joe O’Brien, Shaun Ee, Jam Kraprayoon, Bill Anderson-Samways, Oscar Delaney, and Zoe Williams. 2024.Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI . Technical Report. Institute for AI Policy and Strategy. https://www.iaps.ai/research/c...

  72. [80]

    James Petrie, Onni Aarne, Nora Ammann, and David "davidad" Dalrymple. 2024. Interim Report: Mechanisms for Flexible Hardware-Enabled Guarantees. Technical Report

  73. [81]

    Qianqian Pan, Mianxiong Dong, Kaoru Ota, and Jun Wu. 2022. Device-Bind Key-Storageless Hardware AI Model IP Pro- tection: A PUF and Permute-Diffusion Encryption-Enabled Approach. doi:10.48550/arXiv.2212.11133 arXiv:2212.11133 [cs]

  74. [82]

    Jarrett Renshaw and Trevor Hunnicutt. 2024. Biden, Xi agree that humans, not AI, should control nuclear arms. Reuters (Nov. 2024). https://www.reuters.com/world/biden-xi-agreed-that-humans-not-ai-should-control-nuclear- weapons-white-house-2024-11-16/

  75. [83]

    Hadrien Pouget, Claire Dennis, Jon Bateman, Robert F. Trager, Renan Araujo, Haydn Belfield, Belinda Cleeland, Malou Estier, Gideon Futerman, Oliver Guest, Carlos Ignacio Gutierrez, Vishnu Kannan, Casey Mahoney, Matthijs Maas, Charles Martinet, Jakob Mökander, Kwan Yee Ng, Seán...

  76. [84]

    Angelika Romanou, Negar Foroutan, Anna Sotnikova, Zeming Chen, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Mohamed A. Haggag, Snegha A, Alfonso Amayuelas, Azril Hafizi Amirudin, Viraat Aryabumi, Danylo Boiko, Michael Chang, Jenny Chim, Gal Cohen,...

  77. [85]

    Kochenderfer, and Robert Trager

    Anka Reuel, Ben Bucknall, Stephen Casper, Tim Fist, Lisa Soder, Onni Aarne, Lewis Hammond, Lujain Ibrahim, Alan Chan, Peter Wills, Markus Anderljung, Ben Garfinkel, Lennart Heim, Andrew Trask, Gabriel Mukobi, Rylan Schaeffer, Mauricio Baker, Sara Hooker, Irene Solaiman, Alexan...

  78. [86]

    Hadfield, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Janet Egan, Robert F

    Girish Sastry, Lennart Heim, Haydn Belfield, Markus Anderljung, Miles Brundage, Julian Hazell, Cullen O’Keefe, Gillian K. Hadfield, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Janet Egan, Robert F. Trager, Shahar Avin, Adrian Weller, Yoshua Bengio, an...

  79. [87]

    Tim Rühlig. 2023. The Geopolitics of Technical Standardization. https://dgap.org/en/research/publications/ geopolitics-technical-standardization

  80. [88]

    Yonadav Shavit. 2023. What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. doi:10.48550/arXiv.2303.11341 arXiv:2303.11341 [cs]

  81. [89]

    Jonas Schuett, Markus Anderljung, Alexis Carlier, Leonie Koessler, and Ben Garfinkel. 2024. From Principles to Rules: A Regulatory Approach for Frontier AI. doi:10.48550/arXiv.2407.07300 arXiv:2407.07300 [cs]

  82. [90]

    Scott Singer. 2024. How the UK Should Engage China at AI’s Frontier. https://carnegieendowment.org/posts/2024/ 10/lammy-china-ai-safety-cooperation?lang=en

  83. [91]

    Matt Sheehan and Jacob Feldgoise. 2023. What Washington Gets Wrong About China and Technical Stan- dards. https://carnegieendowment.org/research/2023/02/what-washington-gets-wrong-about-china-and-technical- standards?lang=en

  84. [92]

    Arthur A. Stein. 1982. Coordination and collaboration: regimes in an anarchic world. International Organization 36, 2 (1982), 299–324. doi:10.1017/S0020818300018968

  85. [93]

    Tobin South, Alexander Camuto, Shrey Jain, Shayla Nguyen, Robert Mahari, Christian Paquin, Jason Morton, and Alex ’Sandy’ Pentland. 2024. Verifiable evaluations of machine learning models using zkSNARKs. doi:10.48550/arXiv. 2402.02675 arXiv:2402.02675 [cs]

  86. [94]

    Merlin Stein and Connor Dunlop. 2024. Safe beyond sale: post-deployment monitoring of AI. https://www. adalovelaceinstitute.org/blog/post-deployment-monitoring-of-ai/

  87. [95]

    Merlin Stein, Jamie Bernardi, and Connor Dunlop. 2024. The Role of Governments in Increasing Interconnected Post-Deployment Monitoring of AI. doi:10.48550/arXiv.2410.04931 arXiv:2410.04931 [cs]

  88. [96]

    Andrew Trask, Aziz Berkay Yesilyurt, Bennett Farkas, Callis Ezenwaka, Carmen Popa, Dave Buckley, Eelco van der Wel, Francesco Mosconi, Grace Han, Ionesio Junior, Irina Bejan, Ishan Mishra, Khoa Nguyen, Koen van der Veen, Kyoko Eng, Lacey Strahm, Logan Graham, Madhava Jay, Mate...

  89. [97]

    Mengqi Sun. 2024. U.S., China to Cooperate in the Fight Against Dirty Money. Wall Street Journal (April 2024). https://www.wsj.com/articles/u-s-china-to-cooperate-in-the-fight-against-dirty-money-1edb9a25

  90. [98]

    UN HLAB. 2024. Govering AI for Humanity: Final Report . Technical Report. United Nations, New York, NY. https: //www.un.org/ai-advisory-body

  91. [99]

    UKRI. 2022. Managing risks in international research and innovation: An overview of higher education sector guidance. https://www.ukri.org/wp-content/uploads/2022/07/UKRI-07072022-managing-risks-in-international- In Which Areas of Technical AI Safety Could Geopolitical Rivals ...

  92. [100]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou

  93. [101]

    Caterina Urban and Antoine Miné. 2021. A Review of Formal Methods applied to Machine Learning. doi:10.48550/ arXiv.2104.02466 arXiv:2104.02466 [cs]

  94. [102]

    JoAnne Yates and Craig N. Murphy. 2019.Engineering Rules. Johns Hopkins University Press. doi:10.1353/book.66187

  95. [103]

    Zhang, Kevin Klyman, Yifan Mai, Yoav Levine, Yian Zhang, Rishi Bommasani, and Percy Liang

    Andy K. Zhang, Kevin Klyman, Yifan Mai, Yoav Levine, Yian Zhang, Rishi Bommasani, and Percy Liang. 2024. Language model developers should report train-test overlap. doi:10.48550/arXiv.2410.08385 arXiv:2410.08385 [cs]

  96. [104]

    Zico Kolter, Jakob Foerster, and Martin Strohmeier

    Christian Schroeder de Witt, Samuel Sokota, J. Zico Kolter, Jakob Foerster, and Martin Strohmeier. 2023. Perfectly Secure Steganography Using Minimum Entropy Coupling. doi:10.48550/arXiv.2210.14889 arXiv:2210.14889 [cs]

  97. [105]

    Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han. 2023. Don’t Make Your LLM an Evaluation Benchmark Cheater. doi:10.48550/arXiv.2311.01964 arXiv:2311.01964 [cs]

  98. [106]

    Nicholas Zúñiga, Saheli Datta Burton, Filippo Blancato, and Madeline Carr. 2024. The geopolitics of technology standards: historical context for US, EU and Chinese approaches. International Affairs 100, 4 (July 2024), 1635–1652. doi:10.1093/ia/iiae124 22 Bucknall, Siddiqui et ...

  99. [107]

    Daniel Zhang, Nestor Maslej, Erik Brynjolfsson, John Etchemendy, Terah Lyons, James Manyika, Helen Ngo, Juan Carlos Niebles, Michael Sellitto, Ellie Sakhaee, Yoav Shoham, Jack Clark, and Raymond Perrault. 2022. The AI Index 2022 Annual Report . Technical Report. Stanford Insti...

  100. [2023]

    doi:10.48550/arXiv.2201.11903 arXiv:2201.11903 [cs]

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. doi:10.48550/arXiv.2201.11903 arXiv:2201.11903 [cs]

  101. [2024]

    https://www.gov.uk/government/publications/frontier-ai-safety-commitments-ai-seoul-summit-2024/frontier- ai-safety-commitments-ai-seoul-summit-2024

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.