{"id":"513d40c6-9f39-4e13-88f9-103c1facb7fe","arxiv_id":"2504.12914","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Based on a four-risk typology, the paper concludes that verification mechanisms and codified protocols are the least risky areas for cooperation between geopolitical rivals on technical AI safety.","lead":"This paper argues that international rivals can most safely cooperate on AI verification tools and shared safety protocols, while cooperation on shared infrastructure and model evaluations carries higher risks. It offers a framework for weighing capability, information, and sabotage risks before governments choose where to cooperate.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Suitability ranking relies on a risk-only rubric that omits the political/process risks the paper itself discusses; Protocols' 'Minimal' harmful-action coding is internally inconsistent.","rationale":"The reader identified the risk-only treatment of suitability as the weakest assumption. I agree that this is load-bearing, and I extend it with a specific internal inconsistency: the paper's own Section 4.2.1 presents evidence that standardisation processes are used to advance national interests, yet Table 1 codes Protocols as 'Minimal' on all four risk dimensions, including harmful action. This makes the protocols pillar of the central claim unsupported by the text. The paper is otherwise careful and well-referenced, with appropriately hedged language in the body, and the verification claim may well survive. But the combination of an omitted-benefits caveat and an inconsistent risk coding means the conclusion requires revision or re-scoring before it can be accepted as a policy recommendation. The proposed concrete test—adding a political/process risk dimension or re-coding the harmful-action cell—would settle whether the ranking changes. This does not overturn the reader's CONDITIONAL verdict; it sharpens the conditions under which the paper's conclusion would hold.","tokens_in":21800,"tokens_out":5048,"duration_ms":52740,"concrete_test":"Produce a revised Table 1 that adds a fifth risk dimension, 'Political/process capture risk,' defined via refs [62], [85], and [106] as the risk that a cooperation venue becomes a vehicle for advancing one party's interests (e.g., standards capture, politicization of technical outputs). Re-score all four areas using only evidence already cited in the paper. If Protocols' total risk (the four original dimensions plus the fifth) exceeds Evaluations' total, the paper's ranking of protocols as less challenging is not supported. As a minimal sensitivity check, re-code only the Protocols row's 'Provides opportunity for harmful action' cell from 'Minimal' to 'Moderate' (consistent with Section 4.2.1's standards-capture discussion) and re-derive the qualitative ordering; if the ordering changes, the central claim fails to robustly follow from the paper's own analysis.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim (verification and protocols are 'less-challenging' and 'may be suitable' for cooperation) is derived from Table 1, which scores four risk dimensions. This is load-bearing for the policy conclusion, but the table and the text diverge. The Introduction explicitly disclaims any analysis of benefits: 'We do not aim to definitively identify the most suitable areas for cooperation, nor to investigate specific benefits of cooperation in different AI safety research areas.' Suitability is thus equated with low risk on four technical dimensions, leaving out political, reputational, and process risks. More concretely, Section 4.2.1 states that 'both states and industry actors tend to use the international standardisation process to advance their own interests, potentially to the detriment of other actors' (citing [62], [85], [106]), and the protocols subsection warns that 'protocols... could be more politicised... if the process becomes co-opted and overly politicised.' Yet Table 1 codes Protocols as 'Minimal' on all four dimensions, including 'Provides opportunity for harmful action.' This is internally inconsistent: a venue documented as a vehicle for unilateral interest advancement cannot credibly receive the lowest possible harmful-action score. If that cell were re-coded as 'Moderate' (as the text's own evidence suggests), the aggregate risk of Protocols would be comparable to or greater than that of Evaluations, whose harmful-action risk is rated 'Minimal/moderate.' The headline conclusion that protocols are a less-challenging area for cooperation therefore rests on an unsupported coding, not merely on an omitted-benefits caveat. The paper's own evidence undermines one of the two pillars of the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks where US-China rivals can safely cooperate on technical AI safety. It defines four risk dimensions relevant to cooperation—advancing the global capabilities frontier, differentially advancing a rival's capabilities, exposing sensitive information, and providing opportunities for harmful action—and applies them to four candidate areas: verification mechanisms, codified protocols and best practices, shared infrastructure, and evaluation methodologies. Using historical analogies, survey evidence on US-China AI collaboration, and a qualitative assessment table, the paper concludes that verification and protocols are less-challenging areas for cooperation than infrastructure and evaluations and 'may be well-suited' for international cooperation. The paper explicitly disclaims any analysis of the benefits of cooperation and notes that its risk categories are not comprehensive, so the conclusion is best read as a risk-based suitability judgment rather than a full cost-benefit analysis.","tokens_in":22056,"tokens_out":5645,"duration_ms":59467,"significance":"If the ranking holds, the paper gives policymakers a usable starting point for directing rival cooperation toward verification and protocols while treating infrastructure and evaluations as riskier. The paper's strengths are its clear definitions, its structured four-dimensional typology, its use of concrete historical analogues (Open Skies, PALs, the Joint Verification Experiment), and the explicit hedged language ('may,' 'less-challenging'). Table 1 makes the underlying qualitative judgments checkable, and the paper is transparent that the list of areas and risks is non-exhaustive. The central limitation is that suitability is inferred from low risk alone, without weighing benefits or political/process risks; this does not invalidate the paper but does bound the strength of the policy conclusion.","major_comments":[{"comment":"The paper's central conclusion—that verification and protocols 'may be well-suited' for cooperation—is derived from a risk-only assessment, but the Introduction explicitly states that the paper does 'not aim to definitively identify the most suitable areas for cooperation, nor to investigate specific benefits.' If benefits differ across areas (e.g., protocols may be easy to agree on but weak in changing behavior, while verification may require deep system access to be valuable), a higher-risk, higher-benefit area could be more suitable than a lower-risk, lower-benefit one. I recommend either reframing the conclusion as 'less risky areas for cooperation' or adding a qualitative benefit discussion for each area to support the suitability claim.","section":"Introduction, §4, Table 1"},{"comment":"The Protocols row codes 'Provides opportunity for harmful action' as 'Minimal,' yet the same subsection states that 'both states and industry actors tend to use the international standardisation process to advance their own interests, potentially to the detriment of other actors' and that protocols 'could be more politicised... leading to a degradation in scientific rigour.' The table's own note concedes that 'standardisation has sometimes been used to advance unilateral interests.' At minimum, the coding and the text are in tension; if the harmful-action category is meant to include harm through process capture, a 'Minimal/moderate' or 'Moderate' coding would be more consistent, and this would materially weaken the paper's ranking of Protocols as uniformly lowest-risk.","section":"§4.2.1, Table 1"}],"minor_comments":[{"comment":"'We begin by why nations historically cooperate' should read 'We begin by examining why nations historically cooperate.'","section":"Abstract"},{"comment":"The caption does not define the denominator for the percentage (e.g., percentage of US-authored AI-safety papers with at least one co-author from the indicated country), and the phrase 'co-authorship instances' is ambiguous; please clarify.","section":"Figure 1"},{"comment":"Footnote 9 states 'we do not take a position on whether existing measures are sufficient,' but the Conclusion says the four risk sources are 'under-addressed in current risk mitigation strategies'; please reconcile these statements.","section":"§2.3, footnote 9 vs. §5"},{"comment":"The TBT/ITU example would be clearer if it distinguished 'international standards developed in the ITU' from the legal requirement in the WTO TBT Agreement, since the current phrasing could be read as claiming the TBT itself endorses ITU standards.","section":"§4.2.1"},{"comment":"The table is hard to parse because row labels and ratings are interleaved in the provided rendering; consider presenting a standard matrix with explicit row and column headers.","section":"Appendix A, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on the authors' own prior outputs (e.g., [10], [83]) and on reports from affiliated institutions. This is not inappropriate given the topic, but a more diverse citation base—particularly on international standardisation politics and on empirical studies of international research collaboration—would strengthen the appearance of balance. I do not see this as warranting rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Sarah, quick take: this is one of the few papers that tells governments where cooperation between AI rivals is less likely to blow up. The four-risk typology applied to verification, protocols, infrastructure, and evaluations is a useful synthesis; the historical parallels (Open Skies, PALs, Joint Verification Experiment) are apt. The paper is honest about scope: it does not claim to weigh benefits, and the body hedges with 'may'. A policymaker asking for a starting point on US–China technical AI safety cooperation would get value from this.\n\nData and citation pattern are fine: no formal math to check, the co-authorship figure is descriptive, and the references are relevant and wide. There is substantial self-citation to the authors' Open Problems in Technical AI Governance and International AI Safety Report, but those are real prior works and the central claim does not depend on them.\n\nSoft spots are real but fixable. First, Table 1 is a set of qualitative judgments without a rubric or any account of who scored and how, so the ranking is hard to audit. Second, the paper explicitly sets aside benefits yet draws a suitability conclusion; low risk is not well-suited unless benefits are roughly equal across areas, and nothing supports that. Third, an internal inconsistency: the Protocols row codes harmful action as Minimal, but Section 4.2.1 says standardisation processes have been used by states and industry to advance their own interests to others' detriment. That is not Minimal; re-coding to at least Minimal/moderate pulls Protocols close to Evaluations and weakens the headline claim. Fourth, the abstract states the conclusion more strongly than the body's caveats. Fifth, the candidate areas are preselected from advocacy, so the menu is not systematic.\n\nNone of this sinks the paper. The central observation—verification and protocols are comparatively lower-risk, with benefits honestly left out—is plausible and well argued. It deserves serious peer review, not desk rejection. I'd send it to a referee with instructions to focus on the risk-rating methodology and the internal consistency of Table 1. With a rubric, a re-coded Protocols row, and a conclusion that tracks the benefits caveat, this would be a solid policy-facing paper.\n\nRecommendation: engage, bring to reading group, cite, and referee if asked.","headline":"A genuinely useful risk-mapping for US–China AI safety cooperation, but the headline ranking leans on a qualitative table whose Protocol row is coded against the paper's own evidence.","tokens_in":22699,"tokens_out":3368,"would_cite":true,"duration_ms":34785,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Verification mechanisms and shared protocols are the least risky areas for AI safety cooperation between geopolitical rivals; shared infrastructure and joint evaluations carry higher risks of capability transfer, sensitive-information…","keywords":["technical AI safety","international cooperation","geopolitical rivalry","AI verification mechanisms","AI protocols and standards","AI evaluation","AI infrastructure","US-China AI cooperation"],"falsifier":"A concrete test would be to run the same four-risk analysis on documented cooperation projects—for example, a bilateral compute-attestation prototype or the joint US–UK evaluation exercise—and find that a verification or protocol project exposed sensitive information, advanced a rival's capabilities, or enabled a covert backdoor; a structured expert elicitation that ranks verification or protocols riskier than infrastructure or evaluations would likewise falsify the paper's ordering.","tokens_in":21628,"feed_emoji":"🤝","tokens_out":12128,"duration_ms":115630,"temperature":0.7,"pith_summary":"The paper asks which areas of technical AI safety research are safe enough for geopolitical rivals—chiefly the United States and China—to cooperate on, and it argues that riskiness is area-specific rather than uniform. It builds a four-part risk typology—advancing the global capabilities frontier, differentially advancing a rival's capabilities, exposing sensitive information, and creating opportunities for harmful action—and applies it to four proposed cooperation areas: verification mechanisms, shared protocols and best practices, shared infrastructure, and evaluation methodologies. Its central finding is that research on AI verification mechanisms and on shared protocols is less challenging for rival cooperation than infrastructure and evaluations, because these areas mostly attest to or codify existing knowledge instead of extending it. The paper is best read as a risk-screening device: it identifies where the downside of cooperation is smallest, while explicitly setting aside whether the benefits in each area justify the effort.","feed_headline":"Verification and protocols carry the least AI-safety cooperation risk","feed_subtitle":"Comparing four AI safety research areas shows where rivals can cooperate with minimal downside.","key_machinery":"The analytical core is a two-way risk matrix. Four risk categories—advancing the global frontier of AI capabilities, differentially advancing a rival's capabilities, exposing nationally strategic information, and enabling harmful action by motivated actors—are crossed with four candidate cooperation areas: verification mechanisms, codified protocols and best practices, shared infrastructure, and evaluation methodologies. Each area is unpacked into subareas such as formal verification, verifiable audits, compute attestation, watermarking, safety frameworks, incident-reporting standards, secure-weight standards, evaluation infrastructure, shared compute clusters, and capability evaluations. The matrix does the argumentative work: the paper's verdict that verification and protocols are 'less-challenging areas for international cooperation' is a summary of the risk columns of that matrix, with historical precedents such as the Open Skies Treaty and Permissive Action Links supplying the reason to believe verification cooperation can stay low-risk.","core_discovery":"The paper's central claim is that the risks of international cooperation on technical AI safety vary systematically by area, and that verification mechanisms and shared protocols are the two areas where such risks are lowest. Verification is comparatively low-risk, the authors argue, because it predominantly certifies claims about systems rather than improving them, and the properties verified can be restricted in advance to those all parties already know, as with the certified sensors of the Open Skies Treaty. Protocols are low-risk because they codify existing knowledge into agreed procedures rather than pushing the research frontier, and they need not involve direct access to live AI systems or sensitive infrastructure. Infrastructure and evaluations, by contrast, carry higher risk of misuse, backdoor insertion, leakage of sensitive domain knowledge, and differential capability gain. The paper does not say the riskier areas should be avoided, nor that the safer areas are the most valuable; it limits its conclusion to a risk-based assessment of which areas are less challenging, and explicitly leaves benefit analysis to future work.","pith_inferences":["Because the paper assesses only risks, its 'suitable' verdict should be read as a risk ranking rather than a priority list; a low-risk area could still be low-value, and verifying the benefits of verification and protocol work is the natural next step.","The historical analogies suggest a sharper test than the paper states: verification cooperation is easiest when the verified object is standardized, mutually observable, and low-complexity, so early compute attestation or watermarking is a more realistic first target than formal verification of frontier models.","The protocols category's low technical risk may be offset by political risk, since standards bodies have repeatedly been used to advance national advantage; the paper's verdict for protocols therefore depends on institutional design that resists capture.","A finer-grained risk map—scoring the paper's named subareas across the four risks through expert elicitation—would turn its four-area ranking into a portfolio tool for choosing which specific projects to propose to rivals."],"forward_implications":["Governments can treat AI verification mechanisms—compute attestation, watermarks, verifiable audits—as the first candidates for structured cooperation between rival AI safety institutes and in track-II dialogues.","Protocol and standards development (safety frameworks, incident-reporting definitions, secure-weight standards) can be pursued without requiring parties to disclose sensitive details about their own systems.","Cooperation on shared infrastructure and joint evaluations of dangerous capabilities should be paired with explicit mitigations, such as open-source development, bug bounties, restricted access, and secure evaluation enclaves.","Vetting processes for international research collaborations can be supplemented with the paper's four technology-specific risk categories, rather than relying only on due diligence and sanctions screening.","The finding gives intergovernmental cooperation on AI safety a concrete starting agenda: verification and protocol work can build trust before harder questions of shared compute or capability evaluation are attempted."],"supporting_citations":[{"why":"The International AI Safety Report identifies verification mechanisms as a proposed area for cooperation, supplying the paper's leading low-risk candidate.","marker":"[10]"},{"why":"Defines verification mechanisms as technical procedures for supporting verifiable claims about AI systems, grounding the paper's lowest-risk category.","marker":"[15]"},{"why":"Extends the definition of verification and the case for guaranteed safe AI, cited alongside [15] to frame what verification research covers.","marker":"[28]"},{"why":"Supplies the contrast between assessment and verification problems and grounds verifiable audits, a subarea used in the verification-versus-evaluation comparison.","marker":"[83]"},{"why":"Source for the capability-externality risk that safety research can advance AI capabilities, the first category in the paper's risk typology.","marker":"[45]"},{"why":"Game-theoretic account of why arms-control verification is rare and what makes it possible, used to frame rival cooperation on verification.","marker":"[25]"},{"why":"Open Skies Treaty precedent showing verification can be restricted to certified, pre-agreed sensors, a core analogue for low-risk verification.","marker":"[8]"},{"why":"Documents the transfer of Permissive Action Links between rivals, showing safety technology can be shared without leaking sensitive weapons information.","marker":"[31]"},{"why":"Argues for outcomes-based regulation and shared rules, supporting the protocols category's premise that codification need not expose sensitive details.","marker":"[87]"}],"fun_headline_variants":["Verification and shared protocols are lowest-risk AI safety cooperation","Where can AI rivals cooperate safely? Verification and protocols","AI safety cooperation: verification and protocols carry low risk","For rivals, verification and protocols are the safest AI safety areas"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper equates suitability for cooperation with low exposure to its four risks and explicitly leaves the benefits of cooperation unexamined; if benefits differ by area, a low-risk area could still be a poor place to cooperate, and the ranking could change.","fun_headline_variants_meta":{"raw":{"variants":["Verification and shared protocols are lowest-risk AI safety cooperation","Where can AI rivals cooperate safely? Verification and protocols","AI safety cooperation: verification and protocols carry low risk","For rivals, verification and protocols are the safest AI safety areas"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000615,"raw_usage":{"total_tokens":2852,"prompt_tokens":938,"completion_tokens":1914,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":1848}},"tokens_in":554,"tokens_out":1914,"duration_ms":13802,"temperature":1.0,"reasoning_tokens":1848,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:19:09.850027+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be to run the same four-risk analysis on documented cooperation projects—for example, a bilateral compute-attestation prototype or the joint US–UK evaluation exercise—and find that a verification or protocol project exposed sensitive information, advanced a rival's capabilities, or enabled a covert backdoor; a structured expert elicitation that ranks verification or protocols riskier than infrastructure or evaluations would likewise falsify the paper's ordering.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the contrast between assessment and verification problems and grounds verifiable audits, a subarea used in the verification-versus-evaluation comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argues for outcomes-based regulation and shared rules, supporting the protocols category's premise that codification need not expose sensitive details."}],"review_version":1}