REVIEW 2 major objections 5 minor 11 references
Against racing to AGI: Cooperation, deterrence, and catastrophic risks
T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper argues that racing to build AGI first is contrary to every major actor's own self-interest, since the expected costs are higher and the prize of domination less attainable than commonly assumed.
desk verdict A serious, well-written argument against AGI racing that grants the proponent's assumptions and still finds the race unjustifiable; the payoff matrices encode the conclusion, but the paper is honest about its limits and deserves review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a sequence of ordinal payoff matrices for two agents deciding whether to race or cooperate, updated as each section revises the expected value of unilateral racing and mutual cooperation. The matrices show the race shifting from a prisoner's dilemma, in which racing dominates, to a stag hunt, in which cooperation is both better and undominated. Supporting this machinery are two central concepts: a decisive strategic advantage (DSA), the level of technological superiority sufficient for world domination, which the paper argues is hard to obtain; and capability restraint, the ability to stop, inhibit, or steer AI development, which the paper argues racing erodes and yet much technical safety research presupposes.
What would settle it
Show that a racing state can actually secure a decisive strategic advantage, for example by protecting frontier model weights at the highest cybersecurity level and withstanding sabotage, or show that verification and enforcement cannot credibly raise the cost of defection; either result would undercut the claim that cooperation is the self-interested alternative.
Extended reading notes
Core claim
The central claim is that AGI Racing, the doctrine that a major actor should accelerate frontier AI development to get AGI first, is not in anyone's self-interest. Even granting that AGI is feasible, that it could confer a decisive strategic advantage, that national security depends on winning, and that rivals may race, the expected costs of racing outweigh its benefits. Racing multiplies catastrophic risks, including misaligned takeover, catastrophic misuse, preventive war, systemic collapse, gradual disempowerment, and nuclear instability, and it undermines the technical safety research that is supposed to justify it. At the same time, the upside of obtaining a decisive strategic advantage is doubtful: weak cybersecurity enables model theft, rivals can respond with sabotage or kinetic attack, and capability gains may be incremental rather than sudden. The paper therefore transforms the strategic picture from a prisoner's dilemma, where racing is the dominant strategy, into a stag hunt, where mutual cooperation is both preferable and strategically viable.
Load-bearing premise
The argument depends on the possibility of making international cooperation robust through credible verification and enforcement, so that defection becomes costly enough to remove racing's appeal; the authors themselves leave the likelihood of that cooperation open.
Editorial extensions
If this is right
- A race to AGI raises catastrophic risks, including nuclear instability, and erodes the time, transparency, and social adaptation needed to control them.
- The expected prize of racing, a decisive strategic advantage, is less likely than assumed because model theft, sabotage, and the difficulty of maintaining a large lead make domination uncertain.
- Because a state's own racing signals make rivals more likely to race, the belief that a race is inevitable can become a self-fulfilling prophecy.
- International cooperation with verification and enforcement can deliver most of the security benefits racing promises while lowering catastrophic risk and, unlike deterrence, need not increase escalation risk.
- Deterrence of the mutually assured AI malfunction kind may buy time, but its escalation risks and the difficulty of stopping algorithmic or hardware progress limit its usefulness.
Reading between the lines
- If the payoff-shifting argument is right, then public 'race' declarations are not neutral descriptions: they are moves that alter competitors' incentives, so governments should evaluate race rhetoric as an intervention rather than a forecast.
- The paper leaves open whether verification and enforcement can be made credible; a natural next step is to study enforcement mechanisms as seriously as alignment, because without credible defection costs the stag-hunt conclusion would fail.
- The model-theft and sabotage considerations imply that security of frontier AI development is as important as capability: robust cybersecurity could invert the paper's skepticism about decisive strategic advantages, making the argument's empirical fork testable.
- If racing degrades capability restraint, then safety research and restraint agreements are complements rather than substitutes: safety research is more valuable when institutions can actually pause or redirect development.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues against AGI Racing, the view that states should accelerate frontier AI development to beat competitors to AGI. It grants several assumptions of racing proponents (technical feasibility, DSA potential, existential stakes, opponent racing risk) and then argues that racing nonetheless is not in any actor's self-interest. The argument has three parts: (i) racing substantially increases catastrophic risks (misalignment, misuse, preventive war, nuclear instability, social fragility) and undermines technical safety research; (ii) the expected benefit of racing—achieving a decisive strategic advantage—is less likely than commonly assumed because of diminishing returns, model theft, and sabotage/kinetic risks; (iii) alternatives exist, particularly international cooperation and possibly deterrence, which avoid the worst risks while addressing the security motivations for racing. The conclusion is that racing to AGI is not in anyone's self-interest because cooperation is preferable. The paper uses a series of 2x2 payoff matrices (Tables 1–4) to illustrate the strategic structure, moving from a prisoner's dilemma to a stag hunt as the payoffs are adjusted based on the preceding arguments.
Significance. If the conclusion were established, the paper would make a valuable contribution to AI governance debates by reframing the race as a coordination problem rather than a competitive one. The paper is unusually fair-minded: it grants many assumptions of its opponents and then shows that the argument for racing still fails. It also provides a useful and well-referenced survey of catastrophic AI risks and a nuanced discussion of deterrence (MAIM) and its limitations. The paper is explicitly transparent about what it does not establish, notably the feasibility of international cooperation. Because the central claim is conditional on that feasibility, the significance of the paper as a decisive argument is limited, but it remains a strong and useful policy analysis.
major comments (2)
- [Section 6.2, Table 4] The central conclusion that cooperation is preferable to racing rests on the payoff ordering N > UR in Table 4, but this ordering is asserted rather than derived. The paper explicitly states in Section 6.2 that 'It is outside the scope of this paper to assess the likelihood that the requisite extent of international coordination and cooperation comes to pass' and acknowledges 'a paucity of concrete proposals for enforcement methods.' Without a credible mechanism that makes cooperation robust against defection, a state considering cooperation risks the worst outcome (UN = 0) if its adversary defects, and the game may remain a prisoner's dilemma. The categorical conclusion in Section 7 ('racing to AGI is not in anyone's self-interest') is therefore stronger than the analysis supports. The authors should either supply an argument for why robust enforcement is feasible or revise the conclusion to a conditional form, e.g., 'if robust international cooperation is achievable, then racing is not in anyone's self-interest.'
- [Section 5] The claim that 'if one refrains from racing, one's opponents are substantially less likely to race' is load-bearing, because it is used to argue that the conditional scenario where the opponent races is not the relevant baseline. However, this causal claim is asserted without direct evidence. The paper cites Ó hÉigeartaigh (2025) to argue that we are not currently in a race, which supports the claim that racing is not inevitable, but it does not establish that one's own choice to refrain actually reduces the opponent's probability of racing. If the opponent's likelihood of racing is exogenous or perversely affected by one's restraint, the paper's conclusion against racing loses its force. The authors should provide a more careful probabilistic argument or at least acknowledge the empirical uncertainty explicitly in the formal structure of the decision.
minor comments (5)
- [Abstract and general presentation] The abstract contains several typographical artifacts ('competit ors', 'He nce') that should be corrected in the published version.
- [Section 3.3] The assumption that racing makes AI development 'twice as fast' is a simplification; the argument would benefit from noting that the relationship is illustrative and depends on a speed multiplier that is itself uncertain.
- [Table 2] The text refers to 'the scenario in which both agents race (in bold)', but the table is not formatted in bold; this is a minor presentation inconsistency.
- [References] The in-text citation 'Kokotaljo et al. 2025' should be 'Kokotajlo et al. 2025' to match the reference list.
- [Section 6.2] The list of cooperation areas is said to be 'surely incomplete,' but the paper would be strengthened by mentioning at least one concrete successful precedent of verification/enforcement in another domain, even if imperfect.
Circularity Check
No circularity found: the paper's central argument rests on external evidence and qualitative reasoning; the payoff matrices are explicitly illustrative summaries, not fitted inputs presented as predictions.
full rationale
This is an argumentative policy paper, not a formal derivation, and its central claims are supported by external empirical and philosophical sources such as Nevo et al. (2024), Rivera et al. (2024), Brundage & Werner (2025), Carlsmith (2025), and Hendrycks et al. (2025). The payoff matrices in Tables 1-4 are explicitly illustrative: the Table 1 caption says 'One should not place much weight on specific numbers beyond their ordering,' and the subsequent revisions of the payoff values are presented as crude summaries of the preceding qualitative arguments, not as empirically fitted parameters or as predictions. The conclusion that cooperation is preferable is not derived from the matrix alone; the matrix is offered after the arguments for high racing costs, low racing benefits, and viable alternatives. The paper's self-citations (Dung 2023, Dung 2024a, Dung 2024b, Friederich & Dung 2025, and the forthcoming Hellrigel-Holderbaum & Dung) supply definitions and background assumptions but are not load-bearing for the central policy conclusion. The paper also candidly disclaims scope in Section 6.2, stating that it will not assess the likelihood of the requisite cooperation, which is a limitation rather than a circular move. Consequently, there is no specific equation or fitted parameter that reduces a claimed result to its inputs, and no significant circularity is present.
Assumptions & free parameters
free parameters (2)
- Payoff matrix utility values =
N=5, UR=3, R=1, UN=0 (Table 4)
- Racing speed multiplier =
2
assumptions (5)
- domain assumption AGI is technically feasible soon and can provide a decisive strategic advantage (DSA).
- domain assumption Racing to AGI substantially increases catastrophic risks from AI.
- domain assumption The relevant actors are rational, self-interested unitary agents (primarily states).
- ad hoc to paper International cooperation can be made robust through verification and enforcement.
- domain assumption Model theft by well-resourced state actors cannot currently be prevented.
Cite this review
Pith. "Pith review of Against racing to AGI: Cooperation, deterrence, and catastrophic risks." pith.science (2026). https://pith.science/paper/3AIISITC
@misc{pith2026250721839,
author = {Pith},
title = {Pith review of: Against racing to AGI: Cooperation, deterrence, and catastrophic risks},
year = {2026},
howpublished = {\url{https://pith.science/paper/3AIISITC}},
note = {Machine review of arXiv:2507.21839}
}
read the original abstract
AGI Racing is the view that it is in the self-interest of major actors in AI development, especially powerful nations, to accelerate their frontier AI development to build highly capable AI, especially artificial general intelligence (AGI), before competitors have a chance. We argue against AGI Racing. First, the downsides of racing to AGI are much higher than portrayed by this view. Racing to AGI would substantially increase catastrophic risks from AI, including nuclear instability, and undermine the prospects of technical AI safety research to be effective. Second, the expected benefits of racing may be lower than proponents of AGI Racing hold. In particular, it is questionable whether winning the race enables complete domination over losers. Third, international cooperation and coordination, and perhaps carefully crafted deterrence measures, constitute viable alternatives to racing to AGI which have much smaller risks and promise to deliver most of the benefits that racing to AGI is supposed to provide. Hence, racing to AGI is not in anyone's self-interest as other actions, particularly incentivizing and seeking international cooperation around AI issues, are preferable.
Reference graph
Works this paper leans on
-
[1]
We already talked about takeover through misaligned AGI. To rehearse, the idea is roughly that (by default) we may create AGI systems with goals that are at odds with human values. If so, such systems may in service of their goals learn to pursue dangerous subgoals like seeking power or accumulating resources since those are generally useful for most goal...
-
[2]
Catastrophic aligned AGI misuse . If AGI provides a DSA, then a group or an individual which obtains control over AGI could use it to achieve world domination and use this for malevolent ends (Friederich 2023; Katzke and Futerman 2024; Yum 2024)
work page 2023
-
[3]
The downsides of an AGI race 3.1 The many risks argument For concreteness, let us call a risk “catastrophic” if it is the chance of an event, or a set of related events, which causes the death of at least 100 million human lives, or something which has comparable negative moral significance (Dung 2024b). We hold that there are multiple credible independen...
work page 2023
-
[4]
Preventive war. If AGI is believed to provide a DSA, then nations have an incentive to wage war to prevent their adversaries from obtaining AGI (Hendrycks et al. 2025; Katzke and Futerman 2024). This is particularly plausible if their adversaries will not, or cannot, cr edibly pre-commit to not using AGI to achieve domination over them
work page 2025
-
[5]
Accumulative catastrophic risk . Rapid advances in AI may cause successive disruptions which lower the stability and response capacity of different societal systems until a cascading breakdown of these systems leads to irrecoverable collapse (Kasirzadeh 2025; Bales 2025)
work page 2025
-
[6]
Gradual aligned AGI takeover. Even if individual AI systems behave in line with individual human intentions, competitive incentives and coordination failures may drive humans to gradually increase their reliance on AI and thereby to transfer power to AI systems, until societal systems are decoupled from human feedback and control (Kulveit et al. 2025). 9 ...
work page 2024
-
[7]
The (putative) benefits of racing: Will there be a decisive strategic advantage? In this section, we analyze an assumption we have granted thus far. There are some reasons to think the chance of realizing the upside that motivates proponents of AGI Racing – achieving a sufficiently large lead ahead of competitors to institute a decisiv e strategic advanta...
work page 2023
-
[8]
The comparative danger of racing We have argued in section 3, using three separate arguments, that the expected downsides of racing to AGI are dire, and much worse than commonly appreciated by proponents of AGI Racing. In section 4, we further argued that the expected upsides of racing to AGI are smaller than commonly portrayed. The upshot of this is that...
work page 2025
Show all 11 references
-
[9]
Verification
Alternatives to racing In the preceding sections, we argued against racing to AGI. What should be done instead? Table 3 (and table 4 below) reflect that the scenario in which neither agent races is overall preferable and that not racing oneself is a viable strategic choice. Ho...
2025
-
[10]
First, we have shown in section 3 that the downsides of racing to AGI are much higher than proponents of AGI Racing hold
Conclusion Our argument against AGI Racing has proceeded as follows. First, we have shown in section 3 that the downsides of racing to AGI are much higher than proponents of AGI Racing hold. Racing would substantially increase catastrophic risks from AI, including nu clear ins...
-
[104]
https://doi.org/10.1007/s13347-024-00794-0
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.