REVIEW 3 major objections 5 minor 1 cited by
Inference Scaling Reshapes AI Governance
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read AI governance would need an overhaul if compute scaling shifts from training to inference.
desk verdict A clear, honest scenario analysis of governance under inference scaling, built on an empirical premise the paper itself flags as uncertain. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a distinction between two destinations for scaled inference compute. Inference-at-deployment spends extra compute per user query, and the paper's rule of thumb is that each order of magnitude of inference compute adds roughly 0.7 orders of magnitude of effective pre-training equivalence. Inference-during-training spends the compute inside the lab, using an inference-amplified model to generate synthetic training data or to guide a search whose outputs are distilled into a new model. The second mechanism, iterated distillation and amplification, is the paper's most consequential piece of machinery: start from a system-1 model, amplify it with inference-time search, distill the amplified behaviour into a new model, and repeat — the loop that carried AlphaGo Zero past world-champion play in forty days, and which the paper argues is a plausible pathway for general LLMs. Against these, the paper sets the existing governance machinery: compute thresholds such as the EU AI Act's $10^{25}$ FLOP and the US executive order's $10^{26}$ FLOP, which draw a bright regulatory line by training compute alone, and which inference scaling threatens to make unworkable.
What would settle it
Track the compute and capability of successive frontier pre-training runs over the next few years. If leading labs show that orders of magnitude more pre-training compute still yield the capability gains they did from 2020 to 2024, the premise that the pre-training era is over collapses, and with it the urgency of the paper's governance analysis. In the other direction, if inference-scaling techniques reliably plateau well below a 100,000x multiplier for general tasks, the claim that thresholds such as the EU AI Act's $10^{25}$ FLOP can be breached by inference amplification weakens.
Extended reading notes
Core claim
The central claim is that the marginal compute driving frontier AI is shifting from pre-training to inference, and that this is not an implementation detail but a change of regime. If the scaled compute is spent at deployment, then the number of simultaneous copies of a new frontier model falls by roughly a factor of 100 per two orders of magnitude of inference scaling; the first human-level systems may cost more than equivalent human labour; model weights become less worth stealing; open-weight models become less attractive and less dangerous; the software-like business model of frontier AI erodes; monolithic data centres lose strategic centrality; and governance via training-compute thresholds (the EU AI Act's $10^{25}$ FLOP, the US executive order's $10^{26}$ FLOP) breaks, because a model trained below the threshold can be amplified by inference to perform at the level of a far larger training run. If the scaled compute is spent during training, the effects are more ambiguous: it could feed high-quality synthetic data back into pre-training, or — in the more consequential variant — it could power iterated distillation and amplification in the manner of AlphaGo Zero, a ladder in which each rung distills an inference-amplified model into a stronger base model. The paper treats this loop as a plausible form of recursive self-improvement for general LLMs that could shorten timelines to transformative AI while remaining invisible to outside observers.
Load-bearing premise
The load-bearing premise is that pre-training scaling has plateaued near GPT-4 level, so that future frontier progress must come substantially from inference compute; if pre-training resumes rapid scaling, the described governance disruptions lose their force.
Editorial extensions
If this is right
- Inference-at-deployment could let a model trained below a $10^{25}$ FLOP threshold perform at the level of a $10^{27}$ FLOP training run, breaking threshold-based regulation and pushing governance toward use-based rules.
- Inference-at-deployment would make open-weight releases less attractive to users, since the heavy inference costs fall on the deployer rather than the trainer.
- Inference-at-deployment could make the first human-level systems more expensive to run than equivalent human labour, creating a window to study or demonstrate them before transformative deployment.
- Inference-during-training, if it works through iterated distillation and amplification, could shorten timelines to transformative AGI while keeping the best models hidden from outside observers.
- Both scenarios argue for policies that require disclosure of current capabilities and immediate plans, and for monitoring which kind of inference scaling is actually happening.
Reading between the lines
- My inference: the threshold problem the paper identifies may partially self-correct, because inference costs pass to users; if only well-resourced actors can afford a dangerous level of amplification, the governance target shifts from model weights to who can pay for compute.
- My inference: the two scenarios are asymmetric — the deployment scenario is the near-term default, while the distillation–amplification scenario is the higher-consequence tail; a governance regime that tracks the ratio of training to deployment compute at frontier labs would be the natural early-warning metric.
- My inference: because inference scaling is tunable per task, a dangerous capability could be concentrated on a few high-value targets at enormous compute multipliers even when average deployment compute stays low, so safety cases should weigh worst-case concentration of inference compute rather than averages.
- My inference: the argument's empirical engine is data scarcity, so the thesis is testable by watching whether frontier labs resume rapid pre-training scaling after absorbing synthetic-data techniques; a resumption would weaken the urgency of the governance overhaul without refuting the mechanics of the deployment scenario.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that the apparent plateau in pre-training compute scaling and the rise of inference-time compute may transform AI governance. It distinguishes two scenarios: inference-at-deployment, where extra compute is spent on reasoning during use, and inference-during-training, where extra compute is spent inside the training process (for synthetic data or iterated distillation and amplification). The paper derives a list of governance consequences for each scenario, including reduced value of closed-model weights, reduced strategic importance of open-weight models, higher cost of first human-level systems, breakage of compute-threshold regulation, changes to the frontier-lab business model, and less transparency into state-of-the-art capabilities in the training-internal scenario. The essay is explicitly exploratory, with many hedged claims and caveats.
Significance. The paper is a timely and clearly written mapping of an important potential shift in AI development. Its main value is in laying out a structured set of governance implications that differ sharply between deployment-side and training-side inference scaling, and in flagging the threshold problem in compute-based regulation. The treatment of iterated distillation and amplification, while speculative, is a useful contribution because it connects an existing AI-alignment idea to current frontier-lab practices. The essay's strengths are its conceptual clarity, honest acknowledgment of uncertainty, and the explicit scenario structure. Its main weakness is that the central empirical premise that pre-training scaling has plateaued near GPT-4 level rests on thin evidence, so the governance conclusions are conditional on an assumption that could fail.
major comments (3)
- [Section 2 (The end of an era) and Section 3] The claim that pre-training scaling has plateaued near GPT-4 level is supported only by anonymous Reuters reporting (Hu & Tong 2024) and a single OpenAI chart, and the paper itself later concedes in Section 3 that it is 'not yet clear' whether the pre-training growth rate has fallen to zero or to some fraction of its previous rate. This premise is load-bearing for every governance consequence in Sections 4–6: threshold breakage, the diminished value of model weights, the reduced strategic importance of open-weight models, the business-model shift, and the inference-during-training scenarios. If pre-training continues at something close to the historical 5x/year (Epoch AI), most of the described effects would be substantially delayed or would not occur. The paper would be much stronger if it either presented this as one explicit scenario among two and analyzed the continuation case in parallel, or provided additional systematic evidence (e.g., recent training-run compute records, algorithmic-efficiency trends) for the plateau. As written, the central argument rests on an under-evidenced empirical premise that is acknowledged as uncertain but is then treated as the basis for the rest of the essay.
- [Appendix (cost comparison)] The statement that 'deployment compute exceeding total training compute on commercial frontier systems' (footnote ‡‡‡) is asserted without supporting data or a citation, and the footnote text appears to be missing from the manuscript. This claim is load-bearing for the appendix's conclusion that when deployment compute dominates, scaling inference by 10x increases total costs by nearly 10x while scaling pre-training by 10x increases costs by only about 3x. That cost comparison in turn motivates the essay's claim that inference scaling changes the industry's business model. Please either add a verifiable citation or clearly label this as an assumption rather than an established fact, and restore the missing footnote.
- [Section 4 (Breaking the strategy of AI governance via compute thresholds)] The threshold-breakage example uses a linear effective-OOM conversion (0.7 × OOMs of inference) to claim that a 10^24 FLOP model with 4 OOM of inference scaling would perform at the level of a 10^27 FLOP model. This extrapolates to a 10,000x inference multiplier even though the paper later acknowledges that current inference-scaling techniques hit performance plateaus that cannot be exceeded by any level of compute. The argument would be more convincing if the threshold-breakage conclusion were framed as a growing risk that depends on breakthrough research in inference scaling, with the plateau limitation discussed in the same section rather than in the later paragraph. As it stands, the example overstates the certainty of the threshold problem.
minor comments (5)
- [Global/typography] Several footnote markers (‡, ‡‡‡) appear in the text but the corresponding footnote text is missing or not rendered in the manuscript; this should be fixed before publication.
- [Section 6] The word 'rapdily' should be 'rapidly'.
- [Section 6] The phrase 'overton window' should be capitalized as 'Overton window'.
- [Section 4] The 'Effective orders of magnitude' equation would benefit from a citation to the source of the 0.7 coefficient and a note that the original source gives a range (0.5 to 1.0) rather than a fixed value.
- [References] The reference to 'Epoch AI (2024)' for the 5x/year pre-training growth rate is generic; please specify the exact page or dataset.
Circularity Check
No circular derivation; the essay's conclusions are scenario-conditional and rest on external, explicitly uncertain premises.
full rationale
This is an analytic essay, not a fitted or self-referential derivation. Each governance consequence is conditional on a scenario ('if inference-at-deployment is scaled by two orders of magnitude...'), with the uncertainty about the pre-training growth rate explicitly flagged: the paper says it is 'not yet clear if it has gone to zero... or to some fraction of its previous rate.' The effective-OOMs heuristic in Section 4 is introduced as a 'rule of thumb' from prior external work (Jones 2021, Villalobos & Atkinson 2023), and the compute-threshold example (10^24 FLOP plus 4 OOM of inference reaching 10^27-level performance) is illustrative arithmetic rather than a parameter fitted to make any conclusion come out. The AlphaGo Zero discussion cites Silver et al. (2017) as independent, established external evidence, and the paper explicitly states that it is 'not at all clear whether this would work' for LLMs. There are no load-bearing self-citations: the author's own prior work is not invoked to justify any premise. The skeptical concern about the empirical support for the pre-training plateau is a correctness or evidence concern, not a circularity concern under the specified rubric.
Assumptions & free parameters
free parameters (3)
- Inference-to-pretraining OOM conversion factor =
0.7
- Scenario scale-up factor for copies and price =
100 (2 OOM)
- Scenario scale-up factor for compute-threshold example =
10,000 (4 OOM)
assumptions (5)
- domain assumption Pre-training scaling has plateaued near GPT-4 level or will soon decelerate enough that inference compute becomes the main scaling axis.
- domain assumption Compute cost is approximately linear in parameters, training tokens, deployment calls, and inference steps: C ≈ ND + Cpost + dNI.
- domain assumption Deployment compute is currently the dominant lifetime compute cost for commercial frontier systems.
- ad hoc to paper Iterated distillation and amplification, demonstrated in AlphaGo Zero, transfers to general LLM training.
- domain assumption Inference scaling benefits verifiable and long-horizon tasks more than intuitive tasks.
Cite this review
Pith. "Pith review of Inference Scaling Reshapes AI Governance." pith.science (2026). https://pith.science/paper/NRV7OMGK
@misc{pith2026250305705,
author = {Pith},
title = {Pith review of: Inference Scaling Reshapes AI Governance},
year = {2026},
howpublished = {\url{https://pith.science/paper/NRV7OMGK}},
note = {Machine review of arXiv:2503.05705}
}
read the original abstract
The shift from scaling up the pre-training compute of AI systems to scaling up their inference compute may have profound effects on AI governance. The nature of these effects depends crucially on whether this new inference compute will primarily be used during external deployment or as part of a more complex training programme within the lab. Rapid scaling of inference-at-deployment would: lower the importance of open-weight models (and of securing the weights of closed models), reduce the impact of the first human-level models, change the business model for frontier AI, reduce the need for power-intense data centres, and derail the current paradigm of AI governance via training compute thresholds. Rapid scaling of inference-during-training would have more ambiguous effects that range from a revitalisation of pre-training scaling to a form of recursive self-improvement via iterated distillation and amplification.
Figures
Forward citations
Cited by 1 Pith paper
-
How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements
Verification of international AI agreements will fail first at detecting hidden compute facilities, around the 10,000-H100-equivalent scale, before other enforcement mechanisms break.
Reference graph
Works this paper leans on
-
[1]
1 Inference Scaling Reshapes AI Governance Toby Ord* Oxford Martin AI Governance Initiative University of Oxford The shift from scaling up the pre-training compute of AI systems to scaling up their inference compute may have profound effects on AI governance. The nature of these effects depends crucially on whether this new inference compute will primaril...
work page 2024
-
[2]
How o1’s performance on AIME scales with post-training compute and with inference compute, from OpenAI (2024). This has led to intense speculation that the previous era of scaling pre-training compute could be followed by an era of scaling up inference-compute. In this essay, I explore the implications of this possibility for AI governance. In some ways a...
work page 2024
-
[5]
tasks that typically take humans a long time (as this shows these tasks can benefit from a lot of thinking before diminishing marginal returns kick in). Because some tasks benefit more from additional inference than others, it is possible to tailor the amount of inference compute to the task, spending 1,000x the normal amount for a hard, deep maths proble...
work page 2024
-
[10]
Training Compute Thresholds: Features and Functions in AI Regulation. arXiv:2405.10799v2 [cs.CY]. J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al
-
[14]
Mastering the game of Go without human knowledge. Nature 550, 354–359. https://doi.org/10.1038/nature24270 Pablo Villalobos and David Atkinson
-
[100]
Then the value of stealing model weights hasn’t increased over time — it is just the value of not having to train a GPT-4 level model (which has been decreasing over time by about 4x per year due to algorithmic efficiency improvements and Moore’s law (Epoch AI 2024)). And even if the weights were stolen, the thief would still have to pay the high inferenc...
work page 2024
-
[1010]
Apparently, this is usually the case, with deployment compute exceeding total training compute on commercial frontier systems.‡‡‡ The most standard way of training LLMs to minimise training compute involves scaling up N and D by the same factor (Hoffmann et al., 2022). For example, if you scale up training compute by 1 OOM, that means 0.5 OOMs more parame...
work page 2022
-
[2017]
Thinking Fast and Slow with Deep Learning and Tree Search, arXiv:1705.08439 [cs.AI]. Paul Christiano
Show all 14 references
-
[2020]
17 OpenAI, 12 Sep
Scaling Laws for Neural Language Models, arXiv:2001.08361 [cs.LG]. 17 OpenAI, 12 Sep
2001 arXiv
-
[2021]
arXiv:2104.03113v2 [cs.LG]
Scaling Scaling Laws with Board Games. arXiv:2104.03113v2 [cs.LG]. Kadrey v. Meta Platforms, Inc
-
[2022]
Krystal Hu and Anna Tong
Training compute-optimal large language models, arXiv:2203.15556 [cs.CL]. Krystal Hu and Anna Tong
-
[2023]
have revealed that Meta’s Llama3 team decided to train on an illegal Russian repository of copyrighted books, LibGen, because they were unable to reach GPT-4 level without it.: 9 ‘Libgen is essential to meet SOTA [state-of-the-art] numbers, across all categories, and it is kno...
2017
-
[2024]
A second — and ultimately more important — question concerns the nature of inference-scaling
It seems like that rate has now fallen, but it is not yet clear if it has gone to zero (with AI progress coming from things other than pre-training compute) or to some fraction of its previous rate. A second — and ultimately more important — question concerns the nature of inf...
2024
-
[2025]
Lennart Heim and Leonie Koessler
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, arXiv:2501.12948 [cs.CL]. Lennart Heim and Leonie Koessler
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.