Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Inference Scaling Reshapes AI Governance

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read AI governance would need an overhaul if compute scaling shifts from training to inference.

desk verdict A clear, honest scenario analysis of governance under inference scaling, built on an empirical premise the paper itself flags as uncertain. read the letter →

arxiv 2503.05705 v1 pith:NRV7OMGK submitted 2025-02-12 cs.CY cs.AI

classification cs.CYcs.AI
keywords inferencescalingAIgovernancepre-trainingcomputetrainingthresholdsiterateddistillationandamplificationopen-weightmodelsrecursiveself-improvementfrontier
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the era of scaling pre-training compute — the basis of most current AI governance — may be ending, and that a shift to scaling inference compute would upend standard governance assumptions. Whether the new compute is spent at deployment (extra reasoning per user query, as in o1 and R1) or during training (using reasoning models to build better models) leads to sharply different futures. Inference-at-deployment would lower the importance of open-weight models, blunt the impact of the first human-level systems, change the frontier business model, and break regulation keyed to training-compute thresholds. Inference-during-training could either revive pre-training by supplying high-quality synthetic data or, through iterated distillation and amplification, produce rapid recursive self-improvement with far less outside visibility. In either case, the author concludes, the future becomes less predictable, and governance should track which form of inference scaling is actually unfolding.

What carries the argument

The argument is carried by a distinction between two destinations for scaled inference compute. Inference-at-deployment spends extra compute per user query, and the paper's rule of thumb is that each order of magnitude of inference compute adds roughly 0.7 orders of magnitude of effective pre-training equivalence. Inference-during-training spends the compute inside the lab, using an inference-amplified model to generate synthetic training data or to guide a search whose outputs are distilled into a new model. The second mechanism, iterated distillation and amplification, is the paper's most consequential piece of machinery: start from a system-1 model, amplify it with inference-time search, distill the amplified behaviour into a new model, and repeat — the loop that carried AlphaGo Zero past world-champion play in forty days, and which the paper argues is a plausible pathway for general LLMs. Against these, the paper sets the existing governance machinery: compute thresholds such as the EU AI Act's $10^{25}$ FLOP and the US executive order's $10^{26}$ FLOP, which draw a bright regulatory line by training compute alone, and which inference scaling threatens to make unworkable.

What would settle it

Track the compute and capability of successive frontier pre-training runs over the next few years. If leading labs show that orders of magnitude more pre-training compute still yield the capability gains they did from 2020 to 2024, the premise that the pre-training era is over collapses, and with it the urgency of the paper's governance analysis. In the other direction, if inference-scaling techniques reliably plateau well below a 100,000x multiplier for general tasks, the claim that thresholds such as the EU AI Act's $10^{25}$ FLOP can be breached by inference amplification weakens.

Watch

Extended reading notes

Core claim

The central claim is that the marginal compute driving frontier AI is shifting from pre-training to inference, and that this is not an implementation detail but a change of regime. If the scaled compute is spent at deployment, then the number of simultaneous copies of a new frontier model falls by roughly a factor of 100 per two orders of magnitude of inference scaling; the first human-level systems may cost more than equivalent human labour; model weights become less worth stealing; open-weight models become less attractive and less dangerous; the software-like business model of frontier AI erodes; monolithic data centres lose strategic centrality; and governance via training-compute thresholds (the EU AI Act's $10^{25}$ FLOP, the US executive order's $10^{26}$ FLOP) breaks, because a model trained below the threshold can be amplified by inference to perform at the level of a far larger training run. If the scaled compute is spent during training, the effects are more ambiguous: it could feed high-quality synthetic data back into pre-training, or — in the more consequential variant — it could power iterated distillation and amplification in the manner of AlphaGo Zero, a ladder in which each rung distills an inference-amplified model into a stronger base model. The paper treats this loop as a plausible form of recursive self-improvement for general LLMs that could shorten timelines to transformative AI while remaining invisible to outside observers.

Load-bearing premise

The load-bearing premise is that pre-training scaling has plateaued near GPT-4 level, so that future frontier progress must come substantially from inference compute; if pre-training resumes rapid scaling, the described governance disruptions lose their force.

Editorial extensions

If this is right

  • Inference-at-deployment could let a model trained below a $10^{25}$ FLOP threshold perform at the level of a $10^{27}$ FLOP training run, breaking threshold-based regulation and pushing governance toward use-based rules.
  • Inference-at-deployment would make open-weight releases less attractive to users, since the heavy inference costs fall on the deployer rather than the trainer.
  • Inference-at-deployment could make the first human-level systems more expensive to run than equivalent human labour, creating a window to study or demonstrate them before transformative deployment.
  • Inference-during-training, if it works through iterated distillation and amplification, could shorten timelines to transformative AGI while keeping the best models hidden from outside observers.
  • Both scenarios argue for policies that require disclosure of current capabilities and immediate plans, and for monitoring which kind of inference scaling is actually happening.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: the threshold problem the paper identifies may partially self-correct, because inference costs pass to users; if only well-resourced actors can afford a dangerous level of amplification, the governance target shifts from model weights to who can pay for compute.
  • My inference: the two scenarios are asymmetric — the deployment scenario is the near-term default, while the distillation–amplification scenario is the higher-consequence tail; a governance regime that tracks the ratio of training to deployment compute at frontier labs would be the natural early-warning metric.
  • My inference: because inference scaling is tunable per task, a dangerous capability could be concentrated on a few high-value targets at enormous compute multipliers even when average deployment compute stays low, so safety cases should weigh worst-case concentration of inference compute rather than averages.
  • My inference: the argument's empirical engine is data scarcity, so the thesis is testable by watching whether frontier labs resume rapid pre-training scaling after absorbing synthetic-data techniques; a resumption would weaken the urgency of the governance overhaul without refuting the mechanics of the deployment scenario.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper argues that the apparent plateau in pre-training compute scaling and the rise of inference-time compute may transform AI governance. It distinguishes two scenarios: inference-at-deployment, where extra compute is spent on reasoning during use, and inference-during-training, where extra compute is spent inside the training process (for synthetic data or iterated distillation and amplification). The paper derives a list of governance consequences for each scenario, including reduced value of closed-model weights, reduced strategic importance of open-weight models, higher cost of first human-level systems, breakage of compute-threshold regulation, changes to the frontier-lab business model, and less transparency into state-of-the-art capabilities in the training-internal scenario. The essay is explicitly exploratory, with many hedged claims and caveats.

Significance. The paper is a timely and clearly written mapping of an important potential shift in AI development. Its main value is in laying out a structured set of governance implications that differ sharply between deployment-side and training-side inference scaling, and in flagging the threshold problem in compute-based regulation. The treatment of iterated distillation and amplification, while speculative, is a useful contribution because it connects an existing AI-alignment idea to current frontier-lab practices. The essay's strengths are its conceptual clarity, honest acknowledgment of uncertainty, and the explicit scenario structure. Its main weakness is that the central empirical premise that pre-training scaling has plateaued near GPT-4 level rests on thin evidence, so the governance conclusions are conditional on an assumption that could fail.

major comments (3)
  1. [Section 2 (The end of an era) and Section 3] The claim that pre-training scaling has plateaued near GPT-4 level is supported only by anonymous Reuters reporting (Hu & Tong 2024) and a single OpenAI chart, and the paper itself later concedes in Section 3 that it is 'not yet clear' whether the pre-training growth rate has fallen to zero or to some fraction of its previous rate. This premise is load-bearing for every governance consequence in Sections 4–6: threshold breakage, the diminished value of model weights, the reduced strategic importance of open-weight models, the business-model shift, and the inference-during-training scenarios. If pre-training continues at something close to the historical 5x/year (Epoch AI), most of the described effects would be substantially delayed or would not occur. The paper would be much stronger if it either presented this as one explicit scenario among two and analyzed the continuation case in parallel, or provided additional systematic evidence (e.g., recent training-run compute records, algorithmic-efficiency trends) for the plateau. As written, the central argument rests on an under-evidenced empirical premise that is acknowledged as uncertain but is then treated as the basis for the rest of the essay.
  2. [Appendix (cost comparison)] The statement that 'deployment compute exceeding total training compute on commercial frontier systems' (footnote ‡‡‡) is asserted without supporting data or a citation, and the footnote text appears to be missing from the manuscript. This claim is load-bearing for the appendix's conclusion that when deployment compute dominates, scaling inference by 10x increases total costs by nearly 10x while scaling pre-training by 10x increases costs by only about 3x. That cost comparison in turn motivates the essay's claim that inference scaling changes the industry's business model. Please either add a verifiable citation or clearly label this as an assumption rather than an established fact, and restore the missing footnote.
  3. [Section 4 (Breaking the strategy of AI governance via compute thresholds)] The threshold-breakage example uses a linear effective-OOM conversion (0.7 × OOMs of inference) to claim that a 10^24 FLOP model with 4 OOM of inference scaling would perform at the level of a 10^27 FLOP model. This extrapolates to a 10,000x inference multiplier even though the paper later acknowledges that current inference-scaling techniques hit performance plateaus that cannot be exceeded by any level of compute. The argument would be more convincing if the threshold-breakage conclusion were framed as a growing risk that depends on breakthrough research in inference scaling, with the plateau limitation discussed in the same section rather than in the later paragraph. As it stands, the example overstates the certainty of the threshold problem.
minor comments (5)
  1. [Global/typography] Several footnote markers (‡, ‡‡‡) appear in the text but the corresponding footnote text is missing or not rendered in the manuscript; this should be fixed before publication.
  2. [Section 6] The word 'rapdily' should be 'rapidly'.
  3. [Section 6] The phrase 'overton window' should be capitalized as 'Overton window'.
  4. [Section 4] The 'Effective orders of magnitude' equation would benefit from a citation to the source of the 0.7 coefficient and a note that the original source gives a range (0.5 to 1.0) rather than a fixed value.
  5. [References] The reference to 'Epoch AI (2024)' for the 5x/year pre-training growth rate is generic; please specify the exact page or dataset.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation; the essay's conclusions are scenario-conditional and rest on external, explicitly uncertain premises.

full rationale

This is an analytic essay, not a fitted or self-referential derivation. Each governance consequence is conditional on a scenario ('if inference-at-deployment is scaled by two orders of magnitude...'), with the uncertainty about the pre-training growth rate explicitly flagged: the paper says it is 'not yet clear if it has gone to zero... or to some fraction of its previous rate.' The effective-OOMs heuristic in Section 4 is introduced as a 'rule of thumb' from prior external work (Jones 2021, Villalobos & Atkinson 2023), and the compute-threshold example (10^24 FLOP plus 4 OOM of inference reaching 10^27-level performance) is illustrative arithmetic rather than a parameter fitted to make any conclusion come out. The AlphaGo Zero discussion cites Silver et al. (2017) as independent, established external evidence, and the paper explicitly states that it is 'not at all clear whether this would work' for LLMs. There are no load-bearing self-citations: the author's own prior work is not invoked to justify any premise. The skeptical concern about the empirical support for the pre-training plateau is a correctness or evidence concern, not a circularity concern under the specified rubric.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central argument rests on uncertain empirical premises: the end of pre-training scaling, the transferability of AlphaGo Zero style amplification to LLMs, and the cost structure of deployed frontier systems. These are scenario assumptions rather than fitted parameters. The paper is transparent about most of them, but they are load-bearing for its conclusions.

free parameters (3)
  • Inference-to-pretraining OOM conversion factor = 0.7
    Used in the 'Effective orders of magnitude' rule to estimate capability equivalence; imported from Jones (2021) and Villalobos & Atkinson (2023), not fitted or validated in this paper, and applied to multi-OOM extrapolations.
  • Scenario scale-up factor for copies and price = 100 (2 OOM)
    Illustrative scenario assumption used to argue that the number of simultaneous deployed copies falls by 100x and that the first human-level systems could cost more than human labour.
  • Scenario scale-up factor for compute-threshold example = 10,000 (4 OOM)
    Used to argue that a 10^24 FLOP model with 4 OOM of inference could perform like a 10^27 FLOP model; based on o3's reported 10,000x inference multiplier, extrapolated to a governance claim.
assumptions (5)
  • domain assumption Pre-training scaling has plateaued near GPT-4 level or will soon decelerate enough that inference compute becomes the main scaling axis.
    Introduced in Section 'The end of an era' based on anonymous employee reports (Hu & Tong, 2024). This is the premise that makes the governance analysis relevant.
  • domain assumption Compute cost is approximately linear in parameters, training tokens, deployment calls, and inference steps: C ≈ ND + Cpost + dNI.
    Appendix cost model. This ignores batching, hardware utilisation, memory, and constant factors, but the qualitative conclusions rely on the linear trade-off structure.
  • domain assumption Deployment compute is currently the dominant lifetime compute cost for commercial frontier systems.
    Stated in the appendix as 'apparently, this is usually the case' with no citation; used to argue that scaling inference has immediate cost consequences.
  • ad hoc to paper Iterated distillation and amplification, demonstrated in AlphaGo Zero, transfers to general LLM training.
    The inference-during-training scenario depends on this analogy; the paper itself acknowledges it is 'not at all clear whether this would work'.
  • domain assumption Inference scaling benefits verifiable and long-horizon tasks more than intuitive tasks.
    Used to conclude unequal performance across tasks and users; a plausible heuristic supported by the o1/o3 examples but not systematically evidenced.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Inference Scaling Reshapes AI Governance." pith.science (2026). https://pith.science/paper/NRV7OMGK

@misc{pith2026250305705,
  author       = {Pith},
  title        = {Pith review of: Inference Scaling Reshapes AI Governance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NRV7OMGK}},
  note         = {Machine review of arXiv:2503.05705}
}
read the original abstract

The shift from scaling up the pre-training compute of AI systems to scaling up their inference compute may have profound effects on AI governance. The nature of these effects depends crucially on whether this new inference compute will primarily be used during external deployment or as part of a more complex training programme within the lab. Rapid scaling of inference-at-deployment would: lower the importance of open-weight models (and of securing the weights of closed models), reduce the impact of the first human-level models, change the business model for frontier AI, reduce the need for power-intense data centres, and derail the current paradigm of AI governance via training compute thresholds. Rapid scaling of inference-during-training would have more ambiguous effects that range from a revitalisation of pre-training scaling to a form of recursive self-improvement via iterated distillation and amplification.

Figures

Figures reproduced from arXiv: 2503.05705 by the authors.

Figure 1
Figure 1. How o1’s performance on AIME scales with post-training compute and with inference compute, from OpenAI (2024). This has led to intense speculation that the previous era of scaling pre-training compute could be followed by an era of scaling up inference-compute. In this essay, I explore the implications of this possibility for AI governance. In some ways a move to scaling of inference compute may be a continuation of… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements

    cs.CY 2026-06 conditional novelty 6.0 of 10

    Verification of international AI agreements will fail first at detecting hidden compute facilities, around the 10,000-H100-equivalent scale, before other enforcement mechanisms break.

Reference graph

Works this paper leans on

14 extracted references · 7 canonical work pages · cited by 1 Pith paper

  1. [1]

    1 Inference Scaling Reshapes AI Governance Toby Ord* Oxford Martin AI Governance Initiative University of Oxford The shift from scaling up the pre-training compute of AI systems to scaling up their inference compute may have profound effects on AI governance. The nature of these effects depends crucially on whether this new inference compute will primaril...

  2. [2]

    This has led to intense speculation that the previous era of scaling pre-training compute could be followed by an era of scaling up inference-compute

    How o1’s performance on AIME scales with post-training compute and with inference compute, from OpenAI (2024). This has led to intense speculation that the previous era of scaling pre-training compute could be followed by an era of scaling up inference-compute. In this essay, I explore the implications of this possibility for AI governance. In some ways a...

  3. [5]

    tasks that typically take humans a long time (as this shows these tasks can benefit from a lot of thinking before diminishing marginal returns kick in). Because some tasks benefit more from additional inference than others, it is possible to tailor the amount of inference compute to the task, spending 1,000x the normal amount for a hard, deep maths proble...

  4. [10]

    arXiv:2405.10799v2 [cs.CY]

    Training Compute Thresholds: Features and Functions in AI Regulation. arXiv:2405.10799v2 [cs.CY]. J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al

  5. [14]

    Nature 550, 354–359

    Mastering the game of Go without human knowledge. Nature 550, 354–359. https://doi.org/10.1038/nature24270 Pablo Villalobos and David Atkinson

  6. [100]

    And even if the weights were stolen, the thief would still have to pay the high inference-at-deployment costs

    Then the value of stealing model weights hasn’t increased over time — it is just the value of not having to train a GPT-4 level model (which has been decreasing over time by about 4x per year due to algorithmic efficiency improvements and Moore’s law (Epoch AI 2024)). And even if the weights were stolen, the thief would still have to pay the high inferenc...

  7. [1010]

    For example, if you scale up training compute by 1 OOM, that means 0.5 OOMs more parameters and 0.5 OOMs more data

    Apparently, this is usually the case, with deployment compute exceeding total training compute on commercial frontier systems.‡‡‡ The most standard way of training LLMs to minimise training compute involves scaling up N and D by the same factor (Hoffmann et al., 2022). For example, if you scale up training compute by 1 OOM, that means 0.5 OOMs more parame...

  8. [2017]

    Paul Christiano

    Thinking Fast and Slow with Deep Learning and Tree Search, arXiv:1705.08439 [cs.AI]. Paul Christiano

Show all 14 references
  1. [2020]

    17 OpenAI, 12 Sep

    Scaling Laws for Neural Language Models, arXiv:2001.08361 [cs.LG]. 17 OpenAI, 12 Sep

  2. [2021]

    arXiv:2104.03113v2 [cs.LG]

    Scaling Scaling Laws with Board Games. arXiv:2104.03113v2 [cs.LG]. Kadrey v. Meta Platforms, Inc

  3. [2022]

    Krystal Hu and Anna Tong

    Training compute-optimal large language models, arXiv:2203.15556 [cs.CL]. Krystal Hu and Anna Tong

  4. [2023]

    have revealed that Meta’s Llama3 team decided to train on an illegal Russian repository of copyrighted books, LibGen, because they were unable to reach GPT-4 level without it.: 9 ‘Libgen is essential to meet SOTA [state-of-the-art] numbers, across all categories, and it is kno...

  5. [2024]

    A second — and ultimately more important — question concerns the nature of inference-scaling

    It seems like that rate has now fallen, but it is not yet clear if it has gone to zero (with AI progress coming from things other than pre-training compute) or to some fraction of its previous rate. A second — and ultimately more important — question concerns the nature of inf...

  6. [2025]

    Lennart Heim and Leonie Koessler

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, arXiv:2501.12948 [cs.CL]. Lennart Heim and Leonie Koessler

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.