REVIEW 3 major objections 7 minor 48 references
Assessing confidence in frontier AI safety cases
T0 review · 3 major / 7 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read To hold 95 percent confidence in a seven-part frontier AI cyber-safety case, each part must run about 99.3 percent confident.
desk verdict A useful, honest application of Assurance 2.0 to a frontier AI cyber safety case: the LLM Delphi pipeline is the real contribution, and the 99.3% propagation result is conditional on assumptions the paper mostly discloses. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a pair of arithmetic propagation rules applied to a conjunctive argument tree. In the product method, the confidence of a claim is the product of the confidences of its side-claim and all its sub-claims; in the sum-of-doubts method, doubt in a claim is bounded above by the sum of doubts in its supports, so confidence becomes the sum of component confidences minus the number of components, rounded up to zero if negative. A second piece of machinery is an LLM-based Delphi pipeline: 50 instances of a large language model act as expert forecasters, iterating over up to five rounds until their estimates converge below a standard-deviation threshold, with final probabilities weighted by each expert's consistency across rounds. This pipeline is used to estimate how likely each safety-case defeater is to be sustained, producing reproducible input probabilities for the propagation rules.
What would settle it
Recompute the required per-component confidence for the same C2.2.1 fragment after replacing the independence assumption with realistic positive correlations among the three sub-claims, or after allowing the top claim to be supported by any one of several diverse sub-arguments; if a 95% top-level confidence is then reached with per-component confidences well below about 99.3%, the challenge result does not generalise beyond its stated assumptions.
Extended reading notes
Core claim
Working from a cyber-misuse inability safety case, the paper isolates the C2.2.1 fragment in which a 7-day maximum time to detect a novel AI-enabled attack and take the system offline is decomposed into three sub-claims plus a side-claim, each supported by evidence. It assigns illustrative posterior probabilities $P(C|E)$ and side-claim strengths, then propagates them through the argument using Assurance 2.0's product and sum-of-doubts rules. Both rules give the same threshold for 95% top-level confidence: if all assigned probabilities are equal, each must be about $0.993$ ($99.3\%$). Because the fragment is only a small part of a full safety case, the required per-element confidence for the whole case would be even higher, and the paper argues this makes reliable absolute probabilistic valuation of frontier AI safety cases very difficult. The paper therefore presents probabilistic valuation as most useful for comparative 'what if' analysis rather than as an absolute go/no-go number.
Load-bearing premise
The numerical threshold assumes each claim is supported only by one fixed conjunction of independent sub-claims and a side-claim, so that confidence multiplies or doubts add; if the sub-claims overlap, share causes, or offer alternative support paths, the 99.3% figure changes.
Editorial extensions
If this is right
- A full frontier AI safety case, being much larger than the seven-component fragment, would require even higher per-element confidences to reach any fixed top-level threshold under these propagation rules.
- The sum-of-doubts rule can return zero confidence on realistic fragments, so an absolute numerical verdict from these methods alone should not be treated as a deploy/no-deploy test.
- The LLM Delphi method provides a reproducible, auditable procedure for generating leaf-node probabilities, which is useful for third-party evaluation even though the probabilities remain estimates.
- Prioritising defeaters by their probability of being sustained, their impact on top-level confidence, and the estimated effort to resolve them can shorten the path to finding decisive flaws in a safety case.
- Using probabilistic claims inside the argument itself, or structuring the case with diverse redundant sub-arguments, may lower the required per-element confidence, but the paper leaves that as future work.
Reading between the lines
- The 99.3% figure depends on an independence assumption that the paper itself flags as false for its own fragment; modelling the shared budget, team, and information among the three C2.2.1 sub-claims is the natural next computation and could move the threshold in either direction.
- A regulator who simply demands 95% top-level probability on a large conjunctive safety case may be demanding near-impossible per-leaf numbers; a more workable requirement would target the argument's structure, such as diverse independent sub-arguments or a probabilistic top-level claim.
- The LLM Delphi results on generic forecasting questions suggest a reproducible route to leaf probabilities, but the paper does not test that pipeline on actual safety-case claims about future AI behaviour; a hybrid human-LLM panel would be the direct next test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper applies the Assurance 2.0 methodology to confidence assessment for a frontier AI cyber-misuse inability safety case. It introduces a purely LLM-implemented Delphi method for eliciting leaf-node probabilities, validates this pipeline against 100 resolved Metaculus forecasting questions, and then uses Assurance 2.0's product and sum-of-doubts propagation methods on a seven-component fragment of the case. The paper's central quantitative finding is that achieving 95% confidence in the fragment's top-level claim requires approximately 99.3% confidence in each assigned component probability under both propagation methods. It also proposes a defeater-prioritisation scoring scheme and recommendations for communicating confidence to executives. The authors conclude that absolute probabilistic confidence is very difficult to achieve for frontier AI safety cases, while arguing that the process of producing such valuations can still improve safety cases.
Significance. If the central result is taken at face value, the paper provides a concrete, policy-relevant warning: even a small conjunctive safety case with only seven components demands near-perfect confidence in every leaf and side claim before a 95% top-level confidence can be reported. This is a useful counterweight to casual uses of quantitative safety-case confidence. The paper's strengths include its explicit statement of the propagation assumptions, its worked examples, the availability of code for the Delphi pipeline, and its unusually candid discussion of limitations, including the known violation of the independence assumption in the worked fragment. The two weak points are load-bearing: the 99.3% headline figure is conditional on mutual independence, and the calibration evidence for the LLM Delphi method rests on a confounded, post-hoc benchmark. The paper is a valuable applied contribution, but these two issues need to be addressed before the quantitative claims can be regarded as fully supported.
major comments (3)
- [Executive Summary and Section 7.1 (Eqs. 7, 9); Section 7.3] The headline claim that 95% confidence in the C2.2.1 fragment requires ~99.3% confidence in each of the seven components is derived under the explicit assumption that the side-claim and sub-claims are mutually independent. Section 7.3 concedes that C2.2.1.1, C2.2.1.2, and C2.2.1.3 share budget, team, and information, so they are positively correlated. With positive correlation, the probability of the conjunction is higher than the product of the marginal probabilities, so the per-component confidence needed to reach 95% top-level confidence can be substantially lower than 99.3%. Because this number is presented in the Executive Summary and Section 10.1 as the paper's main challenge result, the authors should quantify the sensitivity to plausible correlations (e.g., a simple common-factor model) or clearly reframe the 99.3% figure as an independence-conditional upper bound rather than a general property of conjunctive safety cases.
- [Section 4.2 (Figure 5E)] The claim that the LLM-based Delphi pipeline is 'better calibrated' than Metaculus is not supported by the evidence as presented. The 100 questions are a post-hoc selected set of resolved Metaculus questions; Metaculus's forecasts incorporate information up to question close, whereas the LLM is restricted to an October 2023 knowledge cutoff, creating a systematic information asymmetry. No confidence interval or statistical test is reported for the calibration difference, and the selection of questions is not pre-registered. The paper acknowledges some of these limitations, but the stated conclusion still overstates the strength of the evidence. The authors should either temper the calibration claim to 'preliminary and not directly comparable' or provide a benchmark with matched information sets and a pre-specified question set.
- [Section 7.1 (Eq. 9 and following)] The statement that 'for both methods, the confidence for assigned probabilities must be around 99.3%' conflates a sufficient condition with a necessary condition in the sum-of-doubts method. Equation (9) is a lower bound on the confidence of the top-level claim, so setting that lower bound to 0.95 gives a sufficient condition on the reported bound, not a necessary condition on the actual probability of the top-level claim. The product method, by contrast, gives an exact value under the independence assumption. The paper should clarify this logical asymmetry, since a reader may otherwise conclude that the sum-of-doubts method imposes a hard requirement that it does not actually establish.
minor comments (7)
- [Section 2.1] In the paragraph beginning 'Where safety engineering standards exist', the phrase 'the developers’ of frontier AI systems' contains a typo and should read 'the developers of frontier AI systems'.
- [Section 4.2 (Figure 5C)] The text cites p = 0.053 from a Mann-Whitney U test as a 'strong tendency', but this is not significant at the conventional 0.05 level. Please report an effect size and confidence interval, and avoid language that implies a robust finding.
- [Section 7.1 (Figure 7)] The sentence 'we simply round such results up to 0' is imprecise: rounding -0.5 up to 0 is not standard rounding. The text should say that the lower-bound result is truncated or set to 0 when it is negative.
- [Section 9 (Figures 9 and 10)] Figures 9 and 10 are referenced in the text but appear only as placeholders in the manuscript; please ensure the actual visualisations are included in the final version.
- [Section 4.2 (Figure 5D)] The choice of pseudo-counts (=10) for the credible intervals is not justified and is not varied in a sensitivity analysis. A short robustness check would strengthen the presentation of the uncertainty estimates.
- [Section 4.2 (Figure 5F caption)] The text says 'we show 3 examples of what these questions look like', but the figure caption lists five example questions. Please align the text and the figure.
- [Section 8.4 (Eq. 10)] The prioritisation formula depends on weighting factors W_Probability, W_Impact, and W_Effort, but the paper does not suggest how these weights should be set beyond noting that they vary by team. Although Section 8.6 acknowledges the limitation, a brief discussion of calibration or sensitivity analysis would make the proposed method more actionable.
Circularity Check
No significant circularity: the §7.1 99.3% figure is a transparent algebraic consequence of the cited Assurance 2.0 propagation formulas, not a fitted or self-referential input.
full rationale
After walking the paper's derivation chain, no circular step was found. The central quantitative claim in §7.1—that achieving 95% confidence in the seven-component C2.2.1 fragment requires roughly 99.3% confidence in each component when all assigned probabilities are equal—is an explicit algebraic consequence of the external Assurance 2.0 propagation formulas (Eq. 7 product method and Eq. 9 sum-of-doubts method) under the paper's stated equal-p assumption. The paper shows the arithmetic directly (0.95 = 7p − 6 and 0.95 = p^7), so the result is not a hidden fit, a renamed empirical pattern, or a conclusion smuggled into its premises. The LLM-based Delphi probabilities in Section 4 are benchmarked against external Metaculus outcomes and are not fitted to the safety-case target; they are then used only as illustrative leaf-node inputs. The paper's reliance on Bloomfield and Rushby [2] is substantive but external to the present authors, and the assumptions—mutual independence, conjunctive entailment, equal importance—are stated explicitly rather than imported through self-citation or an unexamined uniqueness theorem. The independence limitation is acknowledged in §7.3 for the very fragment used in the headline calculation, which is a robustness concern about the model, not circularity. No equation in the paper is equivalent to its input by construction in a way that would make the derivation self-supporting.
Assumptions & free parameters
free parameters (5)
- LLM temperature =
1.5
- Delphi consensus threshold =
sigma < 10%
- Pseudo-counts for credible intervals =
10
- Prioritisation weighting factors =
set by practitioner
- Illustrative leaf and side-claim probabilities =
e.g., 0.6, 0.8, 0.9, 0.7
assumptions (6)
- domain assumption The safety case is logically sound before probabilistic valuation is applied.
- domain assumption Sub-claims and side-claims are mutually independent.
- domain assumption Each claim is entailed by the conjunction of its direct sub-claims and side-claim, and that conjunction is necessary and sufficient.
- ad hoc to paper LLM expert responses can substitute for human expert judgment in safety-case probability elicitation.
- domain assumption Metaculus community probabilities are a suitable calibration benchmark for forecasting ability.
- domain assumption Natural Language Deductivism is the appropriate standard for logical soundness in safety cases.
invented entities (1)
-
Contextual Doubt
Cite this review
Pith. "Pith review of Assessing confidence in frontier AI safety cases." pith.science (2026). https://pith.science/paper/6MG7DKG4
@misc{pith2026250205791,
author = {Pith},
title = {Pith review of: Assessing confidence in frontier AI safety cases},
year = {2026},
howpublished = {\url{https://pith.science/paper/6MG7DKG4}},
note = {Machine review of arXiv:2502.05791}
}
read the original abstract
Powerful new frontier AI technologies are bringing many benefits to society but at the same time bring new risks. AI developers and regulators are therefore seeking ways to assure the safety of such systems, and one promising method under consideration is the use of safety cases. A safety case presents a structured argument in support of a top-level claim about a safety property of the system. Such top-level claims are often presented as a binary statement, for example "Deploying the AI system does not pose unacceptable risk". However, in practice, it is often not possible to make such statements unequivocally. This raises the question of what level of confidence should be associated with a top-level claim. We adopt the Assurance 2.0 safety assurance methodology, and we ground our work by specific application of this methodology to a frontier AI inability argument that addresses the harm of cyber misuse. We find that numerical quantification of confidence is challenging, though the processes associated with generating such estimates can lead to improvements in the safety case. We introduce a method for better enabling reproducibility and transparency in probabilistic assessment of confidence in argument leaf nodes through a purely LLM-implemented Delphi method. We propose a method by which AI developers can prioritise, and thereby make their investigation of argument defeaters more efficient. Proposals are also made on how best to communicate confidence information to executive decision-makers.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Marie Davidsen Buhl, Gaurav Sett, Leonie Koessler, Jonas Schuett, and Markus Anderljung. Safety cases for frontier AI. arXiv preprint arXiv:2410.21572, 2024
arXiv 2024
-
[2]
Assessing Confidence with Assurance 2.0
Robin Bloomfield and John Rushby. Assessing confidence with assurance 2.0. arXiv preprint arXiv:2205.04522, 2022
work page Pith review arXiv 2022
-
[3]
Safety case template for frontier AI: A cyber inability argument
Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett, Tomek Korbak, Jessica Wang, Benjamin Hilton, and Geoffrey Irving. Safety case template for frontier AI: A cyber inability argument. arXiv preprint arXiv:2411.08088, 2024
arXiv 2024
-
[4]
Safety cases: How to justify the safety of advanced AI systems
Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen. Safety cases: How to justify the safety of advanced AI systems. arXiv preprint arXiv:2403.10462, 2024
arXiv 2024
-
[5]
Anthropic’s responsible scaling policy. Accessed 7 December 2024. Available at: https: //www.anthropic.com/news/anthropics-responsible-scaling-policy
work page 2024
-
[6]
P. J. Graydon and C. M. Holloway. An investigation of proposed techniques for quantifying confidence in Assurance arguments. Safety Science, 92:53–65, 2017
work page 2017
-
[7]
Confidence in assurance 2.0 cases
Robin Bloomfield and John Rushby. Confidence in assurance 2.0 cases. In The Practice of Formal Methods: Essays in Honour of Cliff Jones, Part I, pages 1–23. Springer, 2024
work page 2024
-
[8]
How do practitioners gain confidence in assurance cases?
Simon Diemert, Caleb Shortt, and Jens H Weber. How do practitioners gain confidence in assurance cases? arXiv preprint arXiv:2411.03657, 2024
work page Pith review arXiv 2024
Show all 48 references
-
[9]
L. Groarke. Deductivism within pragma-dialectics. Argumentation, 13(1):1–16, 1999
1999
-
[10]
Towards guar- anteed safe ai: A framework for ensuring robust and reliable ai systems
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, et al. Towards guar- anteed safe ai: A framework for ensuring robust and reliable ai systems. arXiv preprint arXiv:2405.06624, 2024
2024 arXiv
-
[11]
Assurance 2.0: A manifesto
Robin Bloomfield and John Rushby. Assurance 2.0: A manifesto. arXiv preprint arXiv:2004.10474, 2020
2004 arXiv
-
[12]
R. E. Bloomfield and K. Netkachova. Building blocks for Assurance cases. In Proceedings of the 2nd International Workshop on Assurance Cases for Software-intensive Systems (ASSURE). International Symposium on Software Reliability Engineering (ISSRE), 2014
2014
-
[13]
J. L. Pollock. Defeasible reasoning. Cognitive Science, 11(4):481–518, 1987
1987
-
[14]
Eliminative argumentation: A basis for arguing confidence in system properties
John B Goodenough, Charles B Weinstock, and Ari Z Klein. Eliminative argumentation: A basis for arguing confidence in system properties. Software Engineering Institute, Carnegie Mellon University, Pittsburgh, PA, Tech. Rep. CMU/SEI-2015-TR-005, 2015
2015
-
[15]
Pegasus (spyware). 2024. In Wikipedia. https://en.wikipedia.org/w/index.php? title=Pegasus_(spyware)&oldid=1261562226
2024
-
[16]
Stuxnet. 2024. In Wikipedia. https://en.wikipedia.org/w/index.php?title= Stuxnet&oldid=1257264443
2024
-
[17]
Dalkey and O
N. Dalkey and O. Helmer. An experimental application of the Delphi method to the use of experts. Management Science, 9(3):458–467, 1963
1963
-
[18]
Khodyakov, S
D. Khodyakov, S. Grant, J. Kroger, and M. Bauman. RAND methodological guidance for conducting and critically appraising delphi panels. Technical report, RAND Corporation, 2023
2023
-
[19]
Lewis, E
P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. Küttler, et al. Retrieval- augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474. Curran Associates, Inc., 2020. 46
2020
-
[20]
Y . Qin, S. Liang, Y . Ye, K. Zhu, L. Yan, Y . Lu, Y . Lin, et al. ToolLLM: Facilitating Large Language Models to master 16000+ real-world APIs, 2023
2023
-
[21]
H. L. Harney. Bayesian Inference. Springer, Berlin, Heidelberg, 2003
2003
-
[22]
Z. Kunda. The case for motivated reasoning. Psychological Bulletin, 108(3):480–498, 1990
1990
-
[23]
Confirmation bias: A ubiquitous phenomenon in many guises
Raymond S Nickerson. Confirmation bias: A ubiquitous phenomenon in many guises. Review of general psychology, 2(2):175–220, 1998
1998
-
[24]
P. C. Wason. On the failure to eliminate hypotheses in a conceptual task. Quarterly Journal of Experimental Psychology, 12(3):129–140, 1960
1960
-
[25]
Dunning, D
D. Dunning, D. W. Griffin, J. D. Milojkovic, and L. Ross. The overconfidence effect in social prediction. Journal of Personality and Social Psychology, 58(4):568–581, 1990
1990
-
[26]
Tversky and D
A. Tversky and D. Kahneman. Judgment under uncertainty: Heuristics and biases. Science, 185(4157):1124–1131, 1974
1974
-
[27]
Tversky and D
A. Tversky and D. Kahneman. Availability: A heuristic for judging frequency and probability. Cognitive Psychology, 5(2):207–232, 1973
1973
-
[28]
I. L. Janis. Victims of Groupthink: A Psychological Study of Foreign-Policy Decisions and Fiascoes. Houghton Mifflin, Oxford, England, 1972
1972
-
[29]
S. Milgram. Behavioral study of obedience. The Journal of Abnormal and Social Psychology, 67(4):371–378, 1963
1963
-
[30]
Gohar, M
U. Gohar, M. C. Hunter, R. R. Lutz, and M. B. Cohen. Codefeater: Using LLMs to find defeaters in Assurance cases. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering (ASE ’24), pages 2262–2267, New York, NY , USA, 2024. Association fo...
2024
-
[31]
N. G. Leveson. The use of safety cases in certification and regulation. Technical report, Massachusetts Institute of Technology, Engineering Systems Division, 2011. Working Paper
2011
-
[32]
R. Uuk, A. Brouwer, N. Dreksler, V . Pulignano, and R. Bommasani. Effective mitigations for systemic risks from general-purpose AI, 2024. SSRN Scholarly Paper. Rochester, NY: Social Science Research Network
2024
-
[33]
Cârlan, F
C. Cârlan, F. Gomez, Y . Mathew, K. Krishna, R. King, P. Gebauer, and B. R. Smith. Dynamic safety cases for frontier AI, 2024
2024
-
[34]
Varadarajan, R
S. Varadarajan, R. Bloomfield, J. Rushby, G. Gupta, A. Murugesan, R. Stroud, K. Netkachova, and I. H. Wong. CLARISSA: Foundations, tools & automation for Assurance cases. In 2023 IEEE/AIAA 42nd Digital Avionics Systems Conference (DASC), pages 1–10, Barcelona, Spain,
2023
-
[35]
M. J. Eppler and M. Aeschimann. A systematic framework for risk visualization in risk management and communication. Risk Management, 11(2):67–89, 2009
2009
-
[36]
L. J. Trevena, B. J. Zikmund-Fisher, A. Edwards, W. Gaissmaier, M. Galesic, P. K. J. Han, J. King, et al. Presenting quantitative information about decision outcomes: A risk commu- nication primer for patient decision aid developers. BMC Medical Informatics and Decision Making...
2013
-
[37]
Hausmann and D
D. Hausmann and D. Läge. Sequential evidence accumulation in decision making: The individ- ual desired level of confidence can explain the extent of information acquisition. Judgment and Decision Making, 3(3):229–243, 2008
2008
-
[38]
Spiegelhalter
D. Spiegelhalter. Risk and uncertainty communication. Annual Review of Statistics and Its Application, 4(1):31–60, 2017
2017
-
[39]
Fischhoff, N
B. Fischhoff, N. T. Brewer, and J. S. Downs. Communicating risks and benefits: An evidence- based user’s guide. Technical report, FDA, 2011. 47
2011
-
[40]
Reuel, L
A. Reuel, L. Soder, B. Bucknall, and T. A. Undheim. Position paper: Technical research and talent is needed for effective AI governance, 2024
2024
-
[41]
N. Guha, C. M. Lawrence, L. A. Gailmard, K. T. Rodolfa, F. Surani, R. Bommasani, I. D. Raji, et al. The AI regulatory alignment problem. https://dho.stanford.edu/wp-content/ uploads/AI_Regulation.pdf
-
[42]
Krook et al
J. Krook et al. Trustworthy autonomous systems hub – written evidence. Technical report, House of Lords Communications and Digital Select Committee, 2023. Inquiry: Large language models
2023
-
[43]
Accessed 20 December 2024
AI index report 2024 – artificial intelligence index. Accessed 20 December 2024. Available at: https://aiindex.stanford.edu/report/
2024
-
[44]
Peters, P
E. Peters, P. S. Hart, and L. Fraenkel. Informing patients: The influence of numeracy, framing, and format of side effect information on risk perceptions. Medical Decision Making, 31(3):432– 436, 2011
2011
-
[45]
Janzwood
S. Janzwood. Confident, likely, or both? the implementation of the uncertainty language framework in IPCC special reports. Climatic Change, 162(3):1655–1675, 2020
2020
-
[46]
M. C. Duke. Probability and confidence: How to improve communication of uncertainty about uncertainty in intelligence analysis. Journal of Behavioral Decision Making, 37(1):e2364, 2024
2024
-
[47]
Chozos and R
N. Chozos and R. Bloomfield. Mini-guide 6: Summarising and communication. Part of Declare CAE guidance document set
-
[48]
Treweek, A
DECIDE Consortium, S. Treweek, A. D. Oxman, P. Alderson, P. M. Bossuyt, L. Brandt, J. Bro˙zek, et al. Developing and evaluating communication strategies to support informed decisions and practice based on evidence (DECIDE): Protocol and preliminary results. Imple- mentation Sc...
2013
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.