REVIEW 3 major objections 4 minor 1 cited by
AI Safety Frameworks Should Include Procedures for Model Access Decisions
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read To govern who gets what access to frontier AI models, companies should adopt explicit 'Responsible Access Policies' built on empirical evaluation, user risk profiling, and pre-commitments.
desk verdict A credible, clearly argued policy proposal that names a framework (RAPs) but leans on an unproven feasibility assumption; worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's central instrument is the Responsible Access Policy, which is operationalized through an 'Access Assessment Matrix' that maps access styles against user groups. Access styles include chat, fine-tuning, weights inspection, and weights modification; user groups range from the general public to researchers, AI safety institutes, and governments. The matrix is meant to force explicit, evidence-backed reasoning about the risks and benefits of each combination, and to make those decisions visible to external stakeholders. The three pillars—empirical evaluation, user profiling, and pre-commitments—are the machinery that gives the matrix its content.
What would settle it
A systematic comparison of pre-release access-style evaluations with post-release outcomes across several frontier models would settle the empirical pillar: if predicted capability uplifts from fine-tuning or open-weights access do not correlate with observed misuse or beneficial use, the core justification for RAPs fails.
Extended reading notes
Core claim
The central claim is that frontier AI companies should build on existing safety frameworks by adding formal procedures for model access decisions. The authors propose Responsible Access Policies with three minimum components: i) processes for empirically evaluating model capabilities given different styles of access, ii) processes for assessing the risk profiles of different categories of user, and iii) clear pre-commitments regarding when to grant or revoke specific types of access under specified conditions. They argue that these components would make access governance accountable, legible, and empirically grounded, and that companies have an opportunity to set a standard for the industry and for regulators.
Load-bearing premise
The proposal assumes that the risks and benefits of different access styles for different user groups can be reliably measured before access is granted, and that those measurements remain meaningful as usage patterns and technologies change.
Editorial extensions
If this is right
- Safety frameworks at frontier companies would expand to include explicit access-governance procedures, making release decisions more predictable.
- Regulators would gain a clearer basis for evaluating whether companies are managing access risk responsibly.
- Researchers and downstream users would have greater certainty about the stability of their access to model capabilities.
- A new research agenda would emerge around measuring how different access styles change model capabilities and misuse potential.
Reading between the lines
- The Access Assessment Matrix could evolve into a shared industry template or regulatory reporting standard, much like model cards.
- The empirical-evaluation requirement implies that companies will need to invest in adversary-aware evaluation methods, since many capability uplifts may only manifest after users adapt.
- The pre-commitment element could be strengthened by making some commitments legally binding or independently auditable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that frontier AI companies should extend their existing safety frameworks with explicit procedures for model access decisions, which the authors call Responsible Access Policies (RAPs). The proposal has three pillars: (i) empirical evaluation of model capabilities under different access styles, (ii) assessment of the risk profiles of different user categories, and (iii) clear, robust pre-commitments governing when specific access styles are granted or revoked for particular groups. The paper reviews related work on structured access, open versus closed source debates, and current safety frameworks, and motivates RAPs by highlighting both the risks of incautious access (jailbreak bypass, irreversible spread, reduced oversight) and the opportunity costs of overly restrictive access (slowed safety research, underutilization, inequitable power concentration). It also argues that developers cannot be assumed to make good access decisions by default, and that governments and society need transparency about future access regimes. The paper is a normative/conceptual contribution and contains no empirical data or formal modeling.
Significance. The paper addresses a real and timely gap: existing safety frameworks such as Anthropic's RSP, OpenAI's Preparedness Framework, and Google DeepMind's Frontier Safety Framework say little about the procedures for deciding who gets which style of access. The proposal to require transparent, empirically grounded, pre-committed access policies is a plausible and useful policy recommendation for developers and regulators. The paper's strength is its clear conceptual structure: separating access styles from user groups, distinguishing reversible from irreversible access styles, and emphasizing pre-commitment and transparency as governance tools. It is also honest about the difficulty of the empirical evaluations it recommends. The main limitation is that the feasibility of the core evaluation pillar is asserted rather than demonstrated; the paper provides no protocol, evidence standard, or pilot example, and it concedes that access-style use cases can shift unpredictably. As a position paper the contribution is worthwhile, but the central mechanism remains under-specified.
major comments (3)
- [§4.2, 'Evaluating Access Styles' and §5] The load-bearing assumption of RAPs is that the risks and benefits of different access styles can be empirically evaluated before an irreversible release, yet the paper provides no account of what such an evaluation would look like, what evidence would suffice, or how uncertainty should be handled. The paper itself concedes in §4.2 that 'modelling different malicious actors using new technologies over different time frames with different resources will be a significant challenge' and in §5 that 'the different use cases afforded by different access styles may be unclear and change substantially over time.' If these concessions are accurate, then pillar (i) cannot reliably support irreversible decisions such as downloadable weights release, and the pre-commitments of §4.1 become vacuous because the conditions to be specified cannot be reliably anticipated. The paper needs at least a worked example of a plausible evaluation protocol or a clear statement of epistemic standards—e.g., what level of evidence justifies a 'safety case' for an access style—to make the central recommendation actionable.
- [§2.2 and §4.3, definitions of 'access style', 'user groups', and 'Access Assessment Matrix'] Key definitions are deferred to the first author's forthcoming work [5], making it difficult to assess how much of the framework is new and whether it is operationalizable. In particular, 'access style' is defined informally, 'user groups' are listed only by example, and the 'Access Assessment Matrix' mentioned in the Executive Summary and §4.3 is never actually defined or illustrated in the text. Since the paper's proposal depends on these terms being clear and consistently applied, the authors should either provide precise definitions inline or summarize the relevant parts of [5]. Without this, the transparency requirement in §4.3—which demands 'precise definitions which avoid unfairness'—cannot be evaluated.
- [§4.1, 'Specified Procedure'] The paper recommends that companies pre-commit to 'detailed protocols for conducting evaluations, including defined significance levels,' but it does not discuss how significance levels should be chosen or how they would apply to the kind of uncertain, fast-moving risk assessments described elsewhere. A naive use of significance levels in frontier AI evaluations could create a false impression of scientific rigor while leaving substantial modeler discretion. The paper should address this risk, for example by discussing how pre-registration, independent audits, or confidence intervals might be used, or by acknowledging that statistical significance is only one input into a broader safety case.
minor comments (4)
- [Title and Abstract] The title reads 'AI Safety Frameworks Should Include Procedure for Model Access Decisions'; 'Procedure' should be 'Procedures' to match the plural content of the paper.
- [§2.1, 'Related Literature'] There is a typo: 'Comapnies like Meta' should be 'Companies like Meta'. Also, the sentence 'these extent to which these frameworks are comprehensive or feasible is unclear' contains a redundant 'these' and should be rephrased.
- [Reference [31]] The reference for the EU AI Act cites 'Regulation 2024/1689' but the URL points to CELEX:32014R0269, which is an older directive, not the AI Act. The reference should be corrected or replaced with the correct legal citation.
- [§4.3, 'Transparency'] The paper refers to the 'Access Assessment Matrix' as 'a useful way to build on existing data representation techniques,' but no example or template of the matrix is provided anywhere in the manuscript. Since the figure is referenced in the Executive Summary, it should be included or the text should describe its structure.
Circularity Check
No circularity: the paper is a normative policy proposal whose central recommendation is argued from external literature and qualitative risk/benefit reasoning, not derived from its own definitions, and its only self-citation is not load-bearing.
full rationale
The paper makes no quantitative derivation and contains no equations, fitted parameters, or empirical predictions that could reduce to its inputs. Its central claim - that frontier AI companies should adopt Responsible Access Policies (RAPs) involving empirical evaluation of access styles, user-profile risk assessment, and pre-commitments - is a normative proposal defended through cited external work (e.g., structured access, open vs. closed model governance) and qualitative arguments about societal risk and opportunity cost. The definitions in Section 2.2 (model aspects, access styles, access regimes) are stipulated conceptual vocabulary, not premises from which the recommendation is derived. The only self-citation, [5] (Kembery, Bucknall, and Simpson, forthcoming), is used in the introduction alongside [6,7] to note that safety frameworks help guide release decisions, and in Section 4.3 to recommend careful definitions of terms like 'research organisations' and 'trusted third-parties'. In neither place is a load-bearing result imported from [5]; the surrounding claims are independently supported, and the recommendation does not depend on any specific content of the forthcoming paper. The paper's own concessions about the difficulty of empirical evaluation (Section 4.2 and Conclusion) are feasibility concerns, not circularity. The appended errata about the SoLaR workshop acceptance is unrelated to the argument's validity. Accordingly, no circular step meets the quoted-reduction standard.
Assumptions & free parameters
assumptions (3)
- domain assumption Existing frontier AI safety frameworks lack procedures for responsible model access decisions.
- domain assumption Transparent procedures and pre-commitments increase the likelihood of responsible access governance.
- domain assumption Model access risks and benefits can be empirically evaluated before release.
invented entities (2)
-
Responsible Access Policies (RAPs)
-
Access Assessment Matrix
Cite this review
Pith. "Pith review of AI Safety Frameworks Should Include Procedures for Model Access Decisions." pith.science (2026). https://pith.science/paper/STOXL36A
@misc{pith2026241110547,
author = {Pith},
title = {Pith review of: AI Safety Frameworks Should Include Procedures for Model Access Decisions},
year = {2026},
howpublished = {\url{https://pith.science/paper/STOXL36A}},
note = {Machine review of arXiv:2411.10547}
}
read the original abstract
The downstream use cases, benefits, and risks of AI models depend significantly on what sort of access is provided to the model, and who it is provided to. Though existing safety frameworks and AI developer usage policies recognise that the risk posed by a given model depends on the level of access provided to a given audience, the procedures they use to make decisions about model access are ad hoc, opaque, and lacking in empirical substantiation. This paper consequently proposes that frontier AI companies build on existing safety frameworks by outlining transparent procedures for making decisions about model access, which we term Responsible Access Policies (RAPs). We recommend that, at a minimum, RAPs should include the following: i) processes for empirically evaluating model capabilities given different styles of access, ii) processes for assessing the risk profiles of different categories of user, and iii) clear and robust pre-commitments regarding when to grant or revoke specific types of access for particular groups under specified conditions.
Forward citations
Cited by 1 Pith paper
-
Position Paper: Model Access should be a Key Concern in AI Governance
Model access decisions should be studied and coordinated through a dedicated research field, with recommendations for evaluators, companies, governments, and international bodies.
Reference graph
Works this paper leans on
-
[5]
Towards Model Access Governance
Edward Kembery, Ben Bucknall, and Morgan Simpson. Towards Model Access Governance. [Accessed 13-09-2024]. Forthcoming
work page 2024
-
[1]
Responsible Scaling Policy Updates — anthropic.com
Anthropic. Responsible Scaling Policy Updates — anthropic.com. https://www.anthropic. com/rsp-updates. [Accessed 09-11-2024]
work page 2024
-
[2]
Responsible Scaling Policies (https://metr.org/blog/2023-09-26-rsp/)
METR. Responsible Scaling Policies (https://metr.org/blog/2023-09-26-rsp/) . Tech. rep. METR, 26 September 2023
2023
-
[3]
OpenAI Preparedness Framework
OpenAI. OpenAI Preparedness Framework . https://cdn.openai.com/openai-preparedness- framework-beta.pdf. [Accessed 17-09-2024]. 2023
2024
-
[4]
Introducing the Frontier Safety Framework — deepmind.google
Anca Dragan, Helen King, and Allan Dafoe. Introducing the Frontier Safety Framework — deepmind.google. https://deepmind.google/discover/blog/introducing- the- frontier-safety-framework/. [Accessed 17-09-2024]. 17 May 2024
2024
-
[6]
Elizabeth Seger et al. Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives . 2023. arXiv: 2311.09227 [cs.CY]. URL: https://arxiv.org/abs/2311.09227
arXiv 2023
-
[7]
Near to Mid-term Risks and Opportunities of Open Source Generative AI
Francisco Eiras et al. “Near to Mid-term Risks and Opportunities of Open Source Generative AI”. In: arXiv preprint arXiv:2404.17047 (2024)
arXiv 2024
-
[8]
Open-Source
David Evan Harris. How to Regulate Unsecured “Open-Source” AI: No Exemptions. Tech. rep. Tech Policy Press, December 4 2023
2023
Show all 38 references
-
[9]
On the Societal Impact of Open Foundation Models
Sayash Kapoor et al. On the Societal Impact of Open Foundation Models . 2024. arXiv: 2403.07918 [cs.CY]. URL: https://arxiv.org/abs/2403.07918
2024 arXiv
-
[11]
Mark Zuckerberg - Llama 3, Open Sourcing $10b Models, & Caesar Au- gustus — dwarkeshpatel.com
Dwarkesh Patel. Mark Zuckerberg - Llama 3, Open Sourcing $10b Models, & Caesar Au- gustus — dwarkeshpatel.com. https://www.dwarkeshpatel.com/p/mark-zuckerberg. [Accessed 26-09-2024]. 2024
2024
-
[12]
A Grading Rubric for AI Safety Frame- works
Jide Alaga, Jonas Schuett, and Markus Anderljung. A Grading Rubric for AI Safety Frame- works. 2024. arXiv: 2409.08751 [cs.CY]. URL: https://arxiv.org/abs/2409.08751
2024 arXiv
-
[13]
Open Source AI Is the Path Forward | Meta
Mark Zuckerberg. Open Source AI Is the Path Forward | Meta. https://about.fb.com/ news/2024/07/open- source- ai- is- the- path- forward/ . [Accessed 27-07-2024]. 2024
2024
-
[14]
NTIA AI Open Model Weights Request for Comments
Misc. NTIA AI Open Model Weights Request for Comments. https://www.regulations. gov/document/NTIA-2023-0009-0001/comment . [Accessed 17-09-2024]. 26 February 2024
2023
-
[15]
Beyond Privacy Trade-offs with Structured Transparency
Andrew Trask et al. Beyond Privacy Trade-offs with Structured Transparency. 2024. arXiv: 2012.08347 [cs.CR]. URL: https://arxiv.org/abs/2012.08347
2024 arXiv
-
[16]
Structured access: an emerging paradigm for safe AI deployment
Toby Shevlane. “Structured access: an emerging paradigm for safe AI deployment”. In: arXiv preprint arXiv:2201.05159 (2022)
2022 arXiv
-
[17]
Bucknall and Robert F
Benjamin S. Bucknall and Robert F. Trager. Structured access for third-party research on frontier AI models: Investigating researchers’ model access requirements. Tech. rep. Oxford Martin School of Governance, Oct. 2023. 8
2023
-
[18]
A safe harbor for ai evaluation and red teaming
Shayne Longpre et al. “A safe harbor for ai evaluation and red teaming”. In: arXiv preprint arXiv:2403.04893 (2024)
2024 arXiv
-
[19]
Auditing large language models: a three-layered approach
Jakob Mökander et al. “Auditing large language models: a three-layered approach”. In: AI and Ethics (2023), pp. 1–31
2023
-
[20]
Conformity assessments and post-market monitoring: a guide to the role of auditing in the proposed European AI regulation
Jakob Mökander et al. “Conformity assessments and post-market monitoring: a guide to the role of auditing in the proposed European AI regulation”. In: Minds and Machines 32.2 (2022), pp. 241–268
2022
-
[21]
Ethics-based auditing to develop trustworthy AI
Jakob Mökander and Luciano Floridi. “Ethics-based auditing to develop trustworthy AI”. In: Minds and Machines 31.2 (2021), pp. 323–327
2021
-
[22]
Operationalising AI governance through ethics-based auditing: an industry case study
Jakob Mökander and Luciano Floridi. “Operationalising AI governance through ethics-based auditing: an industry case study”. In: AI and Ethics 3.2 (2023), pp. 451–468
2023
-
[23]
Stealing Machine Learning Models via Prediction APIs
Florian Tramèr et al. Stealing Machine Learning Models via Prediction APIs. 2016. arXiv: 1609.02943 [cs.CR]. URL: https://arxiv.org/abs/1609.02943
2016 arXiv
-
[24]
Stealing part of a production language model
Nicholas Carlini et al. “Stealing part of a production language model”. In: arXiv preprint arXiv:2403.06634 (2024)
2024 arXiv
-
[25]
Vulnerability Detection in Open Source Software: An Introduction
Stuart Millar. Vulnerability Detection in Open Source Software: An Introduction. 2022. arXiv: 2203.16428 [cs.CR]. URL: https://arxiv.org/abs/2203.16428
2022 arXiv
-
[26]
Towards a Framework for Openness in Foundation Models: Pro- ceedings from the Columbia Convening on Openness in Artificial Intelligence
Adrien Basdevant et al. Towards a Framework for Openness in Foundation Models: Pro- ceedings from the Columbia Convening on Openness in Artificial Intelligence. 2024. arXiv: 2405.15802 [cs.SE]. URL: https://arxiv.org/abs/2405.15802
2024 arXiv
-
[27]
The Open Source AI Definition – draft v
Open Source Initiative. The Open Source AI Definition – draft v. 0.0.3. [Accessed 01-08-2024]. Oct. 2023
2024
-
[28]
Beyond Open vs
Jon Bateman et al. Beyond Open vs. Closed: Emerging Consensus and Key Questions for Foundation AI Model Governance. https://carnegieendowment.org/research/2024/ 07 / beyond - open - vs - closed - emerging - consensus - and - key - questions - for - foundation-ai-model-governan...
2024
-
[29]
The Gradient of Generative AI Release: Methods and Considerations
Irene Solaiman. The Gradient of Generative AI Release: Methods and Considerations. 2023. arXiv: 2302.04844 [cs.CY]. URL: https://arxiv.org/abs/2302.04844
2023 arXiv
-
[30]
Considerations for Governing Open Foundation Models
Rishi Bommasani et al. Considerations for Governing Open Foundation Models. Tech. rep. Stanford HAI, 2024
2024
-
[31]
Regulation 2024/1689 of the European Parliament and of the Council
Council of European Union. Regulation 2024/1689 of the European Parliament and of the Council . http : / / eur - lex . europa . eu / legal - content / EN / TXT / ?qid = 1416170084502&uri=CELEX:32014R0269. June 2024
2024
-
[32]
Frontier AI Regulation: Managing Emerging Risks to Public Safety
Markus Anderljung et al. Frontier AI Regulation: Managing Emerging Risks to Public Safety
-
[33]
Ethical and social risks of harm from language models
Laura Weidinger et al. “Ethical and social risks of harm from language models”. In: arXiv preprint arXiv:2112.04359 (2021)
2021 arXiv
-
[34]
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! 2023
Xiangyu Qi et al. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! 2023. arXiv: 2310.03693 [cs.CL]. URL: https://arxiv.org/ abs/2310.03693
2023 arXiv
-
[35]
Black-Box Access is Insufficient for Rigorous AI Audits
Stephen Casper et al. Black-Box Access is Insufficient for Rigorous AI Audits . 2024. DOI: 10.1145/3630106.3659037 . arXiv: 2401.14446 [cs.CY]. URL: https://arxiv.org/ abs/2401.14446
2024
-
[36]
Typhoon: Thai Large Language Models
Kunat Pipatanakul et al. Typhoon: Thai Large Language Models. 2023. arXiv: 2312.13951 [cs.CL]. URL: https://arxiv.org/abs/2312.13951
2023 arXiv
-
[37]
Systematic Inequalities in Language Technology Performance across the World’s Languages
Damian Blasi, Antonios Anastasopoulos, and Graham Neubig. “Systematic Inequalities in Language Technology Performance across the World’s Languages”. In: arXiv preprint arXiv:2110.06733 (2021)
2021 arXiv
-
[38]
Building an early warning system for LLM-aided biological threat creation
OpenAI. Building an early warning system for LLM-aided biological threat creation. Tech. rep. OpenAI, 2024. 9
2024
-
[2023]
URL: https://arxiv.org/abs/2307.03718
arXiv: 2307.03718 [cs.CY]. URL: https://arxiv.org/abs/2307.03718
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.