Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

AI Safety Frameworks Should Include Procedures for Model Access Decisions

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read To govern who gets what access to frontier AI models, companies should adopt explicit 'Responsible Access Policies' built on empirical evaluation, user risk profiling, and pre-commitments.

desk verdict A credible, clearly argued policy proposal that names a framework (RAPs) but leans on an unproven feasibility assumption; worth refereeing. read the letter →

arxiv 2411.10547 v2 pith:STOXL36A submitted 2024-11-15 cs.CY

classification cs.CY
keywords AIsafetyframeworksmodelaccessgovernanceresponsiblepoliciesstylesusercategoriespre-commitmentsempiricalevaluationfrontier
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Frontier AI companies currently decide who gets what kind of access to their models through ad hoc, opaque processes that are not anchored in safety frameworks. This paper argues that those decisions should be governed by explicit, transparent procedures, which the authors call Responsible Access Policies (RAPs). A RAP would require empirical evaluation of what different access styles (such as chat, fine-tuning, or open weights) enable, risk assessment of user categories, and pre-commitments about when access will be granted or revoked. The motivation is that miscalibrated access can either amplify misuse risks or create opportunity costs, and that current frameworks lack the procedural machinery to manage this trade-off responsibly.

What carries the argument

The paper's central instrument is the Responsible Access Policy, which is operationalized through an 'Access Assessment Matrix' that maps access styles against user groups. Access styles include chat, fine-tuning, weights inspection, and weights modification; user groups range from the general public to researchers, AI safety institutes, and governments. The matrix is meant to force explicit, evidence-backed reasoning about the risks and benefits of each combination, and to make those decisions visible to external stakeholders. The three pillars—empirical evaluation, user profiling, and pre-commitments—are the machinery that gives the matrix its content.

What would settle it

A systematic comparison of pre-release access-style evaluations with post-release outcomes across several frontier models would settle the empirical pillar: if predicted capability uplifts from fine-tuning or open-weights access do not correlate with observed misuse or beneficial use, the core justification for RAPs fails.

Watch

Extended reading notes

Core claim

The central claim is that frontier AI companies should build on existing safety frameworks by adding formal procedures for model access decisions. The authors propose Responsible Access Policies with three minimum components: i) processes for empirically evaluating model capabilities given different styles of access, ii) processes for assessing the risk profiles of different categories of user, and iii) clear pre-commitments regarding when to grant or revoke specific types of access under specified conditions. They argue that these components would make access governance accountable, legible, and empirically grounded, and that companies have an opportunity to set a standard for the industry and for regulators.

Load-bearing premise

The proposal assumes that the risks and benefits of different access styles for different user groups can be reliably measured before access is granted, and that those measurements remain meaningful as usage patterns and technologies change.

Editorial extensions

If this is right

  • Safety frameworks at frontier companies would expand to include explicit access-governance procedures, making release decisions more predictable.
  • Regulators would gain a clearer basis for evaluating whether companies are managing access risk responsibly.
  • Researchers and downstream users would have greater certainty about the stability of their access to model capabilities.
  • A new research agenda would emerge around measuring how different access styles change model capabilities and misuse potential.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The Access Assessment Matrix could evolve into a shared industry template or regulatory reporting standard, much like model cards.
  • The empirical-evaluation requirement implies that companies will need to invest in adversary-aware evaluation methods, since many capability uplifts may only manifest after users adapt.
  • The pre-commitment element could be strengthened by making some commitments legally binding or independently auditable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper argues that frontier AI companies should extend their existing safety frameworks with explicit procedures for model access decisions, which the authors call Responsible Access Policies (RAPs). The proposal has three pillars: (i) empirical evaluation of model capabilities under different access styles, (ii) assessment of the risk profiles of different user categories, and (iii) clear, robust pre-commitments governing when specific access styles are granted or revoked for particular groups. The paper reviews related work on structured access, open versus closed source debates, and current safety frameworks, and motivates RAPs by highlighting both the risks of incautious access (jailbreak bypass, irreversible spread, reduced oversight) and the opportunity costs of overly restrictive access (slowed safety research, underutilization, inequitable power concentration). It also argues that developers cannot be assumed to make good access decisions by default, and that governments and society need transparency about future access regimes. The paper is a normative/conceptual contribution and contains no empirical data or formal modeling.

Significance. The paper addresses a real and timely gap: existing safety frameworks such as Anthropic's RSP, OpenAI's Preparedness Framework, and Google DeepMind's Frontier Safety Framework say little about the procedures for deciding who gets which style of access. The proposal to require transparent, empirically grounded, pre-committed access policies is a plausible and useful policy recommendation for developers and regulators. The paper's strength is its clear conceptual structure: separating access styles from user groups, distinguishing reversible from irreversible access styles, and emphasizing pre-commitment and transparency as governance tools. It is also honest about the difficulty of the empirical evaluations it recommends. The main limitation is that the feasibility of the core evaluation pillar is asserted rather than demonstrated; the paper provides no protocol, evidence standard, or pilot example, and it concedes that access-style use cases can shift unpredictably. As a position paper the contribution is worthwhile, but the central mechanism remains under-specified.

major comments (3)
  1. [§4.2, 'Evaluating Access Styles' and §5] The load-bearing assumption of RAPs is that the risks and benefits of different access styles can be empirically evaluated before an irreversible release, yet the paper provides no account of what such an evaluation would look like, what evidence would suffice, or how uncertainty should be handled. The paper itself concedes in §4.2 that 'modelling different malicious actors using new technologies over different time frames with different resources will be a significant challenge' and in §5 that 'the different use cases afforded by different access styles may be unclear and change substantially over time.' If these concessions are accurate, then pillar (i) cannot reliably support irreversible decisions such as downloadable weights release, and the pre-commitments of §4.1 become vacuous because the conditions to be specified cannot be reliably anticipated. The paper needs at least a worked example of a plausible evaluation protocol or a clear statement of epistemic standards—e.g., what level of evidence justifies a 'safety case' for an access style—to make the central recommendation actionable.
  2. [§2.2 and §4.3, definitions of 'access style', 'user groups', and 'Access Assessment Matrix'] Key definitions are deferred to the first author's forthcoming work [5], making it difficult to assess how much of the framework is new and whether it is operationalizable. In particular, 'access style' is defined informally, 'user groups' are listed only by example, and the 'Access Assessment Matrix' mentioned in the Executive Summary and §4.3 is never actually defined or illustrated in the text. Since the paper's proposal depends on these terms being clear and consistently applied, the authors should either provide precise definitions inline or summarize the relevant parts of [5]. Without this, the transparency requirement in §4.3—which demands 'precise definitions which avoid unfairness'—cannot be evaluated.
  3. [§4.1, 'Specified Procedure'] The paper recommends that companies pre-commit to 'detailed protocols for conducting evaluations, including defined significance levels,' but it does not discuss how significance levels should be chosen or how they would apply to the kind of uncertain, fast-moving risk assessments described elsewhere. A naive use of significance levels in frontier AI evaluations could create a false impression of scientific rigor while leaving substantial modeler discretion. The paper should address this risk, for example by discussing how pre-registration, independent audits, or confidence intervals might be used, or by acknowledging that statistical significance is only one input into a broader safety case.
minor comments (4)
  1. [Title and Abstract] The title reads 'AI Safety Frameworks Should Include Procedure for Model Access Decisions'; 'Procedure' should be 'Procedures' to match the plural content of the paper.
  2. [§2.1, 'Related Literature'] There is a typo: 'Comapnies like Meta' should be 'Companies like Meta'. Also, the sentence 'these extent to which these frameworks are comprehensive or feasible is unclear' contains a redundant 'these' and should be rephrased.
  3. [Reference [31]] The reference for the EU AI Act cites 'Regulation 2024/1689' but the URL points to CELEX:32014R0269, which is an older directive, not the AI Act. The reference should be corrected or replaced with the correct legal citation.
  4. [§4.3, 'Transparency'] The paper refers to the 'Access Assessment Matrix' as 'a useful way to build on existing data representation techniques,' but no example or template of the matrix is provided anywhere in the manuscript. Since the figure is referenced in the Executive Summary, it should be included or the text should describe its structure.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a normative policy proposal whose central recommendation is argued from external literature and qualitative risk/benefit reasoning, not derived from its own definitions, and its only self-citation is not load-bearing.

full rationale

The paper makes no quantitative derivation and contains no equations, fitted parameters, or empirical predictions that could reduce to its inputs. Its central claim - that frontier AI companies should adopt Responsible Access Policies (RAPs) involving empirical evaluation of access styles, user-profile risk assessment, and pre-commitments - is a normative proposal defended through cited external work (e.g., structured access, open vs. closed model governance) and qualitative arguments about societal risk and opportunity cost. The definitions in Section 2.2 (model aspects, access styles, access regimes) are stipulated conceptual vocabulary, not premises from which the recommendation is derived. The only self-citation, [5] (Kembery, Bucknall, and Simpson, forthcoming), is used in the introduction alongside [6,7] to note that safety frameworks help guide release decisions, and in Section 4.3 to recommend careful definitions of terms like 'research organisations' and 'trusted third-parties'. In neither place is a load-bearing result imported from [5]; the surrounding claims are independently supported, and the recommendation does not depend on any specific content of the forthcoming paper. The paper's own concessions about the difficulty of empirical evaluation (Section 4.2 and Conclusion) are feasibility concerns, not circularity. The appended errata about the SoLaR workshop acceptance is unrelated to the argument's validity. Accordingly, no circular step meets the quoted-reduction standard.

Assumptions & free parameters 0 free parameters · 3 assumptions · 2 invented entities

The paper contributes a policy framework, not an empirical result. Free parameters: none. Axioms are domain assumptions about the state of safety frameworks and the feasibility of empirical access evaluation. The proposed RAPs and Access Assessment Matrix are conceptual constructs with no independent falsifiable evidence.

assumptions (3)
  • domain assumption Existing frontier AI safety frameworks lack procedures for responsible model access decisions.
    Section 2 and Section 3.2 assert that frameworks say little about access style and user category; this is an empirical claim about current industry practice, supported only by examples.
  • domain assumption Transparent procedures and pre-commitments increase the likelihood of responsible access governance.
    Section 4.1 asserts that pre-commitments increase compliance and inform governments; this is a social-science assumption with no cited evidence.
  • domain assumption Model access risks and benefits can be empirically evaluated before release.
    Section 4.2 recommends empirical evaluations; the conclusion acknowledges this is difficult and that use cases change over time.
invented entities (2)
  • Responsible Access Policies (RAPs)
    purpose: A proposed governance framework for making model access decisions transparent and procedure-based.
    Introduced as a new named policy construct; no empirical validation is provided.
  • Access Assessment Matrix
    purpose: A matrix format for documenting access styles, user groups, and risk/benefit evaluations.
    Proposed as a concrete tool, but not tested or implemented.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI Safety Frameworks Should Include Procedures for Model Access Decisions." pith.science (2026). https://pith.science/paper/STOXL36A

@misc{pith2026241110547,
  author       = {Pith},
  title        = {Pith review of: AI Safety Frameworks Should Include Procedures for Model Access Decisions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/STOXL36A}},
  note         = {Machine review of arXiv:2411.10547}
}
read the original abstract

The downstream use cases, benefits, and risks of AI models depend significantly on what sort of access is provided to the model, and who it is provided to. Though existing safety frameworks and AI developer usage policies recognise that the risk posed by a given model depends on the level of access provided to a given audience, the procedures they use to make decisions about model access are ad hoc, opaque, and lacking in empirical substantiation. This paper consequently proposes that frontier AI companies build on existing safety frameworks by outlining transparent procedures for making decisions about model access, which we term Responsible Access Policies (RAPs). We recommend that, at a minimum, RAPs should include the following: i) processes for empirically evaluating model capabilities given different styles of access, ii) processes for assessing the risk profiles of different categories of user, and iii) clear and robust pre-commitments regarding when to grant or revoke specific types of access for particular groups under specified conditions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Position Paper: Model Access should be a Key Concern in AI Governance

    cs.CY 2024-12 accept novelty 4.0 of 10

    Model access decisions should be studied and coordinated through a dedicated research field, with recommendations for evaluators, companies, governments, and international bodies.

Reference graph

Works this paper leans on

38 extracted references · 5 canonical work pages · cited by 1 Pith paper

  1. [5]

    Towards Model Access Governance

    Edward Kembery, Ben Bucknall, and Morgan Simpson. Towards Model Access Governance. [Accessed 13-09-2024]. Forthcoming

  2. [1]

    Responsible Scaling Policy Updates — anthropic.com

    Anthropic. Responsible Scaling Policy Updates — anthropic.com. https://www.anthropic. com/rsp-updates. [Accessed 09-11-2024]

  3. [2]

    Responsible Scaling Policies (https://metr.org/blog/2023-09-26-rsp/)

    METR. Responsible Scaling Policies (https://metr.org/blog/2023-09-26-rsp/) . Tech. rep. METR, 26 September 2023

  4. [3]

    OpenAI Preparedness Framework

    OpenAI. OpenAI Preparedness Framework . https://cdn.openai.com/openai-preparedness- framework-beta.pdf. [Accessed 17-09-2024]. 2023

  5. [4]

    Introducing the Frontier Safety Framework — deepmind.google

    Anca Dragan, Helen King, and Allan Dafoe. Introducing the Frontier Safety Framework — deepmind.google. https://deepmind.google/discover/blog/introducing- the- frontier-safety-framework/. [Accessed 17-09-2024]. 17 May 2024

  6. [6]

    Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives

    Elizabeth Seger et al. Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives . 2023. arXiv: 2311.09227 [cs.CY]. URL: https://arxiv.org/abs/2311.09227

  7. [7]

    Near to Mid-term Risks and Opportunities of Open Source Generative AI

    Francisco Eiras et al. “Near to Mid-term Risks and Opportunities of Open Source Generative AI”. In: arXiv preprint arXiv:2404.17047 (2024)

  8. [8]

    Open-Source

    David Evan Harris. How to Regulate Unsecured “Open-Source” AI: No Exemptions. Tech. rep. Tech Policy Press, December 4 2023

Show all 38 references
  1. [9]

    On the Societal Impact of Open Foundation Models

    Sayash Kapoor et al. On the Societal Impact of Open Foundation Models . 2024. arXiv: 2403.07918 [cs.CY]. URL: https://arxiv.org/abs/2403.07918

  2. [11]

    Mark Zuckerberg - Llama 3, Open Sourcing $10b Models, & Caesar Au- gustus — dwarkeshpatel.com

    Dwarkesh Patel. Mark Zuckerberg - Llama 3, Open Sourcing $10b Models, & Caesar Au- gustus — dwarkeshpatel.com. https://www.dwarkeshpatel.com/p/mark-zuckerberg. [Accessed 26-09-2024]. 2024

  3. [12]

    A Grading Rubric for AI Safety Frame- works

    Jide Alaga, Jonas Schuett, and Markus Anderljung. A Grading Rubric for AI Safety Frame- works. 2024. arXiv: 2409.08751 [cs.CY]. URL: https://arxiv.org/abs/2409.08751

  4. [13]

    Open Source AI Is the Path Forward | Meta

    Mark Zuckerberg. Open Source AI Is the Path Forward | Meta. https://about.fb.com/ news/2024/07/open- source- ai- is- the- path- forward/ . [Accessed 27-07-2024]. 2024

  5. [14]

    NTIA AI Open Model Weights Request for Comments

    Misc. NTIA AI Open Model Weights Request for Comments. https://www.regulations. gov/document/NTIA-2023-0009-0001/comment . [Accessed 17-09-2024]. 26 February 2024

  6. [15]

    Beyond Privacy Trade-offs with Structured Transparency

    Andrew Trask et al. Beyond Privacy Trade-offs with Structured Transparency. 2024. arXiv: 2012.08347 [cs.CR]. URL: https://arxiv.org/abs/2012.08347

  7. [16]

    Structured access: an emerging paradigm for safe AI deployment

    Toby Shevlane. “Structured access: an emerging paradigm for safe AI deployment”. In: arXiv preprint arXiv:2201.05159 (2022)

  8. [17]

    Bucknall and Robert F

    Benjamin S. Bucknall and Robert F. Trager. Structured access for third-party research on frontier AI models: Investigating researchers’ model access requirements. Tech. rep. Oxford Martin School of Governance, Oct. 2023. 8

  9. [18]

    A safe harbor for ai evaluation and red teaming

    Shayne Longpre et al. “A safe harbor for ai evaluation and red teaming”. In: arXiv preprint arXiv:2403.04893 (2024)

  10. [19]

    Auditing large language models: a three-layered approach

    Jakob Mökander et al. “Auditing large language models: a three-layered approach”. In: AI and Ethics (2023), pp. 1–31

  11. [20]

    Conformity assessments and post-market monitoring: a guide to the role of auditing in the proposed European AI regulation

    Jakob Mökander et al. “Conformity assessments and post-market monitoring: a guide to the role of auditing in the proposed European AI regulation”. In: Minds and Machines 32.2 (2022), pp. 241–268

  12. [21]

    Ethics-based auditing to develop trustworthy AI

    Jakob Mökander and Luciano Floridi. “Ethics-based auditing to develop trustworthy AI”. In: Minds and Machines 31.2 (2021), pp. 323–327

  13. [22]

    Operationalising AI governance through ethics-based auditing: an industry case study

    Jakob Mökander and Luciano Floridi. “Operationalising AI governance through ethics-based auditing: an industry case study”. In: AI and Ethics 3.2 (2023), pp. 451–468

  14. [23]

    Stealing Machine Learning Models via Prediction APIs

    Florian Tramèr et al. Stealing Machine Learning Models via Prediction APIs. 2016. arXiv: 1609.02943 [cs.CR]. URL: https://arxiv.org/abs/1609.02943

  15. [24]

    Stealing part of a production language model

    Nicholas Carlini et al. “Stealing part of a production language model”. In: arXiv preprint arXiv:2403.06634 (2024)

  16. [25]

    Vulnerability Detection in Open Source Software: An Introduction

    Stuart Millar. Vulnerability Detection in Open Source Software: An Introduction. 2022. arXiv: 2203.16428 [cs.CR]. URL: https://arxiv.org/abs/2203.16428

  17. [26]

    Towards a Framework for Openness in Foundation Models: Pro- ceedings from the Columbia Convening on Openness in Artificial Intelligence

    Adrien Basdevant et al. Towards a Framework for Openness in Foundation Models: Pro- ceedings from the Columbia Convening on Openness in Artificial Intelligence. 2024. arXiv: 2405.15802 [cs.SE]. URL: https://arxiv.org/abs/2405.15802

  18. [27]

    The Open Source AI Definition – draft v

    Open Source Initiative. The Open Source AI Definition – draft v. 0.0.3. [Accessed 01-08-2024]. Oct. 2023

  19. [28]

    Beyond Open vs

    Jon Bateman et al. Beyond Open vs. Closed: Emerging Consensus and Key Questions for Foundation AI Model Governance. https://carnegieendowment.org/research/2024/ 07 / beyond - open - vs - closed - emerging - consensus - and - key - questions - for - foundation-ai-model-governan...

  20. [29]

    The Gradient of Generative AI Release: Methods and Considerations

    Irene Solaiman. The Gradient of Generative AI Release: Methods and Considerations. 2023. arXiv: 2302.04844 [cs.CY]. URL: https://arxiv.org/abs/2302.04844

  21. [30]

    Considerations for Governing Open Foundation Models

    Rishi Bommasani et al. Considerations for Governing Open Foundation Models. Tech. rep. Stanford HAI, 2024

  22. [31]

    Regulation 2024/1689 of the European Parliament and of the Council

    Council of European Union. Regulation 2024/1689 of the European Parliament and of the Council . http : / / eur - lex . europa . eu / legal - content / EN / TXT / ?qid = 1416170084502&uri=CELEX:32014R0269. June 2024

  23. [32]

    Frontier AI Regulation: Managing Emerging Risks to Public Safety

    Markus Anderljung et al. Frontier AI Regulation: Managing Emerging Risks to Public Safety

  24. [33]

    Ethical and social risks of harm from language models

    Laura Weidinger et al. “Ethical and social risks of harm from language models”. In: arXiv preprint arXiv:2112.04359 (2021)

  25. [34]

    Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! 2023

    Xiangyu Qi et al. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! 2023. arXiv: 2310.03693 [cs.CL]. URL: https://arxiv.org/ abs/2310.03693

  26. [35]

    Black-Box Access is Insufficient for Rigorous AI Audits

    Stephen Casper et al. Black-Box Access is Insufficient for Rigorous AI Audits . 2024. DOI: 10.1145/3630106.3659037 . arXiv: 2401.14446 [cs.CY]. URL: https://arxiv.org/ abs/2401.14446

  27. [36]

    Typhoon: Thai Large Language Models

    Kunat Pipatanakul et al. Typhoon: Thai Large Language Models. 2023. arXiv: 2312.13951 [cs.CL]. URL: https://arxiv.org/abs/2312.13951

  28. [37]

    Systematic Inequalities in Language Technology Performance across the World’s Languages

    Damian Blasi, Antonios Anastasopoulos, and Graham Neubig. “Systematic Inequalities in Language Technology Performance across the World’s Languages”. In: arXiv preprint arXiv:2110.06733 (2021)

  29. [38]

    Building an early warning system for LLM-aided biological threat creation

    OpenAI. Building an early warning system for LLM-aided biological threat creation. Tech. rep. OpenAI, 2024. 9

  30. [2023]

    URL: https://arxiv.org/abs/2307.03718

    arXiv: 2307.03718 [cs.CY]. URL: https://arxiv.org/abs/2307.03718

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.