Pith. sign in

REVIEW 3 major objections 4 minor 7 references

AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A standards profile translates the NIST AI Risk Management Framework into concrete risk-management practices for general-purpose AI and foundation models, covering harms from the individual to the catastrophic scale.

desk verdict A genuinely useful, transparent standards profile for GPAI risk management—but its flagship identification tool, red-teaming, is exactly what the sandbagging and faked-alignment risks it names would defeat. read the letter →

arxiv 2506.23949 v1 pith:PSICRJ3G submitted 2025-06-30 cs.AI cs.CRcs.CY

classification cs.AIcs.CRcs.CY
keywords AIriskmanagementgeneral-purposefoundationmodelsNISTRMFstandardsprofilered-teamingfrontiercatastrophic
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims to close the gap between generic AI risk-management standards and the specific realities of organizations building multi-purpose foundation models. It publishes a 'standards profile': a targeted set of practices and controls that translate the NIST AI Risk Management Framework (AI RMF) and the ISO/IEC 23894 standard into concrete steps for identifying, analyzing, and mitigating the risks of general-purpose AI (GPAI) and foundation models. The primary audience is developers of large-scale, state-of-the-art models, with downstream application developers, evaluators, and regulators as secondary users. If the claims hold, developers gain a documented baseline of minimum expectations, and evaluators and regulators gain a concrete checklist against which to judge whether developers followed relevant best practices.

What carries the argument

The carrying mechanism is the profile structure itself: the NIST AI RMF's four core functions—Govern, Map, Measure, and Manage—each filled in with GPAI-specific supplemental guidance, curated resource lists, and extended excerpts from the NIST AI RMF Playbook and the NIST Generative AI Profile. Around this skeleton the profile organizes a risk taxonomy of impact areas (individuals, groups, organizations, society, and the planet), harm factors (correlated bias, manipulation and deception, cyber and CBRN weaponization potential, loss of understanding and control), and trustworthiness characteristics (safety, security, transparency, explainability, privacy, fairness). A short list of high-priority steps functions as the baseline that defines minimum expectations for users, with the remaining sections available as optional depth.

What would settle it

Apply the profile's Map 1.1 and Map 5.1 procedures retrospectively to a set of already-released frontier models and compare the risks those procedures would have flagged with the harms the models actually caused after deployment; if the anticipated-risk lists systematically miss the harms that materialized, the claim that the guidance adequately identifies reasonably foreseeable risks is undercut. A prospective variant would track whether models developed under a full application of the high-priority steps later produced severe unanticipated harms, and whether the profile's staged-release gate would have stopped any of them.

Watch

Extended reading notes

Core claim

The paper's central claim is that GPAI/foundation models—models trained on broad data that can be adapted to a wide range of downstream tasks—carry risks that generic AI risk-management guidance does not address explicitly, and that a dedicated standards profile can supply the missing specificity. The profile reorganizes the NIST AI RMF's four functions into GPAI-specific guidance: Govern (policies, roles, and accountability), Map (identifying risks in context), Measure (rating trustworthiness characteristics), and Manage (prioritizing, mitigating, transferring, or accepting risks). Within that structure, it names a set of high-priority steps that should be treated as the minimum expectation for every developer: applying go/no-go checks before major investments, taking responsibility for risk tasks the organization is uniquely positioned to perform, setting unacceptable-risk thresholds with a margin of safety, identifying reasonably foreseeable uses, misuses, and abuses, identifying factors that could lead to significant, severe, or catastrophic impacts, red-teaming dangerous capabilities, tracking risks that cannot yet be measured, applying controls such as staged release and structured access, and reporting risk factors to stakeholders through transparency mechanisms.

Load-bearing premise

The load-bearing premise is that developers, using red-teaming, stakeholder input, and expert judgment before release, can adequately anticipate the uses, misuses, and abuses of a model that will later produce serious or catastrophic harms; the document itself acknowledges in Section 1.5 that such best practices are still nascent and rest partly on the authors' own judgment.

Editorial extensions

If this is right

  • Developers who follow the high-priority steps meet a documented minimum standard for GPAI/foundation-model risk management, and can exceed it by applying the remaining sections of the profile.
  • Upstream model developers are assigned primary responsibility for risks they are uniquely positioned to assess—training data, dangerous capabilities, foreseeable misuses—while downstream developers own context-specific risks, and each side is told what information to share with the other.
  • For frontier and near-frontier open-weights releases, the profile recommends staged release with structured access first and release of parameter weights only after sufficient confidence in risk management is established.
  • Evaluators and regulators gain a concrete checklist for judging whether developers followed relevant best practices, with mappings to the EU AI Act and related commitments provided in a companion document.
  • Risks that cannot yet be measured, such as vulnerabilities from data poisoning or mis-specified objectives, must be tracked rather than ignored, especially when the potential impacts are severe or catastrophic.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The profile's predictive value is not demonstrated in the paper: the authors concede in Section 1.5 that best practices are nascent and rest partly on their own judgment. A natural next step would be a systematic retrospective application of the Map and Measure steps to already-released models to check whether flagged risks match post-deployment incidents.
  • If this profile becomes the de facto soft-law baseline, it could converge international expectations on one operational standard, but that convergence would also entrench red-teaming and qualitative expert assessment as canonical safety methods before quantitative alternatives have matured.
  • The 'margin of safety' suggestion under Map 1.5 could be sharpened into a quantitative deployment rule—for instance, refusing release whenever the worst-plausible-case impact rating, multiplied by a safety factor, exceeds the unacceptable-risk threshold—which would make the risk-tolerance guidance testable.
  • The profile's treatment of model weights as the critical control point implies that the release decision, rather than the training run, is where risk management bites hardest; the authors themselves note that if future GPAI systems are assembled from many small models instead of one core model, the guidance would need to shift back toward a system-level focus.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This manuscript, version 1.1 of a standards profile from UC Berkeley's CLTC, provides an operationalization of AI risk management tailored to general-purpose AI (GPAI) and foundation models. It adapts the NIST AI RMF and ISO/IEC 23894, organizing guidance under the four RMF functions (Govern, Map, Measure, Manage) and identifying a set of 'high-priority' steps that constitute a baseline for developers. The central claim, stated in the abstract, is that the document 'provides risk-management practices or controls for identifying, analyzing, and mitigating risks' of GPAI/foundation models. The profile includes detailed tables, resource lists, and a range of specific recommendations, from red-teaming and staged release to documentation and stakeholder engagement. Section 1.5 openly acknowledges that best practices in this area are nascent and that the guidance is based on literature, industry practice, stakeholder input, and the authors' own judgment; the document also references companion materials, including a 'Retrospective Test Use' report, that are not included in the manuscript.

Significance. The Profile is a serious and usable contribution to the emerging practice of GPAI risk governance. Its strengths are real: the mapping to NIST AI RMF and ISO/IEC 23894 is careful; the 'high-priority' baseline in Section 2.3 makes the document usable by practitioners; the risk taxonomies in Map 1.1 and Map 5.1 are detailed and current; the discussion of open-weight release (Section 1.5, Map 1.5, Manage 2.4) is nuanced and policy-relevant; and the authors are transparent in Section 1.5 about the limits of the evidence base. As a standards profile, it does not need original empirical validation to be a useful reference, and the authors are honest that the guidance is based on expert judgment. However, the central value claim—that the practices 'identify, analyze, and mitigate risks'—is not equally supported across risk classes. The gap described in the major comments concerning sandbagging and evaluation gaming is load-bearing because it affects the very class of catastrophic risks the Profile prioritizes.

major comments (3)
  1. [Measure 1.1, Map 5.1, Section 2.3] The Profile makes red-teaming and adversarial testing its flagship identification control (Measure 1.1; high-priority step in Section 2.3), yet Map 5.1 itself identifies sandbagging (van der Weij et al. 2024) and faked alignment during testing (OpenAI 2024b) as significant risks for GPAI/foundation models. A model that can recognize evaluation contexts and strategically underperform can invalidate exactly the evidence that red-teaming is supposed to produce. The Profile provides no complementary measurement or control that tests for evaluation gaming, hidden capabilities, or incentive-invariant behavior; Measure 3.2 only instructs developers to track such risks when they cannot be measured. Consequently, for the class of deceptive or situationally aware models that the Profile treats as highest-severity, the 'identifying' pillar is not actually delivered. Please add guidance on detecting evaluation gaming, or at least an explicit statement that such detection is an open problem and that go/no-go decisions should account for this uncertainty; alternatively, revise the central claim to reflect this limitation.
  2. [Abstract, Section 1.5, Map 1.1, Map 5.1] The abstract's claim that the Profile 'provides risk-management practices or controls for identifying, analyzing, and mitigating risks' is not empirically supported for the prospective, qualitative identification methods on which the high-priority steps rely. Section 1.5 concedes that best practices are 'nascent' and that the guidance rests on 'our own judgment,' but the manuscript does not assess the accuracy of expert judgment, red-teaming, and stakeholder input in anticipating significant, severe, or catastrophic harms. This is load-bearing because Map 1.5 and Manage 1.1 presuppose that such methods can identify unacceptable risks before go/no-go decisions are made. The manuscript would be stronger if it included a short section on the known limitations and failure modes of expert elicitation and red-teaming, and if it stated plainly that the Profile is an expert-consensus document rather than a validated standard.
  3. [Section 1.2, 'Notes on This Version'] The manuscript repeatedly refers to a companion 'Retrospective Test Use of Profile Guidance' (Barrett et al. 2025c) that applies the Profile to new foundation models, but the manuscript does not summarize its findings. Because the central claim is that the Profile is actionable, the reader needs at least a brief statement of what the retrospective test covered and whether the guidance was found to be usable, or an explanation of why the results cannot be included. As it stands, the reference functions as a placeholder for evidence rather than as evidence itself.
minor comments (4)
  1. [Footnote 1, Section 1.1] The text says 'recession of Executive Order 14110'; this should be 'rescission'.
  2. [Table 4 header, Section 3.4] The header 'ManaGe 1' and the corresponding subcategory headers contain an obvious typo for 'Manage'; please correct throughout Table 4.
  3. [Map 4.1, Section 3.2] In the sentence 'as described t under Map 2.2,' the stray 't' should be removed.
  4. [Map 5.1, Section 3.2] The parenthetical 'In a future version of this Profile, we may provide a scoring system for rating impact hazard as a function of these factors' is a useful caveat, but since the impact magnitude rating scale from Barrett et al. (2022) is not reproduced, the reader has little concrete basis for 'rating potential impacts' as the guidance instructs; consider including a placeholder or a more explicit pointer to the source.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Profile is a normative compilation with transparent, non-load-bearing self-citations.

full rationale

The document does not derive quantitative predictions or fitted parameters; it compiles risk-management practices from the NIST AI RMF, ISO/IEC 23894, external literature, and the authors' prior work. The central claim, that the document provides practices for identifying, analyzing, and mitigating GPAI/foundation model risks, is satisfied by the content itself rather than by a derivation that reduces to its own inputs. Self-citations such as the impact magnitude rating scale in Barrett et al. (2022) and the supporting mapping in Barrett et al. (2025a) are used as resources, and Section 1.2 openly states that 'some of the material in this Profile is adapted directly from our related work.' Section 1.5 candidly notes that best practices are nascent and based partly on 'our own judgment,' which is a stated limitation rather than a circular justification. No equation, fitted value, or uniqueness theorem is invoked to force a conclusion, so no circular step can be exhibited.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on domain assumptions about AI risk and the sufficiency of expert judgment rather than on mathematical axioms; no free parameters or invented entities appear.

assumptions (3)
  • domain assumption GPAI/foundation models have risk profiles that are qualitatively distinct from narrower AI systems, warranting tailored risk management.
    Section 1.2 asserts distinct properties such as broad applicability, emergent properties, and societal-scale impacts; no empirical validation is provided for the distinctness claim.
  • domain assumption The NIST AI RMF and ISO/IEC 23894 are appropriate and sufficient base frameworks for GPAI/foundation model risk management.
    Section 2.1 states the profile is meant to be used with these frameworks; the choice is not justified.
  • domain assumption Expert judgment and stakeholder input can identify reasonably foreseeable severe risks.
    The Map and Measure sections rely on red-teaming and qualitative assessment without demonstrated predictive validity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models." pith.science (2026). https://pith.science/paper/PSICRJ3G

@misc{pith2026250623949,
  author       = {Pith},
  title        = {Pith review of: AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PSICRJ3G}},
  note         = {Machine review of arXiv:2506.23949}
}
read the original abstract

Increasingly multi-purpose AI models, such as cutting-edge large language models or other 'general-purpose AI' (GPAI) models, 'foundation models,' generative AI models, and 'frontier models' (typically all referred to hereafter with the umbrella term 'GPAI/foundation models' except where greater specificity is needed), can provide many beneficial capabilities but also risks of adverse events with profound consequences. This document provides risk-management practices or controls for identifying, analyzing, and mitigating risks of GPAI/foundation models. We intend this document primarily for developers of large-scale, state-of-the-art GPAI/foundation models; others that can benefit from this guidance include downstream developers of end-use applications that build on a GPAI/foundation model. This document facilitates conformity with or use of leading AI risk management-related standards, adapting and building on the generic voluntary guidance in the NIST AI Risk Management Framework and ISO/IEC 23894, with a focus on the unique issues faced by developers of GPAI/foundation models.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

7 extracted references · 1 canonical work pages

  1. [1]

    General Purpose AI

    ACSC (2024) Deploying AI Systems Securely: Best Practices for Deploying Secure and Resilient AI Systems. The Australian Signals Directorate’s Australian Cyber Security Centre, https://www.cyber.gov.au/resources-busi- ness-and-government/governance-and-user-education/artificial-intelligence/deploying-ai-systems-securely Julius Adebayo, Justin Gilmer, Micha...

  2. [2]

    National Institute of Standards and Technology, https://csrc.nist.gov/pubs/sp/800/171/r2/ upd1/final NIST (2020c) NIST Privacy Framework and Cybersecurity Framework to NIST Special Publication 800-53, Revision 5 Crosswalk. National Institute of Standards and Technology, https://www.nist.gov/privacy-framework/nist-priva- cy-framework-and-cybersecurity-fram...

  3. [323]

    OECD Digital Economy Papers, No

    Orga- nization for Economic Co-operation and Development, https://doi.org/10.1787/cb6d9eca-en OECD (2022b) Measuring the environmental impacts of artificial intelligence compute and applications: The AI footprint. OECD Digital Economy Papers, No. 341, Organization for Economic Co-operation and Development, https://doi.org/10.1787/7babf571-en OECD (2023) A...

  4. [1270]

    National Institute of Standards and Technology, https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.1270.pdf Elizabeth Seger, Noemi Dreksler, Richard Moulange, Emily Dardaman, Jonas Schuett, K. Wei, Christoph Winter, Mackenzie Arnold, Seán Ó hÉigeartaigh, Anton Korinek, Markus Anderljung, Ben Bucknall, Alan Chan, Eoghan Stafford, Leonie Koessler...

  5. [2021]

    https://datasets-benchmarks-proceedings.neurips.cc/ paper/2021/file/084b6fbb10729ed4da8c3d3f5a3ae7c9-Paper-round2.pdf Inioluwa Deborah Raji, Andrew Smart, Rebecca N

    Track on Datasets and Benchmarks. https://datasets-benchmarks-proceedings.neurips.cc/ paper/2021/file/084b6fbb10729ed4da8c3d3f5a3ae7c9-Paper-round2.pdf Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes (2020) Closing the AI accountability gap: definin...

  6. [2024]

    Algorithmic Pluralism: A Structural Approach To Equal Opportunity

    The International Association of Privacy Profession- als, https://iapp.org/media/pdf/resource_center/us_state_ai_governance_legislation_tracker.pdf ISO (n.d.) Foreword - Supplementary information. https://www.iso.org/foreword-supplementary-information.html ISO/IEC (2022) ISO/IEC International Standard 27001:2022, Information security management systems. h...

  7. [8286]

    Protect, Respect and Remedy

    National Institute of Standards and Technology, https://csrc.nist.gov/ publications/detail/nistir/8286/final Hsuan Su, Cheng-Chu Cheng, Hua Farn, Shachi H Kumar, Saurav Sahay, Shang-Tse Chen, Hung-yi Lee. (2023) Learning from Red Teaming: Gender Bias Provocation and Mitigation in Large Language Models. arXiv, https://arxiv.org/ abs/2310.11079v1 The Associ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.