Pith. sign in

REVIEW 3 major objections 4 minor 11 references

Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures

T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Most safety measures in the EU's draft GPAI Code already have industry precedent, report finds.

desk verdict Useful new mapping of industry documents to the draft GPAI Code, but the headline 'mostly aligns' claim reads quote counts as stringency and should be tempered. read the letter →

arxiv 2504.15181 v2 pith:V2SJELV7 submitted 2025-04-21 cs.CY cs.AI

classification cs.CYcs.AI
keywords EUAIActGeneral-PurposeCodeofPracticesystemicriskindustryprecedentfrontiersafetyframeworksgovernanceandsecuritymeasures
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This report tries to establish that the Safety and Security section of the EU AI Act's draft General-Purpose AI Code of Practice is largely a codification of practice the industry already publicizes, not a set of novel regulatory burdens. Its central quantitative finding is that 52 of the 72 measures and commitments in Commitments II.1–II.16 had at least three separate company documents containing relevant quotes, and most measures reached the study's cap of five quotes. The comparison is built by reading the draft's safety and security commitments against public frontier-safety frameworks, model cards, system cards, and policy statements from over a dozen major AI developers. The report explicitly does not assess legal compliance or take a prescriptive position; it offers the quote-level evidence base to inform the dialogue between regulators and model providers.

What carries the argument

The central machinery is a systematic quote-matching procedure. For each measure and numbered sub-section, the authors selected up to five relevant excerpts from public-facing documents, prioritizing formal organization-wide safety frameworks, then model or system cards, then other policy statements and technical reports. A measure counted as having industry precedent when at least three separate companies' documents contributed quotes, and five quotes were treated as saturation; the document universe was limited to AI Seoul Summit 2024 signatories with publicly released frontier safety frameworks as of mid-2025.

What would settle it

Re-run the quote-matching exercise on the same 72 measures with independent coders who are blinded to the report's conclusions and who apply a stricter relevance rule: a quote counts only if it commits the company to a concrete procedure with named owners, thresholds, or escalation steps, and not if it merely states a value or intention. If fewer than 52 measures reach three-quote saturation under that rule, the claim that the Code mostly aligns with existing practice would not survive.

Watch

Extended reading notes

Core claim

The report claims that the draft Code of Practice's Safety and Security measures mostly align with existing industry precedent. Its core evidence is that 52 of the 72 analysed measures and commitments had at least three separate company documents with quotes relevant to each sub-section, with most measures reaching saturation at five relevant quotes. The authors present this as suggesting that the Code is mostly aligning with current voluntary industry practice rather than imposing novel regulatory obligations, while cautioning that the report is not an indication of legal compliance and that quote coverage does not mean full compliance.

Load-bearing premise

The entire alignment conclusion rests on treating public-facing company documents—safety frameworks, system cards, and policy statements—as reliable evidence of what companies actually do, rather than as aspirational or communications-oriented documents.

Editorial extensions

If this is right

  • If the Code mostly codifies existing practice, much of the implementation burden for leading providers may consist of formalizing and documenting what they already do, rather than building new safety functions from scratch.
  • The measures and sub-sections with fewer than three supporting quotes—such as parts of risk acceptance determination, serious incident reporting, and some documentation obligations—mark the places where regulators should expect the least industry precedent and may need to offer the most guidance.
  • Because the final Code's Safety and Security section was streamlined from the Third Draft, the report's evidence base remains relevant to the final version even though it was prepared before finalization.
  • Companies building their own governance frameworks can use the mapped quotes as starting templates, since the report documents how peer organizations phrase commitments on systemic risk assessment, mitigation, and transparency.
  • Regulators can use the gap table as a concrete checklist for where voluntary industry practice is thin, focusing capacity-building and enforcement attention on those specific measures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Publicly published safety frameworks may be aspirational rather than fully operational, so the 52-of-72 alignment figure likely overstates how many measures are already implemented in day-to-day engineering practice; a stronger test would verify quotes against internal procedures with named owners and escalation steps.
  • The alignment conclusion could also reflect that leading companies anticipated EU regulation and wrote their voluntary frameworks with the Code in mind, which would mean the Code shaped industry practice rather than merely matching it.
  • An extension of this method would weight quotes by document type, treating binding operational policies as stronger evidence of precedent than stated values or future commitments; such weighting could materially change which measures are counted as having industry precedent.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper compares the Safety and Security commitments (II.1–II.16) of the Third Draft of the EU AI Act's GPAI Code of Practice with public-facing documents from leading AI companies. The authors collect relevant quotes from company safety frameworks, system cards, and related documents; select up to five quotes per measure or sub-section; and report that 52 of 72 measures had at least three separate company documents with relevant quotes. They interpret this statistic as suggesting that the Code of Practice mostly aligns with existing industry precedent rather than imposing novel regulatory burdens. The paper includes a large appendix of quoted excerpts and explicitly disclaims any assessment of legal compliance.

Significance. If the paper's central inference were supported, it would be a valuable input to the EU AI Act implementation debate, providing a structured evidence base of what leading frontier-lab safety documents already cover. The strengths of the manuscript are its transparency about quote selection, the breadth of the document corpus (12 companies, multiple document types), and its repeated, honest caveats that quote counts do not imply compliance. The quote database itself is a useful resource for policymakers and researchers. However, the headline claim that quote coverage implies regulatory alignment is not supported by the method as described, because the method measures topical relevance rather than stringency, scope, or bindingness of commitments. With a reframed claim or an added stringency analysis, the paper would make a solid contribution; in its current form, the central inference overreaches the evidence.

major comments (3)
  1. [Executive Summary / Our approach] The headline inference—that 52 of 72 measures having at least 3 relevant quotes 'suggests that the Codes of Practice mostly aligns with existing industry precedent, rather than imposing novel regulatory burdens'—does not follow from the method described in 'Our approach'. The method selects quotes by topical relevance and checks only that the evidence is relevant to the corresponding item; it does not rate whether the quoted text matches the scope, thresholds, or bindingness of the Code of Practice measure. The paper itself acknowledges that 'a measure containing 3–5 relevant quotes by no means indicates full compliance' and that 'we leave interpretation up to the reader on how closely aligned quotes are with the Code's proposed items.' This internal tension means the main claim is an overreach. I recommend removing the 'mostly aligns / no novel burdens' sentence and presenting the 52/72 statistic as evidence of topical coverage, or adding a stringency rating that assesses whether each quote satisfies the substantive elements of the measure before drawing the alignment inference.
  2. [Our approach / Measure II.2.1] The sample is restricted to companies that have already published frontier safety frameworks, so the measured coverage rate is partly by construction. The paper itself states that Measure II.2.1 is 'trivially satisfied by these 12 companies' precisely because the companies were selected for having frameworks in place. This admission undercuts the comparison of quote counts across measures as evidence of alignment: measures that mirror the standard components of a frontier safety framework would be expected to have many quotes even if companies' actual commitments are weaker than the Code's requirements. Please either analyze a broader sample that includes companies without published frameworks, or explicitly frame the 52/72 statistic as 'coverage among framework-publishing companies' and adjust the Executive Summary accordingly.
  3. [Measure II.1.2 and Measure II.1.3 example quotes] Several quoted documents are aspirational rather than descriptions of existing implemented practice. For example, xAI's Risk Management Framework is a 'draft' that intends to set thresholds 'in a future version of the risk management framework,' and OpenAI's Preparedness Framework v1 says 'we will be building and continually improving suites of evaluations.' Counting such statements as evidence of industry precedent blurs the distinction between implemented practice and stated intention. At a minimum, the paper should classify each quote as implemented vs. planned (or add a marker for draft/planned documents) and show that the 52/72 statistic is robust when restricted to implemented commitments. Without this distinction, the 'existing industry precedent' language is overstated.
minor comments (4)
  1. [Abstract and Executive Summary] The abstract states that relevant quotes were found 'from at least 5 companies' documents for the majority of the measures,' while the Executive Summary reports that 52 of 72 measures had at least 3 quotes and that 'the majority of the measures we analysed reached saturation of 5 relevant quotes.' These are different statistics; please align the phrasing so the reader can trace the abstract claim to the Executive Summary numbers.
  2. [Executive Summary, Table 1 and footnotes] The footnote numbering and the sentence 'This number was deduced by excluding the 20 measures in Table 1' are confusing because Table 1 lists measures with fewer than 3 quotes rather than a count of 20. Please rewrite the footnote so the derivation of 52 is transparent and the table title and footnotes are in the correct order.
  3. [Measure II.1.1] The entry for Measure II.1.1 simply says 'See Commitments II.1-II.16 for relevant quotes,' which is a placeholder rather than an analysis of the measure. Either provide specific quotes relevant to the content of the Framework or state explicitly that the measure is addressed through the subsequent commitments and why no separate quotes are listed.
  4. [Our approach] The paper includes IBM's Responsible Use Guide with the caveat that it may not be a safety framework. For reproducibility, please add a short appendix or table listing, for each company, the specific document(s) used, their publication dates, and whether they are drafted, final, or described as living documents.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the report is a disclosed quote-mapping exercise, not a derivation chain.

full rationale

The report makes no quantitative derivation or fitted prediction; it is a quote-mapping exercise. The central claim ('This suggests that the Codes of Practice mostly aligns with existing industry precedent, rather than imposing novel regulatory burdens') is an interpretive generalization from the count '52 Commitments and Measures had at least 3 separate company documents that were found to have relevant quotes.' This count is an input-level observation, not an output forced by definition: the method selects quotes by relevance and counts coverage, and the paper repeatedly disclaims that quote coverage is not compliance ('a measure containing 3–5 relevant quotes by no means indicates full compliance' and 'We leave interpretation up to the reader on how closely aligned quotes are with the Code's proposed items'). No parameter is fitted and then renamed as a prediction; no uniqueness theorem or prior result by the same authors is invoked to rule out alternatives. The footnote citing SaferAI and METR resources is positioning only and is not load-bearing. The only genuine by-construction element is the self-declared note under II.2.1: 'Given we are soliciting quotes from companies with frameworks already in place, this measure is trivially satisfied by these 12 companies'; this is an honest, localized sampling tautology. II.2.1 is one of 72 measures, and even setting it aside the majority statistic is unaffected; the report's own caveats prevent the statistic from being presented as a demonstration of equivalence. Consequently no circular step meeting the definitional test is found.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on interpretative mappings between public documents and legal measures. There are no free parameters or invented entities. The main assumptions concern representativeness of public documents and the 3-quote threshold used to conclude 'alignment.'

assumptions (3)
  • domain assumption Public company documents (safety frameworks, model cards) reflect actual industry practices and commitments.
    The paper's evidence base is quotes from these documents. They are treated as evidence of industry precedent without independent verification of whether the stated policies are implemented in practice.
  • ad hoc to paper The presence of 3 or more relevant quotes indicates 'industry precedent' for that measure.
    The Executive Summary uses this threshold to separate covered measures from those with 2 or fewer quotes. The number 3 is not derived from any external standard and is presented post hoc.
  • domain assumption Quote selection by author judgment of relevance is a valid method for comparing policy text to corporate practice.
    The selection process is subjective; a second internal review provides consistency but no inter-rater reliability metric is reported. The paper acknowledges quotes are not exhaustive.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures." pith.science (2026). https://pith.science/paper/V2SJELV7

@misc{pith2026250415181,
  author       = {Pith},
  title        = {Pith review of: Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V2SJELV7}},
  note         = {Machine review of arXiv:2504.15181}
}
read the original abstract

This report provides a detailed comparison between the Safety and Security measures proposed in the EU AI Act's General-Purpose AI (GPAI) Code of Practice (Third Draft) and the current commitments and practices voluntarily adopted by leading AI companies. As the EU moves toward enforcing binding obligations for GPAI model providers, the Code of Practice will be key for bridging legal requirements with concrete technical commitments. Our analysis focuses on the draft's Safety and Security section (Commitments II.1-II.16), documenting excerpts from current public-facing documents that are relevant to each individual measure. We systematically reviewed different document types, such as companies' frontier safety frameworks and model cards, from over a dozen companies, including OpenAI, Anthropic, Google DeepMind, Microsoft, Meta, Amazon, and others. This report is not meant to be an indication of legal compliance, nor does it take any prescriptive viewpoint about the Code of Practice or companies' policies. Instead, it aims to inform the ongoing dialogue between regulators and General-Purpose AI model providers by surfacing evidence of industry precedent for various measures. Nonetheless, we were able to find relevant quotes from at least 5 companies' documents for the majority of the measures in Commitments II.1-II.16.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

11 extracted references · 6 canonical work pages

  1. [3]

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    DeepSeek. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. Retrieved from https://arxiv.org/pdf/2501.12948 (Published January 22,

  2. [4]

    A Framework for Evaluating Emerging Cyberattack Capabilities of AI

    Google DeepMind. A Framework for Evaluating Emerging Cyberattack Capabilities of AI . Retrieved from https://arxiv.org/pdf/2503.11917 (Published April 1,

  3. [5]

    An Approach to Technical AGI Safety and Security

    Google DeepMind. An Approach to Technical AGI Safety and Security . Retrieved from https://arxiv.org/pdf/2504.01849 (Published April 2,

  4. [6]

    Evaluating Frontier Models for Dangerous Capabilities

    164 Mapping Industry Practices to the Code of Practice Google DeepMind. Evaluating Frontier Models for Dangerous Capabilities . Retrieved from https://arxiv.org/pdf/2403.13793 (Published April 8,

  5. [7]

    Frontier AI Framework

    Meta. Frontier AI Framework . Retrieved from https://ai.meta.com/static-resource/meta-frontier-ai-framework Microsoft. Staying ahead of threat actors in the age of AI. Retrieved from https://www.microsoft.com/en-us/security/blog/2024/02/14/staying-ahead-of-threat-actors-in-the-age-of-ai/ (Published February 14,

  6. [8]

    Staying Ahead of Threat Actors in the Age of AI

    Microsoft Threat Intelligence . Staying Ahead of Threat Actors in the Age of AI. Retrieved from https://www.microsoft.com/en-us/security/blog/2024/02/14/staying-ahead-of-threat-actors-in-the-age-of-ai/ (Published February 14,

  7. [9]

    MLE-Bench: Evaluating Machine Learning Agents on Machine Learning Engineering

    OpenAI. MLE-Bench: Evaluating Machine Learning Agents on Machine Learning Engineering. Retrieved from https://arxiv.org/pdf/2410.07095 (Published October 9,

  8. [10]

    OpenAI Model Spec

    OpenAI. OpenAI Model Spec. Retrieved from https://model-spec.openai.com/2025-04-11.html (Published 11 April,

Show all 11 references
  1. [11]

    Researcher Access Program

    OpenAI. Researcher Access Program . Retrieved from https://openai.com/form/researcher-access-program/ xAI. Risk Management Framework (Draft) . Retrieved from https://x.ai/documents/2025.02.20-RMF-Draft.pdf (Published February 20,

  2. [2024]

    Secure AI Frontier Model Framework

    Cohere. Secure AI Frontier Model Framework . Retrieved from https://cohere.com/security/the-cohere-secure-ai-frontier-model-framework-february-2025.pdf (Published February 11,

  3. [2025]

    External Researcher Access Program

    Anthropic. External Researcher Access Program . Retrieved from https://support.anthropic.com/en/articles/9125743-what-is-the-external-researcher-access-program Anthropic. Responsible Disclosure Policy . Retrieved from https://www.anthropic.com/responsible-disclosure-policy (Pu...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.