{"id":"ec9fc02a-a4ef-4316-9422-98ce0d7c9da5","arxiv_id":"2504.15181","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors found that at least 52 of 72 safety and security measures had relevant quotes from three or more company documents, which they read as evidence that the Code mostly aligns with existing industry precedent.","lead":"This report maps the EU AI Act's draft safety and security Code of Practice measures to voluntary public documents from leading AI companies, finding relevant quotes for most measures. It gives regulators and companies an evidence base for assessing whether the Code imposes novel burdens or codifies existing practice.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 52/72 statistic measures topical quote coverage, not equivalence of commitment; supporting 'mostly aligns / no novel burdens' requires a stringency test the paper does not run.","rationale":"The paper is a valuable and careful evidence compilation: it documents quotes systematically, includes a table of low-coverage measures, and repeatedly caveats that it is not a compliance assessment. The reader's verdict of CONDITIONAL is reasonable. However, the single most load-bearing weakness is not just that documents may be aspirational or that quote selection is subjective; it is that the reported statistic is not evidence for the specific conclusion drawn. 'Relevant quote' is a topical-relevance judgment; 'aligns with existing industry precedent' is a substantive equivalence judgment between a draft regulatory obligation and an industry practice. The latter requires showing that the quote commits to the same level of rigor (e.g., quantitative forecasts, specific thresholds, binding decision procedures), not merely that it addresses the same topic. The paper contains internal evidence of this gap: the II.1.3 forecasting measure is supported by quotes that are informal ('we will aim to improve these forecasts') or describe commissioned forecasters without committing to timeline estimates; II.2.1 is admitted to be 'trivially satisfied' because the sample already consists of companies with frameworks. Thus the central claim overreaches the method. This does not undermine the underlying mapping's utility, but it does require a revision of the headline conclusion, which is exactly the conditional the reader placed. I therefore recommend keeping the verdict as CONDITIONAL (UNCHANGED relative to the reader), with the explicit condition that the Executive Summary reframe 'mostly aligns' as 'topical overlap in public-facing safety documentation.' The proposed stringency re-rating would settle whether even the reframed claim is empirically robust.","tokens_in":59397,"tokens_out":3798,"duration_ms":36941,"concrete_test":"Pre-register a stringency rubric with three levels: (A) topical mention only; (B) partial commitment (some elements of the measure, but weaker scope, thresholds, or accountability); (C) commitment matching the measure's core normative content. Draw a stratified random sample of 20 measures from the 52 reported as covered. Three independent raters, blind to the paper's labels, classify every cited quote for each sampled measure. Define 'alignment' as at least three companies having at least one C-level quote per measure. Recompute the headline coverage statistic. If fewer than half of the sampled measures meet the C-level threshold, revise the Executive Summary to 'substantial topical overlap in public documentation' and remove the 'rather than imposing novel regulatory burdens' claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Executive Summary claims: '52 Commitments and Measures had at least 3 separate company documents that were found to have relevant quotes for each sub-section... This suggests that the Codes of Practice mostly aligns with existing industry precedent, rather than imposing novel regulatory burdens.' The method ('Our approach') selects up to 5 quotes per measure by relevance and retains quotes that are contextually related; quality assurance checks relevance, not stringency. The report itself disclaims compliance and says 3–5 quotes 'by no means indicates full compliance.' So the headline inference treats topical overlap as alignment. Two aggravating factors: (1) the sample is limited to companies that already published frontier safety frameworks, so coverage is partly by construction—the paper even calls II.2.1 'trivially satisfied' for these 12 companies; (2) several cited quotes are explicitly aspirational (e.g., xAI's 'draft' framework with thresholds 'in a future version'; OpenAI's Preparedness Framework v1 'we will be building ... evaluations'). The central claim therefore rests on an unmeasured assumption: relevant quotes imply comparable commitments. Without a rating of whether quotes match the scope, threshold, and bindingness of each CoP measure, 52/72 supports 'topical coverage,' not 'mostly aligns / no novel burdens.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares the Safety and Security commitments (II.1–II.16) of the Third Draft of the EU AI Act's GPAI Code of Practice with public-facing documents from leading AI companies. The authors collect relevant quotes from company safety frameworks, system cards, and related documents; select up to five quotes per measure or sub-section; and report that 52 of 72 measures had at least three separate company documents with relevant quotes. They interpret this statistic as suggesting that the Code of Practice mostly aligns with existing industry precedent rather than imposing novel regulatory burdens. The paper includes a large appendix of quoted excerpts and explicitly disclaims any assessment of legal compliance.","tokens_in":59634,"tokens_out":3461,"duration_ms":34164,"significance":"If the paper's central inference were supported, it would be a valuable input to the EU AI Act implementation debate, providing a structured evidence base of what leading frontier-lab safety documents already cover. The strengths of the manuscript are its transparency about quote selection, the breadth of the document corpus (12 companies, multiple document types), and its repeated, honest caveats that quote counts do not imply compliance. The quote database itself is a useful resource for policymakers and researchers. However, the headline claim that quote coverage implies regulatory alignment is not supported by the method as described, because the method measures topical relevance rather than stringency, scope, or bindingness of commitments. With a reframed claim or an added stringency analysis, the paper would make a solid contribution; in its current form, the central inference overreaches the evidence.","major_comments":[{"comment":"The headline inference—that 52 of 72 measures having at least 3 relevant quotes 'suggests that the Codes of Practice mostly aligns with existing industry precedent, rather than imposing novel regulatory burdens'—does not follow from the method described in 'Our approach'. The method selects quotes by topical relevance and checks only that the evidence is relevant to the corresponding item; it does not rate whether the quoted text matches the scope, thresholds, or bindingness of the Code of Practice measure. The paper itself acknowledges that 'a measure containing 3–5 relevant quotes by no means indicates full compliance' and that 'we leave interpretation up to the reader on how closely aligned quotes are with the Code's proposed items.' This internal tension means the main claim is an overreach. I recommend removing the 'mostly aligns / no novel burdens' sentence and presenting the 52/72 statistic as evidence of topical coverage, or adding a stringency rating that assesses whether each quote satisfies the substantive elements of the measure before drawing the alignment inference.","section":"Executive Summary / Our approach"},{"comment":"The sample is restricted to companies that have already published frontier safety frameworks, so the measured coverage rate is partly by construction. The paper itself states that Measure II.2.1 is 'trivially satisfied by these 12 companies' precisely because the companies were selected for having frameworks in place. This admission undercuts the comparison of quote counts across measures as evidence of alignment: measures that mirror the standard components of a frontier safety framework would be expected to have many quotes even if companies' actual commitments are weaker than the Code's requirements. Please either analyze a broader sample that includes companies without published frameworks, or explicitly frame the 52/72 statistic as 'coverage among framework-publishing companies' and adjust the Executive Summary accordingly.","section":"Our approach / Measure II.2.1"},{"comment":"Several quoted documents are aspirational rather than descriptions of existing implemented practice. For example, xAI's Risk Management Framework is a 'draft' that intends to set thresholds 'in a future version of the risk management framework,' and OpenAI's Preparedness Framework v1 says 'we will be building and continually improving suites of evaluations.' Counting such statements as evidence of industry precedent blurs the distinction between implemented practice and stated intention. At a minimum, the paper should classify each quote as implemented vs. planned (or add a marker for draft/planned documents) and show that the 52/72 statistic is robust when restricted to implemented commitments. Without this distinction, the 'existing industry precedent' language is overstated.","section":"Measure II.1.2 and Measure II.1.3 example quotes"}],"minor_comments":[{"comment":"The abstract states that relevant quotes were found 'from at least 5 companies' documents for the majority of the measures,' while the Executive Summary reports that 52 of 72 measures had at least 3 quotes and that 'the majority of the measures we analysed reached saturation of 5 relevant quotes.' These are different statistics; please align the phrasing so the reader can trace the abstract claim to the Executive Summary numbers.","section":"Abstract and Executive Summary"},{"comment":"The footnote numbering and the sentence 'This number was deduced by excluding the 20 measures in Table 1' are confusing because Table 1 lists measures with fewer than 3 quotes rather than a count of 20. Please rewrite the footnote so the derivation of 52 is transparent and the table title and footnotes are in the correct order.","section":"Executive Summary, Table 1 and footnotes"},{"comment":"The entry for Measure II.1.1 simply says 'See Commitments II.1-II.16 for relevant quotes,' which is a placeholder rather than an analysis of the measure. Either provide specific quotes relevant to the content of the Framework or state explicitly that the measure is addressed through the subsequent commitments and why no separate quotes are listed.","section":"Measure II.1.1"},{"comment":"The paper includes IBM's Responsible Use Guide with the caveat that it may not be a safety framework. For reproducibility, please add a short appendix or table listing, for each company, the specific document(s) used, their publication dates, and whether they are drafted, final, or described as living documents.","section":"Our approach"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful evidence base, but the headline inference is being used in policy discussions and needs to be brought in line with the method. I am not asking for a new empirical study; a reframing of the Executive Summary claim, or a simple classification of quotes as implemented vs. planned, would address the core concern. Please also consider that the selection of companies with published frameworks makes the coverage statistic partly self-fulfilling, and that point should be acknowledged directly in the Executive Summary."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is the first direct measure-by-measure mapping I've seen between public industry safety documents and the Safety & Security section of the third draft GPAI Code of Practice. If you need to know what industry precedent exists for a given commitment, this is the fastest route. The work is transparent, systematically organized, and unusually honest about many of its own limits. The authors reproduce each measure, list up to five relevant quotes per sub-section, flag measures with sparse coverage, and explicitly say the report is not a compliance assessment and that 3–5 quotes 'by no means indicates full compliance.' That is real credit where it is due.\n\nThe soft spot is real and load-bearing, but it is confined to the interpretation, not the underlying compilation. The Executive Summary says that finding relevant quotes for 52 of 72 measures 'suggests that the Codes of Practice mostly aligns with existing industry precedent, rather than imposing novel regulatory burdens.' That inference does not follow from the method. The quotes were selected for topical relevance, not for whether the company commitment matches the scope, thresholds, and bindingness of each Code measure. So 52/72 is a count of topical coverage, not a measure of alignment. The 3-quote threshold is arbitrary. The sample is restricted to companies that already published frontier safety frameworks, so coverage is partly by construction—the paper even calls II.2.1 'trivially satisfied' for that group. And some of the supporting quotes are explicitly aspirational, such as xAI's draft framework with thresholds to be set 'in a future version' and OpenAI's Preparedness Framework v1 saying 'we will be building' evaluations. The paper's own caveats partially cover this, but the Executive Summary still overstates.\n\nWho this is for: policymakers and companies needing a rapid evidence base during the Code finalization, and researchers mapping governance discourse. It is a policy evidence compilation, not a scientific claim, and the counting is descriptive. I would cite it for the mapping, but not for the 'mostly aligns' conclusion without qualification. It deserves a serious referee—a good reviewer can help separate coverage from stringency and get the Executive Summary to match the actual method.","headline":"Useful new mapping of industry documents to the draft GPAI Code, but the headline 'mostly aligns' claim reads quote counts as stringency and should be tempered.","tokens_in":60127,"tokens_out":1505,"would_cite":true,"duration_ms":18050,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Most safety measures in the EU's draft GPAI Code already have industry precedent, report finds.","keywords":["EU AI Act","General-Purpose AI","Code of Practice","systemic risk","industry precedent","frontier safety frameworks","AI governance","safety and security measures"],"falsifier":"Re-run the quote-matching exercise on the same 72 measures with independent coders who are blinded to the report's conclusions and who apply a stricter relevance rule: a quote counts only if it commits the company to a concrete procedure with named owners, thresholds, or escalation steps, and not if it merely states a value or intention. If fewer than 52 measures reach three-quote saturation under that rule, the claim that the Code mostly aligns with existing practice would not survive.","tokens_in":59226,"feed_emoji":"⚖️","tokens_out":4231,"duration_ms":42718,"temperature":0.7,"pith_summary":"This report tries to establish that the Safety and Security section of the EU AI Act's draft General-Purpose AI Code of Practice is largely a codification of practice the industry already publicizes, not a set of novel regulatory burdens. Its central quantitative finding is that 52 of the 72 measures and commitments in Commitments II.1–II.16 had at least three separate company documents containing relevant quotes, and most measures reached the study's cap of five quotes. The comparison is built by reading the draft's safety and security commitments against public frontier-safety frameworks, model cards, system cards, and policy statements from over a dozen major AI developers. The report explicitly does not assess legal compliance or take a prescriptive position; it offers the quote-level evidence base to inform the dialogue between regulators and model providers.","feed_headline":"Draft EU AI safety code mostly matches industry practice, report finds","feed_subtitle":"A quote-level read of 72 draft measures finds 52 backed by at least three company documents, easing fears of brand-new rules.","key_machinery":"The central machinery is a systematic quote-matching procedure. For each measure and numbered sub-section, the authors selected up to five relevant excerpts from public-facing documents, prioritizing formal organization-wide safety frameworks, then model or system cards, then other policy statements and technical reports. A measure counted as having industry precedent when at least three separate companies' documents contributed quotes, and five quotes were treated as saturation; the document universe was limited to AI Seoul Summit 2024 signatories with publicly released frontier safety frameworks as of mid-2025.","core_discovery":"The report claims that the draft Code of Practice's Safety and Security measures mostly align with existing industry precedent. Its core evidence is that 52 of the 72 analysed measures and commitments had at least three separate company documents with quotes relevant to each sub-section, with most measures reaching saturation at five relevant quotes. The authors present this as suggesting that the Code is mostly aligning with current voluntary industry practice rather than imposing novel regulatory obligations, while cautioning that the report is not an indication of legal compliance and that quote coverage does not mean full compliance.","pith_inferences":["Publicly published safety frameworks may be aspirational rather than fully operational, so the 52-of-72 alignment figure likely overstates how many measures are already implemented in day-to-day engineering practice; a stronger test would verify quotes against internal procedures with named owners and escalation steps.","The alignment conclusion could also reflect that leading companies anticipated EU regulation and wrote their voluntary frameworks with the Code in mind, which would mean the Code shaped industry practice rather than merely matching it.","An extension of this method would weight quotes by document type, treating binding operational policies as stronger evidence of precedent than stated values or future commitments; such weighting could materially change which measures are counted as having industry precedent."],"forward_implications":["If the Code mostly codifies existing practice, much of the implementation burden for leading providers may consist of formalizing and documenting what they already do, rather than building new safety functions from scratch.","The measures and sub-sections with fewer than three supporting quotes—such as parts of risk acceptance determination, serious incident reporting, and some documentation obligations—mark the places where regulators should expect the least industry precedent and may need to offer the most guidance.","Because the final Code's Safety and Security section was streamlined from the Third Draft, the report's evidence base remains relevant to the final version even though it was prepared before finalization.","Companies building their own governance frameworks can use the mapped quotes as starting templates, since the report documents how peer organizations phrase commitments on systemic risk assessment, mitigation, and transparency.","Regulators can use the gap table as a concrete checklist for where voluntary industry practice is thin, focusing capacity-building and enforcement attention on those specific measures."],"supporting_citations":[{"why":"Supplies the legal backdrop: Articles 52–55 and Article 56 of the AI Act define the obligations for GPAI models with systemic risk and make the Code of Practice the default means of demonstrating compliance.","marker":"[4]"},{"why":"Identifies prior analyses of frontier AI risk-management practice and states that none provide a direct measure-by-measure comparison with the General-Purpose AI Code of Practice, establishing the report's specific contribution.","marker":"[5]"},{"why":"Defines the sample frame: companies covered were AI Seoul Summit 2024 signatories with publicly released frontier safety frameworks, with some named signatories excluded for lacking such frameworks.","marker":"[6]"},{"why":"Describes the quality-assurance step in which each section was independently reviewed to confirm that selected quotes were relevant to the corresponding Code measure.","marker":"[7]"}],"fun_headline_variants":["EU AI code aligns with industry: 52 of 72 measures","52 of 72 EU AI code measures already industry practice","Most EU AI safety code measures mirror industry practice"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire alignment conclusion rests on treating public-facing company documents—safety frameworks, system cards, and policy statements—as reliable evidence of what companies actually do, rather than as aspirational or communications-oriented documents.","fun_headline_variants_meta":{"raw":{"variants":["EU AI code aligns with industry: 52 of 72 measures","52 of 72 EU AI code measures already industry practice","Most EU AI safety code measures mirror industry practice"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000838,"raw_usage":{"total_tokens":3620,"prompt_tokens":881,"completion_tokens":2739,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":2686}},"tokens_in":497,"tokens_out":2739,"duration_ms":16297,"temperature":1.0,"reasoning_tokens":2686,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:30:29.037037+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the quote-matching exercise on the same 72 measures with independent coders who are blinded to the report's conclusions and who apply a stricter relevance rule: a quote counts only if it commits the company to a concrete procedure with named owners, thresholds, or escalation steps, and not if it merely states a value or intention. If fewer than 52 measures reach three-quote saturation under that rule, the claim that the Code mostly aligns with existing practice would not survive.","supporting_citations":[{"cited_title":"Frontier AI Framework","cited_arxiv_id":null,"evidence_quote":"Describes the quality-assurance step in which each section was independently reviewed to confirm that selected quotes were relevant to the corresponding Code measure."}],"review_version":1}