REVIEW 3 major objections 5 minor 1 cited by
Building Trust: Foundations of Security, Safety and Transparency in AI
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read AI safety hazards can be tracked like software vulnerabilities, with centralized identifiers and exposure statements, this paper argues.
desk verdict A pragmatic, honest proposal to adapt CVE/VEX to AI safety hazards; the core gap is the undefined status semantics, which the paper itself acknowledges. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the pair (CFE identifier, HEX statement), working with an extended model card. A CFE number is a unique, centrally assigned identifier for a recognized AI safety hazard, analogous to a CVE number for software flaws; a HEX statement is a VEX-like machine-readable message that communicates the hazard's status (affected, unaffected, fixed, under investigation) and justification relative to a model's intended use. The model card supplies the declared intent and scope that make both possible: reporters check the card before filing, and consumers check HEX statuses against their own use case. Together these objects convert an amorphous safety concern into a referenceable, comparable, and trackable artifact.
What would settle it
Run an inter-rater reliability trial in which model makers, using a public CFE taxonomy and an extended model card, independently assign HEX statuses to a fixed set of previously reported safety hazards; if agreement on "affected" versus "unaffected" for the same intended use is near chance, the ecosystem fails its transparency function.
Extended reading notes
Core claim
The paper's central claim is that AI Security and AI Safety need separate but coordinated management processes, and that the missing piece is a standardized way to name and communicate safety hazards. It defines an AI security vulnerability as an exploitable flaw affecting confidentiality, integrity, or availability of an AI system, and an AI safety hazard as unexpected model behavior outside the model's defined intent and scope that may cause harm varying by culture and context. On that distinction it builds a proposal: extend model cards with required intent-and-use, scope, evaluation data, and governance fields; create a neutral "Coordinated Hazard Disclosure" body that assigns CFE numbers to safety hazards; and introduce HEX statements, a VEX-style format whose status fields (affected, unaffected, fixed, under investigation) tell consumers whether a hazard impacts their operational use. The paper also proposes an adjunct panel to adjudicate contested hazard claims and a reporting workflow that closes out-of-scope reports, tracks accepted hazards to public advisories, and builds industry knowledge over time.
Load-bearing premise
The load-bearing premise is that safety hazards in AI can be identified, measured, and interpreted consistently enough across contexts for CVE-style identifiers and VEX-style status messages to carry meaning; if harm is too context-relative or too statistical to triage reliably, the CFE/HEX scaffolding loses its value.
Editorial extensions
If this is right
- Public AI models would acquire a safety-hazard record analogous to a vulnerability record, giving consumers a stable identifier to reference when a hazard is reported.
- Model makers would publish HEX statements so an organization could determine, before deployment, whether a known hazard touches its intended use and what status it is in.
- Standardized model cards would make "out of scope" reports resolvable: a report contradicting the declared intent and scope is closed as invalid, while reports within scope proceed to triage and CFE assignment.
- Disputes over whether a hazard exists would go to an adjudication panel, with the statistical validity of submitted samples as the explicit deciding criterion.
- The workflow would route security and safety reports to distinct but coordinated teams, so a prompt-injection attack that also produces harmful content is not dropped between two processes.
Reading between the lines
- Inference: If the CFE/HEX ecosystem matures, a natural next step would be machine-readable attestations signed by model makers, letting deployment pipelines block or gate models automatically based on HEX status—something the paper gestures toward with metadata like AIBOM but does not specify.
- Inference: Because the paper assigns statistical validity thresholds, low-prevalence but severe harms could fall below the reporting bar; a separate channel for one-off incidents would be needed to avoid systematically missing rare harms.
- Inference: A testable extension follows from the paper's own logic: measure inter-rater agreement on hazard classification and HEX status assignment across a culturally diverse panel; if agreement is low, the taxonomy and status justifications need more precision before the ecosystem can function.
- Inference: A single global registry may need per-jurisdiction status values, since the paper acknowledges safety judgments vary across cultures and over time; the HEX status for a given hazard could differ by region.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a position paper on AI security and safety for publicly available models. It distinguishes AI security (technical threats to confidentiality, integrity, and availability) from AI safety (unintended harm arising from model behavior), reviews existing frameworks such as NIST's AI RMF and the EU AI Act, and proposes three interconnected interventions: standardized model/system cards with required fields; a CVE-like Common Flaws and Exposures (CFE) identifier for AI safety hazards, assigned by a neutral body; and a Hazards Exposure eXchange (HEX) format that adapts VEX to convey hazard status to model consumers. The paper also proposes an Adjunct Panel to adjudicate contested hazards. The central claim is that this ecosystem, modeled on software vulnerability disclosure, would improve transparency and risk management for public AI models.
Significance. If the CFE/HEX ecosystem were operationalized, it could address a real and growing gap: there is currently no common mechanism for reporting, tracking, and communicating AI safety hazards across model makers and users. The paper usefully distinguishes security from safety and correctly emphasizes that safety hazards are dynamic and culturally situated. It also gives credit to relevant prior work, notably Cattell et al., and surveys ongoing industry efforts, which helps position the proposal. Its strengths are primarily synthetic: it translates established CVE/VEX practices into the AI safety domain and flags the statistical-validity problem. However, the proposal is at the concept stage; the manuscript contains no pilot, schema, or formal model, and its own text repeatedly defers key definitions to future research. The significance is therefore conditional on the unresolved operationalization of hazard status and CFE assignment.
major comments (3)
- [Adapting VEX] The HEX proposal inherits VEX's discrete status semantics ('affected', 'unaffected', 'fixed', 'under_investigation'), yet the paper defines an AI Safety Hazard as an 'unexpected model behavior' whose impact and severity 'will vary greatly from group to group' and as a dynamic condition on a spectrum that 'may evolve' with societal expectations. No operational rule is given for when a hazard counts as 'affected' or 'fixed' for a given model, for a given use, and at a given time. This is load-bearing because CFE/HEX coordination and comparability depend on consistent status assignments across reporters, model makers, and adjudicators. The sentence 'Additional research is required to develop status justification statements and other potential HEX fields' concedes the gap but does not resolve it; a concrete example, such as how a specific bias hazard would be assigned a HEX status, is needed.
- [Common flaws and exposures (CFEs) for Hazard tracking] The CFE assignment workflow states that a hazard must meet 'statistical validity thresholds' and be established as a safety hazard, but the paper does not specify who defines these thresholds, on what benchmark or evaluation data they are computed, or how they account for the model card's intended use. The paper itself notes that trustworthiness and bias 'often extend beyond the scope of security vulnerabilities,' which is precisely the problem: a VEX-like process presupposes a well-defined vulnerability condition, while a safety hazard's existence is relative to an annotator population and evaluation context. Without a defined unit of analysis and threshold-setting procedure, CFE identifiers cannot be assigned consistently, so the ecosystem's core value of coordination and comparability is not secured.
- [Extending model/system cards] The paper calls for a 'consistent set of minimum fields and content that must be present' in model cards and proposes an 'industry accepted format,' but it provides only illustrative examples (intent/use, scope, evaluation data, governance) rather than a specification. Similarly, the conclusion asserts the proposals 'may provide a shortcut without compromise' to managing AI safety, but no pilot, case study, or worked example is offered. For a paper whose central contribution is a standardization proposal, this is not merely a presentation issue: it leaves the core proposal untestable and makes it difficult for a reader to judge whether the proposed fields and workflows are sufficient or coherent. Adding one worked example, such as a real or realistic model card extended with CFE/HEX entries, would materially improve the paper.
minor comments (5)
- [Common flaws and exposures (CFEs) for Hazard tracking] The phrase 'Common and Flaws and Exposure (CFE)' is a typo; the intended name is presumably 'Common Flaws and Exposures.'
- [Scope of AI Safety flaws] The claim that 'AI is however the first time in which the technology and its development are the cause of violations in trust and safety' is too strong and unsupported; prior technologies such as medical devices, automobiles, and algorithmic content-ranking systems have been designed and operated in ways that caused such violations.
- [Definitions] Several definitions are sourced from 'GENAI Commons' and marked 'Modified for RedHat'; since Red Hat is the authors' employer, the provenance and potential vendor-specific framing of these definitions should be acknowledged more prominently and the sources cited in the reference list.
- [References and footnotes] Many key sources, including the prior work by Cattell et al. cited in footnote 39, appear only as URLs in footnotes rather than in the reference list, which makes it difficult to verify the relationship between this proposal and that prior work.
- [Introduction] The term 'public model' is defined as 'publicly available for download and use,' but later in the conclusion the paper speaks of models 'developed according to open source principles'; these are different conditions, and the distinction should be maintained throughout.
Circularity Check
No circular derivation; the CFE/HEX proposal is an external synthesis, not a self-referential prediction.
full rationale
The paper proposes adapting existing security disclosure processes (CVE/VEX) to AI safety hazards. Its load-bearing elements—CFE as a CVE analogue, HEX as a VEX analogue, and the adjunct dispute panel—are explicitly drawn from external prior work (Sven Cattell et al., arXiv:2402.07039) and industry standards (CISA VEX, CVE, Google model cards), not from the authors' own fitted values or theorems. There is no equation, fitted parameter, or derived quantity whose output equals its input by construction. The only Red Hat-affiliated sources are a topical web page and definitions attributed to GenAI Commons 'Modified for RedHat'; these are not used to justify the central proposal and are not load-bearing. The paper itself flags the open problems: 'Additional research is required to develop status justification statements and other potential HEX fields' and CFE requires 'statistical validity thresholds.' Those are acknowledged limitations in an early-stage proposal, not circular reasoning. Consequently, the derivation chain is self-contained as a policy proposal, and any concerns about practicality belong to correctness risk, not circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption CVE/VEX processes are effective and adaptable to AI safety hazards.
- domain assumption A central neutral body for hazard tracking can be established and gain industry adoption.
- domain assumption Model cards can be standardized with the proposed required fields without hampering model makers.
- domain assumption The proposed definitions of AI Security and AI Safety are accepted as a useful basis.
invented entities (4)
-
CFE (Common Flaws and Exposures)
-
HEX (Hazards Exposure eXchange)
-
Coordinated Hazard Disclosure committee
-
Adjunct Panel
Cite this review
Pith. "Pith review of Building Trust: Foundations of Security, Safety and Transparency in AI." pith.science (2026). https://pith.science/paper/IPSFK5A4
@misc{pith2026241112275,
author = {Pith},
title = {Pith review of: Building Trust: Foundations of Security, Safety and Transparency in AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/IPSFK5A4}},
note = {Machine review of arXiv:2411.12275}
}
read the original abstract
This paper explores the rapidly evolving ecosystem of publicly available AI models, and their potential implications on the security and safety landscape. As AI models become increasingly prevalent, understanding their potential risks and vulnerabilities is crucial. We review the current security and safety scenarios while highlighting challenges such as tracking issues, remediation, and the apparent absence of AI model lifecycle and ownership processes. Comprehensive strategies to enhance security and safety for both model developers and end-users are proposed. This paper aims to provide some of the foundational pieces for more standardized security, safety, and transparency in the development and operation of AI models and the larger open ecosystems and communities forming around them.
Forward citations
Cited by 1 Pith paper
-
Catastrophic Liability: Managing Systemic Risks in Frontier AI Development
Voluntary AI safety standards may already bind frontier labs through US tort law, making thorough safety documentation a liability shield.
Reference graph
Works this paper leans on
-
[1]
GenerativeAdversarial Networks, IanJ. Goodfellow, JeanPouget-Abadie, Mehdi Mirza, BingXu, DavidWarde-Farley, Sherjil Ozair, AaronCourville, YoshuaBengio.https://arxiv.org/abs/1406.2661, 2014AttentionIsAll YouNeed, AshishVaswani, NoamShazeer, Niki Parmar, JakobUszkoreit,LlionJones, AidanN. Gomez, LukaszKaiser, IlliaPolosukhin.https://arxiv.org/abs/1706.037...
arXiv 2016
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.