REVIEW 3 major objections 8 minor 1 cited by
A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety
T0 review · 3 major / 8 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Openness can enhance AI safety by enabling independent scrutiny, decentralized mitigation, and culturally plural oversight, this proceedings paper argues through a mapping of the open-safety toolchain.
desk verdict Useful workshop synthesis with practical tool maps, but the core openness-safety thesis rests on an unexamined net-benefit assumption; worth engaging as a position paper. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism that carries the argument is the 'openness triad': transparent weights, interoperable tooling, and public governance, each linked to a distinct safety channel (independent scrutiny, decentralized mitigation, culturally plural oversight). A second carrying structure is the post-training workflow framework, which maps technical interventions by stage—data, model tuning, online and offline filtering, evaluation, monitoring—for each scoped risk domain. The content-safety-filter comparison adds a taxonomy of safeguards (guardrail LLMs, programmable guardrails, non-LLM ML guardrails) used to identify coverage gaps such as missing multilingual and multimodal filters.
What would settle it
An independent, stratified sample of AI safety researchers—including closed-model developers and members of communities most affected by AI harms—could be asked to rank the five proposed priorities; a materially different ranking would undercut the roadmap's claimed representativeness. A matched-deployment study comparing harm rates and mitigation speed between open-weight and closed-weight AI systems would test the core claim that openness actually enables faster, better safety responses.
Extended reading notes
Core claim
The paper's central claim is that safety in AI systems is better served by openness than by closure, provided openness is understood as a bundle of three properties: transparent weights that allow researchers to inspect and probe models, interoperable tooling that lets developers apply safeguards across the stack, and public governance that brings affected communities into oversight. Each property is paired with a safety mechanism: scrutiny, decentralized mitigation, and cultural pluralism. By mapping post-training interventions—data, tuning, filtering, evaluation, monitoring—against risk domains such as child safety, content safety, bias, privacy, and model integrity, the paper shows where tooling exists and where it does not. Its roadmap names five priorities: participatory inputs, future-proof content filters, ecosystem-wide safety infrastructure, rigorous agentic safeguards, and expanded harm taxonomies.
Load-bearing premise
The load-bearing premise is that the roughly forty-five convening participants, most of whom come from organizations invested in open AI, are representative enough that their qualitative consensus is a reliable basis for a global research agenda, especially since no formal selection criteria or diversity metrics are reported.
Editorial extensions
If this is right
- If openness enhances safety through scrutiny, then restricting open weights should be weighed against the loss of independent inspection and external research.
- The roadmap's five priorities—participatory inputs, future-proof content filters, ecosystem-wide safety infrastructure, agentic safeguards, and expanded harm taxonomies—define where the report says investment and research should go.
- Developers deploying open models currently lack standardized, interoperable safety tooling; filling that gap is framed as a safety measure, not just a convenience.
- Content safety filters must become more controllable, transparent, and multilingual and multimodal, or they will keep producing uneven and biased moderation across languages and identities.
- Agentic systems require trajectory-level and compositional safety evaluation, because individually harmless steps can combine into harmful real-world outcomes.
Reading between the lines
- Editorial inference: the report bundles three different kinds of openness, and the safety case may hold more strongly for some (transparent weights, interoperable tooling) than for others (public governance); future work should test each channel separately.
- Editorial inference: representativeness is the natural pressure point; an independent replication with a broader, less open-AI-leaning participant pool would directly test the roadmap's claimed consensus.
- Editorial inference: if 'safety is a system property' is taken seriously, policy attention should shift from model release alone to deployment context, monitoring, and accountability regimes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a proceedings report from the Columbia Convening on AI Openness and Safety (November 2024), authored by a diverse group of researchers, engineers, and policy leaders. It presents a research agenda at the intersection of open-source/weight AI and safety, a mapping of post-training technical interventions and tooling, a mapping of the content-safety-filter ecosystem, a discussion of agentic-system risks, and a set of participatory approaches. The central claim is that openness—understood as transparent weights, interoperable tooling, and public governance—can enhance AI safety by enabling independent scrutiny, decentralized mitigation, and culturally plural oversight. The report identifies gaps such as scarce multimodal/multilingual benchmarks, limited defenses against prompt injection and compositional attacks in agentic systems, and insufficient participatory mechanisms, and it concludes with five priority research directions. The paper is an updated version of a convening output that reportedly informed the February 2025 French AI Action Summit.
Significance. If the central claim holds, the report provides a useful counterweight to closed-model-centric safety discussions and catalogs many concrete open tools, benchmarks, and gaps in a way that could inform both research and policy. The appendices and tables are a resource for practitioners, and the roadmap is actionable. The paper is transparent about many of its limits (e.g., scoped risks, acknowledged gaps in taxonomies, participant composition). However, the paper does not provide empirical validation or a structured net-benefit analysis for its main claim; it is an expert-consensus position piece rather than a scientific demonstration. Its significance is therefore more as a policy- and research-shaping document than as a technical contribution, and its conclusions need to be read with that caveat.
major comments (3)
- [Abstract; Section 1.1; Section 1.3.2; Table 1; Table 3] The central claim that openness 'can enhance safety' is not supported by a net-benefit argument, and the report scopes out the most direct countervailing mechanism. Table 1 excludes 'deliberate actions by motivated, experienced adversaries' from model-integrity risks, and Section 1.3.2 states that the paper focuses on intentional misuse 'by low-capability users.' Yet open weights are precisely what enables more capable adversaries to fine-tune or ablate safety behavior; Table 3 itself lists 'Tampering Attack Resistance' as a needed intervention, implicitly conceding that current safeguards in open weights are removable. The paper offers no comparison of the safety gains from independent scrutiny, decentralized mitigation, and cultural pluralism against the safety losses from increased removal and misuse capability. To make the claim load-bearing, the authors should either qualify the finding as conditional on specific mitigations (e.g., robust tamper resistance, monitoring, and acceptable-use regimes) or provide evidence or a structured argument that the scrutiny benefits exceed the misuse costs. The abstract's unqualified 'We find that openness can enhance safety' is especially problematic given this unexamined tradeoff.
- [Appendix 4; Section 2 (acknowledgements)] The participant pool appears heavily weighted toward organizations with a direct stake in open AI (e.g., Mozilla, HuggingFace, EleutherAI, AI2, Mistral), and the convening was funded and hosted by Mozilla and Columbia (Section 2). The report does not describe participant selection criteria, diversity metrics, or any effort to balance perspectives from more skeptical parts of the AI-safety community, yet the roadmap is presented as 'community-informed.' This is a limitation on the generalizability of the qualitative consensus that underlies the central claim. Please add a description of the selection process and a statement about the participant pool's composition, or reframe the roadmap explicitly as the considered view of this particular convened group rather than a broader community consensus.
- [Section 1.1 (CISA analogy)] The analogy with open-sourcing dual-use cybersecurity tools is invoked as a key support for the open-safety claim, but the analogy does not directly transfer to open-weight foundation models. Weights are not patchable software components; an adversary with full weights can fine-tune, ablate, or otherwise modify the model to remove safety behavior, and the model is a static artifact once released, not a continuously maintained codebase. The report does not address this disanalogy, which is load-bearing for inferring the cybersecurity consensus to open-weight AI. Please add a discussion of this disanalogy and its implications for the central claim, or reduce the weight placed on the CISA analogy.
minor comments (8)
- [Abstract; Section 1.4] The abstract states that 'these recommendations informed the February 2025 French AI Action Summit,' but no evidence is provided about the nature or extent of that influence; please substantiate or soften this assertion.
- [Section 1.1, footnote 14] The claim that Representation Engineering leads to '+18% improvement and SOTA on TruthfulQA' lacks a citation and appears to be a specific empirical assertion that may quickly become outdated; please add a reference and date or remove the SOTA claim.
- [Section 1.3.2, footnote 27] The note on FraudGPT and WormGPT, saying they 'pretend to help to facilitate cyber attacks but were reported to claim capabilities it didn’t have,' is confusing and grammatically unclear; please rewrite for precision.
- [Table 2] In the Child Safety row, the entry 'PJ' (Perverted Justice) is used without defining the abbreviation on first use; please spell it out.
- [Table 3] The acronym TAR (Tampering Attack Resistance) is used without elaboration or reference in the table; please provide a citation or a parenthetical definition.
- [Section 3.2, Table 8] The 'Genie' row describes a video-based agent but the text emphasizes software engineering; please clarify whether these are different systems or correct the description. Also, the 'Harvey' row is marked 'Open Source: Yes' with the caveat 'Limited info'; this is misleading and should be corrected or removed.
- [Section 4.3, Table 10] The comparison of content-safety filters would be more useful if it included columns for supported languages, modalities, license type, and, where available, reported performance metrics; the current table lists features but no evaluation data.
- [General presentation] The manuscript contains numerous typos, inconsistent spellings (e.g., 'decentralised' vs. 'decentralized'), and formatting irregularities in tables and footnotes; a careful copyedit is needed before publication.
Circularity Check
No load-bearing circularity: the openness-enhances-safety claim rests on convening consensus and external evidence; self-citations to prior Columbia Convening and ROOST are contextual, not the derivation.
full rationale
This paper is a qualitative, participatory policy synthesis rather than a formal derivation. It contains no equations, fitted parameters, uniqueness theorems, or quantitative predictions whose outputs could reduce to their inputs. The central claim that openness can enhance safety is supported by convening consensus, an external CISA analogy in Section 1.1, and external academic risk taxonomies in Appendix 1. Self-citations to the first Columbia Convening (footnotes 10 and 13) and to the ROOST initiative (Section 2.1) provide background or describe the authors' own project; they do not carry the logical weight of the safety argument. The paper also explicitly acknowledges a significant scope limitation in Section 1.3.2 and Table 1, excluding deliberate actions by motivated, experienced adversaries; this is a substantive boundary choice, not a circular step. No reduction of the central claim to its inputs is exhibited, and no cited prior work is invoked to forbid alternatives. The minor self-citations are present but not load-bearing, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Openness is an antidote, not a poison, for AI safety (from the Mozilla open letter, cited in Section 1.1).
- domain assumption Safety is a system property, not a model property (Narayanan & Kapoor, cited in Section 1.1).
- domain assumption Participatory input and collective intelligence improve AI safety outcomes (Section V).
- domain assumption The cybersecurity precedent that open-sourcing defensive tools benefits defenders more than adversaries applies to AI (CISA quote, Section 1.1).
Cite this review
Pith. "Pith review of A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety." pith.science (2026). https://pith.science/paper/QU5W5DKM
@misc{pith2026250622183,
author = {Pith},
title = {Pith review of: A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety},
year = {2026},
howpublished = {\url{https://pith.science/paper/QU5W5DKM}},
note = {Machine review of arXiv:2506.22183}
}
read the original abstract
The rapid rise of open-weight and open-source foundation models is intensifying the obligation and reshaping the opportunity to make AI systems safe. This paper reports outcomes from the Columbia Convening on AI Openness and Safety (San Francisco, 19 Nov 2024) and its six-week preparatory programme involving more than forty-five researchers, engineers, and policy leaders from academia, industry, civil society, and government. Using a participatory, solutions-oriented process, the working groups produced (i) a research agenda at the intersection of safety and open source AI; (ii) a mapping of existing and needed technical interventions and open source tools to safely and responsibly deploy open foundation models across the AI development workflow; and (iii) a mapping of the content safety filter ecosystem with a proposed roadmap for future research and development. We find that openness -- understood as transparent weights, interoperable tooling, and public governance -- can enhance safety by enabling independent scrutiny, decentralized mitigation, and culturally plural oversight. However, significant gaps persist: scarce multimodal and multilingual benchmarks, limited defenses against prompt-injection and compositional attacks in agentic systems, and insufficient participatory mechanisms for communities most affected by AI harms. The paper concludes with a roadmap of five priority research directions, emphasizing participatory inputs, future-proof content filters, ecosystem-wide safety infrastructure, rigorous agentic safeguards, and expanded harm taxonomies. These recommendations informed the February 2025 French AI Action Summit and lay groundwork for an open, plural, and accountable AI safety discipline.
Forward citations
Cited by 1 Pith paper
-
Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns
A few open-weight video models and distribution platforms dominate the creation and spread of NSFW AI video, making developer and platform choices the main intervention points for reducing non-consensual deepfake abuse.
Reference graph
Works this paper leans on
-
[1]
Assess Actions Individually: Each action taken by an agent, such as making a booking, sending a command, or interacting with sensitive data, needs to be scrutinized for potential harm. Safety guards should evaluate each action within its context to ensure compliance with both ethical standards and user intent
-
[2]
AI Safety Is Not a Model Property
Evaluate Compositional Risks: Actions that are harmless individually may become problematic when combined. Consider a banking AI agent with three seemingly harmless abilities: checking account balances (just viewing information), creating transaction templates (saving payment details without execution), and setting up automated payment rules (each followi...
-
[3]
Taxonomy of actions, risks, and mitigations. The developer community would benefit from a clear taxonomy of agentic actions (i.e., searching information, sharing user information, purchasing, etc. ) with their associated risks and possible mitigations
-
[4]
Expanding notions of consequence . The increased complexity of the scope of safety requires redefining the current taxonomy of harms to include real-life consequences of apparently innocuous actions
-
[5]
Many real-world risks stem from misspecified user requests
Human-Computer Interaction insights. Many real-world risks stem from misspecified user requests. More research in Human-Computer Interaction is needed to develop user-agent interaction models that reduce the likelihood of harmful real-world actions
-
[6]
Accountability and liability regimes . Advancing policy research on models of stakeholders’ accountability and liability is key to enforcing safety mitigations
-
[7]
Guardrails. More research on guardrails for agentic systems is needed. This includes expanding beyond current filters narrowly focused on preventing specific outputs to a more holistic approach that ensures the agentic system remains aligned with its goal, to include addressing compositional risks. IV) Content Safety Filters This section outlines both the o...
-
[8]
Planning for unreliable performance: Content safety systems are far from perfect and will result in mistakes, including both false positives (e.g. classifying an output as unsafe when it is not) and false negatives (failing to classify an output as unsafe when it is). Developers should evaluate classifiers with metrics like F1, Precision, Recall, and AUC-R...
Show all 28 references
-
[9]
Black,” “Muslim,
Considering discriminatory biases: Like other ML models, safety classifiers can reflect and amplify societal biases, often over-triggering on content related to marginalized 23 identities. For example, early versions of Google’s Perspective API 91 flagged comments mentioning word...
-
[10]
Toxic Bias: Perspective API Misreads German as More Toxic
Classifiers can complicate the system: Integrating classifiers into a system introduces technical challenges, such as added latency and higher memory usage, which can be particularly problematic for devices like low-memory smartphones. Classifiers can increase the system’s comple...
2023
-
[11]
cookbook
Developing a practical “cookbook” could help guide and empower developers implementing safety filters throughout the deployment process. Currently, a lack of clear guidance on basic safety issues creates friction. Developers face challenges in selecting relevant policies, ident...
-
[12]
Toxic Bias
Larger modality coverage is needed. As highlighted in Table 10, most content safety classifiers are currently designed for text, or text and images, 94 with other modalities or agentic systems mostly out of scope. 94 It’s worth noting a helpful exception here, that of Roblox op...
2023
-
[13]
Better multilingual coverage. While some traditional ML classifiers cover a large language spectrum (Perspective API is available in 18 languages for instance 95 ), most LLM safeguards are available in English only, or in a reduced set of languages as indicated in the overview Table
-
[14]
Several emerging approaches support this direction
Developers seek greater control over filters—including the ability to adjust thresholds, customize policies, and set exceptions . Several emerging approaches support this direction. For example, Google’s agile classifiers framework 96 enables developers to create custom safeguar...
-
[15]
aligned,
Developers want more transparency about filters . As working group participants noted, it is often unclear why a filter blocks a specific input or output. Some approaches, such as autoraters 97 (models that provide explanations for their decisions) can offer a useful foundation fo...
-
[16]
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
To surface valuable but hard-to-intuit insights and to gather a wider range of constructive ideas. Many AI safety challenges involve complex sociotechnical systems where critical information is distributed across diverse stakeholders. Local/contextual knowledge and lived exper...
2024 arXiv
-
[17]
default state
Offsetting conflicts of interest or bias. Complex technical systems benefit from multiple independent layers of review to catch potential failure modes. Different stakeholders bring diverse priorities and risk perceptions, leading to more robust evaluations. Institutional incentiv...
-
[18]
Democratic principles assert that those affected by AI systems should have a voice in their development and oversight
Supporting normative goals around agency and input. Democratic principles assert that those affected by AI systems should have a voice in their development and oversight. Collective input fosters legitimacy and trust by enabling meaningful stakeholder participation in safety de...
-
[19]
Collective input strengthens the broader AI landscape, including public labs and independent researchers
Bolstering a broader ecosystem. Collective input strengthens the broader AI landscape, including public labs and independent researchers. This contributes valuable assets and helps prevent the concentration of knowledge and influence within a few large technology companies, who...
-
[20]
Current evaluation often relies on community-driven leaderboards like LMSYS, which may reflect a narrow notion of quality if the user base lacks diversity
Creating a more diverse leaderboard for model evaluation. Current evaluation often relies on community-driven leaderboards like LMSYS, which may reflect a narrow notion of quality if the user base lacks diversity. To mitigate this bias, mechanisms such as reweighting contributi...
-
[21]
Applying collective intelligence principles to incident reporting can help surface risks that might otherwise be overlooked
Building robust incident reporting systems. Applying collective intelligence principles to incident reporting can help surface risks that might otherwise be overlooked. Engaging 98 In 2023, the private AI investment in the US reached $67.2bn ( https://www.statista.com/statisti...
2023
-
[22]
When inserted directly into a system message, this content can help tailor model behavior to better align with the values and needs of a specific community
System prompts could take the form of raw, open-ended text files that include framing, cultural nuance, and community context. When inserted directly into a system message, this content can help tailor model behavior to better align with the values and needs of a specific community
-
[23]
Fine-tuning with prompt/response pairs that reflect region-specific cultural nuances can produce models that communities find more aligned and trustworthy than the default
Prompts and response pairs. Fine-tuning with prompt/response pairs that reflect region-specific cultural nuances can produce models that communities find more aligned and trustworthy than the default. While generating such pairs is more challenging than drafting a constitution, t...
-
[24]
While useful for regional alignment, this approach introduces an intermediary between the community and the model
Model spec / local soft alignment targets for RLHF or RLAIF: This would involve a set of culturally-informed guidelines that human raters or an LLM could use to judge which responses are more likely to resonate with a specific region. While useful for regional alignment, this a...
-
[25]
This would be a reward preference model that can be integrated into existing RLHF pipelines 100 and bypass the reward model training step
Localized reward models. This would be a reward preference model that can be integrated into existing RLHF pipelines 100 and bypass the reward model training step. However, like RLHF guidelines, they introduce an intermediary layer and may limit direct community control over a...
-
[26]
Non-monetary incentives can overcome motivation crowding out
Encourage and incentivize public institutions to contribute to datasets to help reduce biases and improve fairness in AI models 101 . 5.3 Future Work There are several promising research avenues: 1. Research on incentive mechanisms for diverse communities to contribute to part...
2011
-
[27]
A starting point could be to develop technical systems that ensure informed consent, such as clear opt-in and opt-out options
Ensuring ethical participation: Research is needed to prevent participatory mechanisms from becoming extractive or exploitative forms of free labor. A starting point could be to develop technical systems that ensure informed consent, such as clear opt-in and opt-out options
-
[28]
level 2” categories, 45 “level 3
Steering models with minimal data : Research is needed on how to guide models using small, targeted datasets that reflect local harms and issues specific to a community’s geographic or socio-economic context that are often underrepresented in large-scale training data. APPENDIX ...
2021
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.