Pith. sign in

REVIEW 3 major objections 8 minor 1 cited by

A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety

T0 review · 3 major / 8 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Openness can enhance AI safety by enabling independent scrutiny, decentralized mitigation, and culturally plural oversight, this proceedings paper argues through a mapping of the open-safety toolchain.

desk verdict Useful workshop synthesis with practical tool maps, but the core openness-safety thesis rests on an unexamined net-benefit assumption; worth engaging as a position paper. read the letter →

arxiv 2506.22183 v1 pith:QU5W5DKM submitted 2025-06-27 cs.AI

classification cs.AI
keywords open-sourceAIsafetycontentfiltersagenticsystemsparticipatorygovernancerisktaxonomyopen-weightmodelsmultilingualbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This proceedings paper argues that openness in AI—transparent weights, interoperable tooling, and public governance—can make AI systems safer rather than riskier. Drawing on a six-week working-group process and a day-long convening of more than forty-five researchers, engineers, and policy leaders, it maps where safety interventions and open tools exist across the AI development workflow and where they are missing. The central finding is that open ecosystems enable independent scrutiny, decentralized mitigation, and culturally plural oversight, yet major gaps remain: scarce multilingual and multimodal benchmarks, weak defenses against prompt injection and compositional attacks in agentic systems, and insufficient participatory mechanisms for the communities most affected by AI harms. The paper converts these findings into five priority research directions intended to feed into international AI policy discussions and to lay groundwork for an open, plural, and accountable AI safety discipline.

What carries the argument

The mechanism that carries the argument is the 'openness triad': transparent weights, interoperable tooling, and public governance, each linked to a distinct safety channel (independent scrutiny, decentralized mitigation, culturally plural oversight). A second carrying structure is the post-training workflow framework, which maps technical interventions by stage—data, model tuning, online and offline filtering, evaluation, monitoring—for each scoped risk domain. The content-safety-filter comparison adds a taxonomy of safeguards (guardrail LLMs, programmable guardrails, non-LLM ML guardrails) used to identify coverage gaps such as missing multilingual and multimodal filters.

What would settle it

An independent, stratified sample of AI safety researchers—including closed-model developers and members of communities most affected by AI harms—could be asked to rank the five proposed priorities; a materially different ranking would undercut the roadmap's claimed representativeness. A matched-deployment study comparing harm rates and mitigation speed between open-weight and closed-weight AI systems would test the core claim that openness actually enables faster, better safety responses.

Watch

Extended reading notes

Core claim

The paper's central claim is that safety in AI systems is better served by openness than by closure, provided openness is understood as a bundle of three properties: transparent weights that allow researchers to inspect and probe models, interoperable tooling that lets developers apply safeguards across the stack, and public governance that brings affected communities into oversight. Each property is paired with a safety mechanism: scrutiny, decentralized mitigation, and cultural pluralism. By mapping post-training interventions—data, tuning, filtering, evaluation, monitoring—against risk domains such as child safety, content safety, bias, privacy, and model integrity, the paper shows where tooling exists and where it does not. Its roadmap names five priorities: participatory inputs, future-proof content filters, ecosystem-wide safety infrastructure, rigorous agentic safeguards, and expanded harm taxonomies.

Load-bearing premise

The load-bearing premise is that the roughly forty-five convening participants, most of whom come from organizations invested in open AI, are representative enough that their qualitative consensus is a reliable basis for a global research agenda, especially since no formal selection criteria or diversity metrics are reported.

Editorial extensions

If this is right

  • If openness enhances safety through scrutiny, then restricting open weights should be weighed against the loss of independent inspection and external research.
  • The roadmap's five priorities—participatory inputs, future-proof content filters, ecosystem-wide safety infrastructure, agentic safeguards, and expanded harm taxonomies—define where the report says investment and research should go.
  • Developers deploying open models currently lack standardized, interoperable safety tooling; filling that gap is framed as a safety measure, not just a convenience.
  • Content safety filters must become more controllable, transparent, and multilingual and multimodal, or they will keep producing uneven and biased moderation across languages and identities.
  • Agentic systems require trajectory-level and compositional safety evaluation, because individually harmless steps can combine into harmful real-world outcomes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the report bundles three different kinds of openness, and the safety case may hold more strongly for some (transparent weights, interoperable tooling) than for others (public governance); future work should test each channel separately.
  • Editorial inference: representativeness is the natural pressure point; an independent replication with a broader, less open-AI-leaning participant pool would directly test the roadmap's claimed consensus.
  • Editorial inference: if 'safety is a system property' is taken seriously, policy attention should shift from model release alone to deployment context, monitoring, and accountability regimes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. This manuscript is a proceedings report from the Columbia Convening on AI Openness and Safety (November 2024), authored by a diverse group of researchers, engineers, and policy leaders. It presents a research agenda at the intersection of open-source/weight AI and safety, a mapping of post-training technical interventions and tooling, a mapping of the content-safety-filter ecosystem, a discussion of agentic-system risks, and a set of participatory approaches. The central claim is that openness—understood as transparent weights, interoperable tooling, and public governance—can enhance AI safety by enabling independent scrutiny, decentralized mitigation, and culturally plural oversight. The report identifies gaps such as scarce multimodal/multilingual benchmarks, limited defenses against prompt injection and compositional attacks in agentic systems, and insufficient participatory mechanisms, and it concludes with five priority research directions. The paper is an updated version of a convening output that reportedly informed the February 2025 French AI Action Summit.

Significance. If the central claim holds, the report provides a useful counterweight to closed-model-centric safety discussions and catalogs many concrete open tools, benchmarks, and gaps in a way that could inform both research and policy. The appendices and tables are a resource for practitioners, and the roadmap is actionable. The paper is transparent about many of its limits (e.g., scoped risks, acknowledged gaps in taxonomies, participant composition). However, the paper does not provide empirical validation or a structured net-benefit analysis for its main claim; it is an expert-consensus position piece rather than a scientific demonstration. Its significance is therefore more as a policy- and research-shaping document than as a technical contribution, and its conclusions need to be read with that caveat.

major comments (3)
  1. [Abstract; Section 1.1; Section 1.3.2; Table 1; Table 3] The central claim that openness 'can enhance safety' is not supported by a net-benefit argument, and the report scopes out the most direct countervailing mechanism. Table 1 excludes 'deliberate actions by motivated, experienced adversaries' from model-integrity risks, and Section 1.3.2 states that the paper focuses on intentional misuse 'by low-capability users.' Yet open weights are precisely what enables more capable adversaries to fine-tune or ablate safety behavior; Table 3 itself lists 'Tampering Attack Resistance' as a needed intervention, implicitly conceding that current safeguards in open weights are removable. The paper offers no comparison of the safety gains from independent scrutiny, decentralized mitigation, and cultural pluralism against the safety losses from increased removal and misuse capability. To make the claim load-bearing, the authors should either qualify the finding as conditional on specific mitigations (e.g., robust tamper resistance, monitoring, and acceptable-use regimes) or provide evidence or a structured argument that the scrutiny benefits exceed the misuse costs. The abstract's unqualified 'We find that openness can enhance safety' is especially problematic given this unexamined tradeoff.
  2. [Appendix 4; Section 2 (acknowledgements)] The participant pool appears heavily weighted toward organizations with a direct stake in open AI (e.g., Mozilla, HuggingFace, EleutherAI, AI2, Mistral), and the convening was funded and hosted by Mozilla and Columbia (Section 2). The report does not describe participant selection criteria, diversity metrics, or any effort to balance perspectives from more skeptical parts of the AI-safety community, yet the roadmap is presented as 'community-informed.' This is a limitation on the generalizability of the qualitative consensus that underlies the central claim. Please add a description of the selection process and a statement about the participant pool's composition, or reframe the roadmap explicitly as the considered view of this particular convened group rather than a broader community consensus.
  3. [Section 1.1 (CISA analogy)] The analogy with open-sourcing dual-use cybersecurity tools is invoked as a key support for the open-safety claim, but the analogy does not directly transfer to open-weight foundation models. Weights are not patchable software components; an adversary with full weights can fine-tune, ablate, or otherwise modify the model to remove safety behavior, and the model is a static artifact once released, not a continuously maintained codebase. The report does not address this disanalogy, which is load-bearing for inferring the cybersecurity consensus to open-weight AI. Please add a discussion of this disanalogy and its implications for the central claim, or reduce the weight placed on the CISA analogy.
minor comments (8)
  1. [Abstract; Section 1.4] The abstract states that 'these recommendations informed the February 2025 French AI Action Summit,' but no evidence is provided about the nature or extent of that influence; please substantiate or soften this assertion.
  2. [Section 1.1, footnote 14] The claim that Representation Engineering leads to '+18% improvement and SOTA on TruthfulQA' lacks a citation and appears to be a specific empirical assertion that may quickly become outdated; please add a reference and date or remove the SOTA claim.
  3. [Section 1.3.2, footnote 27] The note on FraudGPT and WormGPT, saying they 'pretend to help to facilitate cyber attacks but were reported to claim capabilities it didn’t have,' is confusing and grammatically unclear; please rewrite for precision.
  4. [Table 2] In the Child Safety row, the entry 'PJ' (Perverted Justice) is used without defining the abbreviation on first use; please spell it out.
  5. [Table 3] The acronym TAR (Tampering Attack Resistance) is used without elaboration or reference in the table; please provide a citation or a parenthetical definition.
  6. [Section 3.2, Table 8] The 'Genie' row describes a video-based agent but the text emphasizes software engineering; please clarify whether these are different systems or correct the description. Also, the 'Harvey' row is marked 'Open Source: Yes' with the caveat 'Limited info'; this is misleading and should be corrected or removed.
  7. [Section 4.3, Table 10] The comparison of content-safety filters would be more useful if it included columns for supported languages, modalities, license type, and, where available, reported performance metrics; the current table lists features but no evaluation data.
  8. [General presentation] The manuscript contains numerous typos, inconsistent spellings (e.g., 'decentralised' vs. 'decentralized'), and formatting irregularities in tables and footnotes; a careful copyedit is needed before publication.

Circularity Check

0 steps flagged · score 2.0 of 10

No load-bearing circularity: the openness-enhances-safety claim rests on convening consensus and external evidence; self-citations to prior Columbia Convening and ROOST are contextual, not the derivation.

full rationale

This paper is a qualitative, participatory policy synthesis rather than a formal derivation. It contains no equations, fitted parameters, uniqueness theorems, or quantitative predictions whose outputs could reduce to their inputs. The central claim that openness can enhance safety is supported by convening consensus, an external CISA analogy in Section 1.1, and external academic risk taxonomies in Appendix 1. Self-citations to the first Columbia Convening (footnotes 10 and 13) and to the ROOST initiative (Section 2.1) provide background or describe the authors' own project; they do not carry the logical weight of the safety argument. The paper also explicitly acknowledges a significant scope limitation in Section 1.3.2 and Table 1, excluding deliberate actions by motivated, experienced adversaries; this is a substantive boundary choice, not a circular step. No reduction of the central claim to its inputs is exhibited, and no cited prior work is invoked to forbid alternatives. The minor self-citations are present but not load-bearing, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The report's central claim rests on normative and analogical assumptions about openness, safety, and participation rather than on derived or measured quantities. No free parameters are fitted to data, and no new theoretical entities are introduced. The assumptions are domain-level premises that frame the analysis, and each is supported primarily by citation or stated values. The lack of empirical validation for these assumptions is the main epistemic weakness.

assumptions (4)
  • domain assumption Openness is an antidote, not a poison, for AI safety (from the Mozilla open letter, cited in Section 1.1).
    The paper adopts this stance as a foundational premise rather than establishing it empirically. It is used to frame the entire report.
  • domain assumption Safety is a system property, not a model property (Narayanan & Kapoor, cited in Section 1.1).
    This premise justifies focusing on the broader AI stack rather than just model weights. It is taken as given and used to structure the tool mapping.
  • domain assumption Participatory input and collective intelligence improve AI safety outcomes (Section V).
    The report argues for participatory methods based on democratic principles and anecdotal evidence, but provides no controlled study showing that participation reduces harm.
  • domain assumption The cybersecurity precedent that open-sourcing defensive tools benefits defenders more than adversaries applies to AI (CISA quote, Section 1.1).
    The report generalizes from software security to AI safety without testing whether the analogy holds for foundation models, which have dual-use risks and different attack surfaces.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety." pith.science (2026). https://pith.science/paper/QU5W5DKM

@misc{pith2026250622183,
  author       = {Pith},
  title        = {Pith review of: A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QU5W5DKM}},
  note         = {Machine review of arXiv:2506.22183}
}
read the original abstract

The rapid rise of open-weight and open-source foundation models is intensifying the obligation and reshaping the opportunity to make AI systems safe. This paper reports outcomes from the Columbia Convening on AI Openness and Safety (San Francisco, 19 Nov 2024) and its six-week preparatory programme involving more than forty-five researchers, engineers, and policy leaders from academia, industry, civil society, and government. Using a participatory, solutions-oriented process, the working groups produced (i) a research agenda at the intersection of safety and open source AI; (ii) a mapping of existing and needed technical interventions and open source tools to safely and responsibly deploy open foundation models across the AI development workflow; and (iii) a mapping of the content safety filter ecosystem with a proposed roadmap for future research and development. We find that openness -- understood as transparent weights, interoperable tooling, and public governance -- can enhance safety by enabling independent scrutiny, decentralized mitigation, and culturally plural oversight. However, significant gaps persist: scarce multimodal and multilingual benchmarks, limited defenses against prompt-injection and compositional attacks in agentic systems, and insufficient participatory mechanisms for communities most affected by AI harms. The paper concludes with a roadmap of five priority research directions, emphasizing participatory inputs, future-proof content filters, ecosystem-wide safety infrastructure, rigorous agentic safeguards, and expanded harm taxonomies. These recommendations informed the February 2025 French AI Action Summit and lay groundwork for an open, plural, and accountable AI safety discipline.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns

    cs.CY 2025-11 conditional novelty 6.0 of 10

    A few open-weight video models and distribution platforms dominate the creation and spread of NSFW AI video, making developer and platform choices the main intervention points for reducing non-consensual deepfake abuse.

Reference graph

Works this paper leans on

28 extracted references · 27 canonical work pages · cited by 1 Pith paper

  1. [1]

    Safety guards should evaluate each action within its context to ensure compliance with both ethical standards and user intent

    Assess Actions Individually: Each action taken by an agent, such as making a booking, sending a command, or interacting with sensitive data, needs to be scrutinized for potential harm. Safety guards should evaluate each action within its context to ensure compliance with both ethical standards and user intent

  2. [2]

    AI Safety Is Not a Model Property

    Evaluate Compositional Risks: Actions that are harmless individually may become problematic when combined. Consider a banking AI agent with three seemingly harmless abilities: checking account balances (just viewing information), creating transaction templates (saving payment details without execution), and setting up automated payment rules (each followi...

  3. [3]

    The developer community would benefit from a clear taxonomy of agentic actions (i.e., searching information, sharing user information, purchasing, etc

    Taxonomy of actions, risks, and mitigations. The developer community would benefit from a clear taxonomy of agentic actions (i.e., searching information, sharing user information, purchasing, etc. ) with their associated risks and possible mitigations

  4. [4]

    The increased complexity of the scope of safety requires redefining the current taxonomy of harms to include real-life consequences of apparently innocuous actions

    Expanding notions of consequence . The increased complexity of the scope of safety requires redefining the current taxonomy of harms to include real-life consequences of apparently innocuous actions

  5. [5]

    Many real-world risks stem from misspecified user requests

    Human-Computer Interaction insights. Many real-world risks stem from misspecified user requests. More research in Human-Computer Interaction is needed to develop user-agent interaction models that reduce the likelihood of harmful real-world actions

  6. [6]

    Advancing policy research on models of stakeholders’ accountability and liability is key to enforcing safety mitigations

    Accountability and liability regimes . Advancing policy research on models of stakeholders’ accountability and liability is key to enforcing safety mitigations

  7. [7]

    online” (when the AI system is live) to filter potentially harmful content from users or the model itself, and “offline

    Guardrails. More research on guardrails for agentic systems is needed. This includes expanding beyond current filters narrowly focused on preventing specific outputs to a more holistic approach that ensures the agentic system remains aligned with its goal, to include addressing compositional risks. IV) Content Safety Filters This section outlines both the o...

  8. [8]

    classifying an output as unsafe when it is not) and false negatives (failing to classify an output as unsafe when it is)

    Planning for unreliable performance: Content safety systems are far from perfect and will result in mistakes, including both false positives (e.g. classifying an output as unsafe when it is not) and false negatives (failing to classify an output as unsafe when it is). Developers should evaluate classifiers with metrics like F1, Precision, Recall, and AUC-R...

Show all 28 references
  1. [9]

    Black,” “Muslim,

    Considering discriminatory biases: Like other ML models, safety classifiers can reflect and amplify societal biases, often over-triggering on content related to marginalized 23 identities. For example, early versions of Google’s Perspective API 91 flagged comments mentioning word...

  2. [10]

    Toxic Bias: Perspective API Misreads German as More Toxic

    Classifiers can complicate the system: Integrating classifiers into a system introduces technical challenges, such as added latency and higher memory usage, which can be particularly problematic for devices like low-memory smartphones. Classifiers can increase the system’s comple...

  3. [11]

    cookbook

    Developing a practical “cookbook” could help guide and empower developers implementing safety filters throughout the deployment process. Currently, a lack of clear guidance on basic safety issues creates friction. Developers face challenges in selecting relevant policies, ident...

  4. [12]

    Toxic Bias

    Larger modality coverage is needed. As highlighted in Table 10, most content safety classifiers are currently designed for text, or text and images, 94 with other modalities or agentic systems mostly out of scope. 94 It’s worth noting a helpful exception here, that of Roblox op...

  5. [13]

    Better multilingual coverage. While some traditional ML classifiers cover a large language spectrum (Perspective API is available in 18 languages for instance 95 ), most LLM safeguards are available in English only, or in a reduced set of languages as indicated in the overview Table

  6. [14]

    Several emerging approaches support this direction

    Developers seek greater control over filters—including the ability to adjust thresholds, customize policies, and set exceptions . Several emerging approaches support this direction. For example, Google’s agile classifiers framework 96 enables developers to create custom safeguar...

  7. [15]

    aligned,

    Developers want more transparency about filters . As working group participants noted, it is often unclear why a filter blocks a specific input or output. Some approaches, such as autoraters 97 (models that provide explanations for their decisions) can offer a useful foundation fo...

  8. [16]

    Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

    To surface valuable but hard-to-intuit insights and to gather a wider range of constructive ideas. Many AI safety challenges involve complex sociotechnical systems where critical information is distributed across diverse stakeholders. Local/contextual knowledge and lived exper...

  9. [17]

    default state

    Offsetting conflicts of interest or bias. Complex technical systems benefit from multiple independent layers of review to catch potential failure modes. Different stakeholders bring diverse priorities and risk perceptions, leading to more robust evaluations. Institutional incentiv...

  10. [18]

    Democratic principles assert that those affected by AI systems should have a voice in their development and oversight

    Supporting normative goals around agency and input. Democratic principles assert that those affected by AI systems should have a voice in their development and oversight. Collective input fosters legitimacy and trust by enabling meaningful stakeholder participation in safety de...

  11. [19]

    Collective input strengthens the broader AI landscape, including public labs and independent researchers

    Bolstering a broader ecosystem. Collective input strengthens the broader AI landscape, including public labs and independent researchers. This contributes valuable assets and helps prevent the concentration of knowledge and influence within a few large technology companies, who...

  12. [20]

    Current evaluation often relies on community-driven leaderboards like LMSYS, which may reflect a narrow notion of quality if the user base lacks diversity

    Creating a more diverse leaderboard for model evaluation. Current evaluation often relies on community-driven leaderboards like LMSYS, which may reflect a narrow notion of quality if the user base lacks diversity. To mitigate this bias, mechanisms such as reweighting contributi...

  13. [21]

    Applying collective intelligence principles to incident reporting can help surface risks that might otherwise be overlooked

    Building robust incident reporting systems. Applying collective intelligence principles to incident reporting can help surface risks that might otherwise be overlooked. Engaging 98 In 2023, the private AI investment in the US reached $67.2bn ( https://www.statista.com/statisti...

  14. [22]

    When inserted directly into a system message, this content can help tailor model behavior to better align with the values and needs of a specific community

    System prompts could take the form of raw, open-ended text files that include framing, cultural nuance, and community context. When inserted directly into a system message, this content can help tailor model behavior to better align with the values and needs of a specific community

  15. [23]

    Fine-tuning with prompt/response pairs that reflect region-specific cultural nuances can produce models that communities find more aligned and trustworthy than the default

    Prompts and response pairs. Fine-tuning with prompt/response pairs that reflect region-specific cultural nuances can produce models that communities find more aligned and trustworthy than the default. While generating such pairs is more challenging than drafting a constitution, t...

  16. [24]

    While useful for regional alignment, this approach introduces an intermediary between the community and the model

    Model spec / local soft alignment targets for RLHF or RLAIF: This would involve a set of culturally-informed guidelines that human raters or an LLM could use to judge which responses are more likely to resonate with a specific region. While useful for regional alignment, this a...

  17. [25]

    This would be a reward preference model that can be integrated into existing RLHF pipelines 100 and bypass the reward model training step

    Localized reward models. This would be a reward preference model that can be integrated into existing RLHF pipelines 100 and bypass the reward model training step. However, like RLHF guidelines, they introduce an intermediary layer and may limit direct community control over a...

  18. [26]

    Non-monetary incentives can overcome motivation crowding out

    Encourage and incentivize public institutions to contribute to datasets to help reduce biases and improve fairness in AI models 101 . 5.3 Future Work There are several promising research avenues: 1. Research on incentive mechanisms for diverse communities to contribute to part...

  19. [27]

    A starting point could be to develop technical systems that ensure informed consent, such as clear opt-in and opt-out options

    Ensuring ethical participation: Research is needed to prevent participatory mechanisms from becoming extractive or exploitative forms of free labor. A starting point could be to develop technical systems that ensure informed consent, such as clear opt-in and opt-out options

  20. [28]

    level 2” categories, 45 “level 3

    Steering models with minimal data : Research is needed on how to guide models using small, targeted datasets that reflect local harms and issues specific to a community’s geographic or socio-economic context that are often underrepresented in large-scale training data. APPENDIX ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.