{"id":"774d9bc7-c16f-437d-bd2f-4016d4bd27a4","arxiv_id":"2603.26487","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Open-source GenAI governance clusters into three orientations and 12 strategies, driven by review burden, accountability, provenance, and infrastructure concerns rather than a simple ban-or-not choice.","lead":"Analyzing public governance documents from 67 prominent open-source projects, this paper maps how communities regulate AI-assisted contributions and identifies three governance orientations and 12 concrete strategies. It gives maintainers and platform designers a structured menu for deciding whether, when, and how to restrict or channel generative-AI contributions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Corpus screen for explicit GenAI mentions may under-represent projects that govern AI via general quality rules; the taxonomy's completeness and orientation proportions are at risk.","rationale":"The paper's central contribution is a qualitative taxonomy. For that taxonomy to be valid, the corpus must sample the range of GenAI governance practices, not just practices that explicitly mention GenAI. I examined the methodology (§3.2) and the three-orientation/12-strategy results (§5). The most insecure link is the operationalization of 'publicly encoded governance' as texts that explicitly regulate GenAI-assisted behavior. This screen is understandable for a first study, but it creates a blind spot for silent absorption: high-visibility projects whose general contribution rules already impose tests, accountability, and review limits that apply identically to AI-assisted submissions. Those projects would be O3-like but are excluded by design. This directly threatens the 'reusable strategy space' claim and the orientation percentages, though it does not threaten the weaker, well-supported claim that some projects go beyond banning.\n\nI considered whether the 'requires coordinated responses' wording is an unsupported normative leap. The data show that concerns cluster across workflow stages, but the paper does not measure whether coordinated responses are causally required for successful governance. This is a rhetorical overreach in the abstract, but it is not the weight-bearing part of the contribution; the descriptive taxonomy stands independently. I therefore keep the reader's focus on corpus construction.\n\nThe reported inter-rater kappa (0.742/0.745/0.840) and extensive quotes provide real support for the reliability of the coding among the selected cases. The problem is not coding accuracy but sampling completeness. The proposed test—a blind audit of high-visibility projects without explicit AI mentions—would settle whether an 'implicit absorption' category is missing and whether O3 is undercounted. Until that check is run or the corpus sheet/codebook is released, the percentages should be read as sample-descriptive, exactly as the CONDITIONAL verdict says. Therefore I do not change the reader's verdict.","tokens_in":18546,"tokens_out":9659,"duration_ms":95907,"concrete_test":"Take a random sample of, say, 100 repositories from the same Gitstar top-800 seed that were not included in Phase 1 (i.e., had no explicit GenAI governance text by the authors' screen). Independently code their full CONTRIBUTING/SECURITY/templates for rules that would constrain AI-assisted contributions (test/evidence requirements, human responsibility, issue linkage, queue/review limits), blind to whether 'AI' appears. Map the coded rules onto the 12 strategies and three orientations. If all coded cases fit the existing taxonomy and O3 is the dominant orientation for this group, the explicit-mention screen did not materially bias the strategy space; if new strategies emerge or the orientation distribution shifts substantially (e.g., O3 > 40%), the reported map is incomplete and percentages are sample-descriptive only.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is that the corpus covers the relevant space of GenAI governance. The inclusion screen in §3.1–3.2 required texts that 'explicitly regulated GenAI-assisted contribution behavior'; the corpus is therefore built from projects that name GenAI. This is narrower than the paper's own definition of governance (footnote 1: rules and interfaces that regulate contribution intake). A project can govern AI-mediated contributions entirely through general-purpose gates—mandatory tests, human accountability, issue-linkage, PR-size/queue limits—without ever mentioning AI. Such projects would plausibly be O3 ('quality-first / tool-agnostic'), but they are invisible to the sample. Consequences: (1) the orientation proportions (O1 20.9%, O2 59.7%, O3 19.4%) are conditional on an explicit-AI-mention screen, not on the full population of OSS governance; (2) the 12-strategy map may omit an 'implicit absorption' mode or variants of Scope/Verification/Moderation that exist only as general rules. The central 'beyond banning' claim survives, but the reusable-map contribution and the transferability of the percentages are weaker. §6.1 acknowledges 'less explicit' governance in smaller/non-English communities but does not test this specific selection effect among the same high-visibility projects.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a qualitative document analysis of public GenAI governance materials from 67 highly visible OSS projects. It identifies seven recurring maintainer concerns across contribution workflows, derives three governance orientations—prohibitionist, boundary-and-accountability, and quality-first—and abstracts 12 governance strategies organized into four functional groups. The central claim is that governing GenAI in OSS is not a binary ban-or-allow decision but a multi-surface design problem involving accountability, verification, review capacity, provenance, and platform infrastructure. The methodology uses a two-phase corpus construction (top-800 star seed + snowball sampling), iterative open coding, project-level memoing, constant comparison, and inter-rater reliability assessment (κ = 0.742 for orientations, 0.745 for concerns, 0.840 for strategies). Limitations are acknowledged in §6.","tokens_in":18800,"tokens_out":10409,"duration_ms":99658,"significance":"If the proposed taxonomy is robust, it provides a useful structured map of an emerging and scattered governance practice, with practical value for maintainers and platform designers and a conceptual baseline for future research. The study's strengths include a transparent audit trail, direct quotations from primary sources, inter-rater reliability reporting, and an honest limitation section. The main risk to the contribution is the corpus inclusion screen, which requires texts that explicitly regulate GenAI-assisted contribution behavior; this may limit the completeness and generalizability of the 12-strategy map and the reported orientation proportions.","major_comments":[{"comment":"The corpus screen requires texts that 'explicitly regulated GenAI-assisted contribution behavior,' which is narrower than the paper's own governance definition in footnote 1 (any rules/interfaces regulating contribution intake). Projects that govern AI-mediated contributions through general-purpose quality gates without naming GenAI are systematically excluded. Because the abstract and RQ2 claim a 'reusable strategy space' and report orientation proportions (O1 20.9%, O2 59.7%, O3 19.4%), this selection effect is load-bearing: the taxonomy may be missing an 'implicit absorption' mode, and the percentages cannot be read as characterizing GenAI governance in OSS broadly. Please either scope the claims to explicit GenAI governance or conduct a supplementary check on a sample of high-visibility projects without explicit AI mentions to test whether new strategies or orientations emerge.","section":"§3.1–3.2, §6.1"},{"comment":"The paper states that the three orientations 'explain why projects facing similar GenAI pressures develop markedly different institutional responses.' The design is cross-sectional and descriptive; no measure of 'pressures' is used, and orientations are derived from the same governance texts that define the strategies. This supports an interpretive account or association, not a causal explanation. Please soften the language (e.g., 'account for' or 'are associated with') or add evidence that projects with similar pressure profiles differ systematically by orientation.","section":"§5.1, Contributions (p. 2)"},{"comment":"The corpus mixes 58 repository-hosted sources and 9 project-adjacent policy texts into a single project-level case analysis, but Table 1 appears to report demographics for only 55 projects (the language counts sum to 55). The paper should clarify whether the 9 project-adjacent texts are counted as cases equivalent to whole projects, and report the breakdown by source type. This matters because the orientation percentages and strategy prevalences are case-level counts, and mixing organizational types may affect the reported distributions.","section":"§3.2, §4, Table 1"}],"minor_comments":[{"comment":"The figure is difficult to parse in the provided rendering: 'Capacity & Queue Control' has no prevalence entries, and the column-major layout is unclear. Please ensure the figure renders all 12 strategies with their three orientation values, or provide a machine-readable table.","section":"Figure 1"},{"comment":"The 'Policy adoption' row is garbled, and the demographic counts do not reconcile with the corpus size. Please clean the table and state explicitly which subset of cases it covers.","section":"Table 1"},{"comment":"The coding codebook is not included. Since the 12 strategies and seven concerns are the main results, a supplementary codebook with example quotes per code would improve reproducibility and reader confidence.","section":"§3.4"},{"comment":"Minor typo: 'independent if AI is used or not' should be 'independent of whether AI is used or not'.","section":"§5.1"},{"comment":"The abstract reports orientation percentages without qualification; §6.3 warns they are approximate patternings. Consider adding a pointer to the limitation or softening the abstract's numerical presentation.","section":"§6.3, Abstract"}],"recommendation":"major_revision","confidential_remarks":"This is a worthwhile qualitative study with a solid methodological core. The major issue is the mismatch between the broad governance framing and the explicit-AI-mention corpus, which threatens the completeness of the proposed 'reusable strategy space.' A supplementary robustness check on projects without explicit AI governance text, or a clear scoping of the claims, would resolve the concern. I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the important thing about this paper: it gives the field its first systematic map of how OSS projects govern GenAI contributions. The three orientations (prohibitionist, boundary-and-accountability, quality-first) and the 12 strategies with their functional groups are grounded in a transparent coding process with decent inter-rater agreement (κ around 0.74–0.84) and are richly illustrated with quotes. The central claim that governance is more than a ban-or-not question is well supported. The discussion of front-loading review and infrastructure-level governance is a genuine contribution, and the paper connects to existing OSS-governance literature instead of pretending the topic emerged from nowhere.\n\nThe soft spots are real but contained. The corpus is built on texts that explicitly mention AI and explicitly regulate AI-assisted contributions (Section 3.1). That is a deliberate scope choice, and the limitation section is honest about it, but it does mean the taxonomy and especially the orientation shares (O1 20.9%, O2 59.7%, O3 19.4%) are conditional on a project having named AI. A project that governs AI-mediated contributions entirely through general quality gates—mandatory tests, human accountability, issue-linkage—without ever saying “AI” is invisible to this sample. Those projects would likely land in the O3 quality-first bucket, so the prevalence figures are likely skewed. The paper frames these counts as sample-descriptive, and Section 6.1 backs off from prevalence claims, so this is a boundary condition more than a fatal flaw. Still, if the authors want the 12-strategy map to be the reusable baseline they advertise, they should either release the corpus sheet and codebook so the coding can be inspected, or soften the percentages further and label them explicitly as conditional on explicit-AI-governance texts.\n\nMy other quibble is minor: the paper mentions 'publicly encoded governance' and then talks about projects such as curl and Zig where some policy is in blog posts or issue threads. The line between repository-embedded 'core' texts and surrounding 'justificatory' materials is handled reasonably, but it can be blurry. That does not threaten the analysis.\n\nBottom line: this deserves a serious referee. The authors have done honest qualitative work on a timely problem, and the paper is a useful baseline for maintainers and for researchers studying GenAI governance. I would send it to peer review with a request for the audit materials and a more careful framing of the orientation counts. It is not a desk reject.","headline":"A well-executed qualitative taxonomy of GenAI governance in OSS that delivers a useful 'beyond banning' frame, but the corpus screen for explicit AI mentions means the orientation counts and strategy map speak mainly to projects that name AI.","tokens_in":19289,"tokens_out":2460,"would_cite":true,"duration_ms":22772,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Open source projects are not simply banning or allowing generative-AI contributions; they are building governance regimes that combine admissibility rules, disclosure and accountability requirements, verification gates, workflow protections","keywords":["generative AI","open source software","contribution governance","maintainer workload","review bottleneck","governance orientations","AI disclosure policy","code provenance"],"falsifier":"A replication that applies the same coding to projects without explicit AI-named policies—for example, a random sample of smaller or lower-activity repositories with strict general review rules—and finds that their governance cannot be sorted into the three orientations, or that a fourth orientation emerges, would refute the claim that the taxonomy covers the OSS governance space. Alternatively, a maintainer survey showing that private enforcement contradicts the public policy in a majority of sampled projects would falsify the public-text premise.","tokens_in":18404,"feed_emoji":"🤖","tokens_out":5994,"duration_ms":60067,"temperature":0.7,"pith_summary":"This paper tries to establish that open source software (OSS) projects are not answering a simple 'ban AI or not' question when they regulate generative-AI contributions. Based on a qualitative analysis of public governance texts from 67 highly visible projects, it argues that maintainer concerns fall into seven recurring areas—review bottlenecks, low-value AI text, issue triage costs, security-report noise, adversarial incentives, provenance and licensing uncertainty, and platform or tooling limits—and that project responses cohere into three governance orientations: prohibitionist, boundary-and-accountability, and quality-first. These orientations are implemented through 12 reusable strategies, from disclosure rules and evidence gates to PR queues and platform migration. The value is that it converts scattered community practices into a conceptual map that maintainers can use to match policy to their specific bottleneck, and it redirects research away from 'ban or not' toward a multi-surface governance problem.","feed_headline":"Three AI-governance orientations emerge from 67 open source projects","feed_subtitle":"Prohibition, boundary-and-accountability, quality-first: most projects use disclosure and review gates, not bans.","key_machinery":"The strategy-orientation map, built through iterative qualitative coding of public governance texts from 67 projects, cross-tabulates three governance orientations—prohibitionist, boundary-and-accountability, and quality-first—against 12 strategies in four functional groups: entry admissibility and input qualification, responsibility and evidence restoration, review burden and workflow protection, and infrastructure and institutional adjustment. This map carries the argument by showing that each orientation is realized by a distinct combination of strategies, that the same strategy plays different roles under different orientations, and that no single device such as disclosure or templates s","core_discovery":"The central claim is that GenAI governance in OSS is a multi-surface design problem, not a binary policy decision. Across 67 project-level cases, the paper identifies three governance orientations—prohibitionist (refusing certain AI inputs at the door), boundary-and-accountability (admitting AI only under explicit disclosure, human accountability, and verification), and quality-first (absorbing AI into existing quality and maintainer-cost thresholds)—and shows that boundary-and-accountability is the dominant orientation, appearing in about 60% of cases. These orientations are operationalized through 12 strategies grouped into four functions: entry admissibility, responsibility and evidence,","pith_inferences":["If the taxonomy is sound, a project's primary bottleneck—legal uncertainty, review capacity, or security-channel noise—should predict its orientation; this diagnostic mapping is implicit in the paper and could be tested on new projects before they write policy.","The sampling design, which requires texts that explicitly name GenAI, likely undercounts quality-first governance, since projects that fold AI expectations into general quality rules without naming AI are excluded; the reported 19.4% share for that orientation is probably a lower bound.","A concrete next experiment would track whether mandatory AI disclosure actually changes review time or contributor mix in projects that adopted it, comparing with matched projects that did not; the paper identifies disclosure compliance as an open question.","The generation-versus-review asymmetry suggests a measurable 'attention tax': review hours spent per merged contribution or per issue before and after agentic tools could quantify the pressure the paper describes and evaluate whether governance strategies restore balance."],"forward_implications":["GenAI governance in OSS should be diagnosed by bottleneck: projects worried about fake security reports can reach for verification and evidence gating, while projects worried about long-term quality can reach for accountability reinforcement and scope control; there is no one-size-fits-all AI policy.","The practical mainstream is not banning AI but requiring disclosure, human ownership, and evidence: boundary-and-accountability is the dominant orientation in the corpus.","The load-bearing shift is upstream: projects are moving admission control before review—pre-approved issues, proof-of-concept gates, and PR limits—to protect maintainer attention as the scarcest resource.","Repository-level rules have limits, so governance is beginning to move to infrastructure: projects escalate to suspending external contributions or changing venues, implying platform designers must take on intake control and traceability."],"fun_headline_variants":["Open source AI governance: 3 orientations, not just bans","GenAI rules in OSS: 60% use boundaries, not bans","Beyond the ban: 12 strategies for GenAI in open source","How 67 OSS projects govern GenAI: 3 orientations"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole map rests on treating publicly written, explicitly AI-naming governance texts as the right and sufficient evidence of how a project governs GenAI; projects that govern AI through general quality rules or private maintainer enforcement are invisible to the corpus, so the taxonomy and orientation shares could shift if they were included.","fun_headline_variants_meta":{"raw":{"variants":["Open source AI governance: 3 orientations, not just bans","GenAI rules in OSS: 60% use boundaries, not bans","Beyond the ban: 12 strategies for GenAI in open source","How 67 OSS projects govern GenAI: 3 orientations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000194,"raw_usage":{"total_tokens":1190,"prompt_tokens":743,"completion_tokens":447,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":487,"completion_tokens_details":{"reasoning_tokens":371}},"tokens_in":487,"tokens_out":447,"duration_ms":4733,"temperature":1.0,"reasoning_tokens":371,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T17:15:23.185141+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that applies the same coding to projects without explicit AI-named policies—for example, a random sample of smaller or lower-activity repositories with strict general review rules—and finds that their governance cannot be sorted into the three orientations, or that a fourth orientation emerges, would refute the claim that the taxonomy covers the OSS governance space. Alternatively, a maintainer survey showing that private enforcement contradicts the public policy in a majority of sampled projects would falsify the public-text premise.","supporting_citations":[],"review_version":1}