{"id":"e1966f32-33a1-46d9-8f76-703972c24a71","arxiv_id":"2608.03329","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Adopting AI governance policies in open source projects is associated with more AI disclosure, more maintainer engagement, and better code quality metrics, but the causal estimates rest on measurement and identification assumptions that are partly circular.","lead":"This paper studies 385 open source projects that adopted AI governance policies and reports that adoption is linked to more AI disclosure, more maintainer engagement, and improved code quality, with AI-assisted contributions continuing to grow. It also introduces TRACE, a five dimension framework for classifying policy content, as a tool for future governance research.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Carry-over AI labeling makes the headline 'AI-assisted contributions grow' an artifact of disclosure-policy signals.","rationale":"The paper's most distinctive empirical claim is that AI policy adoption is followed by continued/increased AI-assisted contribution. The reader independently flagged the same assumption. I focus on it because the outcome construction is confounded with treatment: policies that mandate disclosure create the very signals used to define AI-assisted PRs, and the carry-over rule turns one signal into a persistent label. This construct-validity threat is not addressed by parallel-trends tests or window robustness. The proposed sensitivity analyses directly target the mechanism. Other issues (no unit FE, SonarQube missingness) are real but secondary; if the carry-over test is clean, the conditional verdict can be revisited. Verdict remains CONDITIONAL pending the requested re-analysis.","tokens_in":16785,"tokens_out":5616,"duration_ms":57771,"concrete_test":"Recompute Model I for 'AI-assisted PR share' and 'AI-assisted PR throughput' with (a) no carry-over (only PR-level direct signals) and (b) bounded carry-over (e.g., label only PRs within 4 weeks after a direct signal by the same author). If the positive ATTs shrink to ~0 or reverse, the headline claim is a labeling artifact; if they persist, the carry-over assumption is not the driver.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III-G defines an AI-assisted PR via direct structural signals (agent branches/labels, co-author trailers, self-disclosure) and then: 'Assuming developers continue to use AI tools after initial adoption, we labeled all subsequent PRs by the same author as AI-assisted.' The headline evidence for 'AI-assisted contributions continue to grow' is Table II: AI-assisted PR share +6.8pp, AI-assisted PR throughput +10%, plus the family contrasts in Table IV. But the treatment itself can increase the signals used to label the outcome: several policies (vLLM, Ghostty) mandate disclosure phrases or 'Co-authored-by' trailers, and the post-period begins exactly when those mandates apply. One compliant PR triggers the carry-over rule, so every later PR by that author counts as AI-assisted regardless of actual AI use. The measured increase may thus reflect the policy's visibility mandate, not a behavioral rise in AI tool use. The disclosure-count outcomes are partly mechanical, but the AI-assisted activity metrics are supposed to be behavioral; the carry-over contaminates them. Section VI robustness checks windows but never sensitivity-tests the labeling rule. Since the central claim 'governance regulates rather than prohibits' rests on this measurement, the causal conclusion is not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies AI governance policies in open-source software by analyzing 29,624 GitHub repositories and identifying 385 projects that adopted an AI policy between February 2025 and April 2026. It introduces the TRACE framework (Transparency, Responsibility, Attribution, Constraints, Enforcement), classifies policies into five families, and estimates treatment effects using propensity-score matching and staggered difference-in-differences on 19 SPACE-inspired outcome variables. The headline findings are that policy adoption is associated with increased maintainer engagement, more AI disclosure, richer review interactions, improved code-quality metrics, and continued growth of AI-assisted contributions, with specific dimension- and family-level heterogeneity. The paper argues that AI governance mostly regulates rather than prohibits AI-assisted development.","tokens_in":17034,"tokens_out":4817,"duration_ms":48943,"significance":"If the estimates are valid, this is a valuable first large-scale empirical account of AI governance in OSS, with a reusable taxonomy (TRACE) and a plausible causal framework. The paper's strengths include its large corpus, manual validation and inter-rater reliability checks, a replication package, propensity-score matching, staggered DiD modeling, event-study pre-trend tests for all 19 outcomes, and multiple post-treatment windows. However, the central AI-activity and disclosure outcomes are measured in ways that are entangled with the policy treatment itself, so the headline magnitudes are not yet credible. The paper is a solid empirical contribution that needs substantial robustness work before its policy conclusions can be accepted.","major_comments":[{"comment":"The AI-assisted PR outcome is defined by direct structural signals (agent branches/labels, co-author trailers, self-disclosure) plus a carry-forward rule: once an author is identified as using AI, all subsequent PRs by that author are labeled AI-assisted. Several treatment policies (e.g., vLLM, Ghostty) explicitly require disclosure phrases or Co-authored-by trailers, so a compliant PR triggers the carry-over and all later PRs by that author count as AI-assisted regardless of actual AI use. The estimated +6.8pp AI-assisted PR share and +10% throughput (Table II) may therefore reflect increased detector sensitivity rather than a behavioral increase in AI tool use. Section VI varies the post-treatment window but never sensitivity-tests the labeling rule. Please add analyses that (i) restrict the outcome to signals not mandated by policy text, (ii) remove or cap the carry-forward component,","section":"Section III-G; Table II"},{"comment":"The AI-disclosure PR count and share are keyword matches for phrases such as \"generated by\" and \"AI-assisted\" in PR descriptions. For policies that mandate specific disclosure language, this outcome largely measures compliance with the policy itself, not an independent behavioral response. The +11% disclosure count and +4.2pp disclosure share could arise even if developer behavior were unchanged. The paper should reframe this outcome as compliance with transparency mandates and add robustness checks, e.g., excluding policies that prescribe exact disclosure wording, or measuring disclosure with phrases not mentioned in the policy text. As it stands, the disclosure results are partly circular with the treatment definition.","section":"Section III-G; Table II"},{"comment":"There is an inconsistency in the post-treatment window: Section III-B says the sample is restricted to ensure \"an 8-week pre-treatment and a 4-week post-treatment time window,\" but Section V reports the main effects using an 8-week post-treatment window and calls the 4-week version a robustness check. This matters for the adoption cutoff: the data are described as spanning January 2025 to May 2026, and adoptions through April 30, 2026 cannot all have eight post-treatment weeks within that span. Please state the exact post-treatment horizon used in Tables II–IV and reconcile it with the adoption date restriction and data-end date.","section":"Section III-B; Section V"}],"minor_comments":[{"comment":"There are several typos and grammatical errors: \"specilized,\" \"qualtitive,\" \"introduing,\" \"Rs\" for pull requests, and \"The each family ATT\" in Section III-H. A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The event-study figure shows only 10 of the 19 outcome variables, while the text states that all 19 passed the parallel-trends test. Please indicate where the remaining outcomes are shown or add an appendix figure; also clarify whether any multiple-comparison adjustment was applied across the 19 pre-trend tests.","section":"Figure 2"},{"comment":"The header \"/banCommunication\" should be \"/banConstraints\" or similar. Also, the family column header \"Disclosed Soft-Enforced\" should match the family name \"Disclosed but Soft-Enforced\" used in Section IV-B.","section":"Table III"},{"comment":"The abbreviation \"SEART-GHS\" appears inconsistent with the SEART-GitHub name in the reference [27]. Please verify the tool name and capitalization.","section":"Section III-B"},{"comment":"In Observation 1, the text says \"T3–T4, 66.0%,\" which is arithmetically correct (116+138 = 254; 254/385 = 65.97%), but the phrase \"no disclosure required\" at T1 and \"encouraged/voluntary\" at T2 could be more clearly distinguished from the ordinary review process; consider adding example quotes for T2 and T3 as is done for other levels.","section":"Section IV; Table I"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is likely to be of interest to the software-engineering and MSR community, and the TRACE taxonomy is a useful contribution. However, the measurement entanglement between the treatment and the headlined AI-activity/disclosure outcomes is a serious internal-validity issue that the current robustness checks do not address. I would ask the editor to require the proposed sensitivity analyses before acceptance, rather than treating this as a purely editorial revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the take: this is the first large-scale empirical map of AI governance policies on GitHub, and the TRACE taxonomy is genuinely useful. But the causal headline — that policy adoption increases AI-assisted contributions while improving quality — is not yet supported, because the main outcome for AI activity is measured with a rule that the treatment itself can trigger.\n\nWhat's new: they identified 385 policy adopters from ~30k repos, coded them along five dimensions (Transparency, Responsibility, Attribution, Constraints, Enforcement), grouped them into five families, and released a replication package. The descriptive results — most policies regulate rather than ban, transparency is common, attribution is rare — are likely robust and worth citing. The PSM+DiD design is appropriate in spirit, and the pre-trend tests passing for all 19 outcomes is a good sign.\n\nThe soft spots are real. Section III-G defines an AI-assisted PR via structural signals (agent branches, co-author trailers, self-disclosure) and then labels all subsequent PRs by the same author as AI-assisted. Several policies in the sample (vLLM, Ghostty) mandate disclosure phrases or co-author trailers. So the treatment increases the signal used to define the outcome. One compliant PR triggers the carry-over rule, and every later PR by that author is counted as AI-assisted. The +6.8pp AI-assisted share and +10% throughput in Table II may well be artifacts of detector sensitivity, not changes in behavior. The paper does not sensitivity-test the labeling rule; that's the load-bearing gap.\n\nTwo other issues are more minor but need fixing. The design section says the post window is 4 weeks, but the main results use 8 weeks; robustness checks at 4, 6, 10, 12 weeks don't resolve the internal inconsistency. SonarQube scans failed for 8-13% of repos, and that missingness is unmodeled. 'Maintainer' is never defined, though 'maintainer-responded PR share' and 'maintainer response time' are key outcomes. The staggered DiD uses matched-set and calendar-week random intercepts, not unit fixed effects; that's defensible but fragile with a short window.\n\nBottom line: the taxonomy and descriptive findings are solid and will be cited. The causal estimates should be treated as preliminary until the labeling rule is stress-tested. This paper deserves a serious referee — send it to review — but the referee should demand a corrected time window, a sensitivity analysis for the carry-over rule, and a definition of maintainer. I'd bring it to reading group.","headline":"Large-scale map of AI policies is a real contribution, but the causal estimates collapse on the carry-over labeling rule.","tokens_in":17560,"tokens_out":2390,"would_cite":true,"duration_ms":24054,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Open-source AI policies regulate rather than prohibit AI-assisted development, and adopting them improves disclosure, review, and code quality.","keywords":["AI governance","open source software","GitHub","difference-in-differences","developer experience","TRACE framework","AI-assisted development","propensity score matching"],"falsifier":"Recompute the headline outcomes labeling AI-assisted PRs only by explicit structural signals (agent branches, co-author trailers, self-disclosure) and not by carry-over author history. If the post-adoption increase in AI-assisted PR share and throughput shrinks to zero or reverses specifically in repositories whose policies mandate disclosure, the causal claim would be falsified; if the increase persists, the carry-over assumption is not the driver.","tokens_in":16628,"feed_emoji":"🤖","tokens_out":4808,"duration_ms":47186,"temperature":0.7,"pith_summary":"Open-source projects are starting to write policies for AI-assisted contributions, and this paper asks what those policies do. The authors claim that such policies mostly regulate rather than prohibit AI-assisted development, and that adoption itself changes behavior: after a policy lands, maintainers engage more, developers disclose AI use more often, reviews become more interactive, and code quality metrics improve, while AI-assisted contributions keep growing. They reach this using 29,624 GitHub repositories, 385 actual policy adoptions, and a matched-control difference-in-differences design. If true, it means OSS communities can deliberately shape AI use through governance choices, and that transparency-and-responsibility-oriented policies work better than restriction alone.","feed_headline":"AI policies regulate, not ban, AI coding on GitHub","feed_subtitle":"A study of 385 open-source projects finds disclosure, review depth, and code quality rise after policy adoption.","key_machinery":"The TRACE framework, which codes each policy on five ordered dimensions—Transparency, Responsibility, Attribution, Constraints, and Enforcement—plus a rule-based classification of policies into five governance families. The causal estimates come from a staggered difference-in-differences design with propensity-score matched never-treated controls, event-study parallel-trend tests, and 19 outcomes organized by the SPACE developer-productivity lens.","core_discovery":"Policy adoption is best described as governed permission, not prohibition. Across 385 projects, disclosure requirements and human-validation mandates are the most common provisions, while strict bans are a minority family. Quasi-experimental estimates show that after adoption, maintainer-responded PR share rises, AI-disclosure share and AI-assisted PR share both rise, review turns and reviewers per PR increase, and static-analysis quality metrics (vulnerabilities, duplicated lines, code smells, cognitive complexity) improve. The effects vary systematically with policy content: higher Transparency levels amplify disclosure and participation, Responsibility level 4 improves review engagement a","pith_inferences":["Carry-over labeling: the paper labels every later PR by a developer as AI-assisted once that developer shows any AI signal; if policies themselves trigger new disclosure markers, part of the measured increase in AI-assisted share could reflect detector sensitivity rather than behavior. A sensitivity analysis using only explicit structural signals would test this.","The 8-week post-treatment window means the results are short-run; whether policies cool off, harden, or get revised after longer exposure is untested.","The TRACE dimensions map naturally onto enterprise AI governance, so a comparable matched-cohort study inside firms could test whether transparency-plus-responsibility beats restriction there too.","Policy formation is left unstudied; tracing whether policies emerge from maintainer mandates or contributor deliberation would separate legitimacy effects from mere rule effects."],"forward_implications":["Adopting any explicit AI policy appears to increase maintainers' attention: maintainer-responded PR share rises and LOC per reviewer falls.","Requiring AI disclosure does not suppress AI use; AI-assisted PR share and disclosure share rise together after adoption.","Transparency and responsibility provisions produce stronger community and quality gains than restrictive provisions alone.","Strict prohibition reduces AI-assisted throughput but not to zero, and is associated with the largest quality improvements, suggesting bans filter rather than eliminate.","Silence has a cost: quiet policies are the only family with a net reduction in AI-assisted activity, and opaque restriction lowers engagement."],"supporting_citations":[{"why":"Supplies the initial AI-governance framework that TRACE adapts for OSS policies.","marker":"[8]"},{"why":"Documents quality risks of AI-assisted code and provides the natural-experiment template the authors extend.","marker":"[7]"},{"why":"Establishes the governance-as-treatment approach for measuring policy effects in OSS.","marker":"[24]"},{"why":"Provides the canonical difference-in-differences estimator.","marker":"[33]"},{"why":"Provides the staggered-adoption DiD method used to align repositories on their own adoption weeks.","marker":"[38]"},{"why":"Provides propensity-score matching for constructing comparable control groups.","marker":"[39]"},{"why":"Supplies the SPACE framework organizing the 19 developer-experience outcomes.","marker":"[42]"},{"why":"Provides the structural signals (agent branches, labels, co-author trailers) used to identify AI-assisted PRs.","marker":"[29]"},{"why":"Supplies the SEART-GHS dataset used to sample the 29,624 repositories.","marker":"[27]"}],"fun_headline_variants":["AI policies on GitHub: transparency beats prohibition","Open-source AI governance: disclosure improves outcomes","Governed AI: how GitHub projects regulate coding bots","AI policy adoption: better reviews, quality, and disclosure"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing assumption is that a developer who ever shows an AI signal continues to use AI in all later pull requests: without that carry-over label, the measured AI-assisted share and throughput growth after policy adoption could be an artifact of new disclosure markers rather than a real change in AI use.","fun_headline_variants_meta":{"raw":{"variants":["AI policies on GitHub: transparency beats prohibition","Open-source AI governance: disclosure improves outcomes","Governed AI: how GitHub projects regulate coding bots","AI policy adoption: better reviews, quality, and disclosure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1244,"prompt_tokens":693,"completion_tokens":551,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":491}},"tokens_in":437,"tokens_out":551,"duration_ms":6525,"temperature":1.0,"reasoning_tokens":491,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:52:59.928011+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the headline outcomes labeling AI-assisted PRs only by explicit structural signals (agent branches, co-author trailers, self-disclosure) and not by carry-over author history. If the post-adoption increase in AI-assisted PR share and throughput shrinks to zero or reverses specifically in repositories whose policies mandate disclosure, the causal claim would be falsified; if the increase persists, the carry-over assumption is not the driver.","supporting_citations":[{"cited_title":"Speed at the cost of quality: How cursor ai increases short-term velocity and long-term complexity in open-source projects,","cited_arxiv_id":null,"evidence_quote":"Documents quality risks of AI-assisted code and provides the natural-experiment template the authors extend."},{"cited_title":"Beyond adoption: Examining the evolution and impact of codes of conduct on open source communities,","cited_arxiv_id":null,"evidence_quote":"Establishes the governance-as-treatment approach for measuring policy effects in OSS."},{"cited_title":"Minimum wages and employment: A case study of the fast-food industry in New Jersey and Pennsylvania,","cited_arxiv_id":null,"evidence_quote":"Provides the canonical difference-in-differences estimator."},{"cited_title":"Difference-in-differences with multiple time periods,","cited_arxiv_id":null,"evidence_quote":"Provides the staggered-adoption DiD method used to align repositories on their own adoption weeks."},{"cited_title":"The space of developer productivity: There’s more to it than you think","cited_arxiv_id":null,"evidence_quote":"Supplies the SPACE framework organizing the 19 developer-experience outcomes."},{"cited_title":"Sampling projects in github for msr studies,","cited_arxiv_id":null,"evidence_quote":"Supplies the SEART-GHS dataset used to sample the 29,624 repositories."}],"review_version":1}