{"id":"8459de0b-6f2e-4874-a45e-ae86e2be54bd","arxiv_id":"2607.15769","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A repository file called the Agent Governance Manifest, linking risk zones, evidence obligations, human confirmation, and review gates, raised exact risk-label recovery from 40.5% to 97.4% in a 75-output controlled test and enabled 45/45 feasible contributor-side packages.","lead":"This paper proposes a governance manifest that open-source projects can keep in their repository to tell AI coding agents and contributors what evidence, risk checks, and human sign-offs a proposed change must carry before maintainers review it. In small controlled tests, reviewers using agent tools recovered the intended risk label 97% of the time with the manifest versus 40% without — but the manifest materials stated those labels directly.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reviewer-side result may be a manipulation check: AGM materials contain the exact risk labels and gate states used as outcome measures, so 97.4% vs 40.5% may reflect information inclusion rather than AGM's governance design.","rationale":"The reader's weakest_assumption concerns honest package preparation and human inspection, which is a real boundary condition and is explicitly acknowledged in §4.5. However, the more load-bearing problem is the internal validity of the primary evaluation: the AGM condition supplies the exact governance states that the outcome rubric measures, so the 97.4% vs. 40.5% result conflates 'the information is present' with 'AGM's governance design improves review.' This is not an ad hominem or a demand for production evidence; it is a request for a control that isolates the mechanism. The paper has genuine strengths: a transparent design-science sequence, a pinned public prototype, deterministic local validation scripts, pre-specified rubrics, bootstrap uncertainty, and explicit limitations in §8.3. Those strengths support a conditional verdict, not rejection. The audit's zero-repository finding is also weakened by post-hoc construct refinement (§5.3, Table S10), but the chosen concern is the one that most directly undermines the central empirical claim. A minimal-metadata control and an adversarial-evidence probe would settle whether AGM is doing governance work or merely relabeling. Since this concern reinforces the reader's CONDITIONAL verdict rather than overturning it, the verdict should remain UNCHANGED.","tokens_in":1054,"tokens_out":901,"duration_ms":55335,"concrete_test":"Run a three-condition within-participant reviewer study: (a) ordinary materials, (b) a minimal-metadata control that presents the same risk labels, gate states, and evidence statuses in a plain structured table or header, without AGM-specific evidence packages, confirmation declarations, or manifest logic, and (c) full AGM-supported materials. Use the same five task types and rubric. If (b) and (c) produce statistically indistinguishable exact risk-label recovery, missing-evidence detection, gate-state judgment, and final-acceptance decisions, the headline effect is due to information inclusion, not AGM's design. Separately, plant deliberately fabricated test results in otherwise valid evidence packages and measure whether reviewer outputs flag or challenge them; if AGM-supported reviewers rarely detect fabricated evidence, the 'governable' claim is unsupported beyond legibility.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative support for AGM is the reviewer-side contrast in §7.2: exact risk-label recovery of 97.4% under AGM-supported materials vs. 40.5% under ordinary materials. But the AGM-supported condition (Table S30) includes the very governance states the rubric measures: repository-defined risk-zone references, risk summaries, evidence indexes, contributor-confirmation declarations, and explicit gate states. Ordinary materials omit these as 'not structured.' The outcome rubric (Table S31) then codes whether the reviewer output names the same risk label, evidence status, accountability state, and gate state. This is close to a reading-comprehension check: the answer is literally in the input. The 56.8-point gap therefore chiefly demonstrates that explicit labels are recoverable, not that AGM's specific governance design—risk zoning, evidence packaging, confirmation gates—improves governance judgment. Without a control condition that supplies the same governance information in a plain, non-AGM format, the effect cannot be attributed to the manifest's structure rather than to simply telling reviewers the answer. The one AGM-supported error (T5, one-level underclassification) actually shows agents can ignore labels, but it does not solve the control problem. The contributor-side feasibility result (§7.5) is similarly limited: participants were instructed to prepare compliant packages and had every reason to comply; it shows templates can be filled, not that evidence is truthful. The paper itself acknowledges this boundary (§4.5), but the headline claim 'AGM improves governance-state recovery' is not yet distinguished from 'providing explicit metadata improves label recovery.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the governance burden created by AI coding agents in open-source software. It proposes a three-layer framework (agent-readability, traceability, governability), reports a diagnostic audit of 50 GitHub repositories finding fragmented AI-governance cues but no project-wide governability arrangement, and develops the Agent Governance Manifest (AGM) as a repository-hosted governance resource. The evaluation has two parts: a within-participant reviewer-side study (15 participants, 75 outputs) reporting that AGM-supported materials improve exact risk-label recovery (37/38 vs. 15/37) and perceived review support, and a contributor-side feasibility check (15 participants, 45 tasks) reporting that all final packages represent the core governance state correctly with 41/45 passing strict structural validation.","tokens_in":44369,"tokens_out":2985,"duration_ms":30442,"significance":"If the central claim is established, the paper would make a useful conceptual and practical contribution to OSS governance: it names a distinct governability function, gives it a repository-hosted artifact form, and provides a bidirectional contributor/maintainer workflow. The audit is thoughtfully designed, the statistical reporting is transparent (participant-clustered bootstrap CIs, leave-one-out checks, pre-adjudication inter-coder reliability, disclosure of one AGM error), and the replication-package plan is a genuine strength. However, the headline quantitative support is partly circular: the AGM-supported condition contains the very governance states used as outcome measures. The paper's own boundary statements in §4.5 and §8.3 are honest about what AGM does not guarantee, but the title and abstract overstate what the controlled studies can support.","major_comments":[{"comment":"The central objective contrast is confounded with information inclusion. AGM-supported materials include the repository-defined risk-zone reference, risk summaries, evidence indexes, contributor-confirmation declarations, and explicit gate states (Table S30). The outcome rubric codes whether the output names the same risk label, evidence status, and gate state (Table S31). The 97.4% vs. 40.5% difference therefore largely demonstrates that participants can read an answer that is literally present in the stimulus; it does not isolate AGM's specific governance structure (risk zoning, evidence packaging, confirmation gates) as the causal mechanism. A control condition presenting the same governance information in plain, non-AGM prose, or an analysis holding information constant while varying structure, is needed to support the claim that AGM's design, rather than simply telling reviewers the","section":"§7.2, Tables S30–S31"},{"comment":"The text states that the condition difference is located in 'the structured externalization of repository-defined governance states.' This is an interpretation, not a demonstrated mechanism. The alternative explanation—that ordinary materials simply omitted the relevant risk-zone and gate information—is equally consistent with the data. The phrase 'structured externalization' presumes that the manifest's structure carries the effect. Since the design does not vary structure independently of information content, this specific attribution should be removed or explicitly flagged as untested.","section":"§7.2, 'task-level pattern' paragraph"},{"comment":"The contributor-side feasibility check shows that cooperative participants can fill AGM templates and that agent output can be confirmed by humans. It does not test whether contributors prepare evidence honestly or whether maintainers actually inspect what they confirm. The paper appropriately acknowledges this boundary in §4.5 ('Projects seeking stronger guarantees against ignored rules, omitted evidence, or fabricated declarations require additional identity, signing, audit, attestation, cryptographic, or platform-level mechanisms') and §8.3. Given that acknowledgment, the abstract's and conclusion's phrasing—that AGM makes agent-mediated contributions 'governable'—should be tempered to 'governable under cooperative, non-adversarial conditions,' with the controlled feasibility clearly distinguished from field-level assurance.","section":"§7.5 and Conclusion"}],"minor_comments":[{"comment":"The panel reports 'Gate-state availability or correctness: 0.0% vs. 100.0%.' Under ordinary materials, gate states were not available at all, so 0.0% reflects non-observability, not incorrect judgment. The label should distinguish availability from correctness, and the text should note that the comparison conflates these two dimensions.","section":"Figure 4A"},{"comment":"The abstract reports 'exact risk-label recovery (37/38 vs. 15/37)' without noting that AGM-supported materials contained the risk labels in the stimulus. A short qualifier such as 'when the same governance information was supplied through the manifest structure' would prevent misreading.","section":"Abstract"},{"comment":"Several legacy agreement statistics are reported with κ = 0.000, which can occur with low prevalence or skewed margins. Since these legacy variables were abandoned and replaced by the layered coding, the reporting is transparent, but a one-sentence explanation of why κ is uninformative in those rows would help.","section":"§5.3, Table S10"},{"comment":"The finding that 'no repository in the audit satisfies the four-function criterion' is partly a consequence of the criterion's strictness (canonical, repository-visible, coordinating all four functions). The paper explains this, but it should be stated even more explicitly that the audit measures absence of a particular coordinating arrangement, not absence of governance intent or of individual mechanisms.","section":"§6.1, Finding 1"}],"recommendation":"major_revision","confidential_remarks":"The core conceptual framing and the artifact itself are worthwhile, but the principal evaluation metric is a near-manipulation check. The revision should either add a control condition that supplies the same governance information in a plain format, or substantially reframe the contribution to claim only that AGM packages existing information in a form that makes it recoverable—without claiming that the structure itself causes the improvement. The contributor-side results are honest feasibility evidence but not evidence of governance assurance. I would not recommend rejection because the paper's explicit boundary statements show the authors are aware of these limits; however, the abstract and title currently outrun the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a well-constructed design-science paper with a genuinely useful artifact and an honest audit, but the central evaluation overstates what it shows. The 97.4% vs 40.5% risk-label recovery is real, yet it mostly shows that including the answer in the input makes agents output that answer. The AGM condition contained the exact risk labels, gate states, and evidence statuses that the rubric then coded as \"recovered.\" The ordinary condition omitted them. So the 56.8-point gap is closer to a reading-comprehension check than to evidence that AGM's particular governance design improves judgment. The paper is explicit about what its materials contained, so I don't read this as deception, but the headline claim \"AGM improves governance-state recovery\" is too strong.\n\nWhat is genuinely new: the AGM artifact itself, the 50-repository diagnostic audit, and the two controlled evaluations. The audit usefully separates agent-readability, traceability, and governability, and the finding that no audited repository has a project-wide governability arrangement is a real empirical claim, even if the four-function criterion was refined after coding disagreements. The stats are honest: participant-clustered bootstrap CIs, leave-one-out checks, pre-adjudication inter-coder reliability, and the disclosed single AGM error all speak well of the authors' care. The pinned public prototype and the explicit boundary discussion in §4.5 also earn credit.\n\nSoft spots, in proportion: first, the missing control — a condition that supplies the same governance information in plain prose or a simple template would let the effect be attributed to the manifest's structure rather than to information inclusion. That is the load-bearing fix. Second, the zero-repository audit finding is partly definitional; the criterion needs pre-registration or at least a robustness check with a looser threshold. Third, the contributor-side feasibility check used participants with every reason to comply, so it shows templates can be filled, not that evidence is truthful. The paper acknowledges this boundary, but the practical relevance of AGM hinges on it. Fourth, the replication dataset is promised but not yet there; the pinned artifacts help.\n\nWho this is for: people building or studying governance infrastructure for AI-generated code, and OSS maintainers looking for structured ways to handle contribution risk. A careful reader gets a useful artifact and a clear problem statement, plus a cautionary example of evaluation design near a manipulation check.\n\nRecommendation: send it to peer review. It deserves referee time, but expect a major revision — add the minimal-template control, justify or pre-register the audit criterion, and ship the replication data. With those, the contribution stands as a solid artifact study.","headline":"A solid, transparent design-science paper whose headline evaluation is close to a manipulation check: putting the labels in the stimulus makes the labels recoverable, but the artifact and audit are still worth a serious referee.","tokens_in":44910,"tokens_out":1498,"would_cite":true,"duration_ms":17028,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A repository-hosted Agent Governance Manifest can make AI-made contributions reviewable, raising exact risk-label recovery from 40.5% to 97.4% in a controlled test.","keywords":["open-source governance","AI agents","generation-verification asymmetry","Agent Governance Manifest","contribution review","evidence obligations","risk zones","maintainer authority"],"falsifier":"Give AGM a live trial with contributors who have incentives to pad evidence and reviewers who rubber-stamp confirmations; if fabricated or unverified packages pass structural validation and are merged at rates comparable to ordinary materials, the central claim that AGM makes contributions governable would be refuted.","tokens_in":43882,"feed_emoji":"🤖","tokens_out":4160,"duration_ms":33460,"temperature":0.7,"pith_summary":"The paper argues that coding agents have widened a generation–verification asymmetry in open-source software: AI makes contributions cheap to produce but does not make them cheap to verify. Its central claim is that projects need project-side governability infrastructure—a repository-hosted rule set, instantiated as the Agent Governance Manifest (AGM)—that turns risk classification, evidence obligations, human confirmation, and review gates into contribution-level states that reviewers can recover. A diagnostic audit of 50 repositories found widespread general governance and agent-readable guidance but no project-wide arrangement satisfying all four governability functions. In a controlled reviewer-side evaluation, AGM-supported materials recovered the exact risk level in 37 of 38 outputs versus 15 of 37 for ordinary materials; in a contributor-side feasibility check, all 45 final evidence packages represented the core governance state correctly. If the mechanism holds, maintainers can shift from reconstructing a contribution's governance state to verifying it, while keeping final acceptance authority in human hands.","feed_headline":"One manifest lifts risk-label recovery from 41% to 97%","feed_subtitle":"Repository-hosted rules turn AI-assisted pull requests into evidence-backed, human-confirmed reviews.","key_machinery":"The Agent Governance Manifest (AGM), a repository-hosted boundary resource described in a human-readable Markdown document and a structured YAML file. It acts as a bidirectional governance contract: on the contributor side it allocates risk-sensitive evidence obligations and contributor-confirmation declarations; on the maintainer side it defines review gates and review-support artifacts such as risk summaries, missing-evidence reports, test-evidence summaries, and gate states. The load-bearing mechanism is the externalization of governance states—risk classification, evidence status, accountability, and gate state—into inspectable contribution-level artifacts before review.","core_discovery":"The discovery is that governability is a distinct organizational function, separable from agent-readability and traceability, and that it can be externalized before review. AGM carries project rules into contribution-specific evidence packages: risk zones map changed files to evidence obligations; contributor-side agents prepare change summaries, test evidence, provenance notes, and missing-evidence reports; human contributors confirm declarations for high-risk changes; maintainer-side review packets expose risk, evidence, gate, and accountability states. The paper reports that this structured externalization made repository-defined risk levels recoverable in 97.4% of AGM-supported reviewer","pith_inferences":["Editorial inference: the same bidirectional-contract logic could generalize beyond open-source to any workflow where human approval gates machine-generated output—CI/CD pipelines, scientific analysis scripts, or regulated documentation—with risk-zoned evidence obligations configured per domain.","Editorial inference: because AGM validates structure, not truth, its real-world value will depend on complementary assurance such as signing, attestation, or audit in adversarial settings; the paper explicitly leaves those mechanisms outside its scope.","Editorial inference: a testable extension is whether AGM changes maintainer behavior in live repositories—for example, reducing time-to-review or slowing the acceptance of low-quality AI contributions—which the controlled evaluation does not measure.","Editorial inference: the audit's finding that no repository coordinates all four functions suggests governability is not an emergent byproduct of mature governance; it requires deliberate design as a distinct layer."],"forward_implications":["Projects using AGM can expect higher-risk AI-mediated changes to be classified accurately at review time, reducing governance-risk under-classification from 59.5% of ordinary outputs to 2.6% of AGM-supported outputs.","Contributor-side agents can prepare structured evidence packages that humans confirm, while maintainer-review status remains outside the contributor workflow.","Maintainer review shifts from open-ended reconstruction to targeted verification, with missing or placeholder evidence surfaced by review-support output.","The same governance contract can be rendered in human-facing review interfaces such as risk cards, checklists, and gate-status panels without changing the underlying schema.","AGM can complement existing agent-readable instruction files and provenance records, which the paper positions as inputs rather than substitutes."],"fun_headline_variants":["Project-wide rules lift AI-code risk labeling to 97%","Agent governance manifest boosts risk recall to 97%","Externalized rules make AI code contributions review-ready","Governability: the missing layer for AI-assisted OSS reviews","97% risk-label recovery with repository-hosted governance"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that contributors and their agents prepare evidence packages honestly and that human contributors actually inspect what they confirm; a structurally valid package with invented test results or a rubber-stamped confirmation would pass AGM's gates and mislead maintainers.","fun_headline_variants_meta":{"raw":{"variants":["Project-wide rules lift AI-code risk labeling to 97%","Agent governance manifest boosts risk recall to 97%","Externalized rules make AI code contributions review-ready","Governability: the missing layer for AI-assisted OSS reviews","97% risk-label recovery with repository-hosted governance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000589,"raw_usage":{"total_tokens":2620,"prompt_tokens":785,"completion_tokens":1835,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":1756}},"tokens_in":529,"tokens_out":1835,"duration_ms":10193,"temperature":1.0,"reasoning_tokens":1756,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T22:23:56.531467+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give AGM a live trial with contributors who have incentives to pad evidence and reviewers who rubber-stamp confirmations; if fabricated or unverified packages pass structural validation and are merged at rates comparable to ordinary materials, the central claim that AGM makes contributions governable would be refuted.","supporting_citations":[],"review_version":1}