{"id":"8bca88b3-87df-4fb4-96fd-1098aad99c7a","arxiv_id":"2607.01808","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Archer automates agentic code review for LLVM optimizations and reports finding semantic bugs in 21% of recent open PRs and 11% of closed PRs.","lead":"Archer is an agentic AI system that reviews LLVM compiler optimization pull requests by guiding agents with obligations and admitting only findings with executable validation evidence. If it works, it could supplement scarce expert review capacity and reduce semantic bugs entering large compiler codebases.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Bug rates (21%/11%) depend on unshown accuracy of agentic detection plus deterministic guard with no disclosed false-positive audit or reproduction details.","rationale":"Reader's weakest assumption matches the load-bearing point exactly; the abstract-only view already flags it, and the claim cannot be accepted without evidence that the guard and agent produce reliable bug reports rather than plausible but unverified ones.","tokens_in":1662,"tokens_out":306,"duration_ms":14082,"concrete_test":"Select the first 10 reported buggy PRs (5 open, 5 closed) from the evaluation set; for each, extract the exact executable evidence the guard accepted and attempt independent reproduction of the claimed miscompilation or semantic error on the relevant LLVM revision; if reproduction fails for >2 of the 10, the reported rates are unreliable.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The headline percentages require that Archer's obligations + guard correctly surface only genuine semantic bugs (miscompilations etc.) rather than false positives or unconfirmed cases. The abstract states the guard \"admits only findings backed by executable evidence\" but supplies neither the guard's decision procedure, any manual validation of the 70+328 PRs, nor even one concrete example of a reported bug with its evidence. For closed PRs this additionally requires showing that the merged change actually introduced a latent bug still present in LLVM. Absent these, the rates cannot be distinguished from over-detection.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper presents Archer, the first automated agentic code review tool for compiler optimizations in LLVM. It constrains the agentic process using obligations to guide analysis and a deterministic validation guard to admit only findings backed by executable evidence. Evaluation on 70 open PRs and 328 closed PRs from the last two months finds that 21% of open PRs and 11% of closed PRs introduce semantic bugs such as miscompilations.","tokens_in":1799,"tokens_out":376,"duration_ms":19438,"significance":"If the empirical claims are substantiated with transparent validation, the work would be significant for highlighting the limited capacity for critical review in large compiler projects and demonstrating a practical agentic approach that combines obligations with deterministic guards to reduce false positives in complex domains. The scale of the evaluation (398 PRs) and the focus on real LLVM changes provide a concrete testbed for such tools.","major_comments":[{"comment":"Abstract and Evaluation section: the central claims that Archer discovers 21% of open PRs and 11% of closed PRs are buggy rest on the assertion that the deterministic validation guard 'admits only findings backed by executable evidence,' yet the manuscript supplies no description of the guard's decision procedure, no false-positive audit, no manual validation of the 70+328 cases, and no concrete example of any reported bug together with its executable evidence. This directly undermines the headline percentages.","section":"Abstract and Evaluation"}],"minor_comments":[{"comment":"The abstract states the PRs are 'from the last two months' but does not give the exact date range or the selection criteria used to obtain the 398 PRs.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on the transparency of our evaluation. We address the major comment below and will revise the manuscript to strengthen the substantiation of our empirical claims.","responses":[{"response":"We agree that the manuscript would benefit from expanded detail on the deterministic validation guard to better support the reported percentages. In the revised version, we will add a precise description of the guard's decision procedure, including the criteria and mechanisms it uses to verify that each finding is backed by executable evidence. We will also include at least one concrete example of a reported bug, presenting the relevant PR change, the agent's analysis, and the specific executable evidence (e.g., test case or execution result) that caused the guard to admit the finding. Additionally, we will report a manual false-positive audit performed on a random sample of the admitted findings from both the open and closed PR sets. While a complete manual validation of all 398 PRs exceeds the practical scope of this work, the deterministic, evidence-based design of the guard provides the primary substantiation for the results; the added sample audit and example will further address concerns about the headline percentages.","revision_made":"yes","referee_comment":"[Abstract and Evaluation] Abstract and Evaluation section: the central claims that Archer discovers 21% of open PRs and 11% of closed PRs are buggy rest on the assertion that the deterministic validation guard 'admits only findings backed by executable evidence,' yet the manuscript supplies no description of the guard's decision procedure, no false-positive audit, no manual validation of the 70+328 cases, and no concrete example of any reported bug together with its executable evidence. This directly undermines the headline percentages."}],"tokens_in":1290,"tokens_out":369,"duration_ms":17179,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper introduces Archer as the first agentic review system aimed specifically at compiler optimizations in LLVM. It uses obligations to steer the analysis and a deterministic guard that only accepts findings with executable evidence. That framing is a reasonable response to why off-the-shelf code-review agents fall short on compilers.\n\nWhat stands out is the attempt to make the process more constrained and reproducible than a free-form LLM reviewer. The evaluation runs on real recent PRs (70 open, 328 closed), which is the right kind of test set.\n\nThe soft spot is the central claim. The abstract states that Archer found semantic bugs in 21% of open PRs and 11% of closed ones, yet supplies no description of the guard's decision rules, no manual audit of the flagged cases, and no single concrete example with the executable evidence that supposedly backs it. For the closed PRs the additional step of showing the merged change actually introduced a still-present latent bug is also missing. Without those pieces the percentages cannot be separated from possible over-detection.\n\nThe work is aimed at people who maintain large compiler codebases or build automated review tools. A reader looking for ideas on how to add domain-specific guardrails to an agent might pick up useful details on the obligation structure. The empirical headline, however, needs the missing validation steps before it can be used.\n\nI would send it to peer review. The problem is real and the constrained-agent approach is worth referee scrutiny, but the authors should be asked to document the guard, provide at least a few worked examples, and show how they controlled for false positives.","headline":"Archer's 21% and 11% bug rates on LLVM PRs rest on an agentic detector whose accuracy is not shown, so the headline numbers cannot be taken as evidence yet.","tokens_in":2247,"tokens_out":409,"would_cite":false,"duration_ms":13251,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Archer finds semantic bugs in 21% of open LLVM optimization pull requests and 11% of closed ones.","keywords":["agentic code review","compiler optimizations","LLVM","semantic bugs","pull request review","miscompilation","automated review","validation guard"],"falsifier":"Independent manual verification or re-testing of the specific PRs flagged by Archer to confirm whether they actually introduce miscompilations or other semantic changes.","tokens_in":2566,"feed_emoji":"🔍","tokens_out":597,"duration_ms":24942,"temperature":0.7,"pith_summary":"The paper introduces Archer as an automated agentic code review tool designed specifically for compiler optimizations in LLVM. It guides the review process from both ends by using obligations to direct the agent's analysis and a deterministic validation guard that admits only findings supported by executable evidence. When run on 70 open and 328 closed recent LLVM PRs, Archer reports that 21% of the open PRs and 11% of the closed PRs introduce semantic bugs such as miscompilations. The authors conclude that this reveals a critical shortfall in expert review capacity for large compiler projects and positions Archer as a practical additional reviewer.","feed_headline":"Archer flags semantic bugs in 21% of open LLVM PRs","feed_subtitle":"The agent uses obligations and executable evidence checks to review optimization changes and exposes gaps in manual oversight.","key_machinery":"Archer, the agentic review system that applies obligations to guide analysis and a deterministic validation guard to accept only executable-evidence-backed findings.","core_discovery":"Archer constrains agentic review with obligations and a deterministic validation guard that requires executable evidence, and its application to recent LLVM PRs shows that 21% of open PRs and 11% of closed PRs introduce semantic bugs such as miscompilations.","pith_inferences":["The same constrained agentic approach might be adapted to review changes in other large, correctness-critical codebases such as operating system kernels.","The reported bug rates suggest that existing test suites and continuous integration for LLVM may leave certain semantic properties under-checked.","If the validation guard can be made more general, Archer-style review could shorten the time between patch submission and safe merge while reducing bug escape."],"forward_implications":["A substantial fraction of compiler optimization changes may enter the codebase with undetected semantic errors.","Expert review capacity in large compiler projects is insufficient to catch all such issues before integration.","An automated tool using obligations and executable validation can serve as a scalable additional reviewer for optimization PRs."],"fun_headline_variants":["Archer detects bugs in 21% open LLVM PRs","21% open and 11% closed LLVM PRs have semantic bugs per Archer","Archer finds 21% of open PRs buggy in LLVM","LLVM sees 21% open PR bugs flagged by Archer"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The validation guard and agentic analysis correctly identify actual semantic bugs without substantial false positives or missed cases.","fun_headline_variants_meta":{"raw":{"variants":["Archer detects bugs in 21% open LLVM PRs","21% open and 11% closed LLVM PRs have semantic bugs per Archer","Archer finds 21% of open PRs buggy in LLVM","LLVM sees 21% open PR bugs flagged by Archer"]},"model":"grok-4.3","cost_usd":0.007326,"raw_usage":{"total_tokens":3343,"prompt_tokens":610,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":73262000,"prompt_tokens_details":{"text_tokens":610,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2660,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":610,"tokens_out":73,"duration_ms":17303,"temperature":1.0,"reasoning_tokens":2660,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-03T08:54:15.786171+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Independent manual verification or re-testing of the specific PRs flagged by Archer to confirm whether they actually introduce miscompilations or other semantic changes.","supporting_citations":[],"review_version":1}