{"id":"88b81e2e-7204-41b9-b59a-e0c783f7a9f4","arxiv_id":"2607.21652","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Vibe coding evidence from 47 sources describes an intent-driven, iterative evaluation loop whose productivity benefits are conditional on review and validation practices.","lead":"A systematic review of 47 academic and practitioner sources finds vibe coding is an iterative describe-generate-test-repair loop, not a one-shot prompt. It shows reported productivity gains concentrate in prototyping and UI work, with weak evidence on production, safety-critical, and long-term maintainability.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Search vocabulary and codebook jointly risk manufacturing the iterative-loop claim: iterative terms are included while no one-shot code exists.","rationale":"The reader's weakest assumption identified the exclusion of broader terms such as 'LLM-assisted programming' and 'agentic coding' as a potential bias. I agree that this is a plausible concern, but the stronger and more specific issue is the combination of (a) including terms that semantically presuppose iteration and (b) omitting a one-shot/no-evaluation code from the RQ2 codebook. This makes the review unable to register the very contrast its central claim asserts. The paper itself flags the search-recall limitation in Section 5.1, but not the directional bias toward iterative descriptions. The exact-phrase seed corpus provides a natural control: those records were retrieved before iterative synonyms were added, so re-coding that subset with a one-shot category would show whether the iterative conclusion is driven by the search vocabulary. The paper is otherwise well-documented, with traceable extraction and a released replication package, but the central claim should be conditional on this check rather than treated as established. Since the reader already issued a CONDITIONAL verdict focused on related methodological issues, my concern does not move the verdict; it sharpens the condition.","tokens_in":44262,"tokens_out":4600,"duration_ms":43423,"concrete_test":"Re-run the Stage A seed corpus (the 31 records retrieved by the exact phrase 'vibe coding', before any iterative-connoting synonyms were added) and code it with an explicit 'one-shot / no-evaluation / pure delegation' category alongside the existing Table 18 codes. If a substantial share (e.g., >20%) of exact-phrase sources describe accepting generated code without an evaluation/revision step, the 'rather than one-shot' conclusion is a search artifact. If nearly all exact-phrase sources describe an evaluation-revision loop, the concern fails. Ideally also add the excluded broader terms and compare, but the exact-phrase subset is the decisive control.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Takeaway 2, Section 3.3) is that vibe coding is an iterative generation–evaluation–refinement loop 'rather than a one-shot activity.' The review's own vocabulary and coding scheme make this conclusion difficult to falsify. The Stage B search string (Section 2.2.1) includes 'conversational programming' and 'prompt-based development*' — terms that semantically presuppose multi-turn or prompt-iteration workflows — while deliberately excluding 'LLM-assisted programming,' 'AI pair programming,' 'AI coding assistant,' and 'agentic coding' (Section 2.2.1). A corpus assembled this way is enriched for iterative descriptions. More importantly, Table 18's RQ2 codebook contains no code for one-shot generation, pure delegation, or no-evaluation acceptance, even though the abstract itself states that vibe coding is 'often framed as one-shot prompting.' A source that characterizes vibe coding as a single prompt followed by uncritical acceptance cannot be counted, so the contrast 'rather than a one-shot activity' is not an empirical result of the synthesis; it is built into the extraction categories. The paper's own threat analysis (Section 5.1) acknowledges reduced recall but not this directional bias. Because the 'consistently iterative' claim is the headline contribution and the foundation for the 'evaluation loop' framework, this is load-bearing. This is not an external-consensus objection: it is a question of whether the protocol could detect the opposite pattern.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a multivocal literature review (MLR) of 47 sources (28 peer-reviewed, 19 grey) on vibe coding, following Garousi et al.'s MLR guidelines. The review poses eight research questions covering definitions, workflows, developer-role shifts, outcomes, risks/safeguards, usage contexts, tools, and open challenges. The central claim is that vibe coding is consistently described in the literature as an iterative generation–evaluation–refinement loop rather than a one-shot activity, and that developer work shifts from writing code toward specification, supervision, and validation. The paper contributes a documented protocol, a replication package, an evidence-strength labeling scheme, a conceptual framework, and a comparison with prior reviews by Ge et al., Ray, and Fawzy et al.","tokens_in":44600,"tokens_out":5573,"duration_ms":50808,"significance":"If the central claim holds, this is a useful early consolidation of a fast-moving, practitioner-driven topic. The review's methodological transparency is a genuine strength: the search counts, inclusion/exclusion criteria, quality and credibility checklists, item-level scores, traceability from findings to source excerpts, and the public replication package exceed what is typical for an MLR in this area. The explicit cross-stream triangulation, especially in RQ5, and the careful denominator reporting are also commendable. However, the headline conclusion that vibe coding is 'iterative rather than one-shot' is, as argued below, not currently falsifiable given the protocol's search vocabulary and coding scheme. The paper is therefore valuable as a documented synthesis but needs protocol-level revision before its central interpretive claim can be accepted.","major_comments":[{"comment":"The RQ2 codebook contains no code for one-shot generation, pure delegation, or no-evaluation acceptance. Every code in Table 18 (validation pipelines, chat-based iterative loops, iterate–prompt–patch, etc.) presupposes multi-turn or validation activity. Takeaway 2's contrast—'iterative generation–evaluation–refinement loop rather than a one-shot code-generation activity'—is therefore not an empirical result of the synthesis; it is built into the extraction categories. The abstract itself states that vibe coding 'is often framed as one-shot prompting,' yet no RQ2 code records that framing and the Results section never counts or reconciles it. To make the central claim testable, the authors should add codes such as 'one-shot / direct generation without iterative refinement' and 'acceptance without evaluation,' re-code the corpus, and report how many sources characterize vibe coding that wa","section":"Section 3.3 / Table 18 / Section 2.2.4"},{"comment":"The Stage B search string includes 'prompt-based development*,' 'conversational programming,' and 'natural language programming,' terms that semantically presuppose iterative prompting or conversation, while deliberately excluding 'LLM-assisted programming,' 'AI pair programming,' 'AI coding assistant,' and 'agentic coding.' This is not merely a precision/recall trade-off; it is a directional sampling choice that enriches the corpus for iterative descriptions and removes a large body of literature that could contain one-shot or pure-delegation characterizations. Section 5.1 acknowledges reduced recall but not this directional bias. Because the headline claim is that the evidence 'consistently' describes an iterative loop, the authors should either broaden the search vocabulary to include the excluded terms, or provide a sensitivity analysis showing that the excluded literature would not","section":"Section 2.2.1 / Section 5.1"},{"comment":"The evidence-strength labels are assigned by the first author alone, without formal independent double-coding or an inter-rater agreement measure. Section 5.4 discloses this, but the label 'strong' attached to Takeaway 2 is used to support the central claim, and the underlying percentages (72%, 63%) are the product of a single coder's application of a codebook that lacks a one-shot category. For a review whose stated contribution is an 'evidence-weighted account,' this is a load-bearing reliability risk. I recommend independent coding of a sample of sources against the revised codebook, with agreement reported, or at minimum a clearer pre-specified rule for how source counts, quality bands, and consistency translate into strong/moderate/weak/emerging labels.","section":"Section 2.3 / Section 5.4"}],"minor_comments":[{"comment":"The phrase 'White Screening' appears in the figure; this should read 'Title Screening' or similar.","section":"Figure 2"},{"comment":"The sentence 'Validation and testing sit inside this loop rather than as optional additions' is presented as a synthesis claim, but the supporting Table 18 reports only mentions of validation patterns. Consider softening to 'are frequently described as sitting inside this loop' to avoid overstating the strength of the evidence.","section":"Section 3.3"},{"comment":"In the discussion of Ge et al., the sentence beginning 'Based on the information reported, however, its search...' is grammatically incomplete. Please revise.","section":"Section 6"},{"comment":"Table 28 reports 'Security concerns including unsafe code...' at 50% (10 of 20 RQ8-contributing sources). The subsequent sentence says 'A similar proportion of studies (30%, 6 sources)'—'similar proportion' is misleading; consider 'The next most frequent challenges...'.","section":"Section 3.9"},{"comment":"The reported snowballing yield is very low (2 peer-reviewed and 3 grey additions). This is acknowledged, but the reader would benefit from a brief statement of how many references were checked in forward vs. backward snowballing, since the current description gives only total candidate counts.","section":"Section 2.2.2"}],"recommendation":"major_revision","confidential_remarks":"I share the stress-test concern: the iterative-loop claim is protocol-built rather than empirically won. The manuscript is otherwise well-executed and unusually transparent for an MLR, with a replication package and careful cross-stream triangulation. The revision should focus on making the central claim falsifiable: add one-shot/no-evaluation codes, re-code the corpus, and either broaden the search or justify the directional exclusions with a sensitivity check. If the authors decline to re-code, the claim should be substantially weakened. I would not reject outright, since the issue is fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe short version: this is a well-executed, transparent literature review that will be a useful map of the vibe coding landscape — but don't treat its central \"iterative loop\" finding as solid. The paper's own search and coding scheme make it almost impossible for the review to see anything other than iteration, so the claim that vibe coding is \"rather than a one-shot activity\" is a protocol artifact more than an empirical result.\n\nWhat's genuinely good: this is the first MLR I know of that puts peer-reviewed and grey literature on vibe coding through one documented protocol, with search counts, quality/credibility checklists, evidence-strength labels, and a replication package. The traceability from each finding back to source IDs and excerpts is excellent. The developer-role shift, the security/quality trade-offs, and the domain skew (UI/prototyping heavy, data/ML and production thin) are handled with appropriate caution. The framework and research agenda are reasonable syntheses, not overclaims.\n\nThe soft spots, in order of seriousness. First, the one the stress-test flags: the RQ2 codebook (Table 18) has no category for one-shot generation, pure delegation, or no-evaluation acceptance, even though the abstract itself says vibe coding is \"often framed as one-shot prompting.\" The search string, meanwhile, includes terms like \"conversational programming\" and \"prompt-based development\" that presuppose multi-turn interaction, and deliberately excludes broader terms like \"AI pair programming\" or \"agentic coding.\" So a source that says \"you type a prompt and accept whatever comes out\" literally has nowhere to be coded in the workflow analysis. The 72% for \"validation and evaluation pipelines\" and 63% for \"chat-based iterative loops\" are real frequencies in the corpus, but the corpus was assembled and coded so that one-shot descriptions would dissolve into other categories. The paper's own threat analysis (5.1) acknowledges reduced recall but not this directional bias. That is load-bearing, because Takeaway 2 grounds the whole evaluation-loop framework.\n\nSecond, all screening, extraction, coding, and evidence-strength labeling were done by the first author. The paper transparently states this, and senior review of a sample is not the same as independent double-coding with agreement measures. That's fixable. Third, the inclusion of a low-quality source (WL28) from a borderline venue is a minor concern; they weighted it low, so it doesn't materially change anything.\n\nWho should read it: practitioners wanting an evidence-weighted overview of the topic, researchers working on AI-assisted SE, and anyone thinking about MLR methodology. It deserves a serious referee, but it needs revision before publication: add a one-shot/no-evaluation code and report its frequency, or reframe the \"rather than one-shot\" claim to avoid the false contrast; widen or justify the search exclusion; and add inter-rater agreement or at least a documented reliability sample.\n\nTake it seriously, but push on the protocol.","headline":"Well-documented MLR whose headline \"iterative loop\" claim is partly a product of the search string and codebook — treat that finding as provisional rather than established.","tokens_in":45033,"tokens_out":4637,"would_cite":true,"duration_ms":38583,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 47 sources finds that vibe coding is not one-shot prompting but an iterative loop of intent, generation, evaluation, and refinement, with productivity gains that hinge on the surrounding evaluation and governance.","keywords":["vibe coding","multivocal literature review","large language models","AI-assisted software development","human-AI collaboration","code generation","developer role shift","software engineering"],"falsifier":"A concrete falsifier: re-run the review's search with the deliberately excluded terms ('LLM-assisted programming,' 'AI coding assistant,' 'agentic coding') and count how many of the additionally retrieved sources describe vibe-coding-adjacent work as one-shot, fully delegated generation. If a substantial number do, the claim that the literature 'consistently' describes an iterative loop is falsified. A complementary experiment: compare matched teams using vibe coding with and without enforced test-and-review pipelines; if defect rates and maintainability do not differ, the claim that outcomes","tokens_in":44209,"feed_emoji":"🔁","tokens_out":4382,"duration_ms":38340,"temperature":0.7,"pith_summary":"This review synthesizes 28 peer-reviewed and 19 practitioner sources to establish what vibe coding actually is and how it behaves in software development. Its central claim is that vibe coding is consistently described as an iterative generation–evaluation–revision loop, not a one-shot prompt-to-code activity, and that the real bottleneck is the strength of the evaluation loop around generated code rather than generation speed. The review finds that developer work shifts from writing code to specifying intent, supervising output, and validating results, and that reported productivity gains, appearing in 45% of sources, are conditional on expertise, task type, and verification practices. Evidence is strongest for prototyping and UI work and weakest for production, data-intensive, and safety-critical settings, and tool visibility does not imply demonstrated effectiveness. A sympathetic reader should care because this reframes a hype term as an engineering workflow with identifiable controls and testable failure modes.","feed_headline":"Vibe coding is an iterative loop, not one-shot prompting","feed_subtitle":"A 47-source review finds productivity gains hinge on testing and review, with evidence thinnest in production work.","key_machinery":"The central object is the vibe coding loop: a feedback-driven cycle of intent specification, LLM generation, developer evaluation, and prompt refinement, with validation and testing inside the loop rather than optional additions. The paper's conceptual framework adds four interacting layers—workflow, role and experience, outcomes, and risk and governance—moderated by usage context and tool ecosystem, and attaches an evidence-strength label (strong, moderate, weak, emerging) to each finding. That combination of loop, layered model, and evidence weighting is what carries the argument.","core_discovery":"On its own terms, the paper's central discovery is that the scattered literature on vibe coding converges on a single process shape: the developer states intent in natural language, an LLM generates or revises code, the developer evaluates the result by running, inspecting, or testing it, and then refines the next prompt based on observed behavior. Across 43 of 47 sources, validation and evaluation pipelines (72%) and chat-based iterative loops (63%) are the dominant workflow patterns, and across 36 sources the practice is described most often as an intent-driven socio-technical practice (69%) and as natural-language-to-code prompting (67%). The review labels this an iterative control system","pith_inferences":["If the iterative-control-system framing is right, a testable extension is to instrument real development sessions and check whether teams with stronger evaluation loops (CI gates, test coverage, structured review) show lower escaped-defect rates and better maintainability than teams that iterate without them.","The review's exclusion of broader terms like 'LLM-assisted programming' and 'agentic coding' suggests a conservative lower bound: one-shot or pure-delegation characterizations may exist in that excluded literature, so the 'consistently iterative' claim is best read as applying to the vibe-coding-specific corpus.","The novice/expert split implies a design consequence: tools and training should target verification skill (how to evaluate and test generated code) rather than only prompt skill, because overtrust is most dangerous for the novice users the practice is said to empower."],"forward_implications":["Vibe coding should be managed as an engineering workflow: teams that adopt it without automated tests, review routines, and validation pipelines are likely trading short-term speed for long-term defect and maintenance burden.","Reported productivity gains are most credible for prototyping and UI/front-end work; claims of gains in production, data-intensive, or safety-critical settings should not be assumed until studied.","Safeguards such as human-in-the-loop review and validation pipelines are consistently recommended, but since their effectiveness is under-tested, they should be treated as hypotheses to evaluate rather than proven controls.","Session-level dynamics—momentum breaks, context drift, and repetitive prompt–patch cycles—are real but under-measured, pointing to a need for instrumented long-horizon studies.","The review's evidence-strength labels provide a practical map: strong claims are limited to the iterative-loop description and short-term productivity; everything else is moderate, weak, or emerging."],"fun_headline_variants":["Vibe coding is a loop, not a one-shot prompt","Vibe coding's edge: iteration, not one-shot prompting","Review: Vibe coding is iterative generation and review","Vibe coding: validation makes the loop, says 47-source review","Vibe coding works as a feedback loop, says literature review"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The conclusion that vibe coding is 'consistently' iterative rests on a search string and eligibility rules that deliberately excluded literature using broader terms such as 'LLM-assisted programming,' 'AI coding assistant,' and 'agentic coding'; if that excluded literature contains many one-shot or pure-delegation accounts, the consistency claim would overstate the evidence.","fun_headline_variants_meta":{"raw":{"variants":["Vibe coding is a loop, not a one-shot prompt","Vibe coding's edge: iteration, not one-shot prompting","Review: Vibe coding is iterative generation and review","Vibe coding: validation makes the loop, says 47-source review","Vibe coding works as a feedback loop, says literature review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000248,"raw_usage":{"total_tokens":1416,"prompt_tokens":813,"completion_tokens":603,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":527}},"tokens_in":557,"tokens_out":603,"duration_ms":5316,"temperature":1.0,"reasoning_tokens":527,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:54:31.674439+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete falsifier: re-run the review's search with the deliberately excluded terms ('LLM-assisted programming,' 'AI coding assistant,' 'agentic coding') and count how many of the additionally retrieved sources describe vibe-coding-adjacent work as one-shot, fully delegated generation. If a substantial number do, the claim that the literature 'consistently' describes an iterative loop is falsified. A complementary experiment: compare matched teams using vibe coding with and without enforced test-and-review pipelines; if defect rates and maintainability do not differ, the claim that outcomes","supporting_citations":[],"review_version":1}