{"id":"617f17e2-6aa5-46b8-99c4-c030c6faea55","arxiv_id":"2605.24521","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Survey of 162 vibe coders finds perceptions of AI code quality similar across experience levels but motivations, interaction styles, and quality assurance practices diverge, revealing a perception-action gap.","lead":"This paper surveys 162 users of AI code generation tools across non-coder, novice, and professional developer groups to study differences in vibe coding practices. A smart generalist might read it to understand how AI lowers barriers to software creation while leaving verification skills unevenly distributed.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Self-reported QA practices may not validly measure actual verification capacity differences by experience","rationale":"The reader's weakest assumption directly identifies the same methodological vulnerability; the full-text description of the survey method does not add independent validation that would remove it.","tokens_in":1698,"tokens_out":257,"duration_ms":17918,"concrete_test":"Re-administer the survey instrument alongside a standardized verification task (identical buggy AI-generated functions given to all groups) and compare self-reported verification frequency against observed bug-detection rates; if the correlation between experience group and task performance is weaker than the self-report divergence, the gap claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The perception-action gap claim requires that reported differences in quality assurance (e.g., debugging frequency, verification steps) reflect genuine differences in participants' ability to evaluate AI-generated code. The survey design relies entirely on retrospective self-reports from 162 participants grouped by experience; no objective performance measures, think-aloud protocols, or artifact analysis are described that would distinguish reporting bias or differing internal standards from actual capability. If professionals simply apply stricter criteria when self-assessing, the apparent experience-dependence could be an artifact rather than evidence of a capacity gap.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper reports results from a survey of 162 participants grouped into non-coders, novices, and professional developers on their use of AI code generation tools for 'vibe coding' (prompt-based generation evaluated primarily by execution). It claims that perceptions of code quality and risk awareness are broadly similar across groups, while motivations, interaction styles, and quality-assurance practices diverge with experience. The authors synthesize these patterns as a perception-action gap in which general awareness of AI-code limitations is distributed but the capacity to evaluate, debug, and verify remains experience-dependent, implying partial rather than full democratization of software creation.","tokens_in":1807,"tokens_out":613,"duration_ms":31837,"significance":"If the empirical patterns hold after methodological strengthening, the work usefully extends prior AI-assisted programming studies by including non-professional users and by framing experience as a selective rather than uniform shaper of practice. The multi-group design and the explicit articulation of a perception-action gap supply a concrete hypothesis that future work can test with performance measures. The study also supplies descriptive data on how non-developers versus professionals approach the same tools, which is relevant to tool builders and computing-education researchers.","major_comments":[{"comment":"Methods section: the manuscript provides no information on questionnaire item wording, response scales, piloting, or any validation steps for the self-report measures of quality-assurance practices (debugging frequency, verification steps). Because the perception-action gap claim rests on interpreting these self-reports as evidence of differential verification capacity, the absence of instrument details and bias-mitigation procedures is load-bearing.","section":"Methods"},{"comment":"Results/Discussion: the synthesis that 'capacity to evaluate, debug, and verify remains experience-dependent' is drawn solely from retrospective self-reports without objective performance tasks, think-aloud protocols, or artifact analysis. If professionals simply apply stricter internal standards when answering, the reported differences could reflect reporting bias rather than capability; this alternative explanation is not addressed and directly threatens the central claim.","section":"Results and Discussion"},{"comment":"Participant grouping: the criteria used to assign the 162 respondents to the three experience categories (non-coders, novices, professionals) and any checks for overlap or sampling bias are not described. Clean separability of groups is presupposed by all cross-group comparisons yet is not demonstrated.","section":"Methods"}],"minor_comments":[{"comment":"Abstract: the sample size (162) and group breakdown appear only in the body; moving the key demographic numbers into the abstract would improve immediate readability.","section":"Abstract"},{"comment":"Terminology: 'vibe coding' is introduced without an explicit operational definition or citation to prior usage; a short definitional sentence in the introduction would reduce ambiguity.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback, which identifies key areas where methodological transparency can be strengthened. We respond to each major comment below, indicating planned revisions where appropriate.","responses":[{"response":"We agree the manuscript should provide these details. The revised version will include the exact wording of items related to quality-assurance practices, the 5-point response scales employed, a description of the pilot testing with 12 participants, and steps taken to mitigate bias such as anonymous administration and neutral item phrasing.","revision_made":"yes","referee_comment":"[Methods] Methods section: the manuscript provides no information on questionnaire item wording, response scales, piloting, or any validation steps for the self-report measures of quality-assurance practices (debugging frequency, verification steps). Because the perception-action gap claim rests on interpreting these self-reports as evidence of differential verification capacity, the absence of instrument details and bias-mitigation procedures is load-bearing."},{"response":"Our study is a survey capturing self-reported perceptions and practices; objective performance data would require a different design. We will revise the discussion to explicitly acknowledge the possibility of reporting bias as a limitation and to qualify the perception-action gap as reflecting reported divergences in practices rather than directly measured capability. This preserves the descriptive contribution while noting the need for future behavioral studies.","revision_made":"partial","referee_comment":"[Results and Discussion] Results/Discussion: the synthesis that 'capacity to evaluate, debug, and verify remains experience-dependent' is drawn solely from retrospective self-reports without objective performance tasks, think-aloud protocols, or artifact analysis. If professionals simply apply stricter internal standards when answering, the reported differences could reflect reporting bias rather than capability; this alternative explanation is not addressed and directly threatens the central claim."},{"response":"Grouping relied on self-reported screening questions: non-coders reported zero prior coding experience, novices reported 0–2 years, and professionals reported more than 2 years plus current employment in software development. The revised methods section will detail these criteria, report any post-survey checks for group overlap, and discuss potential sampling biases arising from recruitment channels.","revision_made":"yes","referee_comment":"[Methods] Participant grouping: the criteria used to assign the 162 respondents to the three experience categories (non-coders, novices, and professionals) and any checks for overlap or sampling bias are not described. Clean separability of groups is presupposed by all cross-group comparisons yet is not demonstrated."}],"tokens_in":1494,"tokens_out":545,"duration_ms":29720,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper surveys 162 people across non-coders, novices, and professional developers on vibe coding and reports that perceptions of AI code quality are similar across groups while motivations, interaction styles, and quality assurance practices differ by experience. Non-coders cite accessibility, novices focus on learning, and professionals tie it more to work tasks. The authors frame this as a perception-action gap where awareness of risks is common but actual evaluation capacity is experience-dependent.\n\nThe work fills a clear gap by moving beyond the professional-only samples in earlier studies and supplies new empirical distinctions on how experience selectively shapes these practices. That part is useful for anyone thinking about AI tool design or education.\n\nThe main limitation is the exclusive reliance on retrospective self-reports. The stress-test concern holds: without objective measures, think-aloud data, or artifact checks, reported differences in debugging frequency or verification steps could reflect differing internal standards or social desirability rather than genuine capacity gaps. The abstract gives no information on question wording, response rates, or statistical handling, which makes it difficult to assess how cleanly the three groups separate or how much sampling bias affects the results.\n\nThis is for researchers working on AI-assisted software engineering or computing education who want data on non-professional users. A reader already following vibe coding or prompt-based development would find the experience breakdowns worth seeing.\n\nIt deserves peer review so the methods and data can be examined directly; the self-report issue is fixable with clearer limitations or supplementary validation but needs that scrutiny first.","headline":"The survey documents experience-based differences in motivations and reported QA steps for vibe coding but the perception-action gap rests on unvalidated self-reports.","tokens_in":2261,"tokens_out":374,"would_cite":false,"duration_ms":19505,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Experience creates a perception-action gap where awareness of AI code risks is common but the ability to verify outputs depends on prior development experience.","keywords":["vibe coding","AI code generation","experience levels","perception-action gap","software verification","user survey","quality assurance","non-professional developers"],"falsifier":"An observational study that records actual prompting, execution, and verification steps of matched participants from each experience group and compares those behaviours against the survey self-reports.","tokens_in":2607,"feed_emoji":"","tokens_out":640,"duration_ms":16498,"temperature":0.7,"pith_summary":"The paper surveys 162 users of AI code generation tools divided into non-coders, novices, and professional developers to compare their vibe coding practices. Perceptions of code quality and recognition of risks remain similar across all groups, yet motivations, interaction patterns, and quality assurance methods diverge sharply with experience. Non-coders are drawn by accessibility, novices focus on learning through experimentation, and professionals apply the approach more often in work settings. The authors synthesise the results as a perception-action gap in which general risk awareness is widely shared while the capacity to evaluate, debug, and verify AI-generated code stays experience-dependent. This leads to the claim that vibe coding broadens access to software creation without distributing verification expertise equally.","feed_headline":"Experience creates a gap between seeing AI code risks and fixing them","feed_subtitle":"Survey of 162 users finds all groups spot problems in AI output but only experienced developers routinely verify and debug it.","key_machinery":"The perception-action gap, the disconnect between broadly distributed awareness of risks in AI-generated code and the experience-dependent capacity to evaluate, debug, and verify it.","core_discovery":"The central claim is that vibe coding produces a perception-action gap: all three experience groups recognise both the strengths and limitations of AI-generated code, yet only the capacity to evaluate, debug, and verify outputs scales with prior development experience. Motivations and quality-assurance behaviours therefore separate cleanly by group while reported perceptions of quality do not.","pith_inferences":["The gap suggests that training focused on verification techniques rather than prompting alone could narrow differences between groups.","Non-coders may generate larger volumes of unverified code that later requires professional intervention.","Tool designers could add explicit verification scaffolding that reduces reliance on prior experience."],"forward_implications":["Vibe coding partially democratises software creation by widening access while leaving verification skills unevenly distributed.","Quality assurance practices and interaction styles improve selectively with development experience.","Non-coders are motivated primarily by accessibility, novices by learning and experimentation, and professionals by work-related tasks.","General awareness of AI code risks does not translate into equivalent verification behaviour across experience levels."],"fun_headline_variants":["Experience creates verification gap in vibe coding","All spot AI risks but experience enables fixes","Perception action gap shapes vibe coding behaviors","Only experienced users routinely debug AI code output"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Self-reported survey answers from the 162 participants accurately capture their real practices and the three experience groups are cleanly separable without major sampling or response biases.","fun_headline_variants_meta":{"raw":{"variants":["Experience creates verification gap in vibe coding","All spot AI risks but experience enables fixes","Perception action gap shapes vibe coding behaviors","Only experienced users routinely debug AI code output"]},"model":"grok-4.3","cost_usd":0.00483,"raw_usage":{"total_tokens":2369,"prompt_tokens":659,"num_sources_used":0,"completion_tokens":51,"cost_in_usd_ticks":48299500,"prompt_tokens_details":{"text_tokens":659,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1659,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":659,"tokens_out":51,"duration_ms":16455,"temperature":1.0,"reasoning_tokens":1659,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T13:11:12.091717+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An observational study that records actual prompting, execution, and verification steps of matched participants from each experience group and compares those behaviours against the survey self-reports.","supporting_citations":[],"review_version":1}