{"id":"2208ff92-bacf-4985-bd78-2e8415d95cc4","arxiv_id":"2607.29178","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Even without team rules, hackathon participants reported wanting to check generative AI output, but time pressure and limited domain knowledge often stopped them from truly verifying it.","lead":"This paper interviews four hackathon participants about how they used generative AI during a two-day event. It finds that even with no team rules, everyone reported checking AI output, but time pressure and unfamiliar topics limited real verification.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claim that all participants converged on checking GenAI output is unsupported for P2; Section 4.3 quotes only P1, P3, P4, while P2 reports inability to judge.","rationale":"The reader's weakest assumption concerned reliance on one self-reported member per team and retrospective recall. My concern is narrower and more concrete: within the reported data, the universal claim about 'all participants' is not backed by a quote from P2. This does not change the reader's conditional verdict; it strengthens the need for revision. The paper is otherwise honest about its limitations, clearly frames team-level claims as perceptions, and uses qualitative methods appropriately for an exploratory study. However, the abstract and Section 4.3 overstate support for the key finding. A single missing quote is easily fixable, but without it the central convergence claim is not fully evidenced. Thus the correct disposition remains CONDITIONAL, matching the reader's verdict; no change to the verdict is required.","tokens_in":8687,"tokens_out":3311,"duration_ms":33230,"concrete_test":"Request the full interview transcript or a direct quote from P2 concerning checking or adjusting GenAI output before use. Then independently code P2's statements about verification. If P2 made no statement expressing a desire to check/adjust GenAI output, revise the finding to 'three of four participants described wanting to check' and soften the abstract accordingly. If P2 did make such a statement, include the quote in Section 4.3 and the claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central finding, restated in the abstract as 'all participants we studied converged on an unwritten practice of checking GenAI output before using it,' relies on Section 4.3's assertion: 'every participant described wanting to check or adjust GenAI output before using it.' The immediately following quotes support this for P1, P3, and P4, but no P2 quote is provided. The only P2-specific material in that section describes the opposite: 'I often didn't really understand whether it was doing the right thing or not, because I'm not that strong mathematically' and 'couldn't really judge a lot of the time.' Wanting to check is not the same as being unable to judge; the latter could coexist with wanting to check, but the paper does not show that P2 expressed that desire. Since the headline finding is a universal quantifier over all four participants, one unsupported case weakens it from 'all' to 'three of four.' The paper's own Section 6 cautions that team-level claims are perceptions of a single member, but this issue is more specific: even within the self-reported data as presented, the evidence for P2's convergence is missing. This is not a reliability concern about retrospective recall generally; it is an internal-evidence gap in the reported excerpts. The authors may have additional transcript data for P2, but the manuscript as written does not support the universal claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative interview study of four participants from different teams at a two-day AI-themed hackathon in central Europe. It addresses two research questions: how participants use generative AI (GenAI) across tasks and tools, and how they verify GenAI outputs under hackathon time pressure. The main findings are that participants used GenAI for learning, ideation, coding, and documentation; that they combined multiple GenAI and non-GenAI tools according to task fit; and that, despite the absence of explicit team rules, all participants supposedly converged on an unwritten practice of checking GenAI output before using it, albeit constrained by time pressure and limited domain knowledge. The paper explicitly frames the study as exploratory and proposes a follow-up mixed-methods design combining observation, surveys, and prompt-and-response logs.","tokens_in":9019,"tokens_out":4543,"duration_ms":42206,"significance":"If the convergence claim holds, the paper makes a useful empirical contribution: it complicates the assumption that hackathon time pressure automatically leads to uncritical acceptance of GenAI output, and it shifts attention to supporting informal verification practices rather than merely permitting or restricting GenAI use. The paper is transparent about its limitations—small sample, single event, one interviewee per team, perception-based data, single-coder thematic analysis—and provides a valuable interview guide in the appendix. Its contribution is modest but appropriate for an exploratory empirical study. However, the headline universal claim is not fully supported by the evidence actually presented, and the Discussion repeats that unsupported universal as if established. These are load-bearing issues that need to be fixed before the conclusions can be accepted as stated.","major_comments":[{"comment":"The headline finding—'all participants we studied converged on an unwritten practice of checking GenAI output before using it'—is supported by direct quotes only for P1, P3, and P4. P2's only quoted statements in this section are 'I often didn't really understand whether it was doing the right thing or not, because I'm not that strong mathematically' and 'couldn't really judge a lot of the time.' Wanting to check is not established by these quotes; if anything, they report an inability to verify. Since the universal quantifier is load-bearing for RQ2 and the abstract, either a supporting quote from P2 must be added or the claim should be hedged to 'three of four participants described...' or otherwise qualified.","section":"Section 4.3 and Abstract"},{"comment":"The Discussion repeats the unsupported universal: 'all four participants described this expectation as something they held.' This is stronger than the data in §4.3, which supplies explicit statements for P1, P3, and P4 only. The Discussion also moves from 'described wanting to check' to 'described this expectation' without additional data. Please align the Discussion with whatever evidence exists for P2, or narrow the claim to the participants who actually articulated it.","section":"Section 5.1"},{"comment":"The text says 'Two constraints recurred across the cases.' In the reported excerpts, the domain-knowledge constraint is illustrated only by P2 and the time constraint only by P1. 'Recurred' overstates the support: the data show these constraints were salient to at least one participant each, not that they recurred across multiple cases. Either provide evidence from other participants or revise the phrasing to something like 'two constraints appeared in our data.'","section":"Section 4.3, 'Two constraints'"},{"comment":"The claim that 'no team set explicit rules' is presented as a fact in §4.3 and the abstract, but Section 6 limits team-level claims to a single member's perception. For an exploratory study this is an acceptable limitation, but the wording should carry the same hedge that §5.1 uses ('as far as participants reported'), because the absence-of-rules premise is part of the headline contrast.","section":"Section 4.3 and Section 6"}],"minor_comments":[{"comment":"There are spacing/formatting inconsistencies in phrases like 'relevant toRQ2' and 'relevant toRQ 1' (missing space before 'RQ' in one place).","section":"Section 3, first paragraph"},{"comment":"The DOI for the Pe-Than et al. paper appears malformed or inconsistent with the publisher's usual format; please verify it.","section":"References, [17]"},{"comment":"The bullet 'Despite no explicit team rules, participants checked GenAI output' would be more accurate if written as 'participants reported checking GenAI output in this small sample,' consistent with the evidence in the body.","section":"Section 4, Key findings bullet"}],"recommendation":"major_revision","confidential_remarks":"The evidentiary gap for P2 is fixable by adding a supporting quote or by softening the universal claim; I would not reject the paper for it. The manuscript is at the exploratory end of the empirical spectrum, and the abstract's unqualified universal phrasing should be adjusted to match the data before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a small, honest interview study that gives a plausible account of how hackathon teams use GenAI, with one real evidentiary gap in its headline claim.\n\nWhat's new: interviews, rather than surveys or observations, to capture participants' reasoning about tool choice and verification in a hackathon. The finding that participants bring a checking habit from everyday work, and that it frays under time pressure and unfamiliar domains, is a useful complement to Sajja et al. and Gama et al. The paper is careful to describe task-fit tool combinations and boundaries. Credit where due: the limitations section is unusually candid, acknowledging sample size, self-selection, single perspective per team, and perception-based data.\n\nThe soft spot is the universal quantifier. In Section 4.3 the authors write that 'every participant described wanting to check or adjust GenAI output,' but the supporting quotes are from P1, P3, and P4. The only P2-specific quotes express the opposite: not understanding whether the tool was doing the right thing, and not being able to judge. Wanting to check and being unable to judge can certainly coexist, but the paper does not show P2 expressing the wanting. Since the abstract restates convergence as 'all participants we studied,' this is an internal evidence gap, not just a recall concern. It is fixable: quote P2 on wanting to check (if it exists) or soften the claim to 'most participants.' The paper's own limitations section already disclaims team-level claims as perceptions, which covers the one-participant-per-team issue, but the P2 gap is narrower and remains.\n\nThe rest of the analysis is proportionate. The discussion about organizers supporting existing checking practices rather than merely permitting or banning GenAI is reasonable. I don't see a load-bearing flaw elsewhere.\n\nWho should read this: hackathon organizers and researchers studying informal GenAI governance. It's too small to change practice, but it points to a follow-up design (observation, logs) that makes sense.\n\nRecommendation: send it to peer review. The study design is appropriate for an exploratory qualitative paper, and the issues are addressable in revision. A serious referee would ask for the P2 evidence and a more careful abstract.","headline":"Small, honest hackathon interview study with a solid core and one overclaim: the 'all participants' checking finding lacks P2's supporting quote.","tokens_in":9456,"tokens_out":2184,"would_cite":false,"duration_ms":22161,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Despite having no team rules, hackathon participants converged on an unwritten practice of checking generative AI output before use, but time pressure and unfamiliar domains limited how well they could verify.","keywords":["hackathons","generative AI","verification","software engineering","team norms","qualitative interviews","time pressure"],"falsifier":"A direct observation study with screen recordings and interaction logs of hackathon teams: if a notable share of AI-generated code is adopted without any edit, review, or cross-check, the universal-checking claim would be falsified.","tokens_in":8608,"feed_emoji":"🤖","tokens_out":8779,"duration_ms":71336,"temperature":0.7,"pith_summary":"This paper asks how hackathon teams use generative AI and whether they verify its output. Based on interviews with four participants from four teams at a two-day hackathon, it reports that participants used AI for learning, brainstorming, coding, and documentation, and mixed AI and non-AI tools by task fit. The central finding is that despite the absence of team rules, every participant described an unwritten practice of checking or editing AI output before using it. Yet checking was fragile: participants in unfamiliar domains could not always judge correctness, and time pressure pushed them to accept larger, less-scrutinized chunks. If correct, this shifts the question for hackathon design from whether to permit AI to how to support the verification habits participants already sustain.","feed_headline":"Hackathon teams verify AI output even without explicit rules","feed_subtitle":"Interviews at a two-day hackathon show informal checking of AI output - but time pressure and domain gaps limit it.","key_machinery":"The central object is the unwritten checking practice: a habit of reviewing, editing, or cross-validating generative AI output before use, which the paper finds operating even where no team rules exist. The mechanism has two parts—an individual motive (being able to explain and take responsibility for submitted code) and a social one (group pressure not to let all code be AI-written). Its effectiveness is determined by two opposing forces: it persists because participants bring it from everyday work, and it degrades because time pressure and domain-knowledge limits shrink the review each output receives. This object carries the paper's argument that governance can emerge bottom-up in ad hoc","core_discovery":"The paper's core discovery is that informal verification norms can arise in temporary, fast-moving teams even without explicit governance. All four interviewed participants—from different teams—reported that AI output was always edited, cross-checked, or at least read before use, motivated by a sense of individual responsibility or group pressure to understand one's own code. The paper also documents the limits of that practice: participants with weak domain knowledge said they could not reliably judge AI output, and the same participants abandoned their preferred habit of reviewing small chunks of code under the 28-hour deadline. These accounts support the paper's claim that hackathons do n","pith_inferences":["My inference: the convergence on checking may be specific to participants who self-selected into an AI-themed hackathon and already used GenAI; a broader sample with more varied AI experience might show weaker norms.","My inference: the one-interviewee-per-team design leaves open that the 'unwritten rule' was one person's perception; interviewing all team members could reveal whether checking was collective or just individual habit.","My inference: the finding suggests a testable extension—measuring actual verification depth (e.g., edits made to AI code, cross-model checks) across teams with different domain familiarity, to see whether the self-reported constraints correspond to behavioral differences.","My inference: the impostor-syndrome reflection hints that hackathon AI use may affect participants' confidence beyond the event; a longitudinal follow-up could measure whether AI-heavy practice changes self-efficacy over time."],"forward_implications":["Hackathon organizers can treat verification as an activity to support rather than a behavior to mandate—for example, judging 'explainability of contribution' alongside skillful AI use, as the paper suggests.","Because time pressure curtails checking, events that build review steps into the schedule, such as mid-hack code-review checkpoints, could improve the reliability of AI-assisted output.","Participants working outside their domain are least able to verify; pairing them with domain experts or providing domain primers could reduce uncaught AI errors.","The presence of informal checking norms implies that interventions aimed at preventing over-reliance should reinforce existing habits rather than assume no verification happens.","As agentic AI tools take over larger units of work, the human review window shrinks, so future support must target these tools specifically."],"fun_headline_variants":["Hackathon teams verify AI output despite no rules","Quick builds, careful checks: AI in hackathons","Informal AI verification emerges in hackathon teams","Time-pressed hackathon teams still check AI output","Why hackathon teams trust but verify AI code"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that a single interviewed member per team accurately described the team's shared behavior, so if those four people misremembered or were atypical, the convergence on checking may not be real.","fun_headline_variants_meta":{"raw":{"variants":["Hackathon teams verify AI output despite no rules","Quick builds, careful checks: AI in hackathons","Informal AI verification emerges in hackathon teams","Time-pressed hackathon teams still check AI output","Why hackathon teams trust but verify AI code"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000137,"raw_usage":{"total_tokens":970,"prompt_tokens":710,"completion_tokens":260,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":454,"completion_tokens_details":{"reasoning_tokens":185}},"tokens_in":454,"tokens_out":260,"duration_ms":2979,"temperature":1.0,"reasoning_tokens":185,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T12:11:42.317987+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct observation study with screen recordings and interaction logs of hackathon teams: if a notable share of AI-generated code is adopted without any edit, review, or cross-check, the universal-checking claim would be falsified.","supporting_citations":[],"review_version":1}