{"id":"868f7195-3ada-43c0-ba8b-bcee3be9d876","arxiv_id":"2412.02653","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a small U.S. survey, most college STEM students use generative AI for coursework, and over half of those who answered a prompt-writing task would ask the AI to solve the problem directly, often bypassing their own problem-solving.","lead":"A survey of 40 U.S. STEM students and 28 physics instructors finds that most students use generative AI tools like ChatGPT for coursework, mainly to save time, and a majority would ask the AI to solve a physics problem directly rather than use it to guide their own thinking. The findings show a gap between student practice and instructor recommendations, offering practical pointers for AI guidance and tool design in STEM education.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 54% 'direct-solve' estimate is based on a single hypothetical prompt task and an unvalidated inference from prompt form to cognitive engagement; the claim that such use 'falls short' lacks learning-outcome evidence.","rationale":"The reader's weakest-assumption analysis focused on external validity: the small, self-selected Prolific and APS samples may not represent U.S. STEM students and faculty. That is a real concern, and the authors acknowledge it. My read identifies a more internal but equally load-bearing issue: the central claim about how students use genAI for problem-solving is built on a single hypothetical prompt task, and the interpretation of prompt categories as 'bypassing problem-solving' is not validated against actual behavior or learning outcomes. This matters because even a perfectly representative sample would not rescue the inference if the measure does not capture what it claims. The paper has genuine strengths: the survey was iteratively designed and pilot-tested, open coding achieved a mean inter-rater agreement of 83.4%, and the limitations section is candid about sample selection. However, the limitations section does not address the validity of the prompt-task measure or the absence of outcome data, so the strongest claim remains more interpretive than the data can support. I do not regard this as grounds for rejection: the descriptive patterns (adoption rates, use cases, perceived helpfulness, faculty caution) are informative, and the paper is appropriately framed as an early exploration. The concern reinforces the need for a conditional verdict with explicit requirements: share the instrument and coding scheme, validate or recalibrate the prompt-based measure, and treat the causal-sounding conclusion as a hypothesis for future work rather than an established finding.","tokens_in":12709,"tokens_out":5134,"duration_ms":59581,"concrete_test":"Re-analyze the 37 open-ended prompt responses from Table 3 with an additional coding dimension: whether the response requests explanation, worked steps, or verification beyond the final answer. Compare the re-coded 'direct-solve-but-engaged' group with the current 54% direct-solve figure. If a substantial share of the 20 responses coded as direct-solve also signal post-output engagement, the claim that these students prefer to bypass their own problem-solving is overstated. A stronger follow-up test: have the same participants actually use ChatGPT on the same physics problem, capture their chat logs and think-aloud comments, and test whether the initial prompt category predicts whether they engage with the solution. If prompt category does not predict engagement, the central claim lacks measurement validity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—that a slight majority of students (54%) prefer to use genAI to directly solve problems rather than to support their own problem-solving (Section 4.4, Table 3)—rests entirely on answers to one hypothetical prompt-writing task. Students were given a single physics problem and asked what prompt they would give to a genAI chatbot. This is not a measure of how students actually use genAI in their STEM coursework, and it records only the first prompt, not what students do with the AI's output. A student who types the problem and asks for a solution may subsequently work through the solution, ask follow-up questions, or treat it as a worked example; conversely, a student who asks 'what are the primary factors influencing elevator travel time' may still be offloading central reasoning. The coding in Table 3 treats prompt form as equivalent to cognitive engagement, but that equivalence is asserted rather than validated. The Conclusions then go further, stating that these approaches 'often fall short in enhancing their own STEM problem-solving competencies.' That evaluative claim requires some link between prompt behavior and learning outcomes, yet the survey contains no outcome measure, longitudinal data, or observed interaction logs. Even if the sample were perfectly representative, the headline inference would not follow from the prompt categories as coded. The 54% figure is also fragile: it is 20 out of 37 respondents (54%), but if the 3 non-respondents are included in the full participant pool, the proportion is exactly 20/40, so the abstract's 'over half' is not robust.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a mixed-methods survey of 40 US STEM undergraduates and 28 physics instructors about how and why students use generative AI (genAI) tools in STEM coursework, with a focus on problem-solving. It finds high adoption rates, common use cases (finding explanations, exploring topics, summarizing readings, helping with problem sets), time-saving as the dominant motivation, and a coded prompt-writing task suggesting that 54% of respondents (20 of 37 who answered) would ask a chatbot to directly solve a physics problem rather than ask scaffolded questions. The paper also compares students' perceived helpfulness of genAI for four problem-solving aspects with faculty recommendations, and reports perceived benefits (personalized support, information retrieval) and risks (misinformation, over-reliance, academic integrity). It concludes that students' current approaches 'often fall short in enhancing their own STEM problem-solving competencies.'","tokens_in":12931,"tokens_out":4236,"duration_ms":44573,"significance":"If treated as an exploratory descriptive study with appropriately hedged conclusions, this paper would provide a timely snapshot of student and instructor perspectives in a rapidly evolving area. Its strengths include the two-group design (students and faculty), the use of a concrete physics problem to elicit prompting behavior, qualitative coding with a reported 83.4% inter-rater agreement, and an explicit acknowledgment of selection bias in the Limitations section. The central weakness is that the headline evaluative claim about problem-solving competency is not supported by the survey data, which contain no learning-outcome measure, no longitudinal data, and no observation of actual interaction with genAI outputs. The paper is therefore more convincing as a description of self-reported use and perceptions than as a basis for the conclusion that current use 'falls short.'","major_comments":[{"comment":"The claim that 54% of students 'prefer to use genAI to directly solve problems rather than as a tool to support their own problem-solving' rests entirely on a single hypothetical prompt-writing task, and the coding equates the form of the written prompt with students' cognitive engagement. A student who types the problem and asks for a solution may subsequently work through the output, ask follow-up questions, or treat it as a worked example; conversely, a student who asks 'what are the primary factors influencing elevator travel time' may still be offloading central reasoning. No data on follow-up behavior, actual tool use, or learning outcomes are presented. The evaluative conclusion in the Abstract ('often fall short in enhancing their own STEM problem-solving competencies') therefore goes beyond what the prompt categories can support. Please either validate the prompt-to-engagement inference (e.g., with think-aloud protocols, interaction logs, or a follow-up question about what students do with the AI's output) or substantially soften the claims to describe prompting patterns and perceived helpfulness without asserting that competency is or is not enhanced.","section":"Section 4.4, Table 3; Abstract"},{"comment":"There is a numerical inconsistency in the central result. Table 3 reports 14 students (38%) who would copy/paste and ask AI to solve, and 6 students (16%) who would add instructions for AI to solve, for a total of 20 out of 37 respondents (54%). The Abstract, however, says 'over half of the student participants' reported simply inputting a problem for AI to generate solutions. With N=40, 20 is exactly half, not 'over half,' and the 3 students who did not answer the prompt task are not accounted for in the abstract wording. Please correct the wording, state the denominator explicitly, and clarify why 3 respondents are missing from the analysis.","section":"Abstract and Section 4.4"},{"comment":"The authors acknowledge in Section 6 that the sample of 40 self-selected Prolific volunteers and 28 APS-listserv faculty may not be representative of the broader US college STEM population, yet the Abstract and Conclusions state population-level claims as though they were established facts (e.g., 'students' current approaches to utilizing genAI tools often fall short'). Given the acknowledged selection bias, the reported percentages (e.g., 54% direct-solve, 85% adoption in STEM courses) are not robust estimates of population prevalence. Please hedge all population-level generalizations throughout, and move the representativeness caveat into the Results framing rather than only the Limitations section.","section":"Section 6 (Limitations); Abstract and Conclusions"}],"minor_comments":[{"comment":"The comparison in Figure 4 and the accompanying text contrasts students' 'helpfulness' ratings with instructors' 'recommendation' ratings. These are different constructs, so the 'stark contrast' and 'misalignment' language should be qualified; the gap may reflect differences in the two groups' roles and experiences rather than a direct disagreement about the tools' affordances.","section":"Section 4.5, Figure 4"},{"comment":"Inter-rater reliability is reported only as a mean percent agreement (83.4%). Percent agreement does not account for chance agreement; please report chance-corrected indices (e.g., Cohen's kappa or Krippendorff's alpha) or per-code agreement values.","section":"Section 3.3"},{"comment":"The sentence 'We did not find any significant association between students' year in college and their usage patterns' is based on N=40 and is therefore very low-powered. Please report the test statistic and effect size, or soften the claim to 'no significant association was detected in this small sample.'","section":"Section 4.3"},{"comment":"The full survey instrument is not included in the manuscript or an appendix. For reproducibility and for readers who wish to adapt the prompt-writing task, consider including the complete student and faculty questionnaires as supplementary material.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful exploratory descriptive study in a fast-moving area, but the abstract and conclusions overstate what the data can support. If the authors revise to present the findings as self-reported prompting patterns and perceptions, with explicit caveats about the unvalidated prompt-to-engagement inference and the non-representative sample, the paper could be acceptable for publication. The numerical inconsistency between 'over half' and 20/37 also needs correction. I do not see a need for rejection, but the central claims must be brought in line with the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a small but honest survey paper: 40 STEM students and 28 physics instructors, with a useful direct comparison of how students rate genAI helpfulness for four problem-solving aspects versus what faculty recommend. That comparison, plus the multi-domain STEM sample, is the real contribution. The survey work is transparent and the limitations section is upfront.\n\nThe soft spots are in the headline claims, not the descriptive data. The 54% 'direct-solve' figure comes from a single hypothetical prompt-writing task, and the coding treats prompt form as equivalent to cognitive engagement. A student who pastes the problem might still work through the solution afterwards; a student who asks a pointed question might still be offloading reasoning. The survey has no outcome measure, so the conclusion that current use 'falls short in enhancing problem-solving competency' is an interpretation, not a measured result. Also, 20 of 40 participants is exactly half, not 'over half' as the abstract says; 20 of 37 respondents is 54%, but the denominator matters.\n\nI would not sink the paper over this. The descriptive findings (high adoption, time-saving motive, faculty reservations) are well supported and align with prior work. For peer review, I'd accept it with revisions: soften the abstract, share the instrument and data, and reframe the prompt-coding result as a behavioral observation with unknown learning consequences rather than a judgment of competency. The sample is small and self-selected, but the authors already flag that.\n\nThis paper is for instructors and tool designers who want a current snapshot of student AI use in STEM, and for researchers tracking genAI adoption. It deserves a serious referee, but the evaluative framing needs to be reined in.\n\nRecommendation: send to review, conditional on revisions.","headline":"A transparent small-N survey with a valuable student-faculty comparison, but the 'over half' and 'falls short' claims outrun the data.","tokens_in":13512,"tokens_out":2401,"would_cite":true,"duration_ms":24405,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 40 STEM undergraduates finds that 54% would use generative AI to solve a physics problem outright, leading the authors to argue that current student use often bypasses the problem-solving process rather than scaffolding it.","keywords":["generative AI","STEM education","problem-solving competency","college students","prompting strategies","ChatGPT","higher education","survey research"],"falsifier":"A replication with a large, demographically representative sample of U.S. STEM undergraduates using the same prompt-writing task: if fewer than half of respondents used direct-solution prompts, the claim that the slight majority bypasses their own problem-solving would be falsified. A second falsifier would be a controlled experiment in which direct-solution AI users perform as well as scaffolded users on unassisted post-tests.","tokens_in":12492,"feed_emoji":"🤖","tokens_out":9058,"duration_ms":80395,"temperature":0.7,"pith_summary":"This paper asks whether college STEM students use generative AI tools like ChatGPT as a scaffold for learning or as a crutch that replaces their own thinking. Surveying 40 U.S. STEM undergraduates and 28 physics instructors, it finds nearly all students have adopted genAI and use it mainly to save time. When given a physics problem and asked what prompt they would send, 54% of students said they would have the AI solve the problem directly, while 46% asked scaffolded questions. The authors conclude that current student usage often bypasses the problem-solving process and so falls short of building the competency that STEM education aims to develop. This matters because it suggests concrete mismatches between student practice, faculty recommendations, and the design of AI tools for learning.","feed_headline":"Survey: 54% of STEM students prefer AI that solves problems for them","feed_subtitle":"Students treat ChatGPT as a time-saving answer machine; most instructors would not recommend that use.","key_machinery":"The central instrument is a prompt-writing task embedded in the survey: students were given a typical introductory physics problem and asked what prompt they would give to a ChatGPT-like chatbot for help. Open-ended responses were coded into three categories—copy/paste the problem and ask for a solution, copy/paste with added solving instructions, or ask specific scaffolded questions—and this coding produces the study's headline 54% vs 46% split. The second mechanism is a four-part framework of problem-solving aspects (identifying relevant domain knowledge, collecting needed data, executing the plan, verifying correctness), drawn from the authors' prior work, which is used to compare students' perceived helpfulness of genAI with faculty recommendations for each aspect.","core_discovery":"The study's central finding is that a slight majority of college STEM students, 54% of the 40 surveyed, prefer to use generative AI to produce direct solutions to problems rather than to support their own problem-solving. The prompting task—a physics elevator problem—shows 38% of students would copy/paste or paraphrase the problem and ask the AI to solve it, and another 16% would add instructions for solving; only 46% asked questions that could help them figure it out themselves. The paper interprets this as evidence that students' current approaches often bypass the deeper learning processes needed to develop STEM problem-solving competency. It also documents a sharp gap with faculty: less than one-third of the 28 instructors recommended genAI for any of the four problem-solving aspects surveyed (explaining concepts, gathering missing information, calculations, verifying correctness), and only 7% recommended it for calculations. Both groups named misinformation as a top risk, but students were far less worried than faculty about the quality of learning being damaged.","pith_inferences":["The 54% figure likely understates the prevalence of direct-solution use in real coursework: the prompt task presented a single problem in a low-stakes survey, and the authors themselves note that the self-selected online sample may overrepresent students who are interested in or comfortable with genAI.","The faculty-student gap implies that relying on instructor advice alone will not change student behavior; a testable next step is a randomized course-level comparison between a scaffold-only AI tutor and unrestricted ChatGPT access, measuring unassisted problem-solving performance afterward.","Students' shared skepticism about using genAI to verify solution correctness suggests a possible natural entry point for training: if students doubt the tool's reliability for checking, instruction could leverage that doubt to teach verification as a human responsibility.","The time-saving motive aligns with broader patterns of technology use in higher education, so without structural incentives (assessments that require demonstrated process), students will likely keep defaulting to the fastest route regardless of warnings."],"forward_implications":["If direct-solution prompting dominates, students may experience a false sense of fluency, feeling they have learned a problem type when they have only read an AI's solution; the paper likens this to the known gap between perceived and actual learning from passive lectures.","Students and faculty are misaligned: students rate genAI helpful for explaining concepts (82.5% helpful) and calculations (65%), while fewer than one-third of faculty recommend either use, so courses need explicit guidance on when genAI use supports versus undermines learning.","Because most students rely on free versions of LLMs, colleges that want equitable support should consider providing reliable, education-focused genAI access to all students and teaching critical evaluation of AI output.","A promising design direction, flagged by the paper, is genAI-based tutors that scaffold problem-solving without providing direct solutions, following work showing such tutoring can outperform active learning.","Students' dominant motive is saving time, so any intervention to change prompting behavior must address the efficiency incentive rather than only warn about risks."],"supporting_citations":[{"why":"Provides the experimental evidence that students with GPT-4 access performed 17% worse on exams after access was removed, which the paper uses to argue that direct solution generation can short-circuit learning.","marker":"Bastani et al (2024)"},{"why":"Prior finding that chemistry students predominantly used copy-paste prompting with AI chatbots, the direct precedent for the paper's prompting-strategy coding.","marker":"Tassoti (2024)"},{"why":"Shows that students in an introductory physics course often trusted ChatGPT's answers regardless of accuracy, supporting the paper's concern about over-reliance.","marker":"Ding et al (2023)"},{"why":"The authors' earlier problem-solving framework that supplies the four aspects of problem-solving used to structure the helpfulness ratings.","marker":"Salehi (2018)"},{"why":"Characterizes the expert problem-solving process in science and engineering, underpinning the framework the survey uses.","marker":"Price et al (2021)"},{"why":"Demonstrates that a carefully designed genAI tutor that scaffolds rather than solves can improve engagement and learning, cited as the model for future tool design.","marker":"Kestin et al (2024)"},{"why":"Provides evidence on the gap between actual learning and feeling of learning in passive lectures, used to interpret the risk of direct-solution use.","marker":"Deslauriers et al (2019)"}],"fun_headline_variants":["54% of STEM students use AI as a crutch, not scaffold","Students prefer AI to think for them: 54% skip problem-solving","Faculty vs students: 54% would let AI solve STEM problems","AI as crutch: 54% of STEM students outsource thinking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 40 self-selected online students and 28 volunteer physics instructors are representative enough of U.S. college STEM students and faculty that the reported 54% direct-solution preference and the student–faculty gap generalize beyond this sample.","fun_headline_variants_meta":{"raw":{"variants":["54% of STEM students use AI as a crutch, not scaffold","Students prefer AI to think for them: 54% skip problem-solving","Faculty vs students: 54% would let AI solve STEM problems","AI as crutch: 54% of STEM students outsource thinking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001607,"raw_usage":{"total_tokens":6441,"prompt_tokens":1030,"completion_tokens":5411,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":5333}},"tokens_in":646,"tokens_out":5411,"duration_ms":36926,"temperature":1.0,"reasoning_tokens":5333,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:12:11.058173+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication with a large, demographically representative sample of U.S. STEM undergraduates using the same prompt-writing task: if fewer than half of respondents used direct-solution prompts, the claim that the slight majority bypasses their own problem-solving would be falsified. A second falsifier would be a controlled experiment in which direct-solution AI users perform as well as scaffolded users on unassisted post-tests.","supporting_citations":[{"cited_title":"Available at SSRN 4895486","cited_arxiv_id":null,"evidence_quote":"Provides the experimental evidence that students with GPT-4 access performed 17% worse on exams after access was removed, which the paper uses to argue that direct solution generation can short-circuit learning."},{"cited_title":"Journal of Chemical Education","cited_arxiv_id":null,"evidence_quote":"Prior finding that chemistry students predominantly used copy-paste prompting with AI chatbots, the direct precedent for the paper's prompting-strategy coding."},{"cited_title":"International Journal of Educational Technology in Higher Education 20(1):63","cited_arxiv_id":null,"evidence_quote":"Shows that students in an introductory physics course often trusted ChatGPT's answers regardless of accuracy, supporting the paper's concern about over-reliance."},{"cited_title":"Stanford University","cited_arxiv_id":null,"evidence_quote":"The authors' earlier problem-solving framework that supplies the four aspects of problem-solving used to structure the helpfulness ratings."},{"cited_title":"CBE—Life Sciences Education 20(3):ar43","cited_arxiv_id":null,"evidence_quote":"Characterizes the expert problem-solving process in science and engineering, underpinning the framework the survey uses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates that a carefully designed genAI tutor that scaffolds rather than solves can improve engagement and learning, cited as the model for future tool design."},{"cited_title":"Proceedings of the National Academy of Sciences 116(39):19251--19257","cited_arxiv_id":null,"evidence_quote":"Provides evidence on the gap between actual learning and feeling of learning in passive lectures, used to interpret the risk of direct-solution use."}],"review_version":1}