REVIEW 3 major objections 4 minor 51 references
Scaffold or Crutch? Examining College Students' Use and Views of Generative AI Tools for STEM Education
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A survey of 40 STEM undergraduates finds that 54% would use generative AI to solve a physics problem outright, leading the authors to argue that current student use often bypasses the problem-solving process rather than scaffolding it.
desk verdict A transparent small-N survey with a valuable student-faculty comparison, but the 'over half' and 'falls short' claims outrun the data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central instrument is a prompt-writing task embedded in the survey: students were given a typical introductory physics problem and asked what prompt they would give to a ChatGPT-like chatbot for help. Open-ended responses were coded into three categories—copy/paste the problem and ask for a solution, copy/paste with added solving instructions, or ask specific scaffolded questions—and this coding produces the study's headline 54% vs 46% split. The second mechanism is a four-part framework of problem-solving aspects (identifying relevant domain knowledge, collecting needed data, executing the plan, verifying correctness), drawn from the authors' prior work, which is used to compare students' perceived helpfulness of genAI with faculty recommendations for each aspect.
What would settle it
A replication with a large, demographically representative sample of U.S. STEM undergraduates using the same prompt-writing task: if fewer than half of respondents used direct-solution prompts, the claim that the slight majority bypasses their own problem-solving would be falsified. A second falsifier would be a controlled experiment in which direct-solution AI users perform as well as scaffolded users on unassisted post-tests.
Extended reading notes
Core claim
The study's central finding is that a slight majority of college STEM students, 54% of the 40 surveyed, prefer to use generative AI to produce direct solutions to problems rather than to support their own problem-solving. The prompting task—a physics elevator problem—shows 38% of students would copy/paste or paraphrase the problem and ask the AI to solve it, and another 16% would add instructions for solving; only 46% asked questions that could help them figure it out themselves. The paper interprets this as evidence that students' current approaches often bypass the deeper learning processes needed to develop STEM problem-solving competency. It also documents a sharp gap with faculty: less than one-third of the 28 instructors recommended genAI for any of the four problem-solving aspects surveyed (explaining concepts, gathering missing information, calculations, verifying correctness), and only 7% recommended it for calculations. Both groups named misinformation as a top risk, but students were far less worried than faculty about the quality of learning being damaged.
Load-bearing premise
The 40 self-selected online students and 28 volunteer physics instructors are representative enough of U.S. college STEM students and faculty that the reported 54% direct-solution preference and the student–faculty gap generalize beyond this sample.
Editorial extensions
If this is right
- If direct-solution prompting dominates, students may experience a false sense of fluency, feeling they have learned a problem type when they have only read an AI's solution; the paper likens this to the known gap between perceived and actual learning from passive lectures.
- Students and faculty are misaligned: students rate genAI helpful for explaining concepts (82.5% helpful) and calculations (65%), while fewer than one-third of faculty recommend either use, so courses need explicit guidance on when genAI use supports versus undermines learning.
- Because most students rely on free versions of LLMs, colleges that want equitable support should consider providing reliable, education-focused genAI access to all students and teaching critical evaluation of AI output.
- A promising design direction, flagged by the paper, is genAI-based tutors that scaffold problem-solving without providing direct solutions, following work showing such tutoring can outperform active learning.
- Students' dominant motive is saving time, so any intervention to change prompting behavior must address the efficiency incentive rather than only warn about risks.
Reading between the lines
- The 54% figure likely understates the prevalence of direct-solution use in real coursework: the prompt task presented a single problem in a low-stakes survey, and the authors themselves note that the self-selected online sample may overrepresent students who are interested in or comfortable with genAI.
- The faculty-student gap implies that relying on instructor advice alone will not change student behavior; a testable next step is a randomized course-level comparison between a scaffold-only AI tutor and unrestricted ChatGPT access, measuring unassisted problem-solving performance afterward.
- Students' shared skepticism about using genAI to verify solution correctness suggests a possible natural entry point for training: if students doubt the tool's reliability for checking, instruction could leverage that doubt to teach verification as a human responsibility.
- The time-saving motive aligns with broader patterns of technology use in higher education, so without structural incentives (assessments that require demonstrated process), students will likely keep defaulting to the fastest route regardless of warnings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a mixed-methods survey of 40 US STEM undergraduates and 28 physics instructors about how and why students use generative AI (genAI) tools in STEM coursework, with a focus on problem-solving. It finds high adoption rates, common use cases (finding explanations, exploring topics, summarizing readings, helping with problem sets), time-saving as the dominant motivation, and a coded prompt-writing task suggesting that 54% of respondents (20 of 37 who answered) would ask a chatbot to directly solve a physics problem rather than ask scaffolded questions. The paper also compares students' perceived helpfulness of genAI for four problem-solving aspects with faculty recommendations, and reports perceived benefits (personalized support, information retrieval) and risks (misinformation, over-reliance, academic integrity). It concludes that students' current approaches 'often fall short in enhancing their own STEM problem-solving competencies.'
Significance. If treated as an exploratory descriptive study with appropriately hedged conclusions, this paper would provide a timely snapshot of student and instructor perspectives in a rapidly evolving area. Its strengths include the two-group design (students and faculty), the use of a concrete physics problem to elicit prompting behavior, qualitative coding with a reported 83.4% inter-rater agreement, and an explicit acknowledgment of selection bias in the Limitations section. The central weakness is that the headline evaluative claim about problem-solving competency is not supported by the survey data, which contain no learning-outcome measure, no longitudinal data, and no observation of actual interaction with genAI outputs. The paper is therefore more convincing as a description of self-reported use and perceptions than as a basis for the conclusion that current use 'falls short.'
major comments (3)
- [Section 4.4, Table 3; Abstract] The claim that 54% of students 'prefer to use genAI to directly solve problems rather than as a tool to support their own problem-solving' rests entirely on a single hypothetical prompt-writing task, and the coding equates the form of the written prompt with students' cognitive engagement. A student who types the problem and asks for a solution may subsequently work through the output, ask follow-up questions, or treat it as a worked example; conversely, a student who asks 'what are the primary factors influencing elevator travel time' may still be offloading central reasoning. No data on follow-up behavior, actual tool use, or learning outcomes are presented. The evaluative conclusion in the Abstract ('often fall short in enhancing their own STEM problem-solving competencies') therefore goes beyond what the prompt categories can support. Please either validate the prompt-to-engagement inference (e.g., with think-aloud protocols, interaction logs, or a follow-up question about what students do with the AI's output) or substantially soften the claims to describe prompting patterns and perceived helpfulness without asserting that competency is or is not enhanced.
- [Abstract and Section 4.4] There is a numerical inconsistency in the central result. Table 3 reports 14 students (38%) who would copy/paste and ask AI to solve, and 6 students (16%) who would add instructions for AI to solve, for a total of 20 out of 37 respondents (54%). The Abstract, however, says 'over half of the student participants' reported simply inputting a problem for AI to generate solutions. With N=40, 20 is exactly half, not 'over half,' and the 3 students who did not answer the prompt task are not accounted for in the abstract wording. Please correct the wording, state the denominator explicitly, and clarify why 3 respondents are missing from the analysis.
- [Section 6 (Limitations); Abstract and Conclusions] The authors acknowledge in Section 6 that the sample of 40 self-selected Prolific volunteers and 28 APS-listserv faculty may not be representative of the broader US college STEM population, yet the Abstract and Conclusions state population-level claims as though they were established facts (e.g., 'students' current approaches to utilizing genAI tools often fall short'). Given the acknowledged selection bias, the reported percentages (e.g., 54% direct-solve, 85% adoption in STEM courses) are not robust estimates of population prevalence. Please hedge all population-level generalizations throughout, and move the representativeness caveat into the Results framing rather than only the Limitations section.
minor comments (4)
- [Section 4.5, Figure 4] The comparison in Figure 4 and the accompanying text contrasts students' 'helpfulness' ratings with instructors' 'recommendation' ratings. These are different constructs, so the 'stark contrast' and 'misalignment' language should be qualified; the gap may reflect differences in the two groups' roles and experiences rather than a direct disagreement about the tools' affordances.
- [Section 3.3] Inter-rater reliability is reported only as a mean percent agreement (83.4%). Percent agreement does not account for chance agreement; please report chance-corrected indices (e.g., Cohen's kappa or Krippendorff's alpha) or per-code agreement values.
- [Section 4.3] The sentence 'We did not find any significant association between students' year in college and their usage patterns' is based on N=40 and is therefore very low-powered. Please report the test statistic and effect size, or soften the claim to 'no significant association was detected in this small sample.'
- [Methods] The full survey instrument is not included in the manuscript or an appendix. For reproducibility and for readers who wish to adapt the prompt-writing task, consider including the complete student and faculty questionnaires as supplementary material.
Circularity Check
No circularity found: this is a descriptive survey study with no derivation chain, fitted parameters, or predictions that reduce to inputs.
full rationale
The paper is an empirical survey study, not a derivation or modeling paper. There are no equations, fitted parameters, or predictions that could be equivalent to the inputs by construction. The central claims—high adoption rates, reported use cases, prompting strategies, perceived helpfulness, and perceived benefits/risks—are direct summaries of self-report data from 40 students and 28 instructors. The 54% 'preferred to use genAI to directly solve problems' figure is a descriptive count from coded open-ended responses in Table 3, not a fitted or predicted quantity. The only notable self-citation is the use of the authors' prior problem-solving framework (Salehi, 2018; Price et al., 2021) to define the four aspects of problem-solving used in the survey items. This is a construct-definition input to the survey design, not a result derived from the current data, and the survey findings do not depend on that prior framework being true: the prompting behaviors, usage frequencies, and ratings stand as descriptive survey results regardless of how one defines problem-solving aspects. The evaluative conclusion that direct-solution prompting 'often falls short' is an interpretive inference, not a derivation, and the paper explicitly acknowledges sample-size and selection-bias limitations in Section 6. No circular step can be quoted because the specific reduction required by the circularity criteria does not appear in the paper.
Assumptions & free parameters
assumptions (4)
- domain assumption The four aspects of problem-solving (identifying domain knowledge, collecting data, executing a plan, verifying correctness) constitute a valid operationalization of STEM problem-solving for measuring helpfulness.
- domain assumption Self-reported survey behavior accurately reflects actual genAI use and prompting behavior.
- domain assumption The recruited Prolific and APS listserv samples are representative of U.S. STEM students and faculty.
- domain assumption Inter-rater agreement of 83.4% is sufficient to treat the qualitative coding as reliable.
Cite this review
Pith. "Pith review of Scaffold or Crutch? Examining College Students' Use and Views of Generative AI Tools for STEM Education." pith.science (2026). https://pith.science/paper/E5VDMWYJ
@misc{pith2026241202653,
author = {Pith},
title = {Pith review of: Scaffold or Crutch? Examining College Students' Use and Views of Generative AI Tools for STEM Education},
year = {2026},
howpublished = {\url{https://pith.science/paper/E5VDMWYJ}},
note = {Machine review of arXiv:2412.02653}
}
read the original abstract
Developing problem-solving competency is central to Science, Technology, Engineering, and Mathematics (STEM) education, yet translating this priority into effective approaches to problem-solving instruction and assessment remain a significant challenge. The recent proliferation of generative artificial intelligence (genAI) tools like ChatGPT in higher education introduces new considerations about how these tools can help or hinder students' development of STEM problem-solving competency. Our research examines these considerations by studying how and why college students use genAI tools in their STEM coursework, focusing on their problem-solving support. We surveyed 40 STEM college students from diverse U.S. institutions and 28 STEM faculty to understand instructor perspectives on effective genAI tool use and guidance in STEM courses. Our findings reveal high adoption rates and diverse applications of genAI tools among STEM students. The most common use cases include finding explanations, exploring related topics, summarizing readings, and helping with problem-set questions. The primary motivation for using genAI tools was to save time. Moreover, over half of student participants reported simply inputting problems for AI to generate solutions, potentially bypassing their own problem-solving processes. These findings indicate that despite high adoption rates, students' current approaches to utilizing genAI tools often fall short in enhancing their own STEM problem-solving competencies. The study also explored students' and STEM instructors' perceptions of the benefits and risks associated with using genAI tools in STEM education. Our findings provide insights into how to guide students on appropriate genAI use in STEM courses and how to design genAI-based tools to foster students' problem-solving competency.
Reference graph
Works this paper leans on
- [1]
-
[2]
://arxiv.org/abs/2304.14415, 2304.14415
Amani S, White L, Balart T, et al (2023) Generative ai perceptions: A survey to measure the perceptions of faculty, staff, and students on generative ai tools in academia. ://arxiv.org/abs/2304.14415, 2304.14415
arXiv 2023
-
[3]
chatgpt seems too good to be true
Baek C, Tate T, Warschauer M (2024) “chatgpt seems too good to be true”: College students’ use and perceptions of generative ai. Computers and Education: Artificial Intelligence 7:100294
work page 2024
-
[4]
Bastani H, Bastani O, Sungu A, et al (2024) Generative ai can harm learning. Available at SSRN 4895486
work page 2024
-
[5]
://arxiv.org/abs/2005.14165, 2005.14165
Brown TB, Mann B, Ryder N, et al (2020) Language models are few-shot llearners. ://arxiv.org/abs/2005.14165, 2005.14165
arXiv 2020
-
[6]
Physical Review Physics Education Research 16(1):010123
Burkholder E, Miles J, Layden T, et al (2020) Template for teaching and assessment of problem solving in introductory physics. Physical Review Physics Education Research 16(1):010123
work page 2020
-
[7]
Economics of Education Review 56:118--132
Carter SP, Greenberg K, Walker MS (2017) The impact of computer usage on academic performance: Evidence from a randomized trial at the united states military academy. Economics of Education Review 56:118--132
work page 2017
-
[8]
International Journal of Educational Technology in Higher Education 20(1):43
Chan CKY, Hu W (2023) Students’ voices on generative ai: Perceptions, benefits, and challenges in higher education. International Journal of Educational Technology in Higher Education 20(1):43
work page 2023
Show all 51 references
-
[9]
Cuban L (2001) Oversold and underused: Computers in the classroom
2001
-
[10]
Proceedings of the National Academy of Sciences 116(39):19251--19257
Deslauriers L, McCarty LS, Miller K, et al (2019) Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. Proceedings of the National Academy of Sciences 116(39):19251--19257
2019
-
[11]
International Journal of Educational Technology in Higher Education 20(1):63
Ding L, Li T, Jiang S, et al (2023) Students’ perceptions of using chatgpt in a physics class as a virtual tutor. International Journal of Educational Technology in Higher Education 20(1):63
2023
-
[12]
Shaking the foundations of Geo-Engineering education pp 9--14
Felder RM (2012) Engineering education: A tale of two paradigms. Shaking the foundations of Geo-Engineering education pp 9--14
2012
-
[13]
Journal of Science Education and Technology pp 1--12
Feldman-Maggor Y, Blonder R, Alexandron G (2024) Perspectives of generative ai in chemistry education within the tpack framework. Journal of Science Education and Technology pp 1--12
2024
-
[14]
In: Society for Information Technology & Teacher Education International Conference, Association for the Advancement of Computing in Education (AACE), pp 757--766
Goldberg D, Sobo E, Frazee J, et al (2024) Generative ai in higher education: Insights from a campus-wide student survey at a large public university. In: Society for Information Technology & Teacher Education International Conference, Association for the Advancement of Comput...
2024
-
[15]
Computers and Education: Artificial Intelligence 5:100170
Hallal K, Hamdan R, Tlais S (2023) Exploring the potential of ai-chatbots in organic chemistry: An assessment of chatgpt and bard. Computers and Education: Artificial Intelligence 5:100170
2023
-
[16]
Journal of Higher Education Policy and Management 37(3):308--319
Henderson M, Selwyn N, Finger G, et al (2015) Students’ everyday engagement with digital technology in university: exploring patterns of use and ‘usefulness’. Journal of Higher Education Policy and Management 37(3):308--319
2015
-
[17]
Vision report, National Science Foundation, nSF Liaison: Robin Wright, Executive Secretary: Alexandra Medina-Borja
Honey M, Alberts B, Bass H, et al (2020) Stem education for the future - 2020 visioning report. Vision report, National Science Foundation, nSF Liaison: Robin Wright, Executive Secretary: Alexandra Medina-Borja
2020
-
[18]
Essential readings in problem-based learning: Exploring and extending the legacy of Howard S Barrows 1741
Jonassen DH, Hung W (2015) All problems are not equal: Implications for problem-based learning. Essential readings in problem-based learning: Exploring and extending the legacy of Howard S Barrows 1741
2015
-
[19]
Learning and individual differences 103:102274
Kasneci E, Se ler K, K \"u chemann S, et al (2023) Chatgpt for good? on opportunities and challenges of large language models for education. Learning and individual differences 103:102274
2023
-
[20]
Kestin G, Miller K, Klales A, et al (2024) Ai tutoring outperforms active learning
2024
-
[21]
Computers and Education: Artificial Intelligence 5:100156
Kohnke L, Moorhouse BL, Zou D (2023) Exploring generative artificial intelligence preparedness among university language instructors: A case study. Computers and Education: Artificial Intelligence 5:100156
2023
-
[22]
Kortemeyer G (2023) Could an artificial-intelligence agent pass an introductory physics course? Physical Review Physics Education Research 19(1):010132
2023
-
[23]
Washington, DC: Third Way NEXT
Levy F, Murnane RJ (2013) Dancing with robots: Human skills for computerized work. Washington, DC: Third Way NEXT
2013
-
[24]
://arxiv.org/abs/2205.05638, 2205.05638
Liu H, Tam D, Muqeeth M, et al (2022) Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. ://arxiv.org/abs/2205.05638, 2205.05638
2022 arXiv
-
[25]
://arxiv.org/abs/2402.01659, 2402.01659
McDonald N, Johri A, Ali A, et al (2024) Generative artificial intelligence in higher education: Evidence from an analysis of institutional policies and guidelines. ://arxiv.org/abs/2402.01659, 2402.01659
2024 arXiv
-
[26]
Physical Review Physics Education Research 20(2):020103
Montgomery BJ, Price AM, Wieman CE (2024) Characterizing decision-making opportunities in undergraduate physics coursework. Physical Review Physics Education Research 20(2):020103
2024
-
[27]
Psychological science 25(6):1159--1168
Mueller PA, Oppenheimer DM (2014) The pen is mightier than the keyboard: Advantages of longhand over laptop note taking. Psychological science 25(6):1159--1168
2014
-
[28]
Task force report, National Education Association
National Education Association (2024) Report of the NEA task force on artificial intelligence in education. Task force report, National Education Association
2024
-
[29]
://arxiv.org/abs/2303.08774, 2303.08774
OpenAI, Achiam J, Adler S, et al (2024) Gpt-4 technical report. ://arxiv.org/abs/2303.08774, 2303.08774
2024 arXiv
-
[30]
Educational psychologist 20(4):167--182
Pea RD (1985) Beyond amplification: Using the computer to reorganize mental functioning. Educational psychologist 20(4):167--182
1985
-
[31]
In: Proceedings of the 2023 Working Group Reports on Innovation and Technology in Computer Science Education
Prather J, Denny P, Leinonen J, et al (2023) The robots are here: Navigating the generative ai revolution in computing education. In: Proceedings of the 2023 Working Group Reports on Innovation and Technology in Computer Science Education. Association for Computing Machinery, ...
2023
-
[32]
JMIR Medical Education 9(1):e48785
Preiksaitis C, Rose C, et al (2023) Opportunities, challenges, and future directions of generative artificial intelligence in medical education: scoping review. JMIR Medical Education 9(1):e48785
2023
-
[33]
International Journal of Science Education 44(13):2061--2084
Price A, Salehi S, Burkholder E, et al (2022) An accurate and practical method for assessing science and engineering problem-solving expertise. International Journal of Science Education 44(13):2061--2084
2022
-
[34]
CBE—Life Sciences Education 20(3):ar43
Price AM, Kim CJ, Burkholder EW, et al (2021) A detailed characterization of the expert problem-solving process in science and engineering: Guidance for teaching and assessment. CBE—Life Sciences Education 20(3):ar43
2021
-
[35]
Harvard University Press
Reich J (2020) Failure to disrupt: Why technology alone can’t transform education. Harvard University Press
2020
-
[36]
Stanford University
Salehi S (2018) Improving problem-solving through reflection. Stanford University
2018
-
[37]
Journal of Chemical Education 101(3):1332--1340
Schwartz Poehlmann JK, Nardo JE, Rojas M, et al (2024) Introducing the problem-solving template as a tool for equity: Addressing incoming preparation disparities. Journal of Chemical Education 101(3):1332--1340
2024
-
[38]
education 2030
Shiohira K (2021) Understanding the impact of artificial intelligence on skills development. education 2030. UNESCO-UNEVOC International Centre for Technical and Vocational Education and Training
2021
-
[39]
In: Proceedings of the Tenth ACM Conference on Learning @ Scale
Smolansky A, Cram A, Raduescu C, et al (2023) Educator and student perspectives on the impact of generative ai on assessments in higher education. In: Proceedings of the Tenth ACM Conference on Learning @ Scale. Association for Computing Machinery, New York, NY, USA, L@S '23, ...
2023
-
[40]
Journal of Chemical Education
Tassoti S (2024) Assessment of students use of generative artificial intelligence: Prompting strategies and prompt engineering in chemistry education. Journal of Chemical Education
2024
-
[41]
://arxiv.org/abs/2201.08239, 2201.08239
Thoppilan R, Freitas DD, Hall J, et al (2022) Lamda: Language models for dialog applications. ://arxiv.org/abs/2201.08239, 2201.08239
2022 arXiv
-
[42]
Physical Review Physics Education Research 20(1):010152
Wan T, Chen Z (2024) Exploring generative ai assisted feedback writing for students’ written responses to a physics conceptual question with prompt engineering and few-shot learning. Physical Review Physics Education Research 20(1):010152
2024
-
[43]
In: Frontiers in Education, Frontiers Media SA, p 1330486
Wang KD, Burkholder E, Wieman C, et al (2024 a ) Examining the potential and pitfalls of chatgpt in science and engineering problem-solving. In: Frontiers in Education, Frontiers Media SA, p 1330486
2024
-
[44]
Wang KD, Chen Z, Wieman C (2024 b ) Can crowdsourcing platforms be useful for educational research? In: Proceedings of the 14th Learning Analytics and Knowledge Conference, pp 416--425
2024
-
[45]
://arxiv.org/abs/2109.01652, 2109.01652
Wei J, Bosma M, Zhao VY, et al (2022) Finetuned language models are zero-shot learners. ://arxiv.org/abs/2109.01652, 2109.01652
2022 arXiv
-
[46]
://arxiv.org/abs/2303.17012, 2303.17012
West CG (2023) Advances in apparent conceptual physics reasoning in gpt-4. ://arxiv.org/abs/2303.17012, 2303.17012
2023 arXiv
-
[47]
In: International conference on machine learning, PMLR, pp 11328--11339
Zhang J, Zhao Y, Saleh M, et al (2020) Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In: International conference on machine learning, PMLR, pp 11328--11339
2020
-
[48]
Review of educational research 86(4):1052--1084
Zheng B, Warschauer M, Lin CH, et al (2016) Learning in one-to-one laptop environments: A meta-analysis and research synthesis. Review of educational research 86(4):1052--1084
2016
-
[49]
write newline
" write newline " cite write " FUNCTION editor.postfix editor num.names #1 > "( )" "( )" if FUNCTION editor.trans.postfix editor num.names #1 > "( )" "( )" if FUNCTION trans.postfix translator num.names #1 > "( )" "( )" if FUNCTION authors.editors.reflist.apa5 'field := 'dot :...
-
[50]
, " * write output.state after.block = add.period write newline
ENTRY address archive author booktitle chapter doi edition editor eid eprint howpublished institution journal key keywords month note number organization pages publisher school series title type url volume year archivePrefix primaryClass adsurl adsnote version label extra.labe...
-
[51]
write newline
" write newline "" before.all 'output.state := FUNCTION add.period duplicate empty 'skip "." * add.blank if FUNCTION if.digit duplicate "0" = swap duplicate "1" = swap duplicate "2" = swap duplicate "3" = swap duplicate "4" = swap duplicate "5" = swap duplicate "6" = swap dupl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.