REVIEW 4 major objections 5 minor 30 references
Structured Prompts, Better Outcomes? Exploring the Effects of a Structured Interface with ChatGPT in a Graduate Robotics Course
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A structured ChatGPT interface can shift prompting behavior while it is in place, but it does not change learning or performance, and the behavioral gains disappear as soon as the interface is removed.
desk verdict A careful exploratory study where the session mismatch between behavior and learning association breaks the causal claim, but the descriptive data and honest reporting deserve review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a form-based structured interface that sits in front of the ChatGPT API and forces students to label each query as Understanding, Implementing, or Debugging and to compose the prompt through blank fields tailored to that category, so the model only receives the prompt after the student has decomposed it. The analysis then rests on a hand-coded taxonomy of prompts—type (Development, Conceptual, Debugging) and attributes (Understanding, Granularity, Clarity)—and on the proportion of each prompt that is student-generated rather than copy-pasted. This taxonomy is the instrument that links the interface to learning: the paper argues the interface raises the rate of clear Development prompts and self-written text, and that those particular behaviors are the ones associated with learning gains in the regression.
What would settle it
Administer the identical pre- and post-test instrument—same items, same format—to a new sample and check whether clear Development prompts still predict normalized gains; if the prediction disappears when the tests are commensurable, the reported regression link rests on incompatible measures. In addition, compare the two groups' session-3 prompt behavior in a larger sample: if a group difference appears, the null transfer result was a power artifact rather than a genuine absence of transfer.
Extended reading notes
Core claim
The authors report a two-session intervention with 58 graduate students in a mobile robotics course. In the second practice session, students using a structured GPT platform—which required selecting a prompt category (Understanding, Implementing, or Debugging) and filling in form fields—submitted a higher proportion of Development prompts that were also coded as having Clarity (clear, explicit requests) than did students using plain ChatGPT, and they wrote a larger fraction of their prompt text themselves. Those two behaviors are the ones the authors identify as productive: in the session-3 regression, clear development prompts ($\beta = 0.77$) and clear understanding-oriented prompts ($\beta = 0.75$) predicted higher normalized learning gains, together with the proportion of understanding prompts ($\beta = 0.23$). However, the intervention produced no group differences in practice task scores or pre-post test gains, and once the structured interface was removed, the intervention group's prompting behavior was indistinguishable from the control group's. The authors conclude that temporarily restructuring the interface can promote productive prompting during use, but that such bottom-up nudges do not transfer and are not enough to move learning or performance.
Load-bearing premise
The analysis assumes that the pre-test (two open-ended questions, TA-graded) and the post-test (a different MCQ with a different number of items), after separate standardization, measure the same construct on a common scale, so that the normalized learning gain $\frac{\mathrm{post}-\mathrm{pre}}{100-\mathrm{pre}}$ is a valid individual-level measure of learning.
Editorial extensions
If this is right
- If the structured interface is used during a session, the intervention group shows a higher share of clear Development prompts and more self-written prompt text than the control group.
- Those prompting behaviors, not the interface itself, are what predict learning in the session-3 regression; without them, no group learning advantage appears.
- Removing the structured interface erases the behavioral differences: in session 3 both groups prompt alike.
- Students largely reject the structured interface: about 75% of the intervention group preferred plain ChatGPT, citing habit and interface preference, which helps explain the absence of transfer.
Reading between the lines
- If replicated, the result suggests that interface design and prompting skill are separate levers: a tool can raise the frequency of good prompts without teaching the underlying skill, so future interventions should measure whether behavior survives on a novel task, not just during scaffolding.
- The regression pattern—clarity amplifying understanding and development prompts—could be tested directly by randomly assigning students to write prompts in a clear-explicit format versus a vague format and comparing learning, which would separate the prompt's form from its content.
- The finding that self-generated prompt text marginally predicts learning ($\beta = 0.43$, $p = .075$) points to a testable extension: logging keystrokes or draft iterations, rather than counting characters, could distinguish 'thinking through the prompt' from 'typing slowly.'
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This mixed-methods classroom study randomly assigns 58 graduate students in a robotics course to either a structured, form-based prompting interface or to free ChatGPT for two practice lab sessions, followed by a third session in which all students use free ChatGPT. The paper analyzes perception surveys, prompt logs, task performance, and pre/post learning tests. The authors report no significant group differences in performance or learning, but they find that during the intervention the structured-interface group produced a higher proportion of clear Development prompts and more self-written text, and that these behaviors are associated with higher learning gains in a Session 3 regression. The paper concludes that structured interfaces can promote productive prompting behaviors during use, that these behaviors do not transfer once scaffolding is removed, and that most students preferred the unconstrained ChatGPT interface.
Significance. If its central claim held, this would be a valuable, ecologically valid contribution to the literature on LLM use in education. The study has clear strengths: a randomized classroom intervention, process-level log data, inter-rater reliability checks for coding, multiple outcome measures, and a mixed-methods design that gives voice to students' resistance. The process data are especially useful for understanding how interface design shapes prompting behavior. However, the inferential chain from interface to behavior to learning is weakened by a session mismatch, by the questionable commensurability of the pre-test and post-test instruments, and by the statistical fragility of the regression results. As reported, the paper is best read as an exploratory analysis with useful descriptive findings; the causal interpretation needs substantial revision.
major comments (4)
- [§4.1, §4.2, Table 2] The paper's central mediation-style claim—that the structured interface improved learning because it increased clear Development prompts—is not actually tested. The regression linking prompt behaviors to learning gains is reported only for Session 3 (Table 2: F(8,45)=2.06, p<.01, R²=.45), while the behavioral difference induced by the intervention appears only in Session 2 (Section 4.2: χ²(3)=8.57, p=.036 for the Development×Clarity interaction; t(56)=2.07, p=.048 for self-written text). The manuscript does not report whether Session 2 clear Development prompts predict Session 2 learning gains, nor does it test the relevant interaction between condition and behavior on learning. Because Session 3 is measured after the intervention was removed and pools both conditions, the observed association may reflect stable individual differences (prior knowledge, prompt-writing skill) rather than a causal effect of the behavior. The sentence in §5.1 that the interface promoted 'prompting behaviors found to be productive for learning' is therefore an extrapolation rather than a result directly supported by the analyses.
- [§3.2, §3.3] The normalized learning gain variable defined in §3.3, (post−pre)/(100%−pre), requires the pre-test and post-test scores to be on a common scale with a common ceiling. The pre-test is two open-ended questions scored by TAs, while the post-test is an MCQ with a different number of items (8 in Session 2, 5 in Session 3), and the two scores are standardized separately. No equating evidence is provided. If the instruments are not commensurable, the learning-gain regressions in §4.1 and Table 2—and the associated claims about which prompting behaviors 'contribute to learning'—are not interpretable as individual-level learning effects. Please re-analyze with raw scores or an explicitly equated scale, or present the gain analysis only as a sensitivity check.
- [§4.1, Table 2] The stepwise forward regression enters eight predictors and interactions on 54 observations, and the final model is presented with unadjusted p-values and very large interaction coefficients (development:clarity β=0.77, clarity:understanding β=0.75). Stepwise selection invalidates the nominal p-values, and the model has a high risk of overfitting given the sample size (adjusted R²=.35). A related marginal effect is treated as a supportive trend: the proportion of student-generated text in prompts has β=0.43, p=.075 in §4.2, and §5.1 describes this as a 'trend' aligning with the hypothesis. Please present a pre-specified model, evaluate stability (bootstrap or cross-validation), and address the multiple-comparisons problem explicitly.
- [§4.1, Limitations] The paper interprets the absence of group differences in performance and learning as 'no direct impact of the intervention' (§4.1). With 58 participants, these null analyses have wide confidence intervals, and the Limitations paragraph acknowledges limited power but does not provide effect sizes or confidence intervals for the null group contrasts. As written, the null results are presented as evidence of absence rather than as inconclusive with respect to learning effects. Please report CIs for the group contrasts and phrase the null findings accordingly.
minor comments (5)
- [§3.2] The post-test description switches between '8 advanced questions' for Session 2 and '5' for Session 3, while the pre-test has two open-ended questions; please present the scoring formula and standardization for all tests in one place for clarity.
- [§4.2] The statement that the structured-interface group 'tended to disengage more (46% did ≤3 prompts, compared to 25% in the control group)' is not accompanied by a statistical test; please add a test or mark it as descriptive only.
- [Table 2] Table 2 uses †, *, **, and *** to denote significance levels, but the table caption does not define all of these symbols; please add a footnote.
- [§4.3] The student quotes are labeled only as '(student a)' through '(student d)'; please provide minimal coding context or thematic labels so that the quotes' provenance is clearer.
- [§3.2, References] The prompt categorization in §3.2 draws on the authors' own prior work (reference [7]); please state explicitly how the coding protocol used here differs from or extends that work, to avoid any appearance of relying on an untested categorization.
Circularity Check
No circular derivation: the central claims are empirical and self-contained; only a minor, non-load-bearing self-citation appears in the prompt taxonomy.
full rationale
The paper's derivation chain is observational rather than formal. The intervention effect on prompt behavior is measured from logs (Session 2: χ²=8.57, p=.036; t(56)=2.07, p=.048) and the association between prompt attributes and learning from regressions on actual learning-gain data (Session 3: R²=.45, F(8,45)=2.06, p<.01). No outcome is defined as the parameter being fit, and no equation reduces a prediction to its own input. The prompt-type taxonomy is taken from the authors' prior work [7], which is a self-citation, but it is used as a coding scheme re-applied with inter-rater reliability (κ=.71; attribute κs .68–.76), not as an unverified uniqueness theorem or ansatz that forces the result. The normalized learning gain combines non-commensurable pre/post instruments, and the Session 3 regression is pooled across conditions, but these are construct-validity and inference concerns, not circularity. The central null results (no learning/performance differences) and the non-transfer finding would be false if the claimed behavior-learning link were forced. Therefore the paper does not exhibit self-definitional or fitted-input-as-prediction circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Pre-test and post-test scores are commensurable after standardization for normalized gain computation.
- domain assumption Prompt attributes (Understanding, Granularity, Clarity) are valid, reliably coded measures of prompt quality.
- domain assumption Volunteer participants (58 of 143) are representative of the course population, and random assignment balanced unmeasured confounders.
Cite this review
Pith. "Pith review of Structured Prompts, Better Outcomes? Exploring the Effects of a Structured Interface with ChatGPT in a Graduate Robotics Course." pith.science (2026). https://pith.science/paper/IWK3VGWC
@misc{pith2026250707767,
author = {Pith},
title = {Pith review of: Structured Prompts, Better Outcomes? Exploring the Effects of a Structured Interface with ChatGPT in a Graduate Robotics Course},
year = {2026},
howpublished = {\url{https://pith.science/paper/IWK3VGWC}},
note = {Machine review of arXiv:2507.07767}
}
read the original abstract
Prior research shows that how students engage with Large Language Models (LLMs) influences their problem-solving and understanding, reinforcing the need to support productive LLM-uses that promote learning. This study evaluates the impact of a structured GPT platform designed to promote 'good' prompting behavior with data from 58 students in a graduate-level robotics course. The students were assigned to either an intervention group using the structured platform or a control group using ChatGPT freely for two practice lab sessions, before a third session where all students could freely use ChatGPT. We analyzed student perception (pre-post surveys), prompting behavior (logs), performance (task scores), and learning (pre-post tests). Although we found no differences in performance or learning between groups, we identified prompting behaviors - such as having clear prompts focused on understanding code - that were linked with higher learning gains and were more prominent when students used the structured platform. However, such behaviors did not transfer once students were no longer constrained to use the structured platform. Qualitative survey data showed mixed perceptions: some students perceived the value of the structured platform, but most did not perceive its relevance and resisted changing their habits. These findings contribute to ongoing efforts to identify effective strategies for integrating LLMs into learning and question the effectiveness of bottom-up approaches that temporarily alter user interfaces to influence students' interaction. Future research could instead explore top-down strategies that address students' motivations and explicitly demonstrate how certain interaction patterns support learning.
Figures
Reference graph
Works this paper leans on
-
[1]
International Journal of Educational Technology in Higher Education 21(1), 10 (2024)
Abbas, M., Jam, F.A., Khan, T.I.: Is it harmful or helpful? examining the causes and consequences of generative ai usage among university students. International Journal of Educational Technology in Higher Education 21(1), 10 (2024)
work page 2024
-
[2]
In: Proceedings of the 17th international conference on mining software repositories
Abdellatif, A., Costa, D., Badran, K., Abdalkareem, R., Shihab, E.: Challenges in chatbot development: A study of stack overflow posts. In: Proceedings of the 17th international conference on mining software repositories. pp. 174–185 (2020)
work page 2020
-
[3]
Sustainability 15(17), 12983 (2023)
Bahroun, Z., Anane, C., Ahmed, V., Zacca, A.: Transforming education: A compre- hensive review of generative artificial intelligence in educational settings through bibliometric and content analysis. Sustainability 15(17), 12983 (2023)
work page 2023
-
[4]
Psychological bulletin 128(4), 612 (2002)
Barnett, S.M., Ceci, S.J.: When and where do we apply what we learn?: A taxon- omy for far transfer. Psychological bulletin 128(4), 612 (2002)
work page 2002
-
[5]
Available at SSRN 4895486 (2024)
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, O., Mariman, R.: Generative ai can harm learning. Available at SSRN 4895486 (2024)
work page 2024
-
[6]
School psychology review 42(1), 39–55 (2013)
Bowman-Perrott, L., Davis, H., Vannest, K., Williams, L., Greenwood, C., Parker, R.: Academic benefits of peer tutoring: A meta-analytic review of single-case re- search. School psychology review 42(1), 39–55 (2013)
work page 2013
-
[7]
In: International Conference on Artificial Intelligence in Education
Brender, J., El-Hamamsy, L., Mondada, F., Bumbacher, E.: Who’s helping who? when students use chatgpt to engage in practice lab sessions. In: International Conference on Artificial Intelligence in Education. pp. 235–249. Springer (2024)
work page 2024
-
[8]
Byun, H., Lee, J., Cerreto, F.A.: Relative effects of three questioning strategies in ill-structured, small group problem solving. Instr. Sci. 42, 229–250 (2014)
work page 2014
Show all 30 references
-
[9]
Choi, W.C., Chang, C.I.: A survey of techniques, key components, strategies, chal- lenges, and student perspectives on prompt engineering for large language models (llms) in education (2025)
2025
-
[10]
Computers & Education p
Deng, R., Jiang, M., Yu, X., Lu, Y., Liu, S.: Does chatgpt enhance student learn- ing? a systematic review and meta-analysis of experimental studies. Computers & Education p. 105224 (2024)
2024
-
[11]
arXiv preprint arXiv:2307.16364 (2023)
Denny, P., Leinonen, J., Prather, J., Luxton-Reilly, A., Amarouche, T., Becker, B.A., Reeves, B.N.: Promptly: Using prompt problems to teach learners how to effectively utilize ai code generators. arXiv preprint arXiv:2307.16364 (2023)
2023 arXiv
-
[12]
BJET 56(2), 489–530 (2025)
Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., Gaˇ sevi´ c, D.: Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. BJET 56(2), 489–530 (2025)
2025
-
[13]
arXiv preprint arXiv:2409.05511 (2024) Structured Prompts, Better Outcomes? 15
Favero, L., P´ erez-Ortiz, J.A., K¨ aser, T., Oliver, N.: Enhancing critical thinking in education by means of a socratic chatbot. arXiv preprint arXiv:2409.05511 (2024) Structured Prompts, Better Outcomes? 15
2024 arXiv
-
[14]
arXiv preprint arXiv:2310.13712 (2023)
Kumar, H., Musabirov, I., Reza, M., Shi, J., Wang, X., Williams, J.J., Kuzminykh, A., Liut, M.: Impact of guidance and interaction strategies for llm use on learner performance and perception. arXiv preprint arXiv:2310.13712 (2023)
2023 arXiv
-
[15]
In: ICER’2023
Lau, S., Guo, P.: From ”Ban It Till We Understand It” to ”Resistance is Futile”: How University Programming Instructors Plan to Adapt as More Students Use AI Code Generation and Explanation Tools such as ChatGPT and GitHub Copilot. In: ICER’2023. vol. 1, pp. 106–121. ACM (Sep 2023)
2023
-
[16]
Lehmann, M., Cornelius, P.B., Sting, F.J.: Ai meets the classroom: When do large language models harm learning? (2025), https://arxiv.org/abs/2409.09047
2025 arXiv
-
[17]
Science 378(6624), 1092–1097 (2022)
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al.: Competition-level code generation with alphacode. Science 378(6624), 1092–1097 (2022)
2022
-
[18]
In: Proceedings of the 55th ACM SIGCSE TS
Liu, R., Zenke, C., Liu, C., Holmes, A., Thornton, P., Malan, D.J.: Teaching cs50 with ai: leveraging generative artificial intelligence in computer science education. In: Proceedings of the 55th ACM SIGCSE TS. pp. 750–756 (2024)
2024
-
[19]
In: 2016 CHI conference on human factors in computing systems
Loksa, D., Ko, A.J., Jernigan, W., Oleson, A., Mendez, C.J., Burnett, M.M.: Pro- gramming, problem solving, and self-awareness: Effects of explicit guidance. In: 2016 CHI conference on human factors in computing systems. pp. 1449–1461 (2016)
2016
-
[20]
Available at SSRN 4300783 (2022)
Mollick, E.R., Mollick, L.: New modes of learning enabled by ai chatbots: Three methods and assignments. Available at SSRN 4300783 (2022)
2022
-
[21]
Physical Review Physics Education Research 14(1), 010115 (2018)
Nissen, J.M., Talbot, R.M., Nasim Thompson, A., Van Dusen, B.: Comparison of normalized gain and cohen’sd for analyzing gains on concept inventories. Physical Review Physics Education Research 14(1), 010115 (2018)
2018
-
[22]
Science 381(6654), 187–192 (2023)
Noy, S., Zhang, W.: Experimental evidence on the productivity effects of generative artificial intelligence. Science 381(6654), 187–192 (2023)
2023
-
[23]
In: Proceedings of the 50th ACM SIGCSE TS (2019)
Prather, J., Pettit, R., Becker, B.A., Denny, P., Loksa, D., Peters, A., Albrecht, Z., Masci, K.: First things first: Providing metacognitive scaffolding for interpreting problem prompts. In: Proceedings of the 50th ACM SIGCSE TS (2019)
2019
-
[24]
Nature 620(7972), 172–180 (2023)
Singhal, K., Azizi, S., Tu, T., Mahdavi, S.S., Wei, J., Chung, H.W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al.: Large language models encode clinical knowledge. Nature 620(7972), 172–180 (2023)
2023
-
[25]
Nature 614(7947), 214–216 (2023)
Stokel-Walker, C., Van Noorden, R.: What chatgpt and generative ai mean for science. Nature 614(7947), 214–216 (2023)
2023
-
[26]
https://doi.org/10.2139/ssrn.4914743
Swargiary, K.: The Impact of ChatGPT on Student Learning Outcomes: A Com- parative Study of Cognitive Engagement, Procrastination, and Academic Perfor- mance (Aug 2024). https://doi.org/10.2139/ssrn.4914743
2024 doi
-
[27]
Educational psychology review 10, 251–296 (1998)
Sweller, J., Van Merrienboer, J.J., Paas, F.G.: Cognitive architecture and instruc- tional design. Educational psychology review 10, 251–296 (1998)
1998
-
[28]
Computers & education 52(2), 302–312 (2009)
Teo, T.: Modelling technology acceptance in education: A study of pre-service teachers. Computers & education 52(2), 302–312 (2009)
2009
-
[29]
Nature 10 (2023)
Yang, H.: How i use chatgpt responsibly in my teaching. Nature 10 (2023)
2023
-
[30]
Computers in Human Behavior: Artificial Humans 1(2), 100005 (2023)
Yilmaz, R., Yilmaz, F.G.K.: Augmented intelligence in programming learning: Ex- amining student views on the use of chatgpt for programming learning. Computers in Human Behavior: Artificial Humans 1(2), 100005 (2023)
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.