REVIEW 3 major objections 5 minor 58 references
How Adding Metacognitive Requirements in Support of AI Feedback in Practice Exams Transforms Student Learning Behaviors
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Structured reflection in practice exams changed students' study behavior and exam strategies, outweighing the type of AI feedback they received.
desk verdict Large-scale null result on AI feedback types is credible, but the claim that metacognitive requirements drove learning is unsupported because those requirements were constant across all conditions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a two-phase practice-exam workflow: in test mode, each multiple-choice question requires a confidence rating and a written explanation before submission; in review mode, students receive condition-specific feedback, optionally augmented by textbook references or personalized feedback built from the student's answer, explanation, and confidence. The design also uses deterministic assignment so the same student-question pair always receives the same feedback condition, and it maps every practice question to course learning objectives to connect practice to real exam questions. The confidence and explanation requirements are the load-bearing components: they feed the feedback generator, focus students' attention on gaps between certainty and correctness, and appear to shape test-taking strategies that outlast the tool.
What would settle it
Randomize one group to answer practice questions with confidence ratings and explanations and another to answer the same questions without them, keeping feedback identical; the central claim fails if the reflection-free group shows equal or larger exam gains and equal transfer of study strategies in post-exam interviews and surveys.
Extended reading notes
Core claim
The paper's discovery is that the mandatory metacognitive requirements—declaring confidence and writing an explanation for every answer—carried the measurable and reported learning benefits, while the AI feedback that was the system's centerpiece did not outperform a simple right-or-wrong control. Students described using the confidence check and explanation habit during the actual midterm: reordering how they attacked questions, writing out reasoning, slowing down, and checking whether their certainty was backed by reasons. High confidence was the strongest consistent predictor of exam performance, and students were most attentive to feedback when confidence and correctness mismatched. The authors conclude that embedding structured reflection requirements in practice exams may be more impactful than the content of the feedback itself.
Load-bearing premise
The load-bearing premise is that students' self-reports correctly identify the confidence and explanation requirements—rather than the act of practicing or the feedback—as the cause of their changed exam strategies, and since every condition included those requirements, the study cannot contrast them.
Editorial extensions
If this is right
- Course designers should treat required confidence ratings and written explanations as primary practice-exam features, not optional add-ons.
- AI feedback systems do not need to be elaborate to produce benefit at scale; simple correctness feedback may suffice once reflection is required, though combined feedback may help lower-performing students.
- Practice tools can lift textbook engagement: roughly 40 percent of students acted on textbook references, suggesting just-in-time links in feedback outperform assigned reading.
- Students' transfer of reflection habits to real exams implies practice tools can change study strategies, not only test scores.
- Systems should track confidence-correctness mismatches as the moments when students are most open to feedback.
Reading between the lines
- A clean causal test the paper does not include would randomize students to receive or not receive the confidence and explanation requirements while holding feedback identical; if the requirements are the active ingredient, the reflection-free arm should show weaker learning-behavior changes.
- Because all four experimental conditions included the requirements, the null feedback-type result likely reflects the requirements' strong baseline effect rather than feedback being irrelevant; removing them might restore a detectable difference between feedback types.
- The confidence ratings students gave could be repurposed as an instructor-facing diagnostic signal, showing at scale which learning objectives are being overestimated or underestimated by the class.
- The explanation requirement could be made adaptive—skipped for very high-confidence correct answers and for pure guesses—to reduce the fatigue that some students reported while preserving the reflection benefit where it matters most.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports on a large-scale deployment of an AI-powered practice exam system in an introductory biology course (1002 enrolled students, 28,313 question-student interactions across three midterms). Students answered multiple-choice questions, were required to provide a confidence rating and a written explanation for each answer, and then received one of four randomly assigned feedback conditions: right/wrong only, textbook references, AI-generated feedback, or both AI feedback and textbook references. The authors find no statistically significant performance differences between feedback conditions, but report that students' confidence ratings strongly predicted performance (β=0.063, p<0.001). Survey and interview data indicate that students perceived the confidence-rating and explanation requirements as valuable, with many reporting that these requirements transferred to their actual exam strategies and study habits. The paper concludes that embedding structured reflection requirements may be more impactful than the type of feedback provided.
Significance. If the feedback-type null result holds, this is a valuable large-scale field result: it suggests that, in a context where students already engage in structured self-explanation and confidence assessment, the marginal benefit of different feedback formats may be smaller than laboratory studies imply. The detailed description of the system and the interaction-log analysis are useful for the learning-at-scale community. The qualitative findings about students' perceived benefits of the metacognitive requirements and their reported transfer of strategies to exams and future courses are suggestive but, as the authors present them, are not sufficient to support the causal claim that the requirements themselves drove the observed changes. The paper does not provide code or data, but the implementation details are sufficiently specific to facilitate replication.
major comments (3)
- [Abstract; §4.2; §5.3.2] The headline claim that 'embedding structured reflection requirements may be more impactful than sophisticated feedback mechanisms' is not causally identifiable from the design. In §4.2, all four experimental conditions require every student to provide a confidence rating and a written explanation before receiving feedback; there is no experimental condition in which these metacognitive requirements are absent. The supporting evidence in §5.3.2 consists entirely of interviews and surveys (self-reported perceptions), and Section 7 acknowledges the missing baseline biology assessment but not this missing control. Please reframe the claim as an exploratory, perceived-impact finding, or add a follow-up design that manipulates the presence of the requirements.
- [§5.2] The statement that the confidence-performance relationship 'reinforces the value of metacognitive elements in the system' over-interprets the regression coefficient (β=0.063, p<0.001). Because confidence is measured after the student has answered the question, the coefficient primarily reflects calibration or prior knowledge, not the causal effect of requiring confidence ratings. The regression also appears to treat question-level observations as independent without student-level or question-level random effects or clustered standard errors, so the reported p-value likely understates uncertainty. Please present a model with student and question random effects (or cluster-robust standard errors) and explicitly describe this coefficient as correlational.
- [§4.3; §5.2] The analysis framework does not account for the repeated-measures structure of the data: the same student contributes many observations and the same learning objective appears across multiple questions. For the null feedback-type results this is a conservative direction, but for the confidence coefficient and the exploratory subgroup analyses (e.g., bottom-20% students, β=0.049, p=0.067) it could produce misleading significance if the unit of analysis is treated as fully independent. Please add mixed-effects models or cluster-robust standard errors and report how many unique students and learning objectives contribute to the n=10,820 observations.
minor comments (5)
- [§5.3.2] There is a stray 'w' in the sentence 'w One recurring suggestion among students was the inclusion of human-verified explanations.'
- [§4.4 vs §5.3.1] Section 4.4 describes a post-midterm survey (n=279) collected across four categories, but §5.3.1 states that this survey was 'following midterm 1' only. Please clarify the timing and whether the same survey was used after each midterm or only after the first.
- [§5.3.1] The comparison of the 28–39% textbook-link click rate to 'traditional reading rates' is not a controlled comparison; the click rate reflects a specific prompted in-system behavior and is not directly comparable to reading-compliance rates reported in prior studies without additional context. Please temper this phrasing.
- [§3.3] The error-handling description says the system 'gracefully degrades to simpler feedback modes (changing the experimental condition).' Please specify whether the analysis is intention-to-treat and whether the incidence of fallback events differed by assigned condition, as this could influence the null feedback-type results.
- [Table 1] The definition for the 'Course Importance & Motivation' theme includes the placeholder '[Anonymous course]'; this appears to be a leftover anonymization artifact and should be removed.
Circularity Check
No circularity; the paper's empirical claims rest on observed data and self-reports, not on a derivation that feeds its conclusion back into its premises.
full rationale
This paper reports an empirical field experiment rather than a formal derivation, so the circularity patterns do not apply. The central claim that required confidence ratings and explanations were impactful is supported by interview and survey self-reports (Sections 5.3.1, 5.3.2) and by the observation that confidence ratings predict performance (Section 5.2, beta = 0.063), none of which is equivalent by construction to the paper's own inputs. No fitted parameter is renamed as a prediction; the randomized comparison is among feedback conditions, and the metacognitive requirements are constant across arms, which is a causal-identification limitation, not a circularity. The one self-citation (Reference [1]) is unrelated to the load-bearing argument. The lack of a control arm without metacognitive requirements and the reliance on self-report are validity concerns acknowledged in Section 7 only partially, but they do not constitute circular reasoning. Therefore no significant circularity is present, and the score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Students' self-reported attributions of behavior change to the confidence and explanation requirements are accurate.
- domain assumption The learning objective mappings and overlap weights between practice and real exam questions are valid measures of topical alignment.
- domain assumption The deterministic hash-based condition assignment distributes conditions evenly and avoids carryover effects across questions within a student.
- standard math Large-sample normality is sufficient for valid inference despite non-normal residuals.
Cite this review
Pith. "Pith review of How Adding Metacognitive Requirements in Support of AI Feedback in Practice Exams Transforms Student Learning Behaviors." pith.science (2026). https://pith.science/paper/TQVTSI77
@misc{pith2026250513381,
author = {Pith},
title = {Pith review of: How Adding Metacognitive Requirements in Support of AI Feedback in Practice Exams Transforms Student Learning Behaviors},
year = {2026},
howpublished = {\url{https://pith.science/paper/TQVTSI77}},
note = {Machine review of arXiv:2505.13381}
}
read the original abstract
Providing personalized, detailed feedback at scale in large undergraduate STEM courses remains a persistent challenge. We present an empirically evaluated practice exam system that integrates AI generated feedback with targeted textbook references, deployed in a large introductory biology course. Our system encourages metacognitive behavior by asking students to explain their answers and declare their confidence. It uses OpenAI's GPT-4o to generate personalized feedback based on this information, while directing them to relevant textbook sections. Through interaction logs from consenting participants across three midterms (541, 342, and 413 students respectively), totaling 28,313 question-student interactions across 146 learning objectives, along with 279 surveys and 23 interviews, we examined the system's impact on learning outcomes and engagement. Across all midterms, feedback types showed no statistically significant performance differences, though some trends suggested potential benefits. The most substantial impact came from the required confidence ratings and explanations, which students reported transferring to their actual exam strategies. About 40 percent of students engaged with textbook references when prompted by feedback -- far higher than traditional reading rates. Survey data revealed high satisfaction (mean rating 4.1 of 5), with 82.1 percent reporting increased confidence on practiced midterm topics, and 73.4 percent indicating they could recall and apply specific concepts. Our findings suggest that embedding structured reflection requirements may be more impactful than sophisticated feedback mechanisms.
Figures
Reference graph
Works this paper leans on
-
[1]
Mak Ahmad and Kwan-Liu Ma. 2024. More Than Chatting: Conversational LLMs for Enhancing Data Visualization Competencies. In EuroVis 2024 - Education How Adding Metacognitive Requirements in Support of AI Feedback in Practice Exams Transforms Student Learning Behaviors L@S ’25, July 21–23, 2025, Palermo, Italy Papers, Elif E. Firat, Robert S. Laramee, and N...
-
[2]
John Biggs and Catherine Tang. 2011. Train-the-trainers: Implementing outcomes- based teaching and learning in Malaysian higher education. Malaysian Journal of Learning and Instruction 8 (2011), 1–19
work page 2011
-
[3]
Brian Brost and Karen Bradley. 2006. Student compliance with assigned reading: A case study. Journal of the Scholarship of Teaching and Learning (2006), 101–111
work page 2006
-
[4]
Andrew C Butler, Jeffrey D Karpicke, and Henry L Roediger III. 2007. The effect of type and timing of feedback on learning from multiple-choice tests. Journal of Experimental Psychology: Applied 13, 4 (2007), 273
work page 2007
-
[5]
Deborah L Butler and Philip H Winne. 1995. Feedback and self-regulated learning: A theoretical synthesis. Review of educational research 65, 3 (1995), 245–281
1995
-
[6]
Shana K Carpenter. 2012. Testing enhances the transfer of learning. Current directions in psychological science 21, 5 (2012), 279–283
work page 2012
-
[7]
Mei-Mei Chang. 2010. Effects of self-monitoring on web-based language learner’s performance and motivation. Calico Journal 27, 2 (2010), 298–310
work page 2010
-
[8]
Kathy Charmaz. 2008. Grounded theory as an emergent method. Handbook of emergent methods 155 (2008), 172
work page 2008
Show all 58 references
-
[9]
Grace Chipperfield, Lauren Butterworth, and Pablo Munguia. 2022. Embedding resources into digital assessment rubrics: Bringing academic support directly to students. Journal of Academic Language and Learning 16, 1 (2022), C1–C11
2022
-
[10]
Michael A Clump, Heather Bauer, and Catherine Bradley. 2004. The extent to which psychology students read textbooks: a multiple class analysis of reading across the psychology curriculum. Journal of Instructional Psychology 31, 3 (2004), 227–233
2004
-
[11]
Patricia A Connor-Greene. 2000. Assessing and promoting student learning: Blurring the line between teaching and testing. Teaching of Psychology 27, 2 (2000), 84–88
2000
-
[12]
Juliet Corbin et al. 1990. Basics of qualitative research grounded theory proce- dures and techniques. (1990)
1990
-
[13]
Wei Dai, Yi-Shan Tsai, Jionghao Lin, Ahmad Aldino, Hua Jin, Tongguang Li, Dragan Gašević, and Guanliang Chen. 2024. Assessing the proficiency of large language models in automatic feedback generation: An evaluation study. Com- puters and Education: Artificial Intelligence 7 (2...
2024
-
[14]
W Pitt Derryberry and Steven R Wininger. 2008. Relationships among Textbook Usage and Cognitive-Motivational Constructs. Teaching Educational Psychology 3, 2 (2008), n2
2008
-
[15]
Cheryl Cisero Durwin and William M Sherman. 2008. Does choice of college textbook make a difference in students’ comprehension? College teaching 56, 1 (2008), 28–34
2008
-
[16]
Michelle French, Franco Taverna, Melody Neumann, Lena Paulo Kushnir, Jason Harlow, David Harrison, and Ruxandra Serbanescu. 2015. Textbook use in the sciences and its relation to course performance. College Teaching 63, 4 (2015), 171–177
2015
-
[17]
Wayne Geerling, G Dirk Mateer, Jadrian Wooten, and Nikhil Damodaran. 2023. Is ChatGPT smarter than a student in principles of economics. A vailable at SSRN 4356034 (2023)
2023
-
[18]
Sarah J Hatteberg and Kody Steffy. 2013. Increasing reading compliance of undergraduates: An evaluation of compliance methods. Teaching Sociology 41, 4 (2013), 346–352
2013
-
[19]
John Hattie and Helen Timperley. 2007. The power of feedback. Review of educational research 77, 1 (2007), 81–112
2007
-
[20]
Cynthia E Heiner, Amanda I Banet, and Carl Wieman. 2014. Preparing students for class: How to get 80% of students reading the textbook before class. American Journal of Physics 82, 10 (2014), 989–996
2014
-
[21]
Mary E Hoeft. 2012. Why university students don’t read: What professors can do to increase compliance. International journal for the scholarship of teaching and learning 6, 2 (2012), 12
2012
-
[22]
Jay R Howard. 2004. Just-in-time teaching in sociology or how I convinced my students to actually read the assignment. Teaching Sociology 32, 4 (2004), 385–390
2004
-
[23]
Pamela J Howard, Meg Gorzycki, Geoffrey Desa, and Diane D Allen. 2018. Aca- demic reading: Comparing students’ and faculty perceptions of its value, practice, and pedagogy. Journal of College Reading and Learning 48, 3 (2018), 189–209
2018
-
[24]
Lasse X Jensen, Alexandra Buhl, Anjali Sharma, and Margaret Bearman. 2024. Generative AI and higher education: a review of claims from the first months of ChatGPT. Higher Education (2024), 1–17
2024
-
[25]
Bethany C Johnson and Marc T Kiviniemi. 2009. The effect of online chapter quizzes on exam performance in an undergraduate social psychology course. Teaching of Psychology 36, 1 (2009), 33–37
2009
-
[26]
Michael Klymkowsky and Melanie M Cooper. 2024. The end of multiple choice tests: using AI to enhance assessment. arXiv preprint arXiv:2406.07481 (2024)
2024 arXiv
-
[27]
Raymond W Kulhavy and William A Stock. 1989. Feedback in written instruction: The place of response certitude. Educational psychology review 1 (1989), 279–308
1989
-
[28]
Raymond W Kulhavy, Mary T White, Bruce W Topp, Ann L Chan, and James Adams. 1985. Feedback complexity and corrective efficiency. Contemporary educational psychology 10, 3 (1985), 285–291
1985
-
[29]
James A Kulik and Chen-Lin C Kulik. 1988. Timing of feedback and verbal learning. Review of educational research 58, 1 (1988), 79–97
1988
-
[30]
Harsh Kumar, Ruiwei Xiao, Benjamin Lawson, Ilya Musabirov, Jiakai Shi, Xinyuan Wang, Huayin Luo, Joseph Jay Williams, Anna N Rafferty, John Stamper, et al
-
[31]
Matthias Lehmann, Philipp B Cornelius, and Fabian J Sting. 2024. AI Meets the Classroom: When Does ChatGPT Harm Learning?arXiv preprint arXiv:2409.09047 (2024)
2024 arXiv
-
[32]
Abe Leite and Saúl A Blanco. 2020. Effects of human vs. automatic feedback on students’ understanding of AI concepts and programming style. In Proceedings of the 51st ACM Technical Symposium on Computer Science Education . 44–50
2020
-
[33]
Tara Lockhart and Mary Soliday. 2016. The critical place of reading in writing transfer (and beyond): A report of student experiences. Pedagogy 16, 1 (2016), 23–37
2016
-
[34]
Richard S Lysakowski and Herbert J Walberg. 1981. Classroom reinforcement and learning: A quantitative synthesis. The Journal of Educational Research 75, 2 (1981), 69–77
1981
-
[35]
JaneMaree Maher and Jennifer Mitchell. 2010. I’m not sure what to do! Learning experiences in the humanities and social sciences. Issues in Educational Research 20, 2 (2010), 137
2010
-
[36]
Gili Marbach-Ad and Phillip G Sokolove. 2000. Can undergraduate biology students learn to ask higher level questions? Journal of Research in Science Teaching: The Official Journal of the National Association for Research in Science Teaching 37, 8 (2000), 854–870
2000
-
[37]
Teresa Murden and Cindy S Gillespie. 1997. The role of textbooks and reading in content area classrooms: What are teachers and students saying. Exploring literacy (1997), 87–96
1997
-
[38]
Sasha Nikolic, Scott Daniel, Rezwanul Haque, Marina Belkina, Ghulam M Hassan, Sarah Grundy, Sarah Lyden, Peter Neal, and Caz Sandison. 2023. ChatGPT versus engineering education assessment: a multidisciplinary and multi-institutional benchmarking and analysis of this generativ...
2023
-
[39]
Susan Bobbitt Nolen. 1996. Why study? How reasons for learning influence strategy selection. Educational Psychology Review 8 (1996), 335–355
1996
-
[40]
Henry L Roediger and Andrew C Butler. 2011. The critical role of retrieval practice in long-term retention. Trends in cognitive sciences 15, 1 (2011), 20–27
2011
-
[41]
Ernst Z Rothkopf. 1988. Perspectives on study skills training in a realistic in- structional economy. In Learning and study strategies . Elsevier, 275–286
1988
-
[42]
John Sappington, Kimberly Kinsey, and Kirk Munsayac. 2002. Two studies of reading compliance among college students. Teaching of psychology 29, 4 (2002), 272–274
2002
-
[43]
Sima Sengupta. 2002. Developing academic reading at tertiary level: A longitudi- nal study tracing conceptual change. The reading matrix 2, 1 (2002)
2002
-
[44]
Betty Lou Smith, William G Holliday, and Homer W Austin. 2010. Students’ com- prehension of science textbooks using a question-based reading strategy. Journal of Research in Science Teaching: The Official Journal of the National Association for Research in Science Teaching 47,...
2010
-
[45]
Adele Smolansky, Andrew Cram, Corina Raduescu, Sandris Zeivots, Elaine Huber, and Rene F Kizilcec. 2023. Educator and student perspectives on the impact of generative AI on assessments in higher education. In Proceedings of the tenth ACM conference on Learning@ Scale . 378–382
2023
-
[46]
Helen St Clair-Thompson, Alison Graham, and Sara Marsham. 2018. Exploring the reading practices of undergraduate students. Education Inquiry 9, 3 (2018), 284–298
2018
-
[47]
Keith Starcher and Dennis Proffitt. 2011. Encouraging Students to Read: What Professors Are (and Aren’t) Doing About It. International Journal of Teaching and Learning in Higher Education 23, 3 (2011), 396–407
2011
-
[48]
Stephanie Stokes-Eley. 2007. Using Kolb’s experiential learning cycle in chapter presentations. Communication Teacher 21, 1 (2007), 26–29
2007
-
[49]
Petra Stutz, Maximilian Elixhauser, Judith Grubinger-Preiner, Vivienne Linner, Eva Reibersdorfer-Adelsberger, Christoph Traun, Gudrun Wallentin, Katharina Wöhs, and Thomas Zuberbühler. 2023. Ch (e) atGPT? An anecdotal approach addressing the impact of ChatGPT on teaching and l...
2023
-
[50]
Fang Tang and Peida Zhan. 2021. Does diagnostic feedback promote learning? Evidence from a longitudinal cognitive diagnostic assessment. AERA Open 7 (2021), 23328584211060804
2021
-
[51]
Andrea Tick. 2024. Exploring ChatGPT’s Potential and Concerns in Higher Education. In2024 IEEE 22nd Jubilee international symposium on intelligent systems and informatics (SISY). IEEE, 000447–000454
2024
-
[52]
Terry Tomasek. 2009. Critical reading: Using reading prompts to promote active engagement with text. International journal of teaching and learning in higher education 21, 1 (2009), 127–132
2009
-
[53]
Mario Vafeas. 2013. Attitudes toward, and use of, textbooks among marketing undergraduates: an exploratory study.Journal of Marketing Education 35, 3 (2013), 245–258. L@S ’25, July 21–23, 2025, Palermo, Italy Mak Ahmad, Prerna Ravi, David Karger, and Marc Facciotti
2013
-
[54]
Dianna L Van Blerkom, Malcolm L Van Blerkom, and Sharon Bertsch. 2006. Study strategies and generative learning: What works? Journal of College Reading and Learning 37, 1 (2006), 7–18
2006
-
[55]
Ben Ward, Deepshikha Bhati, Fnu Neha, and Angela Guercio. 2024. Analyzing the Impact of AI Tools on Student Study Habits and Academic Performance. arXiv preprint arXiv:2412.02166 (2024)
2024 arXiv
-
[56]
Benedikt Wisniewski, Klaus Zierer, and John Hattie. 2020. The power of feed- back revisited: A meta-analysis of educational feedback research. Frontiers in psychology 10 (2020), 487662
2020
-
[57]
Qi Xia, Xiaojing Weng, Fan Ouyang, Tzung Jin Lin, and Thomas KF Chiu. 2024. A scoping review on how generative artificial intelligence transforms assessment in higher education. International Journal of Educational Technology in Higher Education 21, 1 (2024), 40
2024
-
[2024]
In Proceedings of the eleventh ACM conference on learning@ scale
Supporting self-reflection at scale with large language models: Insights from randomized field experiments in classrooms. In Proceedings of the eleventh ACM conference on learning@ scale . 86–97
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.