{"id":"33aaa5bc-226f-4524-b9af-27c3ac0a70d1","arxiv_id":"2505.13381","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a 1,002-student biology course, AI feedback type did not affect midterm performance, while required confidence ratings and explanations were self-reported as the main driver of changed study behavior.","lead":"A large biology course used an AI-powered practice exam tool that required students to rate their confidence and explain each answer; the type of AI feedback made no measurable difference in exam performance. The study suggests that forcing structured reflection, not the sophistication of feedback, is what students say changed their study habits.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Causal claim about metacognitive requirements rests on a condition held constant across all arms; no contrast supports it, so 'most impactful' is only perceived impact.","rationale":"The reader's conditional verdict already rests on the missing causal contrast for the metacognitive requirements, and my independent read converges on the same point, so the verdict should not change. The paper's strongest claim, repeated in the abstract, introduction, and discussion, is that required confidence ratings and explanations, rather than feedback type, transformed student learning behaviors. In the design, however, the requirements were a constant feature of all four randomized conditions (Section 4.2); no arm omitted them, and no pre-intervention measure of metacognitive strategy use was collected. The interview and survey evidence (Section 5.3.2) can establish perceived impact and plausible mechanisms but cannot separate the requirement effect from practice-testing effects, novelty, or self-report bias. The only quantitative link (confidence rating predicting performance) is correlational. I considered other candidate concerns—the exploratory bottom-20% subgroup trend (p=0.067), the 25.8% survey response rate, and the absence of a baseline biology assessment—but these are secondary to the causal attribution problem and are appropriately handled by the conditional verdict. The proposed concrete test is a between-student randomization of the requirements themselves; a lower-cost interim dose-response analysis on existing logs would strengthen or weaken the claim before a new deployment. With that reframing or additional evidence, the paper is valuable; without it, the headline claim should be stated as perceived impact rather than demonstrated causality.","tokens_in":16585,"tokens_out":3124,"duration_ms":33178,"concrete_test":"Run the same practice exam system in the same course with an additional between-student condition that omits (or makes optional) the confidence-rating and explanation requirements, keeping the four feedback arms identical; assign students randomly and compare midterm performance, interaction depth (e.g., time-on-task, textbook-link clicks), and post-exam survey items on strategy transfer. If the no-requirements arm shows similar strategy transfer or exam gains, the 'most substantial impact' claim fails; if the requirements arm shows superior transfer and gains, the claim is supported. As a lower-cost interim check using existing logs, test for a within-condition dose-response: among students in the same feedback condition, does explanation length/detail or confidence-slider use predict gains on overlapping learning objectives after controlling for prior practice performance?","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline conclusion—that required confidence ratings and explanations, not feedback type, drove student learning behavior change—cannot be tested with the deployed design. In Section 4.2, feedback condition is randomized at the student-question level, but the confidence-rating and explanation requirements are present in every condition, for every student, on every question. There is therefore no causal contrast between a practice exam with metacognitive requirements and one without them. The supporting evidence in Section 5.3.2 is interview and survey self-report, and Section 7 acknowledges only the missing biology baseline, not this missing control. Consequently, the strongest quantitative result—confidence ratings predicting performance (β=0.063, p<0.001, Section 5.2)—is correlational and could reflect calibration or prior knowledge rather than a causal benefit of requiring confidence ratings. The conclusion that structured reflection may be 'more impactful than sophisticated feedback mechanisms' overreaches: feedback type was compared and found null, but the metacognitive requirements were not compared with anything. This is not an internal inconsistency, but the central causal claim is underdetermined by the evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on a large-scale deployment of an AI-powered practice exam system in an introductory biology course (1002 enrolled students, 28,313 question-student interactions across three midterms). Students answered multiple-choice questions, were required to provide a confidence rating and a written explanation for each answer, and then received one of four randomly assigned feedback conditions: right/wrong only, textbook references, AI-generated feedback, or both AI feedback and textbook references. The authors find no statistically significant performance differences between feedback conditions, but report that students' confidence ratings strongly predicted performance (β=0.063, p<0.001). Survey and interview data indicate that students perceived the confidence-rating and explanation requirements as valuable, with many reporting that these requirements transferred to their actual exam strategies and study habits. The paper concludes that embedding structured reflection requirements may be more impactful than the type of feedback provided.","tokens_in":16754,"tokens_out":3899,"duration_ms":37290,"significance":"If the feedback-type null result holds, this is a valuable large-scale field result: it suggests that, in a context where students already engage in structured self-explanation and confidence assessment, the marginal benefit of different feedback formats may be smaller than laboratory studies imply. The detailed description of the system and the interaction-log analysis are useful for the learning-at-scale community. The qualitative findings about students' perceived benefits of the metacognitive requirements and their reported transfer of strategies to exams and future courses are suggestive but, as the authors present them, are not sufficient to support the causal claim that the requirements themselves drove the observed changes. The paper does not provide code or data, but the implementation details are sufficiently specific to facilitate replication.","major_comments":[{"comment":"The headline claim that 'embedding structured reflection requirements may be more impactful than sophisticated feedback mechanisms' is not causally identifiable from the design. In §4.2, all four experimental conditions require every student to provide a confidence rating and a written explanation before receiving feedback; there is no experimental condition in which these metacognitive requirements are absent. The supporting evidence in §5.3.2 consists entirely of interviews and surveys (self-reported perceptions), and Section 7 acknowledges the missing baseline biology assessment but not this missing control. Please reframe the claim as an exploratory, perceived-impact finding, or add a follow-up design that manipulates the presence of the requirements.","section":"Abstract; §4.2; §5.3.2"},{"comment":"The statement that the confidence-performance relationship 'reinforces the value of metacognitive elements in the system' over-interprets the regression coefficient (β=0.063, p<0.001). Because confidence is measured after the student has answered the question, the coefficient primarily reflects calibration or prior knowledge, not the causal effect of requiring confidence ratings. The regression also appears to treat question-level observations as independent without student-level or question-level random effects or clustered standard errors, so the reported p-value likely understates uncertainty. Please present a model with student and question random effects (or cluster-robust standard errors) and explicitly describe this coefficient as correlational.","section":"§5.2"},{"comment":"The analysis framework does not account for the repeated-measures structure of the data: the same student contributes many observations and the same learning objective appears across multiple questions. For the null feedback-type results this is a conservative direction, but for the confidence coefficient and the exploratory subgroup analyses (e.g., bottom-20% students, β=0.049, p=0.067) it could produce misleading significance if the unit of analysis is treated as fully independent. Please add mixed-effects models or cluster-robust standard errors and report how many unique students and learning objectives contribute to the n=10,820 observations.","section":"§4.3; §5.2"}],"minor_comments":[{"comment":"There is a stray 'w' in the sentence 'w One recurring suggestion among students was the inclusion of human-verified explanations.'","section":"§5.3.2"},{"comment":"Section 4.4 describes a post-midterm survey (n=279) collected across four categories, but §5.3.1 states that this survey was 'following midterm 1' only. Please clarify the timing and whether the same survey was used after each midterm or only after the first.","section":"§4.4 vs §5.3.1"},{"comment":"The comparison of the 28–39% textbook-link click rate to 'traditional reading rates' is not a controlled comparison; the click rate reflects a specific prompted in-system behavior and is not directly comparable to reading-compliance rates reported in prior studies without additional context. Please temper this phrasing.","section":"§5.3.1"},{"comment":"The error-handling description says the system 'gracefully degrades to simpler feedback modes (changing the experimental condition).' Please specify whether the analysis is intention-to-treat and whether the incidence of fallback events differed by assigned condition, as this could influence the null feedback-type results.","section":"§3.3"},{"comment":"The definition for the 'Course Importance & Motivation' theme includes the placeholder '[Anonymous course]'; this appears to be a leftover anonymization artifact and should be removed.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for L@S and the deployment scale is a clear strength. The central issue is the gap between the causal conclusion in the abstract and the non-experimental status of the metacognitive-requirement evidence. This is fixable by reframing the contribution as an exploratory qualitative finding plus a null result for feedback types, which would make the paper acceptable. No additional concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a genuinely useful large-scale deployment study, but its headline causal claim about metacognitive requirements doesn't follow from the design. The feedback-type null result is solid; the metacognitive claim rests on a condition that was held constant across all arms.\n\nWhat's new: the system combines GPT-4o feedback with textbook references and required confidence ratings/explanations, run in a real 1002-student biology course over three midterms. 28,313 question-student interactions, four within-subject conditions, plus surveys and interviews. The null result across feedback types (basic, textbook links, AI, both) is credible with this sample. The 40% textbook-link engagement rate is notable, and the qualitative data on students transferring reflection strategies to exams and later courses is genuinely interesting.\n\nSoft spots: the paper says \"the most substantial impact came from required confidence ratings and explanations.\" But every condition in Section 4.2 required those ratings and explanations, so there's no contrast between practice with and without them. The supporting evidence is survey and interview self-report. The confidence-performance correlation (β=0.063) could reflect calibration or prior knowledge, not a causal effect of requiring confidence ratings. The limitations section acknowledges missing a biology baseline, but not this missing control. The exploratory bottom-20% subgroup analysis has p=0.067; treat it as a trend, not evidence. Perceived impact is a legitimate finding, but the causal language overreaches.\n\nWho's it for: people deploying AI feedback at scale, and anyone designing metacognitive scaffolds. The null result on feedback types plus the engagement numbers are worth knowing. The metacognitive claim needs reframing as \"students perceived/attributed benefits\" or better, a randomized contrast with and without the requirements.\n\nRecommendation: send it to peer review. The design flaw is addressable, and the empirical base is stronger than most papers at this level. With a revised framing or an added contrast, it would be a solid publication.","headline":"Large-scale null result on AI feedback types is credible, but the claim that metacognitive requirements drove learning is unsupported because those requirements were constant across all conditions.","tokens_in":17262,"tokens_out":2556,"would_cite":true,"duration_ms":23617,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Structured reflection in practice exams changed students' study behavior and exam strategies, outweighing the type of AI feedback they received.","keywords":["AI-enhanced feedback","practice exams","self-regulated learning","metacognition","confidence ratings","student explanations","textbook engagement","learning at scale"],"falsifier":"Randomize one group to answer practice questions with confidence ratings and explanations and another to answer the same questions without them, keeping feedback identical; the central claim fails if the reflection-free group shows equal or larger exam gains and equal transfer of study strategies in post-exam interviews and surveys.","tokens_in":16387,"feed_emoji":"🧠","tokens_out":4897,"duration_ms":46047,"temperature":0.7,"pith_summary":"The paper reports a large deployment of an online practice-exam system in a 1,002-student introductory biology course, pairing multiple-choice practice with feedback generated by a large language model and links to textbook sections. The central claim is that requiring students to rate their confidence and explain their reasoning before seeing feedback was the most impactful design element: students reported transferring those habits to their actual exams, and textbook-link engagement reached about 40 percent, far above typical reading-compliance rates. Across roughly 28,000 question-student interactions from three midterms, feedback type (correctness-only, textbook links, AI feedback, or both) made no statistically significant difference in exam performance. The paper concludes that embedding structured reflection requirements in practice exams may be a more valuable design lever than the sophistication of the feedback itself.","feed_headline":"Requiring reflection changed study habits more than AI feedback","feed_subtitle":"In a 1,002-student biology class, forced confidence ratings and written explanations transferred to real exam strategies.","key_machinery":"The carrying mechanism is a two-phase practice-exam workflow: in test mode, each multiple-choice question requires a confidence rating and a written explanation before submission; in review mode, students receive condition-specific feedback, optionally augmented by textbook references or personalized feedback built from the student's answer, explanation, and confidence. The design also uses deterministic assignment so the same student-question pair always receives the same feedback condition, and it maps every practice question to course learning objectives to connect practice to real exam questions. The confidence and explanation requirements are the load-bearing components: they feed the feedback generator, focus students' attention on gaps between certainty and correctness, and appear to shape test-taking strategies that outlast the tool.","core_discovery":"The paper's discovery is that the mandatory metacognitive requirements—declaring confidence and writing an explanation for every answer—carried the measurable and reported learning benefits, while the AI feedback that was the system's centerpiece did not outperform a simple right-or-wrong control. Students described using the confidence check and explanation habit during the actual midterm: reordering how they attacked questions, writing out reasoning, slowing down, and checking whether their certainty was backed by reasons. High confidence was the strongest consistent predictor of exam performance, and students were most attentive to feedback when confidence and correctness mismatched. The authors conclude that embedding structured reflection requirements in practice exams may be more impactful than the content of the feedback itself.","pith_inferences":["A clean causal test the paper does not include would randomize students to receive or not receive the confidence and explanation requirements while holding feedback identical; if the requirements are the active ingredient, the reflection-free arm should show weaker learning-behavior changes.","Because all four experimental conditions included the requirements, the null feedback-type result likely reflects the requirements' strong baseline effect rather than feedback being irrelevant; removing them might restore a detectable difference between feedback types.","The confidence ratings students gave could be repurposed as an instructor-facing diagnostic signal, showing at scale which learning objectives are being overestimated or underestimated by the class.","The explanation requirement could be made adaptive—skipped for very high-confidence correct answers and for pure guesses—to reduce the fatigue that some students reported while preserving the reflection benefit where it matters most."],"forward_implications":["Course designers should treat required confidence ratings and written explanations as primary practice-exam features, not optional add-ons.","AI feedback systems do not need to be elaborate to produce benefit at scale; simple correctness feedback may suffice once reflection is required, though combined feedback may help lower-performing students.","Practice tools can lift textbook engagement: roughly 40 percent of students acted on textbook references, suggesting just-in-time links in feedback outperform assigned reading.","Students' transfer of reflection habits to real exams implies practice tools can change study strategies, not only test scores.","Systems should track confidence-correctness mismatches as the moments when students are most open to feedback."],"supporting_citations":[{"why":"It supplies the three-question feedback framework and the connection between feedback and self-regulation that motivates the design.","marker":"[19]"},{"why":"It establishes response certitude as a determinant of feedback receptivity, which the AI feedback design explicitly uses.","marker":"[27]"},{"why":"It provides the basis for comparing simple correctness feedback with elaborated feedback, the comparison at the heart of the experiment.","marker":"[28]"},{"why":"It supports practice testing as a learning strategy, the foundation the practice-exam tool builds on.","marker":"[40]"},{"why":"It offers meta-analytic evidence that feedback effectiveness varies, informing the expectation that feedback type should matter.","marker":"[56]"},{"why":"It reports diagnostic feedback outperforming correct-incorrect feedback, the prior result the paper tests and does not replicate.","marker":"[50]"},{"why":"It shows LLM-powered reflection prompts improving learning at scale, a direct precedent for the reflection-focused findings.","marker":"[30]"},{"why":"It documents low textbook reading compliance, providing the baseline against which the 40 percent textbook-link engagement is interpreted.","marker":"[10]"}],"fun_headline_variants":["Reflection beats AI feedback in practice exams","Requiring confidence checks changes study habits","AI feedback no match for forced reflection","Metacognitive prompts, not AI, drive learning gains"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that students' self-reports correctly identify the confidence and explanation requirements—rather than the act of practicing or the feedback—as the cause of their changed exam strategies, and since every condition included those requirements, the study cannot contrast them.","fun_headline_variants_meta":{"raw":{"variants":["Reflection beats AI feedback in practice exams","Requiring confidence checks changes study habits","AI feedback no match for forced reflection","Metacognitive prompts, not AI, drive learning gains"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000165,"raw_usage":{"total_tokens":1244,"prompt_tokens":934,"completion_tokens":310,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":254}},"tokens_in":550,"tokens_out":310,"duration_ms":3447,"temperature":1.0,"reasoning_tokens":254,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:14:17.886919+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Randomize one group to answer practice questions with confidence ratings and explanations and another to answer the same questions without them, keeping feedback identical; the central claim fails if the reflection-free group shows equal or larger exam gains and equal transfer of study strategies in post-exam interviews and surveys.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It establishes response certitude as a determinant of feedback receptivity, which the AI feedback design explicitly uses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the basis for comparing simple correctness feedback with elaborated feedback, the comparison at the heart of the experiment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supports practice testing as a learning strategy, the foundation the practice-exam tool builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It offers meta-analytic evidence that feedback effectiveness varies, informing the expectation that feedback type should matter."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It reports diagnostic feedback outperforming correct-incorrect feedback, the prior result the paper tests and does not replicate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It documents low textbook reading compliance, providing the baseline against which the 40 percent textbook-link engagement is interpreted."}],"review_version":1}