Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

A Practical Guide for Supporting Formative Assessment and Feedback Using Generative AI

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This review argues that LLMs best serve formative assessment when organized around three instructional processes, and that current feedback evaluation misses the levels that drive learning.

desk verdict A pedagogically grounded narrative review of LLMs in formative assessment with useful prompt templates, but the practical guidance rests on untested illustrative outputs and needs framing as a starting point, not a validated toolkit. read the letter →

arxiv 2505.23405 v2 pith:DKJVGKVB submitted 2025-05-29 cs.CY cs.HC

classification cs.CYcs.HC
keywords formativeassessmentgenerativeAIlargelanguagemodelsfeedbacklevelspromptengineeringself-regulatedlearningBlackandWiliamframeworkHattieTimperley
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review argues that the flood of generative-AI classroom tools has drifted from the pedagogy it is meant to serve. The authors re-anchor LLM use in Black and Wiliam's three formative-assessment processes—clarifying where learners are going, establishing where they are, and moving them forward—and show with concrete ChatGPT prompts how free models can support each one. They further claim that existing research evaluates LLM-generated feedback as a single block, failing to distinguish task-level, process-level, and self-regulation feedback, which Hattie and Timperley's framework says serve different learning functions. If the paper is right, educators get a practical, pedagogy-grounded route to AI integration, and researchers get a clear gap: metrics must be rebuilt to judge feedback by level.

What carries the argument

The framework of Black and Wiliam (2009) and Wiliam and Thompson (2007), which crosses the three instructional processes with the three agents (teacher, learner, peer), is the organizing device; the GPS-guided learning-journey metaphor (destination, position, junction, dynamic feedback, self-monitoring, reflection) carries the argument. Hattie and Timperley's (2007) four feedback levels supply the evaluation lens: task, process, self-regulation, and self-as-person. The paper's working mechanism is prompt engineering—zero-shot, one-shot, few-shot, role-prompting—demonstrated through reproducible ChatGPT templates that teachers can adapt.

What would settle it

A field experiment in which teachers follow the paper's prompt templates and leveled-feedback framework finds no improvement in student learning or metacognition relative to generic AI use, or a benchmark study shows LLM feedback cannot be reliably classified into the three levels.

Watch

Extended reading notes

Core claim

The central claim is that formative assessment, properly understood, provides the missing organizing structure for LLM use in classrooms. Treating the LLM as an in-car GPS, the paper maps each formative process to concrete AI tasks: generating student-friendly objectives and exemplars for 'where the learner is going,' generating diagnostic multiple-choice questions for 'where the learner is,' and producing task-, process-, and self-regulation-level feedback for 'how to move learning forward.' The accompanying critique is that most LLM feedback studies evaluate feedback as a unitary component assumed to improve learning, so the field lacks benchmarks for the feedback types that actually drive learning. The review is intended as a foundation: a framework, prompt templates, and an agenda for building better evaluation metrics.

Load-bearing premise

The guide assumes that the capabilities and limits of current cost-free LLMs, demonstrated through a handful of ChatGPT interactions in the paper, are stable and general enough to anchor classroom recommendations across subjects, grades, and cultures.

Editorial extensions

If this is right

  • Educators can immediately use the prompt templates to generate student-friendly objectives, exemplars, and leveled feedback with free LLMs.
  • Researchers gain a specification for new benchmarks that rate LLM feedback separately at task, process, and self-regulation levels.
  • LLM tools trained on this framework would shift from generic answer-giving toward meta-cognitive coaching.
  • Cost-free LLMs could lower systemic and cultural barriers to formative assessment in teacher-centered and exam-driven systems.
  • AI-aware assessment designs that ask students to critique LLM output turn the model's limitations into learning material.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The framework implies that evaluation of any AI feedback tool should name the feedback level it targets; a tool that scores highly on correctness could still be pedagogically empty at the process level.
  • A natural testable extension is a controlled classroom trial comparing the paper's leveled prompts against generic prompting across subjects and grades, since the paper's evidence is illustrative rather than systematic.
  • The GPS metaphor suggests an adaptive prompt schedule in which LLM feedback degrades gracefully into self-regulation prompts as learner competence rises, a claim the paper leaves implicit.
  • The systemic-barrier argument implies that free LLM access may narrow gaps in high-internet, low-resource settings, but only if teachers receive prompt training; untrained use may reproduce the 'poverty of practice' the paper cites.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper is a narrative review and practical guide that maps uses of generative LLMs (e.g., ChatGPT) onto the three instructional processes of Black and Wiliam's formative assessment framework: clarifying where the learner is going, establishing where the learner is, and moving learning forward. It provides illustrative prompts and ChatGPT outputs, argues that current LLM feedback research does not distinguish task-level, process-level, and self-regulation feedback, and offers practical recommendations for educators and researchers, including future directions for metrics, AI-aware assessment design, and overcoming systemic and cultural barriers.

Significance. If the central claims hold, the paper offers a pedagogically grounded organization of LLM applications that researchers and educators can use to position future work, and it makes a plausible case that feedback evaluation must be stratified by Hattie and Timperley's levels. The paper is candid about its evidence: it explicitly disclaims systematic testing of prompting approaches and acknowledges output variability, including a concrete failure of an MCQ prompt. These caveats are to the authors' credit, but they also limit the confidence one can place in the paper's 'practical guide' aspects, as several core recommendations rest on a small number of author-run ChatGPT interactions.

major comments (3)
  1. [Section 3.2.2 / Table 4] The paper states that generic LLMs cannot consistently produce MCQs with distractors that satisfy human experts, and its own example in Section 3.2.2 produced an out-of-scope question about mRNA that violates the prompt's explicit constraint. Yet Table 4, presented as a summary of LLM-driven formative assessment tasks, lists 'Automatic generation of diagnostic MCQs' with a prompt that is similar in structure to the failed prompt and is not shown to have been tested. This is an internal inconsistency in the paper's main practical message. The table should either flag MCQ generation as requiring substantial expert vetting, or the prompt should be relabeled as illustrative with a caution about the failure mode documented in the text.
  2. [Section 4.3] Section 4.3 argues that cost-free LLMs can help overcome systemic and cultural barriers to formative assessment, citing contexts such as India, Saudi Arabia, China, and South Korea and claiming that LLMs can reduce teacher workload and scaffold student-centered learning in these settings. However, the evidence base is limited to a few English-language ChatGPT 4o-mini interactions in a single university context, none from the cited countries, and the paper itself documents output variability. This extrapolation is load-bearing for the paper's practical-guide claim and is not supported by the presented evidence. The authors should either reframe this as a research hypothesis or provide contextual evidence; otherwise the recommendation risks being misleading.
  3. [Section 3 preamble / Section 3.3] The abstract and Section 3.3 claim that current LLM feedback research evaluates feedback as a unitary component and that this is a significant gap. This is a plausible synthesis, but the paper is a narrative review without a reported search protocol or a method for selecting the works cited, as the Section 3 preamble disclaims systematic evaluation. If the authors want the 'most studies treat feedback as unitary' claim to be load-bearing, they should either describe the search and inclusion criteria or soften the claim to 'selected recent studies' or 'the studies reviewed here.' Without this, the generality of the claim cannot be assessed.
minor comments (5)
  1. [Section 2.2] The sentence ending 'generating responses several times can help!' appears truncated; it should probably introduce a recommendation to regenerate and compare outputs, which would be consistent with the paper's later discussion of output variability.
  2. [Section 3.3.2] The name 'Nguyen and Alan' should be 'Nguyen and Allan' to match the reference list entry (Nguyen & Allan, 2024).
  3. [Section 2.1 / References] The in-text citation 'Wiliam and Thompson (2007)' does not match the reference list entry 'Wiliam, D., & Thompson, M. (2017)'; please harmonize the citation and reference.
  4. [Table 1] The formatting of Table 1 is garbled; the one-shot/few-shot column for process-level feedback appears to be cut off mid-sentence, making it hard to verify the template.
  5. [Section 3.2.2] The phrase 'the free unlimited version available as of December 2023' describing ChatGPT 4o-mini appears to be inaccurate; ChatGPT 4o-mini was released later, so the date and model version should be verified.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the formative-assessment framework is imported from external pedagogical sources, and the authors' self-citations are contextual rather than load-bearing.

full rationale

This is a narrative review and practical guide, not a derivation. The organizing framework (Black & Wiliam's three instructional processes; Hattie & Timperley's feedback levels) comes from external, established pedagogical literature, and the paper's central contribution is the mapping of LLM applications onto that framework. No parameter is fitted and no quantity is predicted from data, so the fitted-input-called-prediction and self-definitional patterns do not arise. The authors' own citations (e.g., Prompiengchai et al., 2024; Ali et al., 2024; Ratnayake et al., 2024; Cleverley et al., 2025) appear in contextual discussions about peer assessment, procurement, reflection, and participatory research, not as the justification for the core claim that LLMs can support formative assessment. The claim that current LLM feedback research 'evaluate[s] feedback as a unitary component' is supported by citations to external work (e.g., Nguyen et al., 2024; Koutcheme et al., 2024), not by the authors' own prior results. The paper explicitly disclaims systematic testing of prompts ('the purpose is not to rigorously evaluate or systematically test the efficiency of specific prompting approaches'), so the illustrative ChatGPT outputs are demonstrations, not predictions masquerading as evidence. The skeptical concern about reliability and generalizability of the recommended prompts is a correctness or evidential weakness, not circularity. Accordingly, the circularity score is low.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper contributes no fitted parameters or invented entities. Its load-bearing assumptions are the pedagogical taxonomies it borrows (Black and Wiliam; Hattie and Timperley) and the representativeness of the ChatGPT examples. These are reasonable but unproven.

assumptions (3)
  • domain assumption Black and Wiliam's three-process model is the correct organizing framework for formative assessment
    The entire review is structured around 'where learners are going, where learners are, and how to move forward'; this is adopted without comparing alternative frameworks, Section 2.1.
  • domain assumption Hattie and Timperley's feedback levels (task, process, self-regulation, self) are the appropriate taxonomy for evaluating LLM feedback
    Used to critique existing metrics in Section 3.3; no alternative taxonomy is considered.
  • ad hoc to paper Free-tier ChatGPT outputs illustrated in the paper are representative of current LLM capabilities and are reasonably stable
    The guide's examples rely on ChatGPT 4o-mini screenshots; the paper notes output variability but still draws general guidance from these instances, Section 2.2 and Section 3.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Practical Guide for Supporting Formative Assessment and Feedback Using Generative AI." pith.science (2026). https://pith.science/paper/DKJVGKVB

@misc{pith2026250523405,
  author       = {Pith},
  title        = {Pith review of: A Practical Guide for Supporting Formative Assessment and Feedback Using Generative AI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DKJVGKVB}},
  note         = {Machine review of arXiv:2505.23405}
}
read the original abstract

Formative assessment is a cornerstone of effective teaching and learning, providing students with feedback to guide their learning. While there has been an exponential growth in the application of generative AI in scaling various aspects of formative assessment, ranging from automatic question generation to intelligent tutoring systems and personalized feedback, few have directly addressed the core pedagogical principles of formative assessment. Here, we critically examined how generative AI, especially large-language models (LLMs) such as ChatGPT, can support key components of formative assessment: helping students, teachers, and peers understand "where learners are going," "where learners currently are," and "how to move learners forward" in the learning process. With the rapid emergence of new prompting techniques and LLM capabilities, we also provide guiding principles for educators to effectively leverage cost-free LLMs in formative assessments while remaining grounded in pedagogical best practices. Furthermore, we reviewed the role of LLMs in generating feedback, highlighting limitations in current evaluation metrics that inadequately capture the nuances of formative feedback, such as distinguishing feedback at the task, process, and self-regulatory levels. Finally, we offer practical guidelines for educators and researchers, including concrete classroom strategies and future directions such as developing robust metrics to assess LLM-generated feedback, leveraging LLMs to overcome systemic and cultural barriers to formative assessment, and designing AI-aware assessment strategies that promote transferable skills while mitigating overreliance on LLM-generated responses. By structuring the discussion within an established formative assessment framework, this review provides a comprehensive foundation for integrating LLMs into formative assessment in a pedagogically informed manner.

Figures

Figures reproduced from arXiv: 2505.23405 by the authors.

Figure 5
Figure 5. Screenshot of ChatGPT generating learning objectives for a fourth-year undergraduate neuroimaging (functional MRI) course. Since large language models generate text based on probabilistic predictions of word sequences, variations in the output are inevitable, even for the same prompt, as the model may produce different responses depending on subtle variations in its internal processing or interpretation of the input… view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Feedback That Clicks: Introductory Physics Students' Valued Features in AI Feedback Generated From Self-Crafted and Engineered Prompts

    physics.ed-ph 2025-09 conditional novelty 5.0 of 10

    Introductory physics students rarely use prompt engineering on their own, but they rate AI feedback as most useful when the prompt explicitly requests evaluation, the correct answer, and improvement suggestions.

Reference graph

Works this paper leans on

17 extracted references · 8 canonical work pages · cited by 1 Pith paper

  1. [1]

    hints” or alternative steps to approaching the coding problem. “Hint

    s used to approach the task (Hattie & Timperley, 2007). For instance, the teacher might suggest the learner to think about an alternative strategy to solve a particular math problem set. Nguyen and Alan (2024) engineered prompts to provide formative code feedback for computer science students completing complex programming tasks. Although they did not sho...

  2. [3]

    next steps

    … 40 4.2 Developing Metrics to Evaluate LLM-Generated Feedback Across Different Feedback Types While LLMs have been widely explored for generating feedback, much of the literature does not distinguish between different types of feedback and their unique roles in learning. Formative assessment relies on various feedback types—such as task-level feedback, p...

  3. [5]

    I Am a Parrot

    Conclusion In this review, we have systematically situated the emerging applications of large language models within the established theoretical framework of formative assessment and feedback (Black & Wiliam, 2009; Hattie & Timperley, 2007). By organizing existing LLM research around core principles of formative assessment—clarifying intended learning goa...

  4. [9]

    J., Touchie, C., Pugh, D., Boulais, A.-P., & De Champlain, A

    Lai, H., Gierl, M. J., Touchie, C., Pugh, D., Boulais, A.-P., & De Champlain, A. (2016). Using Automatic Item Generation to Improve the Quality of MCQ Distractors. Teaching and Learning in Medicine, 28(2), 166–173. https://doi.org/10.1080/10401334.2016.1146608 Langdon, J., Botnaru, D. T., Wittenberg, M., Riggs, A. J., Mutchler, J., Syno, M., & Caciula, M....

  5. [13]

    https://www.education.gov.in/sites/upload_files/mhrd/files/NEP_Final_English_0.pdf Mizumoto, A., & Eguchi, M

    Government of India. https://www.education.gov.in/sites/upload_files/mhrd/files/NEP_Final_English_0.pdf Mizumoto, A., & Eguchi, M. (2023). Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics, 2(2), 100050. https://doi.org/10.1016/j.rmal.2023.100050 Mok, J. (2011). A case study of student...

  6. [15]

    Ratnayake, A., Bansal, A., Wong, N., Saseetharan, T., Prompiengchai, S., Jenne, A., Thiagavel, J., & Ashok, A. (2024). All “wrapped” up in reflection: Supporting metacognitive awareness to promote students’ self-regulated learning. Journal of Microbiology & Biology Education, 25(1), e00103-23. https://doi.org/10.1128/jmbe.00103-23 Ray, P. P. (2023). ChatGP...

  7. [36]

    (Shane), Reid, M., Matsuo, Y., & Iwasawa, Y

    https://doi.org/10.1186/2229-0443-1-3-36 Kojima, T., Gu, S. (Shane), Reid, M., Matsuo, Y., & Iwasawa, Y. (2022). Large Language Models are Zero-Shot Reasoners. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Advances in Neural Information 53 Processing Systems (Vol. 35, pp. 22199–22213). Curran Associates, Inc. https://proceedin...

  8. [38]

    Karataş, F., Eriçok, B., & Tanrikulu, L. (2025). Reshaping curriculum adaptation in the age of artificial intelligence: Mapping teachers’ AI ‐driven curriculum adaptation patterns. British Educational Research Journal, 51(1), 154–180. https://doi.org/10.1002/berj.4068 Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser...

Show all 17 references
  1. [63]

    Lee, J., Smith, D., Woodhead, S., & Lan, A. (2024). Math Multiple Choice Question Generation via Human-Large Language Model Collaboration (arXiv:2405.00864). arXiv. https://doi.org/10.48550/arXiv.2405.00864 Lee, U., Lee, S., Koh, J., Jeong, Y., Jung, H., Byun, G., Lee, Y., Moo...

  2. [149]

    https://doi.org/10.1057/s41599-025-04471-1 Yang, J., & Tan, C. (2019). Advancing Student-Centric Education in Korea: Issues and Challenges. The Asia-Pacific Education Researcher, 28(6), 483–493. https://doi.org/10.1007/s40299-019-00449-1 Yang, M., Qu, Q., Tu, W., Shen, Y., Zhao...

  3. [194]

    I., & Mukhi, N

    https://doi.org/10.3390/electronics14010194 Sung, C., Dhamecha, T. I., & Mukhi, N. (2019). Improving Short Answer Grading Using Transformer-Based Pre-training. In S. Isotani, E. Millán, A. Ogan, P. Hastings, B. McLaren, & R. Luckin (Eds.), Artificial Intelligence in Education (...

  4. [410]

    https://doi.org/10.3390/educsci13040410 55 Lo, L. S. (2023). The CLEAR path: A framework for enhancing information literacy through prompt engineering. The Journal of Academic Librarianship, 49(4), 102720. https://doi.org/10.1016/j.acalib.2023.102720 Logan IV, R. L., Balažević...

  5. [1259]

    R., Norbury, R., Selvaraj, S., Cowen, P

    https://doi.org/10.1057/s41599-024-03611-3 Godlewska, B. R., Norbury, R., Selvaraj, S., Cowen, P. J., & Harmer, C. J. (2012). Short-term SSRI treatment normalises amygdala hyperactivity in depressed patients. Psychological Medicine, 42(12), 2609–2617. https://doi.org/10.1017/S...

  6. [1986]

    https://www.education.gov.in/sites/upload_files/mhrd/files/upload_document/npe.pdf Ministry of Human Resource Development

    Government of India. https://www.education.gov.in/sites/upload_files/mhrd/files/upload_document/npe.pdf Ministry of Human Resource Development. (2020). National Education Policy

  7. [2005]

    https://ncert.nic.in/pdf/nc-framework/nf2005-english.pdf Nguyen, H., & Allan, V. (2024). Using GPT-4 to Provide Tiered, Formative Code Feedback. Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1, 958–964. https://doi.org/10.1145/3626252.3630960...

  8. [2020]

    https://journals-sagepub-com.myaccess.library.utoronto.ca/doi/10.1177/0098628320901385 Evans, D

    Teaching of Psychology. https://journals-sagepub-com.myaccess.library.utoronto.ca/doi/10.1177/0098628320901385 Evans, D. J. R., Zeun, P., & Stanier, R. A. (2014). Motivating student learning using a formative assessment journey. Journal of Anatomy, 224(3), 296–303. https://doi...

  9. [2024]

    You already know the key features of the 35 opening of an argument. Check to see whether you have incorporated them in your first paragraph

    3.3.3 Feedback and Self-Regulated Learning Formative assessment encapsulates self-regulated learning (Clark, 2012). Similar to formative assessment, a self-regulated learner will plan their goals and learning strategies, execute their plan while monitoring their actions, and s...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.