REVIEW 4 major objections 4 minor 97 references
CoGrader: Transforming Instructors' Assessment of Project Reports through Collaborative LLM Integration
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that CoGrader's workflow — LLM-assisted metric design, instructor-chosen benchmarks, and benchmark-driven regrading — makes project-report grading more efficient and consistent while producing reliable peer-comparative…
desk verdict A well-designed instructor-in-the-loop LLM grading system whose central efficiency/consistency claim is undercut by the absence of a no-tool baseline and reliance on self-report. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the benchmark-driven regrading loop. In the first pass the LLM drafts scores and comments per metric; the instructor corrects these, then designates one strong submission as Benchmark High and one weak submission as Benchmark Low. The 'Regrade Reports' action then sends every report back through the model with explicit instructions to compare the target report's current evaluation against the two benchmarks and adjust scores and comments accordingly, so each updated comment carries the shape 'this report is stronger than the low benchmark on criterion X because...'. Radar charts comparing the focused report with both benchmarks, plus sortable score-distribution views, let the instructor audit the comparisons at a glance. This loop is what the paper credits for both halves of its claim: consistency, because every report is measured against the same instructor-approved reference points instead of a drifting internal standard, and efficiency, because the most tedious step — revisiting and rescoring earlier reports after benchmarks change — becomes a single automated call. Everything else in CoGrader (metric extraction, feedback synthesis) feeds or consumes this loop.
What would settle it
A controlled experiment would settle it: recruit two matched groups of instructors with no prior connection to the course, have one group grade the same reports with CoGrader and the other with a conventional rubric only, and compare time per report, inter-grader agreement, and agreement with the official grades; the paper reports no control condition, and its efficiency numbers are self-rated Likert scores. A cheaper check targets consistency directly: run the identical regrade prompt on the same five reports several times and measure score variance, since the paper cites evidence that LLM grading output is not stable across attempts.
Extended reading notes
Core claim
The paper's central claim is that the right unit for LLM assistance in grading is the workflow, not the individual grade: a model alone cannot judge design innovation or practical knowledge application, but a model embedded in a loop where the instructor sets the metrics and chooses the reference points can. CoGrader implements this as three connected stages. In metric design, the LLM analyzes the project requirements into objective metrics and suggests extra potential ones, while the instructor selects, edits, adds, and labels each metric as auto-grade or score-reference. In benchmarking-driven grading, the LLM drafts per-metric scores and comments, the instructor reviews and corrects them, selects a Benchmark High and a Benchmark Low, and triggers a regrade in which the model re-evaluates every report explicitly against those two benchmarks, producing comparative comments that cite why a report outperforms or falls short. In feedback generation, the system consolidates the instructor's annotations, edited scores, and AI suggestions into personalized feedback. The evaluation claims that instructors perceived the workflow as efficient and reliable (means of 5.9–6.5 on a 7-point scale across components), that their scores aligned with the original course grades (Kendall's $\tau = 0.799$, Spearman's $\rho = 0.899$, Pearson $r = 0.900$), and that they engaged critically rather than deferring — adjusting or overriding 64.09% of AI scores and comments on the subjective reference metrics. The paper itself notes that LLM outputs were not fully stable across attempts and that comments could come out generic, which is why it argues for iterative calibration as future work.
Load-bearing premise
The claim rests on the premise that the course's official grades — a weighted average of one instructor and two senior teaching assistants' evaluations — are a valid external standard and that agreement with them shows CoGrader improves grading, even though the participants came from the same institution and several had taught or assisted that same course, so the measured agreement could just reflect shared internal standards.
Editorial extensions
If this is right
- When an instructor revises a benchmark mid-way, the regrade function automatically re-scores every report against the updated reference, eliminating the most time-consuming retroactive step in the current manual workflow.
- Students receive feedback anchored to concrete exemplars, with comments that state why a report meets, exceeds, or falls short of the nominated high and low benchmarks on each metric.
- Because instructors retain authority over metric selection, benchmark choice, score edits, and final feedback, the design is a direct counter to documented failure modes of fully automated grading, such as LLM self-favoritism and generic comments.
- If the workflow holds up at larger scale, a high-enrollment project course could give every student comparative, per-metric feedback at a cost that manual grading could not sustain — the drafting and consolidation are automated while the judgment stays human.
Reading between the lines
- The 64.09% override rate on subjective reference metrics cuts against a strong reading of 'reliable': the LLM's draft functions as a scaffold that instructors must correct, so a real deployment should budget human review of most subjective scores rather than expecting to accept AI output wholesale.
- The study has no control arm, and its efficiency numbers are self-reported Likert ratings; a direct with/without comparison measuring time per report and inter-grader spread would be the natural next experiment.
- A side benefit the paper leaves implicit: the instructor-validated benchmark reports and metric sets become reusable course assets, letting a department accumulate calibrated reference corpora that stabilize grading across semesters and sections.
- Scope caution: the evidence covers five reports, one data-visualization course, one model, and participants drawn from the same institution as the ground truth — the workflow may transfer to other subjective assessments, but the effect sizes should not be expected to transfer without re-measurement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents CoGrader, a human-AI collaborative grading workflow for project reports. A formative study with six instructors motivates six design requirements, and CoGrader implements collaborative metric design, LLM-based pre-grading, instructor-selected benchmark reports, benchmark-driven regrading, and feedback generation. The evaluation is a 12-participant user study in which instructors grade five reports from a master-level data visualization course; the paper reports Likert-scale ratings of each workflow component, correlations between participant scores and the course's ground truth grades, letter-grade agreement, and interaction logs. The paper concludes that CoGrader improves grading efficiency and consistency and provides reliable peer-comparative feedback to students.
Significance. If the effectiveness claims were adequately supported, the contribution would be useful: project-report grading is a real and underserved problem, and the proposed division of labor between LLM-generated references and instructor-held authority is well motivated. The formative study, the system design, and the interaction-log analysis (e.g., the 64.09% override rate for AI-generated reference scores and comments) are valuable and worth building on. However, the empirical evaluation does not support the causal claims in the abstract, and the paper's own data undermine the consistency claim. The strengths are real but they are strengths of a system-description and formative-study paper, not of an effectiveness evaluation.
major comments (4)
- [§5.1–5.5] The abstract's causal claim that CoGrader 'improves grading efficiency and consistency' is not supported by the study design. All 12 participants used CoGrader; there is no no-tool or manual-grading control arm, no objective measure of time-to-grade, and no measurement of grading drift before versus after using the tool. Efficiency is measured only by self-report Likert items (Q9–Q20, Figures 7–8). A within-subjects comparison against the instructors' normal grading workflow, or at least objective interaction-time logs against a control condition, is needed before any improvement claim can be made.
- [§5.5, Table 1] The consistency analysis conflates agreement with an external ground truth with grading consistency. The correlations are computed over only five reports, no confidence intervals or inferential statistics are reported, and the label 'internal consistency within individual graders' is misleading because the comparison is to an external standard. More directly, Table 1 undercuts the consistency claim: letter grades for report R03 span A- to C across 12 graders, and for R05 only 2 of 12 participants match the ground-truth A-. These are evidence against inter-grader consistency, not evidence for it.
- [§5.2–5.3] The ground-truth comparison is confounded by recruitment. The ground truth grades come from the same course's instructor and two senior teaching assistants (§5.2), and participants were recruited at the authors' institution, with most having taught or served as teaching assistants for that same data visualization course (§5.3). Agreement with these grades may therefore reflect shared internal standards or prior familiarity with the assignment rather than any effect of CoGrader.
- [§6.2, §6.5] The claim of 'reliable peer-comparative feedback to students' is not tested with students. Section 6.2 reports only informal discussions with students, and Section 6.5 acknowledges that LLM-generated grading was observed to be inconsistent across attempts (P11). To support the feedback and consistency claims, a student-facing study of feedback usefulness and a reliability analysis (e.g., repeated grading runs or inter-rater agreement statistics such as ICC or Krippendorff's alpha) are required.
minor comments (4)
- [Throughout] The term 'reliance' is used where 'reliability' appears to be intended (e.g., RQ2 in §5, Figure 8 heading); please correct this throughout.
- [Table 1] The 'Assigned Grades (Ps)' row lists counts such as '9A 2A- 1B+' without explicitly mapping each count to the ordered grade labels; a legend or table heading would make the counts interpretable.
- [§5.5] Please clarify how the three correlation values were computed: whether they are averaged across participants, pooled across reports, or computed per participant; the current wording makes the sample size and aggregation ambiguous.
- [§5.2] The paper states that ground truth scores were a weighted average of evaluations from one instructor and two senior TAs, but it does not state the weights; since this grade is used as the external standard throughout, the weighting should be specified.
Circularity Check
No significant circularity: CoGrader is an empirical system paper whose evaluation uses external ground-truth grades and self-reported ratings, with no fitted parameter or self-citation chain forcing the claimed outcomes.
full rationale
This paper does not present a derivation chain in which a predicted quantity is defined in terms of, or fitted to, the evidence used to validate it. The central claims about grading efficiency and consistency are supported by an 80-minute user study with 12 participants, Likert-scale questionnaires, interaction logs, and correlations with ground-truth grades. The ground truth itself is external to the system: it is a weighted average of evaluations from one course instructor and two senior teaching assistants (Section 5.2). The reported correlations (Kendall tau = 0.799, Spearman = 0.899, Pearson = 0.900) are descriptive alignments with that external standard, not fitted parameters that make the result true by construction. The only self-citation in the paper, StuGPTViz [19], appears in the related-work discussion of LLM-powered educational support systems and is not load-bearing for any claim about CoGrader's effectiveness. The paper's explicit limitations (Section 6.5), including LLM output instability and lack of support for non-textual components, are honest caveats rather than hidden circular moves. Concerns about the absence of a no-tool control, small sample size, and reliance on self-report are threats to the strength and generalizability of the empirical evidence, but they are not circularity: the evaluation does not reduce to the system's own definitions or to the authors' prior outputs. Accordingly, the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Ground truth grades from one course instructor and two senior TAs are a valid gold standard for grading quality.
- domain assumption Five project reports from one data visualization course represent the range of project-report grading tasks.
- domain assumption Self-reported Likert ratings are valid measures of grading efficiency and consistency.
Cite this review
Pith. "Pith review of CoGrader: Transforming Instructors' Assessment of Project Reports through Collaborative LLM Integration." pith.science (2026). https://pith.science/paper/IX4C25A2
@misc{pith2026250720655,
author = {Pith},
title = {Pith review of: CoGrader: Transforming Instructors' Assessment of Project Reports through Collaborative LLM Integration},
year = {2026},
howpublished = {\url{https://pith.science/paper/IX4C25A2}},
note = {Machine review of arXiv:2507.20655}
}
read the original abstract
Grading project reports are increasingly significant in today's educational landscape, where they serve as key assessments of students' comprehensive problem-solving abilities. However, it remains challenging due to the multifaceted evaluation criteria involved, such as creativity and peer-comparative achievement. Meanwhile, instructors often struggle to maintain fairness throughout the time-consuming grading process. Recent advances in AI, particularly large language models, have demonstrated potential for automating simpler grading tasks, such as assessing quizzes or basic writing quality. However, these tools often fall short when it comes to complex metrics, like design innovation and the practical application of knowledge, that require an instructor's educational insights into the class situation. To address this challenge, we conducted a formative study with six instructors and developed CoGrader, which introduces a novel grading workflow combining human-LLM collaborative metrics design, benchmarking, and AI-assisted feedback. CoGrader was found effective in improving grading efficiency and consistency while providing reliable peer-comparative feedback to students. We also discuss design insights and ethical considerations for the development of human-AI collaborative grading systems.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Alaa Abd-Alrazaq, Rawan AlSaad, Dari Alhuwail, Arfan Ahmed, Padraig Mark Healy, Syed Latifi, Sarah Aziz, Rafat Damseh, Sadam Alabed Alrazak, Javaid Sheikh, et al. 2023. Large language models in medical education: opportunities, challenges, and future directions. JMIR Medical Education 9, 1 (2023), 148–162
2023
-
[2]
Eman A Alasadi and Carlos R Baiz. 2023. Generative AI in education and research: Opportunities, concerns, and solutions. Journal of Chemical Education 100, 8 (2023), 2965–2971
2023
-
[3]
Sharifah Balkish Albar and Jane Elizabeth Southcott. 2021. Problem and project- based learning through an investigation lesson: Significant gains in creative thinking behaviour within the Australian foundation (preparatory) classroom. Thinking Skills and Creativity 41 (2021), 100–118
2021
-
[4]
Anabela C Alves, Francisco Moreira, Rui M Sousa, and Rui M Lima. 2009. Teach- ers’ workload in a project-led engineering education approach. In International Symposium on Innovation and Assessment of Engineering Curricula , Vol. 14
2009
-
[5]
Ian Arawjo, Dongwook Yoon, and François Guimbretière. 2017. Typetalker: A speech synthesis-based multi-modal commenting system. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing. 1970–1981
2017
-
[6]
Stephan Arndt, Carolyn Turvey, and Nancy C Andreasen. 1999. Correlating and predicting psychiatric symptom ratings: Spearmans r versus Kendalls tau correlation. Journal of psychiatric research 33, 2 (1999), 97–104
1999
-
[7]
Stephen Atlas. 2023. ChatGPT for higher education and professional development: A guide to conversational AI. (2023)
2023
-
[8]
Zied Bahroun, Chiraz Anane, Vian Ahmed, and Andrew Zacca. 2023. Transform- ing education: A comprehensive review of generative artificial intelligence in educational settings through bibliometric and content analysis. Sustainability 15, 17 (2023), 129–143
2023
Show all 97 references
-
[9]
Leopold Bayerlein. 2014. Students’ feedback preferences: how do students react to timely and automatically generated assessment feedback? Assessment & Evaluation in Higher Education 39, 8 (2014), 916–931
2014
-
[10]
Gulbahar Beckett. 2002. Teacher and student evaluations of project-based in- struction. TESL Canada journal (2002), 52–66
2002
-
[11]
William N Bender. 2012. Project-based learning: Differentiating instruction for the 21st century. Corwin Press
2012
-
[12]
Margherita Bernabei, Silvia Colabianchi, Andrea Falegnami, and Francesco Costantino. 2023. Students’ use of large language models in engineering educa- tion: A case study on technology acceptance, perceptions, efficacy, and detection chances. Computers and Education: Artificia...
2023
-
[13]
Blumenfeld, Elliot Soloway, Ronald W
Phyllis C. Blumenfeld, Elliot Soloway, Ronald W. Marx, Joseph S. Krajcik, Mark Guzdial, and Annemarie Palincsar. 1991. Motivating Project-Based Learning: Sustaining the Doing, Supporting the Learning. Educational Psychologist 26, 3-4 (1991), 369–398
1991
-
[14]
Mirjam Brassler and Jan Dettmers. 2017. How to Enhance Interdisciplinary Competence—Interdisciplinary Problem-Based Learning versus Interdisciplinary Project-Based Learning. Interdisciplinary Journal of Problem-Based Learning 11, 2 (2017), 369–398
2017
-
[15]
Lyn Brodie and Peter Gibbings. 2009. Comparison of PBL assessment rubrics. In Proceedings of the Research in Engineering Education Symposium (REES 2009)
2009
-
[16]
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking large language models in retrieval-augmented generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 17754–17762
2024
-
[17]
John Chen, Xi Lu, Yuzhou Du, Michael Rejtig, Ruth Bagley, Mike Horn, and Uri Wilensky. 2024. Learning agent-based modeling with LLM companions: Experiences of novices and experts using ChatGPT & NetLogo chat. InProceedings of the 2024 CHI Conference on Human Factors in Computi...
2024
-
[18]
Shih-Yeh Chen, Chin-Feng Lai, Ying-Hsun Lai, and Yu-Sheng Su. 2022. Effect of project-based learning on development of students’ creative thinking. The International Journal of Electrical Engineering & Education 59, 3 (2022), 232–250
2022
-
[19]
Zixin Chen, Jiachen Wang, Meng Xia, Kento Shigyo, Dingdong Liu, Rong Zhang, and Huamin Qu. 2025. StuGPTViz: A Visual Analytics Approach to Understand Student-ChatGPT Interactions. IEEE Transactions on Visualization and Computer Graphics 31, 1 (2025), 908–918
2025
-
[20]
Clary, Robert F
Renee M. Clary, Robert F. Brzuszek, and C. Taze Fulford. 2011. Measuring Creativity: A Case Study Probing Rubric Effectiveness for Evaluation of Project- Based Learning Solutions. Creative Education 02, 04 (2011), 333–340
2011
-
[21]
Derrick Coetzee, Seongtaek Lim, Armando Fox, Bjorn Hartmann, and Marti A Hearst. 2015. Structuring interactions for large-scale synchronous peer learning. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing. 1139–1152
2015
-
[22]
Taufiq Daryanto, Xiaohan Ding, Lance T Wilhelm, Sophia Stil, Kirk McInnis Knutsen, and Eugenia H Rho. 2025. Conversate: Supporting Reflective Learning in Interview Practice Through Interactive Simulation and Dialogic Feedback. Proceedings of the ACM on Human-Computer Interacti...
2025
-
[23]
Joost C. F. de Winter, Dimitra Dodou, and Arno H. A. Stienen. 2023. ChatGPT in Education: Empowering Educators through Methods for Recognition and Assessment. Informatics 10, 4, 87
2023
-
[24]
Haoxiang Fan, Guanzheng Chen, Xingbo Wang, and Zhenhui Peng. 2024. Lesson- Planner: Assisting novice teachers to prepare pedagogy-driven lesson plans with large language models. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–20
2024
-
[25]
Haoxiang Fan, Changshuang Zhou, Hao Yu, Xueyang Wu, Jiangyu Gu, and Zhenhui Peng. 2025. LitLinker: Supporting the Ideation of Interdisciplinary Contexts with Large Language Models for Teaching Literature in Elementary Schools. In Proceedings of the 2025 CHI conference on human...
2025
-
[26]
Tom Farrelly and Nick Baker. 2023. Generative Artificial Intelligence: Implications and Considerations for Higher Education Practice.Education Sciences 13, 11 (2023), 2227–7102
2023
-
[27]
Artemij Fedosejev. 2015. React. js essentials. Packt Publishing Ltd
2015
-
[28]
Elena L Glassman, Aaron Lin, Carrie J Cai, and Robert C Miller. 2016. Learn- ersourcing personalized hints. In Proceedings of the 19th ACM conference on computer-supported cooperative work & social computing . 1626–1636
2016
-
[29]
O’Reilly Media, Inc
Miguel Grinberg. 2018. Flask web development: developing web applications with python. " O’Reilly Media, Inc. "
2018
-
[30]
Pengyue Guo, Nadira Saab, Lysanne S Post, and Wilfried Admiraal. 2020. A review of project-based learning in higher education: Student outcomes and measures. International journal of educational research 102 (2020), 101586
2020
-
[31]
Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, and Chris Kedzie. 2024. LLM-Rubric: A Multidimensional, Calibrated Approach to Auto- mated Evaluation of Natural Language Texts. In Proceedings of the 62nd Annual Meeting of the Association for Computational Lingui...
2024
-
[32]
Steffan Hooper, Burkhard C Wünsche, Andrew Luxton-Reilly, Paul Denny, and Tony Haoran Feng. 2024. Advancing Automated Assessment Tools-Opportunities for Innovations in Upper-level Computing Courses: A Position Paper. In Proceed- ings of the 55th ACM Technical Symposium on Comp...
2024
-
[33]
Bihao Hu, Longwei Zheng, Jiayi Zhu, Lishan Ding, Yilei Wang, and Xiaoqing Gu. 2024. Teaching Plan Generation and Evaluation With GPT-4: Unleashing the Potential of LLM in Instructional Design. IEEE Transactions on Learning Technologies 17 (2024), 1445–1459
2024
-
[34]
Qinjin Jia, Jialin Cui, Haoze Du, Parvez Rashid, Ruijie Xi, Ruochi Li, and Edward Gehringer. 2024. LLM-generated Feedback in Real Classes and Beyond: Per- spectives from Students and Instructors. In Proceedings of the 17th International Conference on Educational Data Mining . 862–867
2024
-
[35]
Hyoungwook Jin, Seonghee Lee, Hyungyu Shin, and Juho Kim. 2024. Teach ai how to code: Using large language models as teachable agents for program- ming education. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–28. CoGrader: Transforming Inst...
2024
-
[36]
Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley, Paul Denny, Michelle Craig, and Tovi Grossman. 2024. Codeaid: Evaluating a classroom deployment of an llm-based programming assistant that balances student and educator needs. In Proceedings of the CHI Conf...
2024
-
[37]
Dimitra Kokotsaki, Victoria Menzies, and Andy Wiggins. 2016. Project-based learning: A review of the literature. Improving Schools 19, 3 (2016), 267–277. https://doi.org/10.1177/1365480216659733 arXiv:https://doi.org/10.1177/1365480216659733
2016 doi
-
[38]
Milan Kostic, Hans Friedrich Witschel, Knut Hinkelmann, and Maja Spahic- Bogdanovic. 2024. LLMs in Automated Essay Evaluation: A Case Study. In Proceedings of the AAAI Symposium Series , Vol. 3. 143–147
2024
-
[39]
Joseph S Krajcik and Phyllis C Blumenfeld. 2006. Project-based learning. na
2006
-
[40]
Harsh Kumar, Ilya Musabirov, Mohi Reza, Jiakai Shi, Xinyuan Wang, Joseph Jay Williams, Anastasia Kuzminykh, and Michael Liut. 2024. Guiding Students in Using LLMs in Supported Learning Environments: Effects on Interaction Dynamics, Learner Performance, Confidence, and Trust. P...
2024
-
[41]
Harsh Kumar, Jonathan Vincentius, Ewan Jordan, and Ashton Anderson. 2025. Human creativity in the age of llms: Randomized experiments on divergent and convergent thinking. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–18
2025
-
[42]
George Leckie and Jo-Anne Baird. 2011. Rater effects on essay scoring: A multi- level analysis of severity drift, central tendency, and rater experience. Journal of Educational Measurement 48, 4 (2011), 399–418
2011
-
[43]
Hao-Ping Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wilson. 2025. The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. In Procee...
2025
-
[44]
Gilly Leshed, Diego Perez, Jeffrey T Hancock, Dan Cosley, Jeremy Birnholtz, Soy- oung Lee, Poppy L McLeod, and Geri Gay. 2009. Visualizing real-time language- based feedback on teamwork behavior in computer-mediated groups. In Proceed- ings of the SIGCHI Conference on Human Fa...
2009
-
[45]
Qingyao Li, Lingyue Fu, Weiming Zhang, Xianyu Chen, Jingwei Yu, Wei Xia, Weinan Zhang, Ruiming Tang, and Yong Yu. 2023. Adapting large language models for education: Foundational capabilities, potentials, and challenges. arXiv preprint arXiv:2401.08664 (2023)
2023 arXiv
-
[46]
Jingxian Liao, Mrinalini Singh, and Hao-Chuan Wang. 2023. Deepthinkingmap: Collaborative video reflection system with graph-based summarizing and com- menting. In Companion Publication of the 2023 Conference on Computer Supported Cooperative Work and Social Computing . 369–371
2023
-
[47]
Ming Liu, Yiling Ren, Lucy Michael Nyagoga, Francis Stonier, Zhongming Wu, and Liang Yu. 2023. Future of education in the era of generative artificial in- telligence: Consensus among Chinese scholars on applications of ChatGPT in schools. Future in Educational Research 1, 1 (2...
2023
-
[48]
Ziyi Liu, Zhengzhe Zhu, Lijun Zhu, Enze Jiang, Xiyun Hu, Kylie A Peppler, and Karthik Ramani. 2024. Classmeta: Designing interactive virtual classmate to promote VR classroom participation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems . 1–17
2024
-
[49]
Carl A Maida. 2011. Project-based learning: A critical pedagogy for the twenty- first century. Policy Futures in Education 9, 6 (2011), 759–768
2011
-
[50]
Alias Masek and Sulaiman Yamin. 2011. The effect of problem based learning on critical thinking ability: a theoretical and empirical review. International Review of Social Sciences and Humanities 2, 1 (2011), 215–221
2011
-
[51]
Kelly McConvey, Shion Guha, and Anastasia Kuzminykh. 2023. A human- centered review of algorithms in decision-making in higher education. In Pro- ceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–15
2023
-
[52]
Victoria Menzies, Catherine Hewitt, Dimitra Kokotsaki, Clare Collyer, and Andy Wiggins. 2016. Project Based Learning: evaluation report and executive summary. (2016)
2016
-
[53]
Jennifer Meyer, Thorben Jansen, Ronja Schiller, Lucas W Liebenow, Marlene Steinbach, Andrea Horbach, and Johanna Fleckenstein. 2024. Using LLMs to bring evidence-based feedback into the classroom: AI-generated feedback increases secondary students’ text revision, motivation, a...
2024
-
[54]
Jesse G Meyer, Ryan J Urbanowicz, Patrick CN Martin, Karen O’Connor, Ruowang Li, Pei-Chen Peng, Tiffani J Bright, Nicholas Tatonetti, Kyoung Jae Won, Gra- ciela Gonzalez-Hernandez, et al. 2023. ChatGPT and large language models in academia: opportunities and challenges. BioDat...
2023
-
[55]
Steven Moore, Richard Tong, Anjali Singh, Zitao Liu, Xiangen Hu, Yu Lu, Joleen Liang, Chen Cao, Hassan Khosravi, Paul Denny, et al. 2023. Empowering educa- tion with llms-the next-gen interface and content generation. In International Conference on Artificial Intelligence in E...
2023
-
[56]
Tricia J Ngoon, C Ailie Fraser, Ariel S Weingarten, Mira Dontcheva, and Scott Klemmer. 2018. Interactive guidance techniques for improving creative feedback. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . 1–11
2018
-
[57]
Linda B Nilson and Claudia J Stanny. 2015. Specifications grading: Restoring rigor, motivating students, and saving faculty time . Routledge
2015
-
[58]
Andrew M Olney. 2023. Generating Multiple Choice Questions from a Textbook: LLMs Match Human Performance on Most Metrics.. In LLM@ AIED. 111–128
2023
-
[59]
Steve Oney, Christopher Brooks, and Paul Resnick. 2018. Creating guided code explanations with chat. codes. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–20
2018
-
[60]
OpenAI. 2024. Assistants API: File Search. https://platform.openai.com/docs/ assistants/tools/file-search. [Online]
2024
-
[61]
OpenAI. 2024. Assistants API Overview. https://platform.openai.com/docs/ assistants/overview. [Online]
2024
-
[62]
OpenAI. 2024. CAPABILITIES: Structured Outputs. https://platform.openai.com/ docs/guides/structured-outputs. [Online]
2024
-
[63]
OpenAI. 2024. GPT-4o System Card. https://openai.com/index/gpt-4o-system- card/, journal=GPT-4O system card | openai. [Online; Accessed 08-August-2024]
2024
-
[64]
Sitong Pan, Robin Schmucker, Bernardo Garcia Bulle Bueno, Salome Aguilar Llanes, Fernanda Albo Alarcón, Hangxiao Zhu, Adam Teo, and Meng Xia. 2025. TutorUp: What If Your Students Were Simulated? Training Tutors to Address En- gagement Challenges in Online Learning. InProceedin...
2025
-
[65]
Ernesto Panadero and Anders Jonsson. 2020. A critical review of the arguments against the use of rubrics. Educational Research Review 30 (2020), 100329
2020
-
[66]
Arjun Panickssery, Samuel Bowman, and Shi Feng. 2024. Llm evaluators recognize and favor their own generations. Advances in Neural Information Processing Systems 37 (2024), 68772–68802
2024
-
[67]
Jungkook Park, Yeong Hoon Park, Suin Kim, and Alice Oh. 2017. Eliph: Effective visualization of code history for peer assessment in programming education. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing. 458–467
2017
-
[68]
Amirhossein Rasooli, Hamed Zandi, and Christopher DeLuca. 2018. Re- conceptualizing classroom assessment fairness: A systematic meta-ethnography of assessment literature and beyond. Studies in Educational Evaluation 56 (2018), 164–181
2018
-
[69]
Prerna Ravi, John Masla, Gisella Kakoti, Grace Lin, Emma Anderson, Matt Taylor, Anastasia Ostrowski, Cynthia Breazeal, Eric Klopfer, and Hal Abelson. 2025. Co-designing Large Language Model Tools for Project-Based Learning with K12 Educators. In Proceedings of the 2025 CHI con...
2025
-
[70]
Steven I Ross, Fernando Martinez, Stephanie Houde, Michael Muller, and Justin D Weisz. 2023. The programmer’s assistant: Conversational interaction with a large language model for software development. InProceedings of the 28th International Conference on Intelligent User Inte...
2023
-
[71]
Liam Saliba, Elisa Shioji, Eduardo Oliveira, Shaanan Cohney, and Jianzhong Qi. 2024. Learning with Style: Improving Student Code-Style Through Better Automated Feedback. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1. 1175–1181
2024
-
[72]
Irit Sasson, Itamar Yehuda, and Noam Malkinson. 2018. Fostering the skills of critical thinking and question-posing in a project-based learning environment. Thinking Skills and Creativity 29 (2018), 203–212
2018
-
[73]
Patricia Schank and Lawrence Hamel. 2004. Collaborative modeling: hiding UML and promoting data examples in NEMo. InProceedings of the 2004 ACM conference on Computer supported cooperative work . 574–577
2004
-
[74]
Philip Sedgwick. 2012. Pearson’s correlation coefficient. Bmj 345 (2012)
2012
-
[75]
Philip Sedgwick. 2014. Spearman’s rank correlation coefficient. Bmj 349 (2014)
2014
-
[76]
Zekai Shao, Siyu Yuan, Lin Gao, Yixuan He, Deqing Yang, and Siming Chen. 2025. Unlocking Scientific Concepts: How Effective Are LLM-Generated Analogies for Student Understanding and Classroom Practice?. In Proceedings of the 2025 CHI conference on human factors in computing systems
2025
-
[77]
Alexander M Sidorkin. 2024. Embracing chatbots in higher education: the use of artificial intelligence in teaching, administration, and scholarship . Taylor & Francis
2024
-
[78]
Yishen Song, Qianta Zhu, Huaibo Wang, and Qinhua Zheng. 2024. Automated Essay Scoring and Revising Based on Open-Source Large Language Models. IEEE Transactions on Learning Technologies (2024)
2024
-
[79]
John Stamper, Ruiwei Xiao, and Xinying Hou. 2024. Enhancing llm-based feed- back: Insights from intelligent tutoring systems and the learning sciences. In International Conference on Artificial Intelligence in Education . Springer, 32–43
2024
-
[80]
Xiaohang Tang, Sam Wong, Kevin Pu, Xi Chen, Yalong Yang, and Yan Chen. 2024. VizGroup: An AI-assisted Event-driven System for Collaborative Programming Learning Analytics. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–22
2024
-
[81]
Shubo Tian, Qiao Jin, Lana Yeganova, Po-Ting Lai, Qingqing Zhu, Xiuying Chen, Yifan Yang, Qingyu Chen, Won Kim, Donald C Comeau, et al. 2024. Opportunities and challenges for ChatGPT and large language models in biomedicine and health. Briefings in Bioinformatics 25, 1 (2024), bbad493
2024
-
[82]
Sue Timmis, Patricia Broadfoot, Rosamund Sutherland, and Alison Oldfield. 2016. Rethinking assessment in a digital age: Opportunities, challenges and risks.British UIST ’25, September 28-October 1, 2025, Busan, Republic of Korea Zixin Chen, Jiachen Wang, Yumeng Li, Haobo Li, C...
2016
-
[83]
Andrew Tran, Kenneth Angelikas, Egi Rama, Chiku Okechukwu, David H Smith, and Stephen MacNeil. 2023. Generating multiple choice questions for comput- ing courses using large language models. In 2023 IEEE Frontiers in Education Conference (FIE). IEEE, 1–8
2023
-
[84]
Meng-Lin Tsai, Chong Wei Ong, and Cheng-Liang Chen. 2023. Exploring the use of large language models (LLMs) in chemical engineering education: Building core course problem models with Chat-GPT. Education for Chemical Engineers 44 (2023), 71–95
2023
-
[85]
Barbara E Walvoord and Virginia Johnson Anderson. 2011. Effective grading: A tool for learning and assessment in college . John Wiley & Sons
2011
-
[86]
Xingbo Wang, Haipeng Zeng, Yong Wang, Aoyu Wu, Zhida Sun, Xiaojuan Ma, and Huamin Qu. 2020. Voicecoach: Interactive evidence-based training for voice modulation skills in public speaking. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–12
2020
-
[87]
Jeremy Warner, Amy Pavel, Tonya Nguyen, Maneesh Agrawala, and Bjoern Hartmann. 2023. Slidespecs: Automatic and interactive presentation feedback collation. In Proceedings of the 28th International Conference on Intelligent User Interfaces. 695–709
2023
-
[88]
Rick Wormeli. 2023. Fair isn’t always equal: Assessment & Grading in the Differ- entiated Classroom. Routledge
2023
-
[89]
Yu-Chun Grace Yen, Isabelle Yan Pan, Grace Lin, Mingyi Li, Hyoungwook Jin, Mengyi Chen, Haijun Xia, Steven P Dow, et al. 2024. When to Give Feedback: Exploring Tradeoffs in the Timing of Design Feedback. (2024)
2024
-
[90]
Dongwook Yoon, Nicholas Chen, Bernie Randles, Amy Cheatle, Corinna E Löck- enhoff, Steven J Jackson, Abigail Sellen, and François Guimbretière. 2016. Richre- view++ deployment of a collaborative multi-modal annotation system for in- structor feedback and peer discussion. In Pr...
2016
-
[91]
Dongwook Yoon and Piotr Mitros. 2015. Multi-modal peer discussion with RichReview on edX. In Adjunct Proceedings of the 28th Annual ACM Symposium on User Interface Software & Technology . 55–56
2015
-
[92]
Alvin Yuan, Kurt Luther, Markus Krause, Sophie Isabel Vennix, Steven P Dow, and Bjorn Hartmann. 2016. Almost an expert: The effects of rubrics and expertise on perceived value of crowdsourced design critiques. In Proceedings of the 19th ACM Conference on Computer-Supported Coo...
2016
-
[93]
Gefei Zhang, Shenming Ji, Yicao Li, Jingwei Tang, Jihong Ding, Meng Xia, Guodao Sun, and Ronghua Liang. 2025. CPVis: Evidence-based Multimodal Learning Analytics for Evaluation in Collaborative Programming. In Proceedings of the 2025 CHI conference on human factors in computin...
2025
-
[94]
Liang Zhang, Tianyi Chen, Yue Zong, and Xiaopeng Gao. 2024. A Peer Grading Approach for Open-ended Programming Projects Based on Binary System and Swiss System. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1. 1484–1490
2024
-
[95]
Chengbo Zheng, Zeyu Huang, Shuai Ma, and Xiaojuan Ma. 2024. SelfGauge: An Intelligent Tool to Support Student Self-assessment in GenAI-enhanced Project- based Learning. In Adjunct Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–3
2024
-
[96]
Chengbo Zheng, Yuheng Wu, Chuhan Shi, Shuai Ma, Jiehui Luo, and Xiaojuan Ma
-
[2023]
In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems
Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–19
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.