REVIEW 3 major objections 5 minor 1 cited by
Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read A transferable iterative reflection module lets LLM-based student agents predict future quiz correctness more accurately than classical deep learning knowledge tracing models.
desk verdict A useful new dataset and a plausible reflection-based method, but the headline accuracy gain over deep learning baselines is not yet protected against split and retrieval variance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the transferable iterative reflection (TIR) module: a two-agent loop in which a reflective LLM explains its wrong predictions on training students, a novice agent re-predicts using that reflection, and the loop continues until 100% accuracy or the maximum iteration count, at which point the best reflection per lecture is logged into a reflection database. At test time, the database supplies example demonstrations from M=4 randomly selected training students in the same lecture, letting a new reflective agent generate one transferable reflection for the test student without seeing ground truth. This compresses long course materials and past history into actionable reasoning, overcoming the token limits of models such as BERT and making prompting-based LLMs more data-efficient.
What would settle it
Hold the model and data fixed but replace the test-time reflections with reflections harvested from a different lecture than the one being simulated; the claim predicts accuracy should fall well below the reported 0.7012 on EduAgent, so little or no drop would indicate that lecture-specific transfer is not the mechanism at work. A second check is to set the number of retrieved training students M to zero and observe whether accuracy falls back toward the non-TIR baseline.
Extended reading notes
Core claim
The paper's central claim is that TIR lets LLM-based student simulation outperform classical deep learning knowledge tracing models at predicting future question-answer correctness, even with limited demonstration data. The decisive comparison is on the EduAgent dataset: BertKT+TIR reaches 0.7012 accuracy versus SimpleKT's 0.6772, and TIR shows a similar or larger margin on the authors' newly collected dataset. TIR works by running an iterative reflection protocol during training—a reflective agent generates reasons for wrong predictions, a novice agent tests them, and the best reflection per lecture is stored—then, at test time, a new reflective agent retrieves reflections from M=4 randomly selected same-lecture training students and produces a single reflection for the unseen student without ground-truth labels. The paper additionally argues that TIR captures the granular dynamism of learning performance—individual, lecture, question, and skill levels—and inter-student correlations better than non-TIR baselines.
Load-bearing premise
The whole gain rests on the assumption that summaries of why a model made wrong predictions for a few randomly chosen training students in the same lecture can be reused for a new student, even though the new student's summary is written without being told the correct answers; if those summaries are tuned to particular students or lectures rather than generalizable, the accuracy improvement disappears.
Editorial extensions
If this is right
- TIR-augmented LLMs, including smaller models, can beat classical deep learning knowledge tracing baselines that train on all training students, using only a handful of example demonstrations.
- The same reflection database can augment prompting-based models (standard and chain-of-thought) and finetuning-based BERT, because reflections compress long course text into short inputs.
- Simulated students track real students' average accuracy across lectures, questions, and slides; in the authors' data, TIR raised the Pearson correlation on individual differences from 0.02 to 0.42 and on question-level accuracy from -0.50 to 0.37.
- The framework points toward a digital twin classroom, where instructors can preview how new course materials and test questions would affect a cohort before teaching real students.
Reading between the lines
- The paper explicitly leaves cross-lecture generalization untested; a direct test would be to build the reflection database from one lecture and apply it to a different lecture, predicting a large accuracy drop if the reflection content is lecture-specific.
- Because test-time reflections are generated without ground-truth feedback, the method's success likely depends on how well the M=4 randomly selected training reflections transfer; ablating M and switching from random to similarity-based retrieval would reveal how sensitive the gain is.
- If the mechanism generalizes, a practical consequence is that educators could use cohort-level simulated accuracy to compare curriculum versions before deployment, not just to predict individual quiz scores; a minimal test would compare simulated aggregate accuracy against a real cohort taught with the same new materials.
- The current pipeline converts slide images to text and excludes lecturer transcripts, so richer material representations could either further improve simulation or show that text-only course content is sufficient for the observed gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Classroom Simulacra, a framework for simulating student learning behavior in online education using LLM-based generative agents. The authors contribute a new fine-grained dataset (CogEdu) collected from a 6-week online workshop with 60 students, and a transferable iterative reflection (TIR) module that augments both prompting-based and finetuning-based LLM student simulators. TIR iteratively refines reflections on training students using ground-truth labels, stores successful reflections in a database, and at test time retrieves reflections from M=4 randomly selected training students in the same lecture to guide a reflective agent. The paper evaluates the approach on the EduAgent public dataset and the new CogEdu dataset, comparing against deep learning knowledge tracing baselines (DKT, AKT, ATKT, DKVMN, SimpleKT). The central claim is that TIR enables LLM-based simulation to outperform classical deep learning models in predicting future question-answer correctness, even with limited demonstration data.
Significance. If the central claim holds, this is a valuable contribution to HCI and educational technology. The paper provides a new dataset with granular course-material annotations, which is rare and useful for contextual student modeling. The TIR method is a practical technique for compressing long course materials and transferring reflections across students, and the authors release both the system and model implementations. The paper also goes beyond aggregate accuracy by examining individual-, lecture-, question-, and skill-level correlations, which is a useful evaluation lens. However, the empirical superiority claim rests on a single data split and a single stochastic instantiation of the retrieval procedure, so the significance of the headline result is not yet firmly established; the evaluation needs additional variance analysis and robustness checks before the claim can be fully accepted.
major comments (3)
- [Section 3.4, Table 1, Figure 9]
- [Section 6.2, Figures 10-12]
- [Section 3.4, Section 6.8, Section 7.2.3]
minor comments (5)
- [Section 3.4]
- [Section 3.2]
- [Table 1]
- [Section 6.2]
- [Figure 9]
Circularity Check
No significant circularity: TIR is a supervised reflection-selection method with a clean train/test split, and the central empirical claim does not reduce to its inputs.
full rationale
The derivation is self-contained with respect to circularity. TIR's training phase (Sec. 3.3) uses ground-truth labels to select the reflection r_best = r_argmax(acc_k) that maximizes a novice agent's accuracy on each training student; this is a supervised prompt-construction procedure, not a prediction. The test phase (Sec. 3.4) explicitly prevents label leakage: 'we use a new reflective agent ... to retrieve reflections from the successful reflection database' and 'this new reflective agent does not experience any other training data,' and the retrieved reflections are used only as demonstrations to generate a fresh reflection for the test student. Test labels enter only as the evaluation target, so the reported accuracies (Table 1, Fig. 9) are not equivalent to the fitted selection objective. The comparison with deep-learning knowledge-tracing baselines uses the same individual-wise train/test split and the same BERT text embeddings (Secs. 3.7, 6.1), so the 'TIR beats DL' claim is an empirical result rather than a definitional consequence. The only self-citations are to the authors' own public EduAgent dataset [86] and prior generative-agent work [85]; EduAgent is used as a public benchmark alongside the newly collected CogEdu dataset, and no theoretical premise or ansatz is imported from these citations. Concerns about a single split, small test sets (df=17), and unreported retrieval seed/variance are correctness/reproducibility risks, not circularity.
Assumptions & free parameters
free parameters (4)
- M (number of retrieved reflections) =
4
- Number of past questions =
5
- Train/test split ratio =
0.7 for CogEdu, 0.8 for EduAgent
- Maximum TIR iterations =
not specified
assumptions (5)
- domain assumption LLMs can generate useful self-reflections about prediction errors when given ground-truth labels.
- domain assumption Reflections distilled from training students generalize to test students in the same lecture.
- domain assumption Course materials can be represented as text; image content is converted to textual descriptions.
- domain assumption Post-test accuracy is a sufficient proxy for student learning behavior.
- domain assumption The first five questions provide enough history to simulate future performance.
Cite this review
Pith. "Pith review of Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation." pith.science (2026). https://pith.science/paper/E6J62DOR
@misc{pith2026250202780,
author = {Pith},
title = {Pith review of: Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation},
year = {2026},
howpublished = {\url{https://pith.science/paper/E6J62DOR}},
note = {Machine review of arXiv:2502.02780}
}
read the original abstract
Student simulation supports educators to improve teaching by interacting with virtual students. However, most existing approaches ignore the modulation effects of course materials because of two challenges: the lack of datasets with granularly annotated course materials, and the limitation of existing simulation models in processing extremely long textual data. To solve the challenges, we first run a 6-week education workshop from N = 60 students to collect fine-grained data using a custom built online education system, which logs students' learning behaviors as they interact with lecture materials over time. Second, we propose a transferable iterative reflection (TIR) module that augments both prompting-based and finetuning-based large language models (LLMs) for simulating learning behaviors. Our comprehensive experiments show that TIR enables the LLMs to perform more accurate student simulation than classical deep learning models, even with limited demonstration data. Our TIR approach better captures the granular dynamism of learning performance and inter-student correlations in classrooms, paving the way towards a ''digital twin'' for online education.
Figures
Figures from the paper (15 more)
Forward citations
Cited by 1 Pith paper
-
Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment
WikiHowAgent generates 114,296 simulated teacher-learner conversations from 14,287 WikiHow tutorials and evaluates their pedagogic quality with LLM and human judges.
Reference graph
Works this paper leans on
-
[1]
Ghodai Abdelrahman, Qing Wang, and Bernardo Nunes. 2023. Knowledge tracing: A survey. Comput. Surveys 55, 11 (2023), 1–37
2023
-
[2]
Badr AlKhamissi, Millicent Li, Asli Celikyilmaz, Mona Diab, and Marjan Ghazvininejad. 2022. A review on language models as knowledge bases. arXiv preprint arXiv:2204.06031 (2022)
arXiv 2022
-
[3]
Riku Arakawa, Hiromu Yakura, and Masataka Goto. 2023. CatAlyst: domain- extensible intervention for preventing task procrastination using large generative models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–19
2023
-
[4]
Ananya Bhattacharjee, Yuchen Zeng, Sarah Yi Xu, Dana Kulzhabayeva, Minyi Ma, Rachel Kornfield, Syed Ishtiaque Ahmed, Alex Mariakakis, Mary P Czerwinski, Anastasia Kuzminykh, et al. 2024. Understanding the Role of Large Language Models in Personalizing and Scaffolding Strategies to Combat Academic Procras- tination. In Proceedings of the CHI Conference on ...
2024
-
[5]
Paul Calle, Ruosi Shao, Yunlong Liu, Emily T Hébert, Darla Kendzor, Jordan Neil, Michael Businelle, and Chongle Pan. 2024. Towards AI-Driven Healthcare: Systematic Optimization, Linguistic Analysis, and Clinicians’ Evaluation of Large Language Models for Smoking Cessation Interventions. In Proceedings of the CHI Conference on Human Factors in Computing Sy...
2024
-
[6]
Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chen Qian, Chi-Min Chan, Yujia Qin, Yaxi Lu, Ruobing Xie, et al. 2023. Agentverse: Facilitat- ing multi-agent collaboration and exploring emergent behaviors in agents. arXiv preprint arXiv:2308.10848 2, 4 (2023), 6
arXiv 2023
-
[7]
Youngduck Choi, Youngnam Lee, Dongmin Shin, Junghyun Cho, Seoyon Park, Seewoo Lee, Jineon Baek, Chan Bae, Byungsoo Kim, and Jaewe Heo. 2020. Ednet: A large-scale hierarchical dataset in education. In Artificial Intelligence in Edu- cation: 21st International Conference, AIED 2020, Ifrane, Morocco, July 6–10, 2020, Proceedings, Part II 21 . Springer, 69–73
2020
-
[8]
Peng Cui and Mrinmaya Sachan. 2023. Adaptive and personalized exercise generation for online language learning. arXiv preprint arXiv:2306.02457 (2023)
work page Pith review arXiv 2023
Show all 114 references
-
[9]
Yang Deng, An Zhang, Yankai Lin, Xu Chen, Ji-Rong Wen, and Tat-Seng Chua
-
[10]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. CoRR abs/1810.04805 (2018). arXiv:1810.04805 http://arxiv.org/abs/1810.04805
2018 arXiv
-
[11]
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022. A survey on in-context learning. arXiv preprint arXiv:2301.00234 (2022)
2022 arXiv
-
[12]
Peitong Duan, Jeremy Warner, Yang Li, and Bjoern Hartmann. 2024. Generating Automatic Feedback on UI Mockups with Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–20
2024
-
[13]
Ralf Engbert and Reinhold Kliegl. 2003. Microsaccades uncover the orientation of covert attention. Vision research 43, 9 (2003), 1035–1045
2003
-
[14]
Raymond Fok, Nedim Lipka, Tong Sun, and Alexa F Siu. 2024. Marco: Supporting Business Document Workflows via Collection-Centric Information Foraging with Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–20
2024
-
[15]
Lingyue Fu, Hao Guan, Kounianhua Du, Jianghao Lin, Wei Xia, Weinan Zhang, Ruiming Tang, Yasheng Wang, and Yong Yu. 2024. SINKT: A Structure- Aware Inductive Knowledge Tracing Model with Large Language Model. arXiv:2407.01245 [cs.AI] https://arxiv.org/abs/2407.01245
2024 arXiv
-
[16]
Jie Gao, Yuchen Guo, Gionnieve Lim, Tianqin Zhang, Zheng Zhang, Toby Jia- Jun Li, and Simon Tangi Perrault. 2024. CollabCoder: a lower-barrier, rigorous workflow for inductive collaborative qualitative analysis with large language models. In Proceedings of the CHI Conference o...
2024
-
[17]
Simret Araya Gebreegziabher, Zheng Zhang, Xiaohang Tang, Yihao Meng, Elena L Glassman, and Toby Jia-Jun Li. 2023. Patat: Human-ai collaborative qualitative coding with explainable interactive rule synthesis. In Proceedings of the 2023 CHI Conference on Human Factors in Computi...
2023
-
[18]
Aritra Ghosh, Neil Heffernan, and Andrew S. Lan. 2020. Context-Aware Attentive Knowledge Tracing. In Proceedings of the 26th ACM SIGKDD International Con- ference on Knowledge Discovery & Data Mining (Virtual Event, CA, USA) (KDD ’20). Association for Computing Machinery, New ...
2020
-
[19]
Arthur C Graesser, Mark W Conley, and Andrew Olney. 2012. Intelligent tutoring systems. (2012)
2012
-
[20]
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024. MiniLLM: Knowledge Distillation of Large Language Models. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=5h0qf7IBZZ
2024
-
[21]
Xiaopeng Guo, Zhijie Huang, Jie Gao, Mingyu Shang, Maojing Shu, and Jun Sun
-
[22]
Muhammad Usman Hadi, Rizwan Qureshi, Abbas Shah, Muhammad Irfan, Anas Zafar, Muhammad Bilal Shaikh, Naveed Akhtar, Jia Wu, Seyedali Mirjalili, et al
-
[23]
Xinying Hou, Zihan Wu, Xu Wang, and Barbara J Ericson. 2024. Codetailor: Llm-powered personalized parsons puzzles for engaging support while learning programming. In Proceedings of the Eleventh ACM Conference on Learning@ Scale . 51–62
2024
-
[24]
Yihan Hou, Manling Yang, Hao Cui, Lei Wang, Jie Xu, and Wei Zeng. 2024. C2Ideas: Supporting Creative Interior Color Design Ideation with a Large Lan- guage Model. InProceedings of the CHI Conference on Human Factors in Computing Systems. 1–18
2024
-
[25]
Wenyang Hui, Yan Wang, Kewei Tu, and Chengyue Jiang. 2024. RoT: Enhanc- ing Large Language Models with Reflection on Search Trees. arXiv preprint arXiv:2404.05449 (2024)
2024 arXiv
-
[26]
It’s the only thing I can trust
JiWoong Jang, Sanika Moharana, Patrick Carrington, and Andrew Begel. 2024. “It’s the only thing I can trust”: Envisioning Large Language Model Use by Autistic Workers for Communication Assistance. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–18
2024
-
[27]
Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung. 2023. Towards mitigating hallucination in large language models via self-reflection. arXiv preprint arXiv:2310.06271 (2023)
2023 arXiv
-
[28]
Hyoungwook Jin, Seonghee Lee, Hyungyu Shin, and Juho Kim. 2024. Teach AI How to Code: Using Large Language Models as Teachable Agents for Pro- gramming Education. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–28
2024
-
[29]
Eunkyung Jo, Daniel A Epstein, Hyunhoon Jung, and Young-Ho Kim. 2023. Under- standing the benefits and challenges of deploying conversational AI leveraging large language models for public health intervention. In Proceedings of the 2023 CHI Conference on Human Factors in Compu...
2023
-
[30]
Eunkyung Jo, Yuin Jeong, SoHyun Park, Daniel A Epstein, and Young-Ho Kim
-
[31]
Heeseok Jung, Jaesang Yoo, Yohaan Yoon, and Yeonju Jang. 2024. CLST: Cold- Start Mitigation in Knowledge Tracing by Aligning a Generative Language Model as a Students’ Knowledge Tracer. arXiv preprint arXiv:2406.10296 (2024)
2024 arXiv
-
[32]
Minsol Kim, Aliea L Nallbani, and Abby Rayne Stovall. 2024. Exploring LLM- based Chatbot for Language Learning and Cultivation of Growth Mindset. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 1–5
2024
-
[33]
Taewan Kim, Seolyeong Bae, Hyun Ah Kim, Su-woo Lee, Hwajung Hong, Chanmo Yang, and Young-Ho Kim. 2024. MindfulDiary: Harnessing Large Language Model to Support Psychiatric Patients’ Journaling. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–20
2024
-
[34]
In Pro- ceedings of the CHI Conference on Human Factors in Computing Systems
Understanding the Impact of Long-Term Memory on Self-Disclosure with Large Language Model-Driven Chatbots for Public Health Intervention. In Pro- ceedings of the CHI Conference on Human Factors in Computing Systems . 1–21
-
[35]
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning. nature 521, 7553 (2015), 436–444
2015
-
[36]
Unggi Lee, Jiyeong Bae, Dohee Kim, Sookbun Lee, Jaekwon Park, Taekyung Ahn, Gunho Lee, Damji Stratton, and Hyeoncheol Kim. 2024. Language Model Can Do Knowledge Tracing: Simple but Effective Method to Integrate Language Model and Knowledge Tracing Task. arXiv preprint arXiv:24...
2024 arXiv
-
[37]
Unggi Lee, Yonghyun Park, Yujin Kim, Seongyune Choi, and Hyeoncheol Kim
-
[38]
Harsh Kumar, Ruiwei Xiao, Benjamin Lawson, Ilya Musabirov, Jiakai Shi, Xinyuan Wang, Huayin Luo, Joseph Jay Williams, Anna N Rafferty, John Stamper, et al
-
[39]
InProceedings of the Eleventh ACM Conference on Learning@ Scale
Supporting Self-Reflection at Scale with Large Language Models: Insights from Randomized Field Experiments in Classrooms. InProceedings of the Eleventh ACM Conference on Learning@ Scale . 86–97
-
[40]
Qingyao Li, Wei Xia, Kounianhua Du, Qiji Zhang, Weinan Zhang, Ruiming Tang, and Yong Yu. 2024. Learning Structure and Knowledge Aware Representation with Large Language Models for Concept Recommendation. arXiv preprint arXiv:2405.12442 (2024)
2024 arXiv
-
[41]
Tianle Li, Ge Zhang, Quy Duc Do, Xiang Yue, and Wenhu Chen. 2024. Long-context llms struggle with long in-context learning. arXiv preprint arXiv:2404.02060 (2024)
2024 arXiv
-
[42]
Zhaoxing Li, Jujie Yang, Jindi Wang, Lei Shi, and Sebastian Stein. 2024. Integrating lstm and bert for long-sequence data analysis in intelligent tutoring systems. arXiv preprint arXiv:2405.05136 (2024)
2024 arXiv
-
[43]
In International Conference on Intelligent Tutoring Systems
Monacobert: Monotonic attention based convbert for knowledge tracing. In International Conference on Intelligent Tutoring Systems . Springer, 107–123
-
[44]
Haoxuan Li, Jifan Yu, Yuanxin Ouyang, Zhuang Liu, Wenge Rong, Juanzi Li, and Zhang Xiong. 2024. Explainable few-shot knowledge tracing. arXiv preprint arXiv:2405.14391 (2024)
2024 arXiv
-
[45]
Moxin Li, Wenjie Wang, Fuli Feng, Fengbin Zhu, Qifan Wang, and Tat-Seng Chua. 2024. Think twice before assure: Confidence estimation for large language models through reflection on multiple answers. arXiv preprint arXiv:2403.09972 (2024)
2024 arXiv
-
[46]
Zitao Liu, Qiongqiong Liu, Jiahao Chen, Shuyan Huang, and Weiqi Luo. 2023. simpleKT: A Simple But Tough-to-Beat Baseline for Knowledge Tracing. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net. http...
2023
-
[47]
Zitao Liu, Qiongqiong Liu, Jiahao Chen, Shuyan Huang, Jiliang Tang, and Weiqi Luo. 2022. pyKT: A Python Library to Benchmark Deep Learning based Knowledge Tracing Models. In Advances in Neural Infor- mation Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Bel- grave,...
2022
-
[48]
Helen E Longino. 2019. Studying human behavior: How scientists investigate aggression and sexuality. University of Chicago Press
2019
-
[49]
Zhenwen Liang, Wenhao Yu, Tanmay Rajpurohit, Peter Clark, Xiangliang Zhang, and Ashwin Kaylan. 2023. Let gpt be a math tutor: Teaching math word problem solvers with customized exercise generation. arXiv preprint arXiv:2305.14386 (2023)
2023 arXiv
-
[50]
What it wants me to say
Michael Xieyang Liu, Advait Sarkar, Carina Negreanu, Benjamin Zorn, Jack Williams, Neil Toronto, and Andrew D Gordon. 2023. “What it wants me to say”: Bridging the abstraction gap between end-user programmers and code- generating large language models. In Proceedings of the 20...
2023
-
[51]
Naiming Liu, Zichao Wang, Richard G Baraniuk, and Andrew Lan. 2022. GPT- based Open-ended Knowledge Tracing. arXiv preprint arXiv:2203.03716 (2022)
2022 arXiv
-
[52]
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al
-
[53]
Julia M Markel, Steven G Opferman, James A Landay, and Chris Piech. 2023. Gpteach: Interactive ta training with gpt-based students. In Proceedings of the tenth acm conference on learning@ scale . 226–236
2023
-
[54]
Jordan K Matelsky, Felipe Parodi, Tony Liu, Richard D Lange, and Konrad P Ko- rding. 2023. A large language model-assisted education tool to provide feedback on open-ended responses. arXiv preprint arXiv:2308.02439 (2023)
2023 arXiv
-
[55]
Xinyi Lu, Simin Fan, Jessica Houghton, Lu Wang, and Xu Wang. 2023. Read- ingQuizMaker: a human-NLP collaborative system that supports instructors to design high-quality reading quiz questions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–18
2023
-
[56]
Xinyi Lu and Xu Wang. 2024. Generative Students: Using LLM-Simulated Student Profiles to Support Question Item Evaluation. In Proceedings of the Eleventh ACM Conference on Learning @ Scale (L@S ’24, Vol. 8). ACM, 16–27. doi:10.1145/3657604. 3662031
2024 doi
-
[57]
Zilin Ma, Yiyang Mei, Yinru Long, Zhaoyuan Su, and Krzysztof Z Gajos. 2024. Evaluating the Experience of LGBTQ+ People Using Large Language Model Based Chatbots for Mental Health Support. In Proceedings of the CHI Conference Classroom Simulacra: Building Contextual Student Gen...
2024
-
[58]
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology . 1–22
2023
-
[59]
Advances in Neural Information Processing Systems 36 (2024)
Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[60]
Chris Piech, Jonathan Spencer, Jonathan Huang, Surya Ganguli, Mehran Sahami, Leonidas Guibas, and Jascha Sohl-Dickstein. 2015. Deep Knowledge Tracing. arXiv:1506.05908 [cs.AI] https://arxiv.org/abs/1506.05908
2015 arXiv
-
[61]
Chen Pojen, Hsieh Mingen, and Tsai Tzuyang. 2020. Junyi Academy On- line Learning Activity Dataset: A large-scale public online learning activity dataset from elementary to senior high school students. Dataset available from https://www.kaggle.com/junyiacademy/learning-activit...
2020
-
[62]
Hiromi Nakagawa, Yusuke Iwasawa, and Yutaka Matsuo. 2019. Graph-based knowledge tracing: modeling student proficiency using graph neural network. In IEEE/WIC/ACM International Conference on Web Intelligence . 156–163
2019
-
[63]
Seyed Parsa Neshaei, Richard Lee Davis, Adam Hazimeh, Bojan Lazarevski, Pierre Dillenbourg, and Tanja Käser. 2024. Towards Modeling Learner Performance with Large Language Models. arXiv:2403.14661 [cs.CY] https://arxiv.org/abs/ 2403.14661
2024 arXiv
-
[64]
Tuan Nguyen. 2015. The effectiveness of online learning: Beyond no significant difference and future horizons. MERLOT Journal of online learning and teaching 11, 2 (2015), 309–319
2015
-
[65]
Fatemeh Sarshartehrani, Elham Mohammadrezaei, Majid Behravan, and Denis Gracanin. 2024. Enhancing E-Learning Experience Through Embodied AI Tutors in Immersive Virtual Environments: A Multifaceted Approach for Personalized Educational Adaptation. In International Conference on...
2024
-
[66]
Savvas Petridis, Nicholas Diakopoulos, Kevin Crowston, Mark Hansen, Keren Henderson, Stan Jastrzebski, Jeffrey V Nickerson, and Lydia B Chilton. 2023. Anglekindling: Supporting journalistic angle ideation with large language models. In Proceedings of the 2023 CHI conference on...
2023
-
[67]
Woosuk Seo, Chanmo Yang, and Young-Ho Kim. 2024. ChaCha: Leveraging Large Language Models to Prompt Children to Share Their Emotions about Personal Events. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–20
2024
-
[68]
Konstantin R Strömel, Stanislas Henry, Tim Johansson, Jasmin Niess, and Paweł W Woźniak. 2024. Narrating Fitness: Leveraging Large Language Models for Reflec- tive Fitness Tracker Data Interpretation. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–16
2024
-
[69]
Hua Xuan Qin, Shan Jin, Ze Gao, Mingming Fan, and Pan Hui. 2024. Char- acterMeet: Supporting Creative Writers’ Entire Story Character Construction Processes Through Conversation with LLM-Powered Chatbot Avatars. In Pro- ceedings of the CHI Conference on Human Factors in Comput...
2024
-
[70]
Niroop Channa Rajashekar, Yeo Eun Shin, Yuan Pu, Sunny Chung, Kisung You, Mauro Giuffre, Colleen E Chan, Theo Saarinen, Allen Hsiao, Jasjeet Sekhon, et al
-
[71]
In Proceedings of the CHI Conference on Human Factors in Computing Systems
Human-Algorithmic Interaction Using a Large Language Model-Augmented Artificial Intelligence Clinical Decision Support System. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–20
-
[72]
Mohi Reza, Nathan M Laundry, Ilya Musabirov, Peter Dushniku, Zhi Yuan “Michael” Yu, Kashish Mittal, Tovi Grossman, Michael Liut, Anastasia Kuzminykh, and Joseph Jay Williams. 2024. ABScribe: Rapid Exploration & Or- ganization of Multiple Writing Variations in Human-AI Co-Writi...
2024
-
[73]
Bryan Wang, Gang Li, and Yang Li. 2023. Enabling conversational interaction with mobile ui using large language models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–17
2023
-
[74]
Frances Scholtz and Suzaan Hughes. 2021. A systematic review of educator interventions in facilitating simulation based learning.Journal of Applied Research in Higher Education 13, 5 (2021), 1408–1435
2021
-
[75]
Tianjia Wang, Ramaraja Ramanujan, Yi Lu, Chenyu Mao, Yan Chen, and Chris Brown. 2024. DevCoach: Supporting Students in Learning the Software Develop- ment Life Cycle at Scale with Generative Agents. In Proceedings of the Eleventh ACM Conference on Learning@ Scale . 351–355
2024
-
[76]
Wei Wang, Lihuan Guo, Ling He, and Yenchun Jim Wu. 2019. Effects of social- interactive engagement on the dropout ratio in online learning: insights from MOOC. Behaviour & Information Technology 38, 6 (2019), 621–636
2019
-
[77]
Sangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li, and Haijun Xia. 2024. Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation. InProceedings of the CHI Conference on Human Factors in Computing Systems . 1–26
2024
-
[78]
Xiangru Tang, Yiming Zong, Yilun Zhao, Arman Cohan, and Mark Gerstein. 2023. Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data? arXiv preprint arXiv:2309.08963 (2023)
2023 arXiv
-
[79]
Annapurna Vadaparty, Daniel Zingaro, David H Smith IV, Mounika Padala, Chris- tine Alvarado, Jamie Gorson Benario, and Leo Porter. 2024. CS1-LLM: Integrating LLMs into CS1 Instruction. InProceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1 . 297–303
2024
-
[80]
Hongyu Wan, Jinda Zhang, Abdulaziz Arif Suria, Bingsheng Yao, Dakuo Wang, Yvonne Coady, and Mirjana Prpa. 2024. Building LLM-based AI Agents in Social Virtual Reality. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–7
2024
-
[81]
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al. 2023. Larger language models do in-context learning differently. arXiv preprint arXiv:2303.03846 (2023)
2023 arXiv
-
[82]
Sitong Wang, Savvas Petridis, Taeahn Kwon, Xiaojuan Ma, and Lydia B Chilton
-
[83]
In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems
PopBlends: Strategies for conceptual blending with large language models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–19
2023
-
[84]
Wanli Xing and Dongping Du. 2019. Dropout prediction in MOOCs: Using deep learning for personalized intervention. Journal of Educational Computing Research 57, 3 (2019), 547–570
2019
-
[85]
Songlin Xu and Xinyu Zhang. 2023. Leveraging generative artificial intelligence to simulate student learning behavior. arXiv preprint arXiv:2310.19206 (2023)
2023 arXiv
-
[86]
Xinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra, and Zhengjie Miao
-
[87]
In Proceedings of the CHI Conference on Human Factors in Computing Systems
Human-LLM collaborative annotation through effective verification of LLM labels. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–21
-
[88]
Yutong Wang, Jiali Zeng, Xuebo Liu, Fandong Meng, Jie Zhou, and Min Zhang
-
[89]
arXiv preprint arXiv:2406.08434 (2024)
TasTe: Teaching Large Language Models to Translate through Self- Reflection. arXiv preprint arXiv:2406.08434 (2024)
2024 arXiv
-
[90]
Zhan Wang, Lin-Ping Yuan, Liangwei Wang, Bingchuan Jiang, and Wei Zeng
-
[91]
In Proceedings of the CHI conference on human factors in computing systems
Virtuwander: Enhancing multi-modal interaction for virtual tour guidance through large language models. In Proceedings of the CHI conference on human factors in computing systems . 1–20
-
[92]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903 [cs.CL] https: //arxiv.org/abs/2201.11903
2023 arXiv
-
[93]
Chao Zhang, Xuechen Liu, Katherine Ziska, Soobin Jeon, Chi-Lin Yu, and Ying Xu. 2024. Mathemyths: leveraging large language models to teach mathematical language through Child-AI co-creative storytelling. In Proceedings of the CHI Conference on Human Factors in Computing Syste...
2024
-
[94]
Ruolan Wu, Chun Yu, Xiaole Pan, Yujia Liu, Ningning Zhang, Yue Fu, Yuhan Wang, Zhi Zheng, Li Chen, Qiaolei Jiang, et al . 2024. MindShift: Leveraging Large Language Models for Mental-States-Based Problematic Smartphone Use Intervention. In Proceedings of the CHI Conference on ...
2024
-
[95]
Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–22
2022
-
[96]
Zheyuan Zhang, Daniel Zhang-Li, Jifan Yu, Linlu Gong, Jinchang Zhou, Zhiyuan Liu, Lei Hou, and Juanzi Li. 2024. Simulating classroom education with llm- empowered agents. arXiv preprint arXiv:2406.19226 (2024)
2024 arXiv
-
[97]
Wu, and Álvaro Tejero-Cantero
Hanqi Zhou, Robert Bamler, Charley M. Wu, and Álvaro Tejero-Cantero. 2024. Predictive, scalable and interpretable knowledge tracing on structured domains. arXiv:2403.13179 [cs.LG] https://arxiv.org/abs/2403.13179
2024 arXiv
-
[98]
Songlin Xu, Xinyu Zhang, and Lianhui Qin. 2024. EduAgent: Generative Student Agents in Learning. arXiv preprint arXiv:2404.07963 (2024)
2024 arXiv
-
[99]
Xiaohan Xu, Ming Li, Chongyang Tao, Tao Shen, Reynold Cheng, Jinyang Li, Can Xu, Dacheng Tao, and Tianyi Zhou. 2024. A Survey on Knowledge Distillation of Large Language Models.ArXiv abs/2402.13116 (2024). https://api.semanticscholar. org/CorpusID:267760021
2024 arXiv
-
[100]
Hanqi Yan, Qinglin Zhu, Xinyu Wang, Lin Gui, and Yulan He. 2024. Mirror: A Multiple-perspective Self-Reflection Method for Knowledge-rich Reasoning. arXiv preprint arXiv:2402.14963 (2024)
2024 arXiv
-
[101]
Jackie Yang, Yingtian Shi, Yuhan Zhang, Karina Li, Daniel Wan Rosli, Anisha Jain, Shuning Zhang, Tianshi Li, James A Landay, and Monica S Lam. 2024. ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models. In Proceedings of the CHI C...
2024
-
[102]
Yang Yu, Yingbo Zhou, Yaokang Zhu, Yutong Ye, Liangyu Chen, and Mingsong Chen. 2024. ECKT: Enhancing Code Knowledge Tracing via Large Language Models. In Proceedings of the Annual Meeting of the Cognitive Science Society , Vol. 46
2024
-
[103]
Michael V Yudelson, Kenneth R Koedinger, and Geoffrey J Gordon. 2013. Individ- ualized bayesian knowledge tracing models. In Artificial Intelligence in Education: 16th International Conference, AIED 2013, Memphis, TN, USA, July 9-13, 2013. Proceedings 16. Springer, 171–180
2013
-
[104]
Murong Yue, Wijdane Mifdal, Yixuan Zhang, Jennifer Suh, and Ziyu Yao. 2024. MathVC: An LLM-Simulated Multi-Character Virtual Classroom for Mathematics Education. arXiv:2404.06711 [cs.CL] https://arxiv.org/abs/2404.06711 CHI ’25, April 26-May 1, 2025, Yokohama, Japan Xu, et al
2024
-
[106]
Jiani Zhang, Xingjian Shi, Irwin King, and Dit-Yan Yeung. 2017. Dynamic Key- Value Memory Networks for Knowledge Tracing. arXiv:1611.08108 [cs.AI] https: //arxiv.org/abs/1611.08108
2017 arXiv
-
[107]
Liang Zhang, Jionghao Lin, Conrad Borchers, John Sabatini, John Hollander, Meng Cao, and Xiangen Hu. 2024. Predicting Learning Performance with Large Language Models: A Study in Adult Literacy. In International Conference on Human-Computer Interaction. Springer, 333–353
2024
-
[110]
Wazeer Deen Zulfikar, Samantha Chan, and Pattie Maes. 2024. Memoro: Using Large Language Models to Realize a Concise Interface for Real-Time Memory Augmentation. In Proceedings of the CHI Conference on Human Factors in Com- puting Systems. 1–18. A Appendix A.1 Deep Learning Ba...
2024
-
[113]
Predict the student's performance: - Based on the analysis from steps 1 and 2, predict whether the student will answer each upcoming question correctly or incorrectly. CoT prompt LLM: Question 6: Correct, Reason: The student has previously answered questions related to the thr...
-
[114]
- Consider the complexity of the questions and the related lecture materials
Analyze the student's past performance: - Identify the topics the student has performed well in and those they have struggled with. - Consider the complexity of the questions and the related lecture materials
-
[115]
- Identify if the upcoming questions are related to the topics the student has previously struggled with or excelled in
Review the topics and related lecture materials of the upcoming questions: - Determine the difficulty level of the upcoming questions based on their topics and related lecture materials. - Identify if the upcoming questions are related to the topics the student has previously ...
-
[116]
Input" while the output of BertKT is denoted as
Predict the student's performance: - Based on the analysis from steps 1 and 2, predict whether the student will answer each upcoming question correctly or incorrectly. CoT prompt with TIR LLM: Question 6: Correct, Reason: The student has previously answered questions related t...
2025
-
[2021]
In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21)
Enhancing Knowledge Tracing via Adversarial Training. In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21). Association for Computing Machinery, New York, NY, USA, 367–375. doi:10.1145/3474085.3475554
-
[2023]
Authorea Preprints (2023)
A survey on large language models: Applications, challenges, limitations, and practical usage. Authorea Preprints (2023)
2023
-
[2024]
In Companion Proceed- ings of the ACM on Web Conference 2024
Large Language Model Powered Agents in the Web. In Companion Proceed- ings of the ACM on Web Conference 2024 . 1242–1245. CHI ’25, April 26-May 1, 2025, Yokohama, Japan Xu, et al
2024
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.