REVIEW 5 major objections 4 minor 4 cited by
Evaluating Gemini in an arena for learning
T0 review · 5 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Blind expert comparisons rank Gemini 2.5 Pro as the top AI tutor among five models.
desk verdict Two-stage arena for learning is a genuine methodological step forward, but the 'leading model' claim overreaches: the instrument is developer-authored, and the evaluated model serves as its own judge in two places. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the two-stage learning arena: first, 189 educators role-play learners from 49 realistic scenarios in blind, multi-turn conversations with pairs of models, yielding 1,333 head-to-head match-ups; second, an independent pool of 206 experts reviews the transcripts, scores them on a 25-item pedagogy rubric, and gives pairwise preferences that are converted into Elo ratings via a Bradley-Terry model. The design lets each model steer the conversation independently, which is what allows it to assess overall tutoring quality rather than single-turn accuracy.
What would settle it
Re-run the same 1,333 match-ups with Gemini 2.5 Pro, ChatGPT-4o, and GPT-4o served from current production APIs with identical thinking settings and compute; if expert preference for Gemini over ChatGPT-4o falls from 61% toward chance, the reported ranking is an artifact of model version or serving infrastructure.
Extended reading notes
Core claim
The paper claims that expert blind preference over multi-turn tutoring conversations places Gemini 2.5 Pro first among the five models evaluated, and that this preference reflects a genuine difference in pedagogical behavior: Gemini scaffolds tasks, resists giving answers away, keeps conversations aimed at the learner's goal, and adapts to the learner, whereas its competitors more often supply solutions and drift off-topic. The authors describe this as establishing Gemini 2.5 Pro as 'a leading model for learning,' with markedly higher performance across key principles of good pedagogy. The paper also reports a notable split: when educators role-played as students, Gemini 2.5 Pro and ChatGPT-4o tied for first on supporting learning goals, but when independent experts assessed the same transcripts, preference shifted decisively to Gemini 2.5 Pro, which the authors attribute to the gap between what feels immediately helpful to a student and what is pedagogically sound.
Load-bearing premise
The comparison is fair across models even though Gemini 2.5 Pro was served on experimental non-production infrastructure, the GPT-4o snapshot was from 2024, and the ChatGPT-4o version is undated.
Editorial extensions
If this is right
- If expert preference tracks pedagogical quality, then multi-turn arena rankings can serve as a general benchmark for AI tutoring that single-turn exam-style evaluations cannot provide.
- Models that give answers away readily will rank below models that scaffold, even when both produce correct content.
- The gap between learners' first-stage preferences, where ChatGPT-4o tied Gemini, and experts' second-stage judgments suggests that user satisfaction alone is an unreliable signal of tutoring quality.
- Targeted tasks in text re-levelling, short-answer grading, and mistake identification corroborate the arena ranking and can be used to diagnose specific pedagogical skills.
Reading between the lines
- A standardized expert-judged learning arena could push model developers to optimize for scaffolding and learning outcomes rather than for user engagement, potentially reversing the current incentive toward answer-bot behavior.
- The arena's 49 scenarios sample a single snapshot of a longer learning journey; extending the design to longitudinal multi-session tutoring could reveal whether expert-preferred pedagogy actually produces better learning outcomes, which the paper flags as an open question.
- If the result is replicated with matched current model snapshots served on production infrastructure, the 73.2% preference figure would be evidence that pedagogical fine-tuning changes tutoring behavior rather than merely that a newer model is stronger.
- The observed divergence between what students say they like and what experts judge as good teaching connects to the paper's cited finding that felt learning and actual learning can differ; future arena designs could add objective learning-gain measures alongside expert judgment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a "learning arena" in which 189 educators role-played learners across 49 scenarios and 206 pedagogy experts subsequently judged 1333 blind head-to-head multi-turn tutoring conversations. The authors report that Gemini 2.5 Pro ranks first on a Bradley-Terry Elo model of expert preference, is preferred in 73.2% of non-tied match-ups, and scores highest on all five dimensions of a 25-item pedagogy rubric. The paper supplements these results with targeted evaluations of text re-levelling, short-answer assessment, and mistake identification. The central claim is that Gemini 2.5 Pro is a leading model for learning.
Significance. If the results hold, the arena would be a valuable complement to existing single-turn and narrow-task benchmarks for AI in education, and the use of role-playing educators, multi-turn interactions, and blind expert raters is a genuinely useful design choice. The paper's strengths include a substantial participant pool, realistic scenarios, and targeted evaluations on real-world data (Ghanaian student responses; Khan Academy's math-tutoring benchmark). However, the evaluation instrument is developer-authored, the model comparison is not contemporaneous, the qualitative corroboration is self-referential, and no data or code are released; these issues materially weaken the general "leading model for learning" claim as currently stated.
major comments (5)
- [§5.1, Models] The model comparison is not contemporaneous. Gemini 2.5 Pro (2025-05-06) is compared with GPT-4o (2024-08-06), an approximately nine-month-old snapshot, ChatGPT-4o has no snapshot date, and Gemini 2.5 Pro was served on non-production infrastructure. Because frontier model capabilities change quickly, the reported win rates (71.3%, 81.8%, 74.2%, 61.0%) cannot be attributed to intrinsic tutoring quality rather than model age or serving conditions. The authors should either retest with contemporaneous production snapshots or explicitly restrict the conclusions to the specific snapshots and infrastructure used.
- [§5.1, Use case coverage and Performance measurement] The evaluation instrument is developer-aligned. The 49-scenario bank includes "archetypal use cases identified by the Gemini product team," and the 25-item pedagogy rubric is "first developed in [10]," the LearnLM technical report, by the same group that integrates LearnLM capabilities into Gemini. The exact system instructions given to all models are not reported. If these prompts and rubric items encode LearnLM's Socratic, scaffolded, anti-answer-giving style, the head-to-head competition measures conformance to a developer-authored specification, not general pedagogical quality. The expert raters are blind and sincere, but the measurement remains potentially circular. The paper should release the full scenario bank and system prompts, detail the rubric's external validation, and include a robustness check with independently authored prompts and rubric items.
- [Appendix B] Appendix B uses Gemini 2.5 Pro (2025-05-06) to summarize qualitative feedback about all five models, including itself. This self-referential step means the qualitative corroboration in §2.1 cannot independently validate the arena results, because the summarizer is the model under evaluation. The authors should replace this with human-coded thematic analysis (with inter-coder agreement) or an independent summarization model.
- [§5.4] The mistake-identification evaluation applies Gemini 2.5 Pro as the classifier for determining correctness of all models' responses, including Gemini's own. This introduces a direct conflict: the target model sets the ground truth for its own evaluation. The paper should either keep the original GPT-4 Turbo classifier or use an independent classifier, and report agreement between classifiers.
- [Statistics and reproducibility] No data or code are released, no inter-rater reliability statistics (e.g., Cohen's kappa, intraclass correlation) are reported for the 206 expert raters, and Tables 2–3 give point estimates without confidence intervals. For example, the short-answer assessment reports 84.1% for both Gemini 2.5 Pro and ChatGPT-4o, and the mistake-identification gap of 1.6 percentage points may be within noise. The authors should provide data, code, rater-agreement measures, and uncertainty intervals for all targeted results.
minor comments (4)
- [Abstract and §2.1] The abstract reports "73.2%" excluding ties, but §2.1 gives per-model win rates (71.3%, 81.8%, 74.2%, 61.0%) without explaining how the aggregate is computed; please clarify the weighting.
- [§5.2] The text re-levelling prompt is rendered with excessive spacing between characters (e.g., "S im pl ify the f o l l o w i n g text ..."), which appears to be a formatting artifact; please fix the typesetting.
- [Table 3] The ChatGPT-4o row shows "79.6 %" with a stray space, and the Rank column is redundant with the row ordering; please clean up the table formatting.
- [General] The paper would benefit from an explicit data-availability statement; the reader is only pointed to a short project URL (goo.gle/LearnLM-May25), not to the actual evaluation materials or anonymous data.
Circularity Check
Mistake-identification benchmark is self-judged by the evaluated model, and Appendix B corroboration is self-generated; the main arena preference claim is not circular.
-
self definitional
[Section 5.4, Mistake identification]
"While the original framework employed GPT-4 Turbo (2024-04-09) to determine whether the tutor model correctly identified student mistakes or accepted accurate work, we updated our methodology to apply Gemini 2.5 Pro (2025-05-06) with dynamic thinking for this classification."
Table 3 reports accuracy scores for all models on the Khan Academy mistake-identification benchmark, including Gemini 2.5 Pro at 87.4%. The correctness labels used to compute those scores are generated by Gemini 2.5 Pro itself, because the paper replaced the original GPT-4 Turbo classifier with Gemini 2.5 Pro. For the Gemini 2.5 Pro row, the model is both the examined tutor and the judge of whether its own responses correctly identified mistakes. The reported accuracy is therefore not an externally grounded measurement but Gemini's self-assessment, making the head-to-head comparison with Claude 3.7 Sonnet, OpenAI o3, ChatGPT-4o, and GPT-4o non-independent by construction.
-
other
[Appendix B, Qualitative feedback (cross-referenced from Section 2.1)]
"For each qualitative dataset, we provided Gemini 2.5 Pro (2025-05-06) with a prompt offering brief context about the feedback and encouraging transparent thematic analysis (rather than artificial balancing of positive and negative points), along with the corresponding dataset."
Section 2.1 states that 'A broader analysis of qualitative feedback from the arena validates these patterns' and directs readers to Appendix B for 'an automated analysis of all qualitative feedback.' The automated analysis is performed by Gemini 2.5 Pro, the same model whose superiority the feedback is said to validate. The Appendix B summaries are therefore self-generated corroboration: the evaluated system is asked to identify themes in comments about itself and its competitors. This does not change the raw expert preference data, but it means the qualitative 'validation' is not independent of the model being promoted.
full rationale
The central arena claim (Gemini 2.5 Pro first by blind expert preference) is not circular: the head-to-head preferences come from 206 external educators and pedagogy experts rating 1333 blind match-ups, and the Elo/preference analysis is a standard Bradley-Terry model applied to those external judgments. The pedagogy rubric is credited to the authors' prior LearnLM report [10], but the paper also states it was developed in consultation with external pedagogy specialists and based on learning-science research [43-48], so the instrument is not defined in terms of the target model's outputs. The scenario bank's provenance from the Gemini product team is a validity/conflict-of-interest concern, not a derivation-level circularity, especially because all models received identical system instructions and the outcome variable is expert preference. Two self-referential steps do undermine specific secondary results. First, Section 5.4 replaces the benchmark's original GPT-4 Turbo classifier with Gemini 2.5 Pro, so the Table 3 accuracy for Gemini 2.5 Pro (87.4%) is computed by Gemini judging its own tutoring responses; this is a by-construction self-evaluation and makes the cross-model comparison non-independent. Second, Appendix B uses Gemini 2.5 Pro to summarize the qualitative feedback that the paper presents as validating Gemini's superiority, so the qualitative corroboration is self-generated. These issues make the targeted corroborations partially circular, but they do not reduce the main arena preference result to its inputs.
Assumptions & free parameters
free parameters (1)
- Claude 3.7 Sonnet thinking budget =
2000 tokens
assumptions (5)
- domain assumption Expert judgments of tutoring quality from transcripts are a valid proxy for learning effectiveness.
- domain assumption Educators role-playing students produce interactions representative of real learner-tutor conversations.
- standard math Bradley-Terry and Elo models are appropriate for ranking from incomplete pairwise preference data.
- domain assumption Flesch-Kincaid grade level and textual entailment accurately measure text re-levelling quality.
- ad hoc to paper Gemini 2.5 Pro is a reliable and unbiased summarizer of qualitative feedback about all models, including itself.
Cite this review
Pith. "Pith review of Evaluating Gemini in an arena for learning." pith.science (2026). https://pith.science/paper/W7TVENUX
@misc{pith2026250524477,
author = {Pith},
title = {Pith review of: Evaluating Gemini in an arena for learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/W7TVENUX}},
note = {Machine review of arXiv:2505.24477}
}
abstract
Artificial intelligence (AI) is poised to transform education, but the research community lacks a robust, general benchmark to evaluate AI models for learning. To assess state-of-the-art support for educational use cases, we ran an "arena for learning" where educators and pedagogy experts conduct blind, head-to-head, multi-turn comparisons of leading AI models. In particular, $N = 189$ educators drew from their experience to role-play realistic learning use cases, interacting with two models sequentially, after which $N = 206$ experts judged which model better supported the user's learning goals. The arena evaluated a slate of state-of-the-art models: Gemini 2.5 Pro, Claude 3.7 Sonnet, GPT-4o, and OpenAI o3. Excluding ties, experts preferred Gemini 2.5 Pro in 73.2% of these match-ups -- ranking it first overall in the arena. Gemini 2.5 Pro also demonstrated markedly higher performance across key principles of good pedagogy. Altogether, these results position Gemini 2.5 Pro as a leading model for learning.
Forward citations
Cited by 4 Pith papers
-
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
A 30-day, knowledge-tracing-grounded simulated learner benchmark for tutoring agents finds that no base model or harness alone determines quality and that almost all tested combinations plateau within days.
-
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
3D-R1 uses a synthetic chain-of-thought cold start plus GRPO reinforcement learning with perception, semantic, and format rewards, and reports best published results across 3D dense captioning, QA, grounding, dialogue...
-
Benchmarking the Pedagogical Knowledge of Large Language Models
The authors release an open benchmark of 1,143 pedagogical knowledge questions from Chilean teacher exams and report accuracy, cost, and size trade-offs for 97 large language models.
-
Nav-R1: Reasoning and Navigation in Embodied Scenes
Nav-R1 uses a 110K synthetic CoT dataset, GRPO with three rewards, and a fast-in-slow system to set new SOTA on R2R-CE, RxR-CE, and HM3D-OVON.
Reference graph
Works this paper leans on
-
[10]
LearnLM: Improving Gemini for learning.arXiv preprint arXiv:2412.16429, 2024
LearnLM Team, Abhinit Modi, Aditya Srikanth Veerubhotla, Aliya Rysbek, Andrea Huber, Brett Wiltshire, Brian Veprek, Daniel Gillick, Daniel Kasenberg, Derek Ahmed, et al. LearnLM: Improving Gemini for learning.arXiv preprint arXiv:2412.16429, 2024
arXiv 2024
-
[1]
Olivia Sidoti, Eugenie Park, and Jeffrey Gottfried. About a quarter of U.S. teens have used ChatGPT for schoolwork – double the share in 2023.https://web.archive.org/web/ 20250430224239/https://www.pewresearch.org/short-reads/2025/01/15/abou t-a-quarter-of-us-teens-have-used-chatgpt-for-schoolwork-double-the-s hare-in-2023/, 2025. Accessed 15 May, 2025
work page 2023
-
[3]
Artificially intelligent? Children’s and parents’ views on generative AI in education
Ali Bissoondath. Artificially intelligent? Children’s and parents’ views on generative AI in education. https://web.archive.org/web/20241119022121/https://www.inte rnetmatters.org/hub/press-release/ai-research-warns-schools-unprepare d-artificial-intelligence/, 2024. Accessed 15 May, 2025
-
[4]
Josh Freeman. Student generative AI survey 2025.https://web.archive.org/web/2025 0513133314/https://www.hepi.ac.uk/2025/02/26/student-generative-ai-sur vey-2025/, 2025. Accessed 15 May, 2025
work page 2025
-
[5]
Salman Khan.Brave new words: How AI will revolutionize education (and why that’s a good thing). Penguin, 2024
work page 2024
-
[6]
Ethan Mollick.Co-intelligence: Living and working with AI. Penguin, 2024
work page 2024
-
[7]
Erin Mote. Artificial intelligence in education: Opportunities, challenges, and policy considera- tions for Congress, 2025
work page 2025
-
[8]
Higher education leaders navigate AI disruption
American Association of Colleges and Universities. Higher education leaders navigate AI disruption. https://web.archive.org/web/20250405093427/https://www.aacu .org/newsroom/higher-education-leaders-navigate-ai-disruption , 2025. Accessed 15 May, 2025. 11 Evaluating Gemini in an Arena for Learning
Show all 51 references
-
[9]
McKee, Daniel Gillick, Shaojian Zhu, Sara Wiltberger, Shubham Milind Phal, Katherine Hermann, Daniel Kasenberg, Avishkar Bhoopchand, et al
Irina Jurenka, Markus Kunesch, Kevin R. McKee, Daniel Gillick, Shaojian Zhu, Sara Wiltberger, Shubham Milind Phal, Katherine Hermann, Daniel Kasenberg, Avishkar Bhoopchand, et al. Towards responsible development of generative AI for education: An evaluation-driven approach. ar...
2024
-
[11]
LearnLM Team. LearnLM. https://ai.google.dev/gemini-api/docs/learnlm/ ,
-
[12]
Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300, 2020
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300, 2020
2009 arXiv
-
[13]
Accessed 15 May, 2025
2025
-
[14]
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021
2021 arXiv
-
[15]
Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021
-
[16]
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. InProceedings of the IEEE/CVF Conference on Co...
2024
-
[17]
Humanity’s last exam.arXiv preprint arXiv:2501.14249, 2025
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al. Humanity’s last exam.arXiv preprint arXiv:2501.14249, 2025
2025 arXiv
-
[18]
Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, et al
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I. Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, et al. Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual...
2024 arXiv
-
[19]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level Google-proof Q&A benchmark. InFirst Conference on Language Modeling, 2024
2024
-
[20]
LLM based math tutoring: Challenges and dataset, 2024
Pepper Miller and Kristen DiCerbo. LLM based math tutoring: Challenges and dataset, 2024
2024
-
[21]
Think you have solved question answering? Try ARC, the AI2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? Try ARC, the AI2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018
2018 arXiv
-
[22]
Gonzalez, et al
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E. Gonzalez, et al. Chatbot Arena: An open platform for evaluating LLMs by human preference. InForty-first International Conference...
2024
-
[23]
The pedagogy benchmark.https://benchmarks.ai-for-educati on.org/, 2024
Ai-for-Education.org. The pedagogy benchmark.https://benchmarks.ai-for-educati on.org/, 2024. Accessed 14 May, 2025
2024
-
[24]
Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom
LouisDeslauriers, LoganS.McCarty, KellyMiller, KristinaCallaghan, andGregKestin. Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. Proceedings of the National Academy of Sciences 116(39), 2019. doi: 10.1073/pnas.1821936116
2019 doi
-
[25]
Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block designs: I. The method of paired comparisons.Biometrika, 39(3/4):324–345, 1952. doi: 10.1093/biomet/39. 3-4.324
1952 doi
-
[26]
Test format and corrective feedback modify the effect of testing on long-term retention.European Journal of Cognitive Psychology, 19:528–558, 07 2007
Sean Kang, Kathleen McDermott, and Henry Roediger. Test format and corrective feedback modify the effect of testing on long-term retention.European Journal of Cognitive Psychology, 19:528–558, 07 2007. doi: 10.1080/09541440601056620
2007 doi
-
[27]
J. P. Kincaid.Derivation of new readability formulas: Automated readability index, fog count and Flesch reading ease formula for Navy enlisted personnel. Research Branch report. Chief of Naval Technical Training, Naval Air Station Memphis, 1975
1975
-
[28]
Roorda, Helma M
Debora L. Roorda, Helma M. Y. Koomen, Jantine L. Spilt, and Frans J. Oort. The influence of affective teacher–student relationships on students’ school engagement and achievement: A meta-analytic approach.Review of Educational Research, 81(4):493–529, 2011. doi: 10.3102/ 00346...
2011
-
[29]
Klem and James P
Adena M. Klem and James P. Connell. Relationships matter: Linking teacher support to student engagement and achievement.Journal of School Health, 74(7), 2004. doi: 10.1111/j.1746-156 1.2004.tb08283.x
2004 doi
-
[30]
Ruzek, Christopher A
Erik A. Ruzek, Christopher A. Hafen, Joseph P. Allen, Anne Gregory, Amori Yee Mikami, and Robert C. Pianta. How teacher emotional support motivates students: The mediating roles of perceived peer relatedness, autonomy support, and competence.Learning and Instruction, 42: 95–10...
2016 doi
-
[31]
Teacher–student relationships and student outcomes: A systematic second-order meta-analytic review.Psychological Bulletin, 2025
ValentinEmslander, DorisHolzberger, SverreBergOfstad, AntoineFischbach, andRonnyScherer. Teacher–student relationships and student outcomes: A systematic second-order meta-analytic review.Psychological Bulletin, 2025. doi: 10.1037/bul0000461
2025 doi
-
[32]
Stone, Charles Underwood, and Jacqueline Hotchkiss
Lynda D. Stone, Charles Underwood, and Jacqueline Hotchkiss. The relational habitus: In- tersubjective processes in learning settings.Human Development, 55(2):65–91, 2012. doi: 10.1159/000337150
2012 doi
-
[33]
Heather A. Davis. Conceptualizing the role and influence of student-teacher relationships on children’s social and cognitive development.Educational Psychologist, 38(4):207–234, 2003. doi: 10.1207/S15326985EP3804_2
2003 doi
-
[34]
AI tutoring outperforms active learning.Research Square, 2024
Gregory Kestin, Kelly Miller, Anna Klales, Timothy Milbourne, and Gregorio Ponti. AI tutoring outperforms active learning.Research Square, 2024. doi: 10.21203/rs.3.rs-4243877/v1
2024 doi
-
[35]
Socratic wisdom in the age of AI: A comparative study of ChatGPT and human tutors in enhancing critical thinking skills.Frontiers in Education, 10, 01
Hoda Fakour and Moslem Imani. Socratic wisdom in the age of AI: A comparative study of ChatGPT and human tutors in enhancing critical thinking skills.Frontiers in Education, 10, 01
-
[36]
From chalkboards to chatbots: Transforming learning in Nigeria, one prompt at a time
Martín De Simone, Federico Tiberti, Wuraola Mosurola, Federico Manolioco, Maria Barron, and Eliott Dikoru. From chalkboards to chatbots: Transforming learning in Nigeria, one prompt at a time. https://web.archive.org/web/20250515142631/https://blogs.worldb ank.org/en/education...
2025
-
[37]
Effective and scalable math support: Experimental evidence on the impact of an AI-math tutor in Ghana
Owen Henkel, Hannah Horne-Robinson, Nessie Kozhakhmetova, and Amanda Lee. Effective and scalable math support: Experimental evidence on the impact of an AI-math tutor in Ghana. InInternational Conference on Artificial Intelligence in Education, pages 373–381. Springer, 2024
2024
-
[38]
Wang, Ana T
Rose E. Wang, Ana T. Ribeiro, Carly D. Robinson, Susanna Loeb, and Dora Demszky. Tutor CoPilot: A human-AI approach for scaling real-time expertise.arXiv preprint arXiv:2410.03017, 2024. 13 Evaluating Gemini in an Arena for Learning
2024 arXiv
-
[39]
Jin Wang and Wenxiang Fan. The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking: Insights from a meta-analysis.Humanities and Social Sciences Communications, 12(1):1–21, 2025. doi: 10.1057/s41599-025-04787-y
2025 doi
-
[40]
Generative AI can harm learning.Available at SSRN, 4895486, 2024
Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Ozge Kabakcı, and Rei Mariman. Generative AI can harm learning.Available at SSRN, 4895486, 2024
2024
-
[41]
Rothschild, Daniel G
Harsh Kumar, David M. Rothschild, Daniel G. Goldstein, and Jake M. Hofman. Math education with large language models: Peril or promise?Available at SSRN 4641653, 2023
2023
-
[42]
World Bank, UNESCO, UNICEF, USAID, FCDO, and the Bill & Melinda Gates Foundation. The state of global learning poverty: 2022 update.https://web.archive.org/web/20250307 163522/https://www.worldbank.org/en/topic/education/publication/state -of-global-learning-poverty, 2022. Acc...
2022
-
[43]
Keith Sawyer ed.The Cambridge handbook of the learning sciences, 2nd ed.Cambridge University Press, 2014
2014
-
[44]
Cornelius, and Fabian J Sting
Matthias Lehmann, Philipp B. Cornelius, and Fabian J Sting. AI meets the classroom: When does ChatGPT harm learning?Available at SSRN 4941259, 2024
2024
-
[45]
The National Academies Press, 2018
National Academies of Sciences, Engineering, and Medicine.How people learn II: Learners, contexts, and cultures. The National Academies Press, 2018
2018
-
[46]
Intrinsic motivation, curiosity, and learning: Theory and applications in educational technologies.Progress in Brain Research, 2016
Pierre-Yves Oudeyer, Jacqueline Gottlieb, and Manuel Lopes. Intrinsic motivation, curiosity, and learning: Theory and applications in educational technologies.Progress in Brain Research, 2016. doi: 10.1016/bs.pbr.2016.05.005
2016 doi
-
[47]
Mayer.Multimedia learning, 2nd ed.Cambridge University Press, 2009
Richard E. Mayer.Multimedia learning, 2nd ed.Cambridge University Press, 2009
2009
-
[48]
Michelene T. H. Chi and Ruth Wylie. The ICAP framework: Linking cognitive engagement to active learning outcomes.Educational Psychologist, 2014. doi: 10.1080/00461520.2014.965823
2014
-
[49]
Kevin R. McKee. Human participants in AI research: Ethics and transparency in practice.IEEE Transactions on Technology and Society, 2024. doi: 10.1109/TTS.2024.3446183
2024
-
[50]
Routledge, 2018
Yana Weinstein, Megan Sumeracki, and Oliver Caviglioli.Understanding how we learn: A visual guide. Routledge, 2018
2018
-
[53]
Strongly disagree
Ido Dagan, Dan Roth, Fabio Zanzotto, and Mark Sammons.Recognizing textual entailment: Models and applications. Springer Nature, 2022. 14 Evaluating Gemini in an Arena for Learning Contributions and Acknowledgments Core Contributors The following individuals made core contribut...
2022
-
[2025]
doi: 10.3389/feduc.2025.1528603
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.