Pith. sign in

REVIEW 5 major objections 4 minor 4 cited by

Evaluating Gemini in an arena for learning

T0 review · 5 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Blind expert comparisons rank Gemini 2.5 Pro as the top AI tutor among five models.

desk verdict Two-stage arena for learning is a genuine methodological step forward, but the 'leading model' claim overreaches: the instrument is developer-authored, and the evaluated model serves as its own judge in two places. read the letter →

arxiv 2505.24477 v1 pith:W7TVENUX submitted 2025-05-30 cs.CY cs.AIcs.LG

classification cs.CYcs.AIcs.LG
keywords learningpedagogyartificialintelligenceevaluationarenamulti-turntutoringblindcomparisonGemini2.5Pro
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to fill a gap the authors see in AI education research: existing benchmarks measure narrow skills like exam accuracy or mistake spotting, but not whether a model can actually guide a learner through a long tutoring conversation. To measure that, they built an 'arena for learning' in which 189 educators role-played students in realistic learning scenarios through blind, multi-turn conversations with pairs of AI models, and then 206 independent experts reviewed the transcripts and judged which tutor better supported the learner. The central claim is that this design produces a general ranking of pedagogical quality, and that in it Gemini 2.5 Pro comes out first: excluding ties, experts preferred it in 73.2% of head-to-head match-ups against GPT-4o, ChatGPT-4o, Claude 3.7 Sonnet, and OpenAI o3. The paper also reports that Gemini 2.5 Pro scores highest on every principle of a 25-item pedagogy rubric, and on targeted tasks of text re-levelling, short-answer assessment, and mistake identification. A fair reader would take away that a human-judged benchmark for AI tutoring is feasible, and that under it Gemini 2.5 Pro is the leading current model for learning.

What carries the argument

The central mechanism is the two-stage learning arena: first, 189 educators role-play learners from 49 realistic scenarios in blind, multi-turn conversations with pairs of models, yielding 1,333 head-to-head match-ups; second, an independent pool of 206 experts reviews the transcripts, scores them on a 25-item pedagogy rubric, and gives pairwise preferences that are converted into Elo ratings via a Bradley-Terry model. The design lets each model steer the conversation independently, which is what allows it to assess overall tutoring quality rather than single-turn accuracy.

What would settle it

Re-run the same 1,333 match-ups with Gemini 2.5 Pro, ChatGPT-4o, and GPT-4o served from current production APIs with identical thinking settings and compute; if expert preference for Gemini over ChatGPT-4o falls from 61% toward chance, the reported ranking is an artifact of model version or serving infrastructure.

Watch

Extended reading notes

Core claim

The paper claims that expert blind preference over multi-turn tutoring conversations places Gemini 2.5 Pro first among the five models evaluated, and that this preference reflects a genuine difference in pedagogical behavior: Gemini scaffolds tasks, resists giving answers away, keeps conversations aimed at the learner's goal, and adapts to the learner, whereas its competitors more often supply solutions and drift off-topic. The authors describe this as establishing Gemini 2.5 Pro as 'a leading model for learning,' with markedly higher performance across key principles of good pedagogy. The paper also reports a notable split: when educators role-played as students, Gemini 2.5 Pro and ChatGPT-4o tied for first on supporting learning goals, but when independent experts assessed the same transcripts, preference shifted decisively to Gemini 2.5 Pro, which the authors attribute to the gap between what feels immediately helpful to a student and what is pedagogically sound.

Load-bearing premise

The comparison is fair across models even though Gemini 2.5 Pro was served on experimental non-production infrastructure, the GPT-4o snapshot was from 2024, and the ChatGPT-4o version is undated.

Editorial extensions

If this is right

  • If expert preference tracks pedagogical quality, then multi-turn arena rankings can serve as a general benchmark for AI tutoring that single-turn exam-style evaluations cannot provide.
  • Models that give answers away readily will rank below models that scaffold, even when both produce correct content.
  • The gap between learners' first-stage preferences, where ChatGPT-4o tied Gemini, and experts' second-stage judgments suggests that user satisfaction alone is an unreliable signal of tutoring quality.
  • Targeted tasks in text re-levelling, short-answer grading, and mistake identification corroborate the arena ranking and can be used to diagnose specific pedagogical skills.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A standardized expert-judged learning arena could push model developers to optimize for scaffolding and learning outcomes rather than for user engagement, potentially reversing the current incentive toward answer-bot behavior.
  • The arena's 49 scenarios sample a single snapshot of a longer learning journey; extending the design to longitudinal multi-session tutoring could reveal whether expert-preferred pedagogy actually produces better learning outcomes, which the paper flags as an open question.
  • If the result is replicated with matched current model snapshots served on production infrastructure, the 73.2% preference figure would be evidence that pedagogical fine-tuning changes tutoring behavior rather than merely that a newer model is stronger.
  • The observed divergence between what students say they like and what experts judge as good teaching connects to the paper's cited finding that felt learning and actual learning can differ; future arena designs could add objective learning-gain measures alongside expert judgment.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper reports a "learning arena" in which 189 educators role-played learners across 49 scenarios and 206 pedagogy experts subsequently judged 1333 blind head-to-head multi-turn tutoring conversations. The authors report that Gemini 2.5 Pro ranks first on a Bradley-Terry Elo model of expert preference, is preferred in 73.2% of non-tied match-ups, and scores highest on all five dimensions of a 25-item pedagogy rubric. The paper supplements these results with targeted evaluations of text re-levelling, short-answer assessment, and mistake identification. The central claim is that Gemini 2.5 Pro is a leading model for learning.

Significance. If the results hold, the arena would be a valuable complement to existing single-turn and narrow-task benchmarks for AI in education, and the use of role-playing educators, multi-turn interactions, and blind expert raters is a genuinely useful design choice. The paper's strengths include a substantial participant pool, realistic scenarios, and targeted evaluations on real-world data (Ghanaian student responses; Khan Academy's math-tutoring benchmark). However, the evaluation instrument is developer-authored, the model comparison is not contemporaneous, the qualitative corroboration is self-referential, and no data or code are released; these issues materially weaken the general "leading model for learning" claim as currently stated.

major comments (5)
  1. [§5.1, Models] The model comparison is not contemporaneous. Gemini 2.5 Pro (2025-05-06) is compared with GPT-4o (2024-08-06), an approximately nine-month-old snapshot, ChatGPT-4o has no snapshot date, and Gemini 2.5 Pro was served on non-production infrastructure. Because frontier model capabilities change quickly, the reported win rates (71.3%, 81.8%, 74.2%, 61.0%) cannot be attributed to intrinsic tutoring quality rather than model age or serving conditions. The authors should either retest with contemporaneous production snapshots or explicitly restrict the conclusions to the specific snapshots and infrastructure used.
  2. [§5.1, Use case coverage and Performance measurement] The evaluation instrument is developer-aligned. The 49-scenario bank includes "archetypal use cases identified by the Gemini product team," and the 25-item pedagogy rubric is "first developed in [10]," the LearnLM technical report, by the same group that integrates LearnLM capabilities into Gemini. The exact system instructions given to all models are not reported. If these prompts and rubric items encode LearnLM's Socratic, scaffolded, anti-answer-giving style, the head-to-head competition measures conformance to a developer-authored specification, not general pedagogical quality. The expert raters are blind and sincere, but the measurement remains potentially circular. The paper should release the full scenario bank and system prompts, detail the rubric's external validation, and include a robustness check with independently authored prompts and rubric items.
  3. [Appendix B] Appendix B uses Gemini 2.5 Pro (2025-05-06) to summarize qualitative feedback about all five models, including itself. This self-referential step means the qualitative corroboration in §2.1 cannot independently validate the arena results, because the summarizer is the model under evaluation. The authors should replace this with human-coded thematic analysis (with inter-coder agreement) or an independent summarization model.
  4. [§5.4] The mistake-identification evaluation applies Gemini 2.5 Pro as the classifier for determining correctness of all models' responses, including Gemini's own. This introduces a direct conflict: the target model sets the ground truth for its own evaluation. The paper should either keep the original GPT-4 Turbo classifier or use an independent classifier, and report agreement between classifiers.
  5. [Statistics and reproducibility] No data or code are released, no inter-rater reliability statistics (e.g., Cohen's kappa, intraclass correlation) are reported for the 206 expert raters, and Tables 2–3 give point estimates without confidence intervals. For example, the short-answer assessment reports 84.1% for both Gemini 2.5 Pro and ChatGPT-4o, and the mistake-identification gap of 1.6 percentage points may be within noise. The authors should provide data, code, rater-agreement measures, and uncertainty intervals for all targeted results.
minor comments (4)
  1. [Abstract and §2.1] The abstract reports "73.2%" excluding ties, but §2.1 gives per-model win rates (71.3%, 81.8%, 74.2%, 61.0%) without explaining how the aggregate is computed; please clarify the weighting.
  2. [§5.2] The text re-levelling prompt is rendered with excessive spacing between characters (e.g., "S im pl ify the f o l l o w i n g text ..."), which appears to be a formatting artifact; please fix the typesetting.
  3. [Table 3] The ChatGPT-4o row shows "79.6 %" with a stray space, and the Rank column is redundant with the row ordering; please clean up the table formatting.
  4. [General] The paper would benefit from an explicit data-availability statement; the reader is only pointed to a short project URL (goo.gle/LearnLM-May25), not to the actual evaluation materials or anonymous data.

Circularity Check

2 steps flagged · score 6.0 of 10

Mistake-identification benchmark is self-judged by the evaluated model, and Appendix B corroboration is self-generated; the main arena preference claim is not circular.

  1. self definitional [Section 5.4, Mistake identification]
    "While the original framework employed GPT-4 Turbo (2024-04-09) to determine whether the tutor model correctly identified student mistakes or accepted accurate work, we updated our methodology to apply Gemini 2.5 Pro (2025-05-06) with dynamic thinking for this classification."

    Table 3 reports accuracy scores for all models on the Khan Academy mistake-identification benchmark, including Gemini 2.5 Pro at 87.4%. The correctness labels used to compute those scores are generated by Gemini 2.5 Pro itself, because the paper replaced the original GPT-4 Turbo classifier with Gemini 2.5 Pro. For the Gemini 2.5 Pro row, the model is both the examined tutor and the judge of whether its own responses correctly identified mistakes. The reported accuracy is therefore not an externally grounded measurement but Gemini's self-assessment, making the head-to-head comparison with Claude 3.7 Sonnet, OpenAI o3, ChatGPT-4o, and GPT-4o non-independent by construction.

  2. other [Appendix B, Qualitative feedback (cross-referenced from Section 2.1)]
    "For each qualitative dataset, we provided Gemini 2.5 Pro (2025-05-06) with a prompt offering brief context about the feedback and encouraging transparent thematic analysis (rather than artificial balancing of positive and negative points), along with the corresponding dataset."

    Section 2.1 states that 'A broader analysis of qualitative feedback from the arena validates these patterns' and directs readers to Appendix B for 'an automated analysis of all qualitative feedback.' The automated analysis is performed by Gemini 2.5 Pro, the same model whose superiority the feedback is said to validate. The Appendix B summaries are therefore self-generated corroboration: the evaluated system is asked to identify themes in comments about itself and its competitors. This does not change the raw expert preference data, but it means the qualitative 'validation' is not independent of the model being promoted.

full rationale

The central arena claim (Gemini 2.5 Pro first by blind expert preference) is not circular: the head-to-head preferences come from 206 external educators and pedagogy experts rating 1333 blind match-ups, and the Elo/preference analysis is a standard Bradley-Terry model applied to those external judgments. The pedagogy rubric is credited to the authors' prior LearnLM report [10], but the paper also states it was developed in consultation with external pedagogy specialists and based on learning-science research [43-48], so the instrument is not defined in terms of the target model's outputs. The scenario bank's provenance from the Gemini product team is a validity/conflict-of-interest concern, not a derivation-level circularity, especially because all models received identical system instructions and the outcome variable is expert preference. Two self-referential steps do undermine specific secondary results. First, Section 5.4 replaces the benchmark's original GPT-4 Turbo classifier with Gemini 2.5 Pro, so the Table 3 accuracy for Gemini 2.5 Pro (87.4%) is computed by Gemini judging its own tutoring responses; this is a by-construction self-evaluation and makes the cross-model comparison non-independent. Second, Appendix B uses Gemini 2.5 Pro to summarize the qualitative feedback that the paper presents as validating Gemini's superiority, so the qualitative corroboration is self-generated. These issues make the targeted corroborations partially circular, but they do not reduce the main arena preference result to its inputs.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central result rests on the arena design and expert judgments rather than on derived equations. The main burdens are domain assumptions about role-play validity and expert judgment as a learning proxy, plus the authors' own rubric and self-referential qualitative analysis.

free parameters (1)
  • Claude 3.7 Sonnet thinking budget = 2000 tokens
    Chosen by the authors to maximize reasoning while limiting latency; this configuration affects Claude's outputs and is not prescribed by any external standard.
assumptions (5)
  • domain assumption Expert judgments of tutoring quality from transcripts are a valid proxy for learning effectiveness.
    Section 2.1 uses second-stage expert preferences as the main ranking; the paper itself notes first-stage student-role preferences tied Gemini and ChatGPT-4o and attributes the gap to a divergence between immediate helpfulness and pedagogical soundness.
  • domain assumption Educators role-playing students produce interactions representative of real learner-tutor conversations.
    Section 5.1 describes stage one as educators role-playing scenarios; this assumes role-play captures real educational use.
  • standard math Bradley-Terry and Elo models are appropriate for ranking from incomplete pairwise preference data.
    Section 5.1 cites Bradley and Terry 1952 for computing Elo ratings from pairwise comparisons.
  • domain assumption Flesch-Kincaid grade level and textual entailment accurately measure text re-levelling quality.
    Section 5.2 uses these metrics to assess grade level and content coverage in the re-levelling evaluation.
  • ad hoc to paper Gemini 2.5 Pro is a reliable and unbiased summarizer of qualitative feedback about all models, including itself.
    Appendix B uses Gemini 2.5 Pro to theme qualitative feedback about its own and competitors' outputs; no independent verification of its neutrality is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating Gemini in an arena for learning." pith.science (2026). https://pith.science/paper/W7TVENUX

@misc{pith2026250524477,
  author       = {Pith},
  title        = {Pith review of: Evaluating Gemini in an arena for learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/W7TVENUX}},
  note         = {Machine review of arXiv:2505.24477}
}
abstract

Artificial intelligence (AI) is poised to transform education, but the research community lacks a robust, general benchmark to evaluate AI models for learning. To assess state-of-the-art support for educational use cases, we ran an "arena for learning" where educators and pedagogy experts conduct blind, head-to-head, multi-turn comparisons of leading AI models. In particular, $N = 189$ educators drew from their experience to role-play realistic learning use cases, interacting with two models sequentially, after which $N = 206$ experts judged which model better supported the user's learning goals. The arena evaluated a slate of state-of-the-art models: Gemini 2.5 Pro, Claude 3.7 Sonnet, GPT-4o, and OpenAI o3. Excluding ties, experts preferred Gemini 2.5 Pro in 73.2% of these match-ups -- ranking it first overall in the arena. Gemini 2.5 Pro also demonstrated markedly higher performance across key principles of good pedagogy. Altogether, these results position Gemini 2.5 Pro as a leading model for learning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

    cs.CY 2026-08 conditional novelty 6.0 of 10

    A 30-day, knowledge-tracing-grounded simulated learner benchmark for tutoring agents finds that no base model or harness alone determines quality and that almost all tested combinations plateau within days.

  2. 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    3D-R1 uses a synthetic chain-of-thought cold start plus GRPO reinforcement learning with perception, semantic, and format rewards, and reports best published results across 3D dense captioning, QA, grounding, dialogue...

  3. Benchmarking the Pedagogical Knowledge of Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The authors release an open benchmark of 1,143 pedagogical knowledge questions from Chilean teacher exams and report accuracy, cost, and size trade-offs for 97 large language models.

  4. Nav-R1: Reasoning and Navigation in Embodied Scenes

    cs.RO 2025-09 reject novelty 4.0 of 10

    Nav-R1 uses a 110K synthetic CoT dataset, GRPO with three rewards, and a fast-in-slow system to set new SOTA on R2R-CE, RxR-CE, and HM3D-OVON.

Reference graph

Works this paper leans on

51 extracted references · 35 canonical work pages · cited by 4 Pith papers

  1. [10]

    LearnLM: Improving Gemini for learning.arXiv preprint arXiv:2412.16429, 2024

    LearnLM Team, Abhinit Modi, Aditya Srikanth Veerubhotla, Aliya Rysbek, Andrea Huber, Brett Wiltshire, Brian Veprek, Daniel Gillick, Daniel Kasenberg, Derek Ahmed, et al. LearnLM: Improving Gemini for learning.arXiv preprint arXiv:2412.16429, 2024

  2. [1]

    About a quarter of U.S

    Olivia Sidoti, Eugenie Park, and Jeffrey Gottfried. About a quarter of U.S. teens have used ChatGPT for schoolwork – double the share in 2023.https://web.archive.org/web/ 20250430224239/https://www.pewresearch.org/short-reads/2025/01/15/abou t-a-quarter-of-us-teens-have-used-chatgpt-for-schoolwork-double-the-s hare-in-2023/, 2025. Accessed 15 May, 2025

  3. [3]

    Artificially intelligent? Children’s and parents’ views on generative AI in education

    Ali Bissoondath. Artificially intelligent? Children’s and parents’ views on generative AI in education. https://web.archive.org/web/20241119022121/https://www.inte rnetmatters.org/hub/press-release/ai-research-warns-schools-unprepare d-artificial-intelligence/, 2024. Accessed 15 May, 2025

  4. [4]

    Student generative AI survey 2025.https://web.archive.org/web/2025 0513133314/https://www.hepi.ac.uk/2025/02/26/student-generative-ai-sur vey-2025/, 2025

    Josh Freeman. Student generative AI survey 2025.https://web.archive.org/web/2025 0513133314/https://www.hepi.ac.uk/2025/02/26/student-generative-ai-sur vey-2025/, 2025. Accessed 15 May, 2025

  5. [5]

    Penguin, 2024

    Salman Khan.Brave new words: How AI will revolutionize education (and why that’s a good thing). Penguin, 2024

  6. [6]

    Penguin, 2024

    Ethan Mollick.Co-intelligence: Living and working with AI. Penguin, 2024

  7. [7]

    Artificial intelligence in education: Opportunities, challenges, and policy considera- tions for Congress, 2025

    Erin Mote. Artificial intelligence in education: Opportunities, challenges, and policy considera- tions for Congress, 2025

  8. [8]

    Higher education leaders navigate AI disruption

    American Association of Colleges and Universities. Higher education leaders navigate AI disruption. https://web.archive.org/web/20250405093427/https://www.aacu .org/newsroom/higher-education-leaders-navigate-ai-disruption , 2025. Accessed 15 May, 2025. 11 Evaluating Gemini in an Arena for Learning

Show all 51 references
  1. [9]

    McKee, Daniel Gillick, Shaojian Zhu, Sara Wiltberger, Shubham Milind Phal, Katherine Hermann, Daniel Kasenberg, Avishkar Bhoopchand, et al

    Irina Jurenka, Markus Kunesch, Kevin R. McKee, Daniel Gillick, Shaojian Zhu, Sara Wiltberger, Shubham Milind Phal, Katherine Hermann, Daniel Kasenberg, Avishkar Bhoopchand, et al. Towards responsible development of generative AI for education: An evaluation-driven approach. ar...

  2. [11]

    LearnLM Team. LearnLM. https://ai.google.dev/gemini-api/docs/learnlm/ ,

  3. [12]

    Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300, 2020

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300, 2020

  4. [13]

    Accessed 15 May, 2025

  5. [14]

    Measuring mathematical problem solving with the math dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021

  6. [15]

    Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021

  7. [16]

    MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. InProceedings of the IEEE/CVF Conference on Co...

  8. [17]

    Humanity’s last exam.arXiv preprint arXiv:2501.14249, 2025

    Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al. Humanity’s last exam.arXiv preprint arXiv:2501.14249, 2025

  9. [18]

    Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, et al

    Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I. Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, et al. Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual...

  10. [19]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level Google-proof Q&A benchmark. InFirst Conference on Language Modeling, 2024

  11. [20]

    LLM based math tutoring: Challenges and dataset, 2024

    Pepper Miller and Kristen DiCerbo. LLM based math tutoring: Challenges and dataset, 2024

  12. [21]

    Think you have solved question answering? Try ARC, the AI2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? Try ARC, the AI2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018

  13. [22]

    Gonzalez, et al

    Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E. Gonzalez, et al. Chatbot Arena: An open platform for evaluating LLMs by human preference. InForty-first International Conference...

  14. [23]

    The pedagogy benchmark.https://benchmarks.ai-for-educati on.org/, 2024

    Ai-for-Education.org. The pedagogy benchmark.https://benchmarks.ai-for-educati on.org/, 2024. Accessed 14 May, 2025

  15. [24]

    Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom

    LouisDeslauriers, LoganS.McCarty, KellyMiller, KristinaCallaghan, andGregKestin. Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. Proceedings of the National Academy of Sciences 116(39), 2019. doi: 10.1073/pnas.1821936116

  16. [25]

    Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block designs: I. The method of paired comparisons.Biometrika, 39(3/4):324–345, 1952. doi: 10.1093/biomet/39. 3-4.324

  17. [26]

    Test format and corrective feedback modify the effect of testing on long-term retention.European Journal of Cognitive Psychology, 19:528–558, 07 2007

    Sean Kang, Kathleen McDermott, and Henry Roediger. Test format and corrective feedback modify the effect of testing on long-term retention.European Journal of Cognitive Psychology, 19:528–558, 07 2007. doi: 10.1080/09541440601056620

  18. [27]

    J. P. Kincaid.Derivation of new readability formulas: Automated readability index, fog count and Flesch reading ease formula for Navy enlisted personnel. Research Branch report. Chief of Naval Technical Training, Naval Air Station Memphis, 1975

  19. [28]

    Roorda, Helma M

    Debora L. Roorda, Helma M. Y. Koomen, Jantine L. Spilt, and Frans J. Oort. The influence of affective teacher–student relationships on students’ school engagement and achievement: A meta-analytic approach.Review of Educational Research, 81(4):493–529, 2011. doi: 10.3102/ 00346...

  20. [29]

    Klem and James P

    Adena M. Klem and James P. Connell. Relationships matter: Linking teacher support to student engagement and achievement.Journal of School Health, 74(7), 2004. doi: 10.1111/j.1746-156 1.2004.tb08283.x

  21. [30]

    Ruzek, Christopher A

    Erik A. Ruzek, Christopher A. Hafen, Joseph P. Allen, Anne Gregory, Amori Yee Mikami, and Robert C. Pianta. How teacher emotional support motivates students: The mediating roles of perceived peer relatedness, autonomy support, and competence.Learning and Instruction, 42: 95–10...

  22. [31]

    Teacher–student relationships and student outcomes: A systematic second-order meta-analytic review.Psychological Bulletin, 2025

    ValentinEmslander, DorisHolzberger, SverreBergOfstad, AntoineFischbach, andRonnyScherer. Teacher–student relationships and student outcomes: A systematic second-order meta-analytic review.Psychological Bulletin, 2025. doi: 10.1037/bul0000461

  23. [32]

    Stone, Charles Underwood, and Jacqueline Hotchkiss

    Lynda D. Stone, Charles Underwood, and Jacqueline Hotchkiss. The relational habitus: In- tersubjective processes in learning settings.Human Development, 55(2):65–91, 2012. doi: 10.1159/000337150

  24. [33]

    Heather A. Davis. Conceptualizing the role and influence of student-teacher relationships on children’s social and cognitive development.Educational Psychologist, 38(4):207–234, 2003. doi: 10.1207/S15326985EP3804_2

  25. [34]

    AI tutoring outperforms active learning.Research Square, 2024

    Gregory Kestin, Kelly Miller, Anna Klales, Timothy Milbourne, and Gregorio Ponti. AI tutoring outperforms active learning.Research Square, 2024. doi: 10.21203/rs.3.rs-4243877/v1

  26. [35]

    Socratic wisdom in the age of AI: A comparative study of ChatGPT and human tutors in enhancing critical thinking skills.Frontiers in Education, 10, 01

    Hoda Fakour and Moslem Imani. Socratic wisdom in the age of AI: A comparative study of ChatGPT and human tutors in enhancing critical thinking skills.Frontiers in Education, 10, 01

  27. [36]

    From chalkboards to chatbots: Transforming learning in Nigeria, one prompt at a time

    Martín De Simone, Federico Tiberti, Wuraola Mosurola, Federico Manolioco, Maria Barron, and Eliott Dikoru. From chalkboards to chatbots: Transforming learning in Nigeria, one prompt at a time. https://web.archive.org/web/20250515142631/https://blogs.worldb ank.org/en/education...

  28. [37]

    Effective and scalable math support: Experimental evidence on the impact of an AI-math tutor in Ghana

    Owen Henkel, Hannah Horne-Robinson, Nessie Kozhakhmetova, and Amanda Lee. Effective and scalable math support: Experimental evidence on the impact of an AI-math tutor in Ghana. InInternational Conference on Artificial Intelligence in Education, pages 373–381. Springer, 2024

  29. [38]

    Wang, Ana T

    Rose E. Wang, Ana T. Ribeiro, Carly D. Robinson, Susanna Loeb, and Dora Demszky. Tutor CoPilot: A human-AI approach for scaling real-time expertise.arXiv preprint arXiv:2410.03017, 2024. 13 Evaluating Gemini in an Arena for Learning

  30. [39]

    Jin Wang and Wenxiang Fan. The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking: Insights from a meta-analysis.Humanities and Social Sciences Communications, 12(1):1–21, 2025. doi: 10.1057/s41599-025-04787-y

  31. [40]

    Generative AI can harm learning.Available at SSRN, 4895486, 2024

    Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Ozge Kabakcı, and Rei Mariman. Generative AI can harm learning.Available at SSRN, 4895486, 2024

  32. [41]

    Rothschild, Daniel G

    Harsh Kumar, David M. Rothschild, Daniel G. Goldstein, and Jake M. Hofman. Math education with large language models: Peril or promise?Available at SSRN 4641653, 2023

  33. [42]

    World Bank, UNESCO, UNICEF, USAID, FCDO, and the Bill & Melinda Gates Foundation. The state of global learning poverty: 2022 update.https://web.archive.org/web/20250307 163522/https://www.worldbank.org/en/topic/education/publication/state -of-global-learning-poverty, 2022. Acc...

  34. [43]

    Keith Sawyer ed.The Cambridge handbook of the learning sciences, 2nd ed.Cambridge University Press, 2014

  35. [44]

    Cornelius, and Fabian J Sting

    Matthias Lehmann, Philipp B. Cornelius, and Fabian J Sting. AI meets the classroom: When does ChatGPT harm learning?Available at SSRN 4941259, 2024

  36. [45]

    The National Academies Press, 2018

    National Academies of Sciences, Engineering, and Medicine.How people learn II: Learners, contexts, and cultures. The National Academies Press, 2018

  37. [46]

    Intrinsic motivation, curiosity, and learning: Theory and applications in educational technologies.Progress in Brain Research, 2016

    Pierre-Yves Oudeyer, Jacqueline Gottlieb, and Manuel Lopes. Intrinsic motivation, curiosity, and learning: Theory and applications in educational technologies.Progress in Brain Research, 2016. doi: 10.1016/bs.pbr.2016.05.005

  38. [47]

    Mayer.Multimedia learning, 2nd ed.Cambridge University Press, 2009

    Richard E. Mayer.Multimedia learning, 2nd ed.Cambridge University Press, 2009

  39. [48]

    Michelene T. H. Chi and Ruth Wylie. The ICAP framework: Linking cognitive engagement to active learning outcomes.Educational Psychologist, 2014. doi: 10.1080/00461520.2014.965823

  40. [49]

    Kevin R. McKee. Human participants in AI research: Ethics and transparency in practice.IEEE Transactions on Technology and Society, 2024. doi: 10.1109/TTS.2024.3446183

  41. [50]

    Routledge, 2018

    Yana Weinstein, Megan Sumeracki, and Oliver Caviglioli.Understanding how we learn: A visual guide. Routledge, 2018

  42. [53]

    Strongly disagree

    Ido Dagan, Dan Roth, Fabio Zanzotto, and Mark Sammons.Recognizing textual entailment: Models and applications. Springer Nature, 2022. 14 Evaluating Gemini in an Arena for Learning Contributions and Acknowledgments Core Contributors The following individuals made core contribut...

  43. [2025]

    doi: 10.3389/feduc.2025.1528603

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.