REVIEW 5 major objections 5 minor 268 references
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
T0 review · 5 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read The paper introduces a benchmark that claims to be the first to measure whether LLMs can apply second-language learning-design principles, reporting a top score of 85.5%.
desk verdict A genuinely useful benchmark for L2 ed, but the leaderboard rests on a same-family judge and should be treated as provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the L2-Bench construct: a validated taxonomy of 12 competencies (course planning, lesson planning, activity planning, classroom management, language presentation, conversational exchange, performance evaluation, feedback, progress tracking, social-emotional management, assessment creation, and professional development) and 31 subcompetencies, each carrying consensus criteria. Around this sit three scoring layers—universal criteria (age/CEFR/cultural/resource/privacy constraints with context-dependent weights), consensus criteria shared by all tasks in a subcompetency, and task-specific criteria—combined into a weighted pass/fail score. Scoring is automated by an
What would settle it
A direct check is the paper's own proposed extension (Appendix A): recompute the leaderboard with judge- and criterion-level uncertainty propagated into the confidence intervals—if the 1.4-percentage-point gap between the top three models falls inside the expanded intervals, the headline ranking is not statistically separable.
Extended reading notes
Core claim
The central claim is that LLM capability in second-language education can and should be measured as performativity—whether a model can apply learning experience design principles in concrete, contextualised tasks—rather than as declarative knowledge of those principles or as downstream learning outcomes. L2-Bench operationalises this with a taxonomy of 12 competencies and 31 subcompetencies validated by 221 practitioners (task authenticity 4.42/5, criteria adequacy 4.18/5), a dataset of 1,000+ single-turn task–response pairs parameterised by 33 context factors, and a three-layer rubric system (universal, consensus, and task-specific criteria) scored by an optimised reference-guided LLM judge
Load-bearing premise
The load-bearing premise is that a reference-guided LLM judge can score open-ended pedagogical quality well enough to rank models, despite expert human raters agreeing only weakly with one another (Krippendorff's α≈0.36 on criterion-level judgments).
Editorial extensions
If this is right
- Stakeholders can compare AIED systems on applied pedagogical performance rather than knowledge recall or anecdote, supporting procurement and governance decisions.
- Model rankings are stable across hard-item, validated-subset, and length-adjusted variants, so the tier ordering is not an artefact of the aggregation rule.
- All models lose ground on open-ended competencies (conversational exchange, feedback, evaluation) and in low-resource or young-learner contexts, pointing adopters to specific risk areas.
- Negative universal criteria (cultural sensitivity, CEFR appropriacy, data privacy) are violated by all models, indicating that safety-relevant failures are not confined to small models.
- Because the benchmark is automated and open, it can be re-run as models and prompts evolve, providing a living baseline for AIED evaluation.
Reading between the lines
- Inference: The paper's own limitation notes imply the headline confidence intervals understate uncertainty: judge and criterion noise are not propagated. Re-reporting the leaderboard with expanded intervals is a natural next step, and the 1.4-point gap between the top three models could plausibly close.
- Inference: The single-turn design means the observed weakness in conversational competencies may be an artefact; a multi-turn extension could reveal different orderings, especially for exchange-partner tasks.
- Inference: The rubric methodology is presented as a generalisable template for other open-ended disciplines, but the paper explicitly leaves that generalization unproven. A testable extension would be to build a parallel benchmark in another subject and check whether expert validation and judge-human agreement replicate.
- Inference: Because practitioners could not reliably distinguish reference answers from frontier-model responses in blind A/B comparison (51.3% preference), the benchmark's headroom for top models may be smaller than the leaderboard suggests; a discrimination-focused test set targeting near-ceiling models would clarify this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. L2-Bench introduces a benchmark for evaluating LLM capabilities in second language (L2) learning experience design. It contributes: (1) a taxonomy of 12 competencies and 31 sub-competencies validated by 221 expert practitioners; (2) a rubric-based, LLM-as-a-judge scoring pipeline; (3) a dataset of 1,000+ single-turn task-response pairs; and (4) a leaderboard of nine models. The central claims are that the taxonomy is practitioner-validated, the automated judge agrees substantially with human expert consensus, and the resulting leaderboard provides reliable signal about model strengths and weaknesses across L2 education contexts. The paper is open about limitations, including low human inter-annotator agreement, the same-family judge/top-model confound, and the absence of judge/criterion uncertainty propagation into leaderboard confidence intervals.
Significance. If the validity arguments hold, L2-Bench would be a valuable contribution: it is one of the first construct-validated, pedagogy-grounded benchmarks for AIED in L2 education, with a substantial practitioner validation study (N=221, 45 countries) and an open dataset and evaluation pipeline. The paper ships reproducible code and data, reports detailed appendices, and honestly discusses limitations. The methodological attention to rater disagreement and to the distinction between sampling variability and measurement uncertainty is a strength. However, the benchmark's central leaderboard claim currently rests on an automated judge whose same-family bias is unquantified, and the low human inter-annotator agreement (Krippendorff's alpha=0.362) raises unresolved questions about what the rubric scores actually measure. These issues are load-bearing for the paper's main empirical contribution.
major comments (5)
- [Table 1 and Appendix A/H.4] The leaderboard is the paper's central empirical claim, but its validity depends on the production judge. The production judge is Claude Sonnet 4.6 (Appendix H.4), and the top-ranked model is Claude Opus 4.7, a same-family model. Table 19 shows DeepSeek V3.2 v1 achieved higher human alignment (F1=0.942, kappa=0.746) than Sonnet 4.6 v1 (F1=0.936, kappa=0.719). The cross-family check in H.4 reports raw verdict agreement kappa=0.764 but per-competency Spearman rho=0.196 (p=0.564), which the paper dismisses as ceiling. With the top three models separated by only 1.4-2.1pp, and Table 1 CIs excluding judge/criterion uncertainty, a same-family bias of a few points could change the headline ranking. Appendix A explicitly lists a self-preference audit as future work, so the leaderboard currently rests on an unquantified confound. A non-Claude judge re-score of at least the top-tier models is need
- [§3.2, Appendix G.7] The validation study reports Krippendorff's alpha=0.362 on criterion-level binary judgments. The paper argues this is acceptable because the judge agrees with a denoised human majority (kappa=0.746), and that disagreement reflects construct complexity rather than noise. This argument is plausible but not fully convincing: the denoised majority is itself derived from a noisy signal, and the paper does not report how stable the majority is under resampling. The claim that aggregate competency-level scores are informative is reasonable (IIC 0.53-0.61 is moderate), but the paper should show sensitivity analyses using alternative consensus thresholds (e.g., 75% or 100%) at the aggregate-score level, not just at the criterion level, to demonstrate that leaderboard ranks are robust to the noise in the human target.
- [Appendix I.2, Eq. (11)] Equation (11) defines the standard error as containing a between-task variance and a within-task term. The text says the intervals capture sampling variability over tasks but do not propagate judge- or criterion-level uncertainty. This is clearly disclosed, but the abstract and Section 4 call the leaderboard 'reliable signal' without this caveat. Since the top-tier gaps are small and the judge uncertainty is unknown, the 95% CIs in Table 1 may mislead readers into interpreting the ranking as more precise than it is. The paper should either include judge uncertainty in the CIs or present the leaderboard with a clear statement that systematic errors could exceed the reported sampling error.
- [Appendix G.8, Table 14] The blind A/B preference study found that practitioners preferred the reference answer only 51.3% of the time, far below the 70% target, and the result was not significant except for Competency 10. The paper interprets this as indicating that frontier models are indistinguishable from expert-authored answers. However, this also suggests that the reference answers may not be 'gold standard' in the eyes of practitioners—a finding that undercuts the use of these references as guides for the automated judge. The paper should discuss how the lack of clear expert preference for references affects the validity of using them as judge anchors.
- [§3.3, Appendix H.2] Judge validation was conducted on only 48 tasks (4 per competency) with 326 matched criteria. This is a small sample for a production judge that will score 1,000 tasks across all models. The paper should report confidence intervals around the judge-human kappa and F1, and ideally validate on a larger held-out set. The current sample size limits the strength of the claim that the judge is 'reliable' across the full dataset.
minor comments (5)
- [Abstract] The abstract says '1,000+ task-response pairs' and later 'reliable signal,' but the limitations in Appendix A are not referenced in the abstract. Consider adding a phrase like 'with caveats regarding judge uncertainty' to avoid overstatement.
- [Figure 1 and Figure 2] Figure 1 and Figure 2 are referenced in the text but are not fully described in the caption; ensure they are readable in print and specify what the error bars represent (if any).
- [§4, Table 1] The table caption says '95% confidence intervals (N=1,000 tasks per model)' but the text correctly notes these are task-sampling CIs, not total uncertainty. Add a footnote to the table stating 'CIs do not include judge or criterion uncertainty.'
- [Appendix H.2, Table 19] Table 19 lists DeepSeek V3.2 as 'v1 (ref-guided)' with F1=0.942, but the main text (Section 3.3, Judge selection) says 'judge vs. rater majority vote'—clarify the exact comparison target (e.g., 67% majority threshold) in the table caption.
- [Appendix I.3] The model configuration says 'Temperature=1.0' for several models, but the judge pipeline uses temperature 0. This asymmetry is not discussed; a comment on whether temperature affects the judge's stability on the 3 repeated runs would be useful.
Circularity Check
No significant circularity; the benchmark's claims rest on external human-rater validation, with the same-family judge issue disclosed as a validity risk rather than a derivational loop.
full rationale
L2-Bench is an empirical benchmark construction, not a mathematical derivation, and its central claims do not reduce to their inputs by construction. The taxonomy is derived from external L2 frameworks (CEFR, Eaquals, CETF, Millin, British Council) and independently rated by 221 practitioners on authenticity and criteria adequacy. The automated scoring pipeline was selected by measuring judge agreement against human majority verdicts on a validation sample (Appendix H.2), and the production judge's verdicts were checked cross-family: "the two judges agreed on the raw Pass/Fail verdict for 92.3% of criteria, with Cohen's κ=0.764 (substantial agreement)" (Appendix H.4). The leaderboard is therefore not a fitted parameter renamed as a prediction. The same-family concern (Claude judge vs. Claude Opus top entrant) is real and is explicitly acknowledged: "The production judge (Claude Sonnet 4.6) shares a model family with the top-ranked entrant (Claude Opus 4.7), raising a potential self-preference confound" (Appendix A); the paper also states it does "not propagate judge- or criterion-level uncertainty into those intervals." These are validity and uncertainty limitations, not circular reductions: nothing in the scoring formula (Eq. 10) or judge prompt defines the target score in terms of the judge's family membership, and the paper's own Table 19 shows an unrelated judge matching or exceeding the chosen judge's human alignment. No self-citation is load-bearing: the cited pilot (Edgell et al. 2026) informed design, but the validation evidence here is new. Therefore no circular step can be exhibited under the required standard.
Assumptions & free parameters
free parameters (5)
- Rubric criterion weights (e.g., +10 essential, +5 important, +2 nice-to-have, −10 severe) =
Range −10 to +10; specific weights defined per criterion (see Appendix E)
- Universal criterion context-conditional weights (e.g., age-appropriateness weight 10 for pre-primary vs 2 for adult) =
See Table 9 (e.g., 10, 5, 2; 9, 2; etc.)
- Human validation targets =
Authenticity target 4.0/5; criteria adequacy target 3.5/5; judge κ target ≥0.60; recall target ≥0.80
- Task difficulty threshold for hard-item variant =
267 tasks where top three models scored below 80%
- Verbosity penalty in length-adjusted scores =
Not specified exactly
assumptions (5)
- domain assumption Expert practitioner ratings of task authenticity and criteria adequacy are valid evidence of benchmark construct validity.
- domain assumption An LLM-as-a-judge with reference answers can reliably evaluate open-ended pedagogical responses.
- domain assumption The single-turn task format is a meaningful proxy for L2 learning-design capability.
- domain assumption The five European-origin frameworks (CEFR, Eaquals, CETF, Millin, British Council CPD) are adequate basis for a global L2 English benchmark.
- domain assumption Scoring with k=3 runs at temperature 1 (for some models) provides stable task-level scores.
invented entities (2)
-
L2-Bench competency taxonomy (12 competencies, 31 subcompetencies, 72 consensus criteria)
independent evidence
-
L2-Bench task dataset (1,000 task-response pairs with rubrics and reference answers)
independent evidence
Cite this review
Pith. "Pith review of L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education." pith.science (2026). https://pith.science/paper/DBPQIA4I
@misc{pith2026260708842,
author = {Pith},
title = {Pith review of: L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education},
year = {2026},
howpublished = {\url{https://pith.science/paper/DBPQIA4I}},
note = {Machine review of arXiv:2607.08842}
}
read the original abstract
Despite rapid AI adoption in education, rigorous evaluation of AI-powered educational (AIED) systems remains critically underdeveloped, particularly in second language (L2) education, one of the most common yet least evaluated AI applications. We introduce L2-Bench, an open-source benchmark of 1,000+ task-response pairs to aid the pedagogy-led evaluation of LLM capabilities relating to language learning and assessment. Crucially, L2-Bench measures model performativity on the application of learning experience design principles rather than mere knowledge of those principles or broad learning outcomes. Our contributions include: (1) a validated taxonomy of 12 competencies and 31 subcompetencies validated by 200+ expert practitioners (task authenticity: 4.42/5.00, criteria adequacy: 4.18/5.00); (2) a rubric-based evaluation methodology that we believe can, if adapted, generalize to similar (open-ended, qualitative) disciplines; (3) an evaluation dataset that produces reliable signal about model strengths, weaknesses, and contextual robustness across diverse L2 education scenarios. We find that, among large models, Claude Opus 4.7 performs best overall (85.5%), though is marginally outperformed on several constituent tasks. We also find that performance drops notably on harder tasks (69.9% to 73.4%). L2-Bench provides education stakeholders better methods to make more informed decisions about real-world AIED adoption, use, and governance, while advancing the maturing science of AI evaluations for education.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2511.21695 , year=
EvalCards: A Framework for Standardized Evaluation Reporting , author=. arXiv preprint arXiv:2511.21695 , year=
-
[2]
AAAI Conference on Artificial Intelligence , year=
Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation , author=. AAAI Conference on Artificial Intelligence , year=
-
[3]
2025 , eprint=
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights , author=. 2025 , eprint=
2025
-
[4]
2025 , eprint=
Audit Cards: Contextualizing AI Evaluations , author=. 2025 , eprint=
2025
-
[5]
2025 , eprint=
Artificial Intelligence Index Report 2025 , author=. 2025 , eprint=
2025
-
[6]
2023 , eprint =
Weidinger, Laura and Rauh, Markus and Marchal, Naomi and Manzini, Anna and Hendricks, Lisa Anne and Mateos-Garcia, Juan and Bergman, Sofia and Kay, John and Griffin, Colin and Bariach, Boris and Gabriel, Iason and Rieser, Verena and Isaac, William , title =. 2023 , eprint =
2023
-
[7]
2025 , eprint=
Fantastic Bugs and Where to Find Them in AI Benchmarks , author=. 2025 , eprint=
2025
-
[8]
2024 , eprint=
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices , author=. 2024 , eprint=
2024
Show all 268 references
-
[9]
2025 Expert consensus on retrospective evaluation of large language model applications in clinical scenarios , journal =
Qing Chang and Fei Chen and Yaolong Chen and Longlong Cheng and Di Dong and Jiahong Dong and Xiaobin Feng and Junbo Ge and Jingjing He and Yihua He and Zhiyang He and Hong Ji and Xue Jiang and Zehua Jiang and Nan Li and Peng Li and Yazi Li and Bing Liu and Junwei Liu and Han L...
2025
-
[10]
Koehler , title =
Punya Mishra and Matthew J. Koehler , title =. Teachers College Record , volume =. 2006 , doi =. https://doi.org/10.1111/j.1467-9620.2006.00684.x , abstract =
2006
-
[11]
Contextual knowledge and TPACK: Evidence from a global south setting , journal =
Van Loi Nguyen and Chung Thi Thanh Hang and Nguyen Trong Nguyen and Huynh Truong Sang , keywords =. Contextual knowledge and TPACK: Evidence from a global south setting , journal =. 2025 , issn =. doi:https://doi.org/10.1016/j.caeo.2025.100290 , url =
2025
-
[12]
SHULMAN , title =
LEE S. SHULMAN , title =. Educational Researcher , volume =. 1986 , doi =
1986
-
[13]
2025 , eprint=
HealthBench: Evaluating Large Language Models Towards Improved Human Health , author=. 2025 , eprint=
2025
-
[14]
2025 , eprint=
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor , author=. 2025 , eprint=
2025
-
[15]
2024 , eprint=
Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations , author=. 2024 , eprint=
2024
-
[16]
2025 , eprint=
Measuring what Matters: Construct Validity in Large Language Model Benchmarks , author=. 2025 , eprint=
2025
-
[17]
2004 , publisher =
Content Analysis: An Introduction to Its Methodology , author =. 2004 , publisher =
2004
-
[18]
Acta Psychologica , volume =
Optimal Number of Response Categories in Rating Scales: Reliability, Validity, Discriminating Power, and Respondent Preferences , author =. Acta Psychologica , volume =. 2000 , doi =
2000
-
[19]
Results-actionability gap: Understanding how practitioners evaluate LLM products in the wild
van der Maden, Willem and Sadek, Malak and Xiao, Ziang and Mottelson, Aske and Vera Liao, Q and Zhu, Jichen. Results-actionability gap: Understanding how practitioners evaluate LLM products in the wild. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems
2026
-
[21]
2025 , eprint=
Establishing Best Practices for Building Rigorous Agentic Benchmarks , author=. 2025 , eprint=
2025
-
[22]
2025 , eprint=
Eval Factsheets: A Structured Framework for Documenting AI Evaluations , author=. 2025 , eprint=
2025
-
[23]
ICML Workshop on Technical AI Governance (TAIG) , year=
Deprecating Benchmarks: Criteria and Framework , author=. ICML Workshop on Technical AI Governance (TAIG) , year=
-
[24]
2026 , eprint=
Measuring What AI Systems Might Do: Towards A Measurement Science in AI , author=. 2026 , eprint=
2026
-
[25]
2019 , isbn =
Mitchell, Margaret and Wu, Simone and Zaldivar, Andrew and Barnes, Parker and Vasserman, Lucy and Hutchinson, Ben and Spitzer, Elena and Raji, Inioluwa Deborah and Gebru, Timnit , title =. 2019 , isbn =. doi:10.1145/3287560.3287596 , booktitle =
2019
-
[26]
Datasheets for datasets , year =
Gebru, Timnit and Morgenstern, Jamie and Vecchione, Briana and Vaughan, Jennifer Wortman and Wallach, Hanna and III, Hal Daum\'. Datasheets for datasets , year =. Commun. ACM , month = nov, pages =. doi:10.1145/3458723 , abstract =
-
[27]
The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks , author=. The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=
-
[28]
2023 , eprint=
Assessing Language Model Deployment with Risk Cards , author=. 2023 , eprint=
2023
-
[29]
and Schade, Sven and O’Sullivan, Declan and Lewis, Dave , title =
Golpayegani, Delaram and Hupont, Isabelle and Panigutti, Cecilia and Pandit, Harshvardhan J. and Schade, Sven and O’Sullivan, Declan and Lewis, Dave , title =. 2024 , isbn =. doi:10.1007/978-3-031-68024-3_3 , booktitle =
2024 doi
-
[30]
It Takes Two to Tango: Navigating Conceptualizations of NLP Tasks and Measurements of Performance
Subramonian, Arjun and Yuan, Xingdi and Daum \'e III, Hal and Blodgett, Su Lin. It Takes Two to Tango: Navigating Conceptualizations of NLP Tasks and Measurements of Performance. Findings of the Association for Computational Linguistics: ACL 2023. 2023. doi:10.18653/v1/2023.fi...
2023 doi
-
[31]
2025 , month = jan, date =
Assessing AI: Surveying the Spectrum of Approaches to Understanding and Auditing AI Systems , author =. 2025 , month = jan, date =
2025
-
[32]
2021 , eprint =
Akula, Ramya and Garibay, Ivan , title =. 2021 , eprint =
2021
-
[33]
Rauh, Maribeth and Mellor, John F.J. and Uesato, Jonathan and Huang, Po-Sen and Welbl, Johannes and Weidinger, Laura and Dathathri, Sumanth and Glaese, Amelia and Irving, Geoffrey and Gabriel, Iason and Isaac, William and Hendricks, Lisa Anne , title =. 2022 , eprint =
2022
-
[34]
Liang, Percy and Bommasani, Rishi and Lee, Tony and Tsipras, Dimitris and Soylu, Dilara and Yasunaga, Michihiro and Zhang, Yian and Narayanan, Deepak and Wu, Yuhuai and Kumar, Ananya and Newman, Benjamin and Yuan, Binhang and Yan, Bobby and Zhang, Ce and Cosgrove, Christian an...
-
[35]
Patterns , volume =
Manheim, David , title =. Patterns , volume =. 2023 , doi =
2023
-
[36]
2024 , eprint =
Hafner, Flavio and Sun, Chang , title =. 2024 , eprint =
2024
-
[37]
2023 , eprint =
Shevlane, Toby and Farquhar, Sebastian and Garfinkel, Ben and Phuong, Mary and Whittlestone, Jess and Leung, Jade and Kokotajlo, Daniel and Marchal, Nahema and Anderljung, Markus and Kolt, Noam and Ho, Lewis and Siddarth, Divya and Avin, Shahar and Hawkins, Will and Kim, Been ...
2023
-
[38]
and Hullman, Jessica and Subramonyam, Hari , title =
Gupta, Neha R. and Hullman, Jessica and Subramonyam, Hari , title =. 2024 , eprint =
2024
-
[39]
and Hu, Xiao , title =
Ding, Cheng and Guo, Zhicheng and Rudin, Cynthia and Xiao, Ran and Nahab, Fadi B. and Hu, Xiao , title =. 2023 , eprint =
2023
-
[40]
and Heim, Lukas and Rodriguez, Marianela and Sandbrink, Jonas B
Kolt, Noam and Anderljung, Markus and Barnhart, Jess and Brass, Imogen and Esvelt, Kevin and Hadfield, Gillian K. and Heim, Lukas and Rodriguez, Marianela and Sandbrink, Jonas B. and Woodside, Tom , title =. Proceedings of the
-
[41]
2022 , eprint =
Bommasani, Rishi , title =. 2022 , eprint =
2022
-
[42]
2022 , eprint =
Hutchinson, Ben and Rostamzadeh, Negar and Greer, Christina and Heller, Katherine and Prabhakaran, Vinodkumar , title =. 2022 , eprint =
2022
-
[43]
2022 , eprint =
Javed, Syed Ashar and Juyal, Dinkar and Shanis, Zahil and Chakraborty, Shreya and Pokkalla, Harsha and Prakash, Aaditya , title =. 2022 , eprint =
2022
-
[44]
and Moons, Karel G
Collins, Gary S. and Moons, Karel G. M. and Dhiman, Paula and Riley, Richard D. and Beam, Andrew L. and. 2024 , doi =
2024
-
[45]
2023 , eprint =
Chandrasekaran, Jaganmohan and Cody, Tyler and McCarthy, Nicola and Lanus, Erin and Freeman, Laura , title =. 2023 , eprint =
2023
-
[46]
Journal of Machine Learning Research , volume =
Leiter, Christoph and Lertvittayakumjorn, Piyawat and Fomicheva, Marina and Zhao, Wei and Gao, Yang and Eger, Steffen , title =. Journal of Machine Learning Research , volume =. 2024 , eprint =
2024
-
[47]
2025 , eprint =
Weidinger, Laura and Raji, Inioluwa Deborah and Wallach, Hanna and Mitchell, Margaret and Wang, Angelina and Salaudeen, Olawale and Bommasani, Rishi and Ganguli, Deep and Koyejo, Sanmi and Isaac, William , title =. 2025 , eprint =
2025
-
[48]
2023 , eprint =
Costanza-Chock, Sasha and Harvey, Emma and Raji, Inioluwa Deborah and Czernuszenko, Martha and Buolamwini, Joy , title =. 2023 , eprint =
2023
-
[49]
, title =
Beddar-Wiesing, Silvia and Moallemy-Oureh, Alice and Kempkes, Marie and Thomas, Josephine M. , title =. 2025 , eprint =
2025
-
[50]
2025 , month =
The General-Purpose. 2025 , month =
2025
-
[51]
2025 , eprint =
Alizadeh, Morteza and Oveisi, Mehrdad and Falahati, Sonya and Mousavi, Ghazal and. 2025 , eprint =
2025
-
[52]
2025 , eprint =
Bilson, Samuel and Cox, Maurice and Pustogvar, Anna and Thompson, Andrew , title =. 2025 , eprint =
2025
-
[53]
Computational Materials Science , volume =
Alampara, Nawaf and Schilling-Wilhelmi, Mara and Jablonka, Kevin Maik , title =. Computational Materials Science , volume =. 2025 , doi =
2025
-
[54]
and Barnes, Elizabeth A
Ullrich, Paul A. and Barnes, Elizabeth A. and Collins, William D. and Dagon, Kate and Duan, Siyuan and Elms, Jacqueline and Lee, Jiwoo and Leung, L. Ruby and Lu, Dan and Molina, Michael J. and O'Brien, Travis A. and Rebassoo, Forrest O. , title =. Journal of Geophysical Resear...
2025
-
[55]
Verifiable Evaluations of Machine Learning Models Using
South, Tobin and Camuto, Alexander and Jain, Shrey and Nguyen, Shayla and Mahari, Robert and Paquin, Christian and Morton, Jason and Pentland, Alex. Verifiable Evaluations of Machine Learning Models Using. 2024 , eprint =
2024
-
[56]
Conformalizing Machine Translation Evaluation , journal =
Zerva, Chrysoula and Martins, Andr. Conformalizing Machine Translation Evaluation , journal =. 2024 , doi =
2024
-
[57]
, title =
Orzechowski, Patryk and Moore, Jason H. , title =. Science Advances , volume =. 2022 , doi =
2022
-
[58]
Evaluation of Machine Learning Algorithms for Health and Wellness Applications: A Tutorial , journal =
Tohka, Jussi and. Evaluation of Machine Learning Algorithms for Health and Wellness Applications: A Tutorial , journal =. 2021 , doi =
2021
-
[59]
Good Practices for Evaluation of Machine Learning Systems , year =
Ferrer, Luciana and Scharenborg, Odette and B. Good Practices for Evaluation of Machine Learning Systems , year =. 2412.03700 , archivePrefix=
-
[60]
2021 , eprint =
Whittlestone, Jess and Clark, Jack , title =. 2021 , eprint =
2021
-
[61]
and Paskov, Patricia and Dev, Sunishchal and Byun, Michael J
Wei, Kevin L. and Paskov, Patricia and Dev, Sunishchal and Byun, Michael J. and Reuel, Anka and Roberts-Gaal, Xavier and Calcott, Rachel and Coxon, Evie and Deshpande, Chinmay , title =. 2025 , eprint =
2025
-
[62]
2023 , month =
Emerging Processes for Frontier. 2023 , month =
2023
-
[63]
2025 , eprint =
Eriksson, Maria and Purificato, Erasmo and Noroozian, Arman and Vinagre, Joao and Chaslot, Guillaume and Gomez, Emilia and Fernandez-Llorca, David , title =. 2025 , eprint =
2025
-
[64]
2024 , eprint =
Longpre, Shayne and Kapoor, Sayash and Klyman, Kevin and Ramaswami, Aviya and Bommasani, Rishi and Blili-Hamelin, Borhane and Huang, Yangsibo and Skowron, Aleksander and Yong, Zheng-Xin and Kotha, Suhas and Zeng, Yi and Shi, Weiyan and Yang, Xianjun and Southen, Reid and Robey...
2024
-
[65]
2025 , month =
A Structured Protocol for Elicitation Experiments: Calibrating. 2025 , month =
2025
-
[66]
Early Insights from Developing Question-Answer Evaluations for Frontier AI , year =
-
[67]
2024 , eprint =
Miller, Evan , title =. 2024 , eprint =
2024
-
[68]
2025 , eprint =
Luko. 2025 , eprint =
2025
-
[69]
and Wei, Kevin and Webster, Toby , title =
Paskov, Patricia and Byun, Michael J. and Wei, Kevin and Webster, Toby , title =. 2025 , month =
2025
-
[70]
2025 , eprint =
Staufer, Leon and Yang, Mick and Reuel, Anka and Casper, Stephen , title =. 2025 , eprint =
2025
-
[71]
NeurIPS 2021 Datasets and Benchmarks Track , year =
Liao, Thomas and Taori, Rohan and Raji, Inioluwa Deborah and Schmidt, Ludwig , title =. NeurIPS 2021 Datasets and Benchmarks Track , year =
2021
-
[72]
Journal of Artificial Intelligence Research , volume =
Gehrmann, Sebastian and Clark, Elizabeth and Sellam, Thibault , title =. Journal of Artificial Intelligence Research , volume =. 2023 , doi =
2023
-
[73]
2025 , month =
Frontier Capability Assessment , institution =. 2025 , month =
2025
-
[74]
2025 , eprint =
McCaslin, Tegan and Alaga, Jide and Nedungadi, Samira and Donoughe, Seth and Reed, Tom and Bommasani, Rishi and Painter, Chris and Righetti, Luca , title =. 2025 , eprint =
2025
-
[75]
Measuring What Matters: Connecting
Rismani, Shalaleh and Shelby, Renee and Davis, Leah and Rostamzadeh, Negar and Moon,. Measuring What Matters: Connecting. 2025 , eprint =
2025
-
[76]
and Peng, Keiran and Pham, Thanh Huy and Bail, Christopher A
Kapoor, Sayash and Cantrell, Ethan M. and Peng, Keiran and Pham, Thanh Huy and Bail, Christopher A. and Gundersen, Odd Erik and Hofman, Jake M. and Hullman, Jessica and Lones, Michael A. and Malik, Meenal M. and Nanayakkara, Priyanka and Poldrack, Russell A. and Raji, Inioluwa...
2024
-
[77]
2024 , eprint =
Mizrahi, Moran and Kaplan, Guy and Malkin, Dan and Dror, Rotem and Shahaf, Dafna and Stanovsky, Gabriel , title =. 2024 , eprint =
2024
-
[78]
Black-Box Access Is Insufficient for Rigorous
Casper, Stephen and Ezell, Carson and Siegmann, Charlotte and Kolt, Noam and Curtis, Taylor Lynn and Bucknall, Benjamin and Haupt, Andreas and Wei, Kevin and Scheurer, J. Black-Box Access Is Insufficient for Rigorous. 2024 , eprint =. doi:10.1145/3630106.3659037 , url =
2024
-
[79]
and Fitz, Stephen and Hendrycks, Dan , title =
Ren, Richard and Basart, Steven and Khoja, Adam and Gatti, Alice and Phan, Long and Yin, Xuwang and Mazeika, Mantas and Pan, Alexander and Mukobi, Gabriel and Kim, Ryan H. and Fitz, Stephen and Hendrycks, Dan , title =. 2024 , eprint =
2024
-
[80]
Feder and Wang, Angelina and Barocas, Solon and Chouldechova, Alexandra and Atalla, Chad and Blodgett, Su Lin and Corvi, Emily and Dow, P
Wallach, Hanna and Desai, Meera and Pangakis, Nicholas and Cooper, A. Feder and Wang, Angelina and Barocas, Solon and Chouldechova, Alexandra and Atalla, Chad and Blodgett, Su Lin and Corvi, Emily and Dow, P. Alex and. Evaluating Generative. 2024 , eprint =
2024
-
[81]
2020 , isbn =
Cave, Stephen , title =. 2020 , isbn =. doi:10.1145/3375627.3375813 , booktitle =
2020
-
[82]
2025 , eprint =
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions , author =. 2025 , eprint =
2025
-
[83]
2025 , eprint =
Measuring what Matters: Construct Validity in Large Language Model Benchmarks , author =. 2025 , eprint =
2025
-
[84]
2504.07971 , archivePrefix=
Ma, Qianou and Zhao, Dora and Zhao, Xinran and Si, Chenglei and Yang, Chenyang and Louie, Ryan and Reiter, Ehud and Yang, Diyi and Wu, Tongshuang , year =. 2504.07971 , archivePrefix=
-
[85]
2025 , eprint =
Dhar, Ruchira and Sanchez Villegas, Danae and Karamolegkou, Antonia and Schiavone, Alice and Yuan, Yifei and Chen, Xinyi and Li, Jiaang and Frank, Stella and. 2025 , eprint =
2025
-
[86]
, year =
Gursoy, Furkan and Kakadiaris, Ioannis A. , year =. System Cards for. 2203.04754 , archivePrefix=
-
[87]
Checklist for Artificial Intelligence in Medical Imaging (
Mongan, John and Moy, Linda and. Checklist for Artificial Intelligence in Medical Imaging (. Radiology: Artificial Intelligence , year =
-
[88]
and Porras, Antonio R
Lekadir, Karim and Frangi, Alejandro F. and Porras, Antonio R. and Glocker, Ben and others , title =. 2025 , doi =
2025
-
[89]
Interactive Journal of Medical Research , volume =
Sallam, Malik and Barakat, Muna and Sallam, Maram , title =. Interactive Journal of Medical Research , volume =. 2024 , month = feb, doi =
2024
-
[90]
[European Construct Validity Checklist -- Reference Missing] , year =
-
[91]
Measurement to Meaning: A Validity-Centered Framework for
Salaudeen, Olawale and Reuel, Anka and Ahmed, Ahmed and Bedi, Suhana and Robertson, Zachary and Sundar, Sudharsan and Domingue, Ben and Wang, Angelina and Koyejo, Sanmi , year =. Measurement to Meaning: A Validity-Centered Framework for. 2505.10573 , archivePrefix=
-
[92]
2025 , url =
Evaluation of. 2025 , url =
2025
-
[93]
Benchmark Profiling: Mechanistic Diagnosis of
Kim, Dongjun and Shim, Gyuho and Chun, Yongchan and Kim, Minhyuk and Park, Chanjun and Lim, Heuiseok , year =. Benchmark Profiling: Mechanistic Diagnosis of. 2510.01232 , archivePrefix=
-
[94]
General Scales Unlock
Zhou, Lexin and Pacchiardi, Lorenzo and. General Scales Unlock. 2025 , eprint =
2025
-
[95]
2023 , eprint =
Extrinsic Evaluation of Machine Translation Metrics , author =. 2023 , eprint =
2023
-
[96]
2024 , eprint =
Stress-Testing Capability Elicitation With Password-Locked Models , author =. 2024 , eprint =
2024
-
[97]
2025 , eprint =
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning , author =. 2025 , eprint =
2025
-
[98]
Evaluating Machine Expertise: How Graduate Students Develop Frameworks for Assessing
Chen, Celia and Leitch, Alex , year =. Evaluating Machine Expertise: How Graduate Students Develop Frameworks for Assessing. 2504.17964 , archivePrefix=
-
[99]
and Soldaini, Luca and Soboroff, Ian and Weller, Orion and Kayi, Efsun and Sanders, Kate and Mason, Marc and Hibbler, Noah , year =
Mayfield, James and Yang, Eugene and Lawrie, Dawn and MacAvaney, Sean and McNamee, Paul and Oard, Douglas W. and Soldaini, Luca and Soboroff, Ian and Weller, Orion and Kayi, Efsun and Sanders, Kate and Mason, Marc and Hibbler, Noah , year =. On the Evaluation of Machine-Genera...
-
[100]
Sandeepa, A. G. R. and Mohottala, Sanka , year =. Evaluation of Machine Learning Models in Student Academic Performance Prediction , url =. doi:10.1109/icarc64760.2025.10963104 , booktitle =
2025
-
[101]
2011 , doi =
Carroll, Christopher and Booth, Andrew and Cooper, Katy , title =. 2011 , doi =
2011
-
[102]
2026 , month = jan, number =
Managing Misuse Risk for Dual-Use Foundation Models , institution =. 2026 , month = jan, number =
2026
-
[103]
2025 , note =
Carro, Mar\'. 2025 , note =
2025
-
[104]
1968 , month = sep, type =
Delphi Process: A Methodology Used for the Elicitation of Opinions of Experts , author =. 1968 , month = sep, type =
1968
-
[105]
2026 , eprint =
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats , author =. 2026 , eprint =
2026
-
[106]
Richard Landis and Gary G
J. Richard Landis and Gary G. Koch , journal =. The Measurement of Observer Agreement for Categorical Data , urldate =
-
[107]
2024 , howpublished =
Digital Education Council Global AI Student Survey 2024 , author =. 2024 , howpublished =
2024
-
[108]
Social Media + Society , volume =
Crystal Abidin , title =. Social Media + Society , volume =. 2021 , doi =. https://doi.org/10.1177/2056305120984458 , abstract =
2021 doi
-
[109]
2025 , eprint=
Configuration Work: Four Consequences of LLMs-in-use , author=. 2025 , eprint=
2025
-
[110]
, title =
Bachman, Lyle F. , title =. 1990 , publisher =
1990
-
[111]
and Han, F
Bahari, A. and Han, F. and Strzelecki, A. , journal =. Integrating. 2025 , doi =
2025
-
[112]
Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems , articleno =
Bai, Jingwen and Cheong, Wei Soon and Muller, Philippe and Lim, Brian Y , title =. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems , articleno =. 2026 , isbn =. doi:10.1145/3772318.3790539 , abstract =
2026
-
[113]
2001 , url=
Pragmatics in Language Teaching: Evaluating the empirical evidence: Grounds for instruction in pragmatics? , author=. 2001 , url=
2001
-
[114]
2024 , note =
Bastani, Hamsa and Bastani, Osbert and Sungu, Alp and Ge, Haoyang and Kabakci, Ozge and Mariman, Rani , title =. 2024 , note =
2024
-
[115]
2024 , eprint=
Lessons from the Trenches on Reproducible Evaluation of Language Models , author=. 2024 , eprint=
2024
-
[116]
Biesta, Gert J. J. , title =. European Journal of Education , volume =
-
[117]
Biesta, Gert J. J. , title =. 2010 , edition =
2010
-
[118]
and Bjork, Robert A
Bjork, Elizabeth L. and Bjork, Robert A. , title =. Psychology and the Real World: Essays Illustrating Fundamental Contributions to Society , editor =. 2011 , publisher =
2011
-
[119]
2007 , publisher =
Block, David , title =. 2007 , publisher =
2007
-
[120]
2026 , eprint=
Who Decides in AI-Mediated Learning? The Agency Allocation Framework , author=. 2026 , eprint=
2026
-
[121]
2026 , eprint=
Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus , author=. 2026 , eprint=
2026
-
[122]
Teaching for Success: Continuing Professional Development (CPD) for Teachers , year =
-
[123]
NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling , year=
Precursors, Proxies, and Predictive Models for Long-Horizon Tasks , author=. NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling , year=
2025
-
[124]
The BEA-2019 Shared Task on Grammatical Error Correction , booktitle =
Bryant, Christopher and Felice, Mariano and Andersen,. The BEA-2019 Shared Task on Grammatical Error Correction , booktitle =. 2019 , publisher =. doi:10.18653/v1/W19-4406 , url =
2019 doi
-
[125]
Applied Linguistics , volume =
Canale, Michael and Swain, Merrill , title =. Applied Linguistics , volume =. 1980 , doi =
1980
-
[126]
Building a Validity Argument for the Test of English as a Foreign Language , year =
-
[127]
2001 , publisher =
Computer Applications in Second Language Acquisition , author =. 2001 , publisher =
2001
-
[128]
2021 , eprint=
Evaluating Large Language Models Trained on Code , author=. 2021 , eprint=
2021
-
[129]
2018 , eprint=
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge , author=. 2018 , eprint=
2018
-
[130]
Regents Science Exams: An Overview of the Aristo Project , author=
From 'F' to 'A' on the N.Y. Regents Science Exams: An Overview of the Aristo Project , author=. 2021 , eprint=
2021
-
[131]
2025 , institution=
Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources , author=. 2025 , institution=
2025
-
[132]
2021 , eprint=
Training Verifiers to Solve Math Word Problems , author=. 2021 , eprint=
2021
-
[133]
1992 , edition =
Cook, Guy , title =. 1992 , edition =
1992
-
[134]
and Campbell, Donald T
Cook, Thomas D. and Campbell, Donald T. , title =. 1979 , publisher =
1979
-
[135]
, year =
Costa-Gomes, Bruno and Chen, Siqi and Hsueh, Chia-Hung and Morgan, David and Schoenegger, Philipp and Shah, Yash and Way, Sam and Zhu, Yiming and Spielman, Seth and Suleyman, Mustafa and Bhaskar, M. , year =. It’s about time: The
-
[136]
The Common European Framework of Reference for Languages: Learning, Teaching, Assessment---Companion Volume with New Descriptors , year =
-
[137]
Crossley and Perpetual Baffour and Yu Tian and Aigner Picou and Meg Benner and Ulrich Boser , keywords =
Scott A. Crossley and Perpetual Baffour and Yu Tian and Aigner Picou and Meg Benner and Ulrich Boser , keywords =. The persuasive essays for rating, selecting, and understanding argumentative and discourse elements (PERSUADE) corpus 1.0 , journal =. 2022 , issn =. doi:https://...
2022
-
[138]
2001 , publisher =
Cuban, Larry , title =. 2001 , publisher =
2001
-
[139]
2007 , publisher =
DeKeyser, Robert , title =. 2007 , publisher =
2007
-
[140]
and Paskov, P
Dev, S. and Paskov, P. and Sloan, A. and Wei, K. and Nascimento de Lima, P. and Chowdhury, S. and Johnson, J. and Marcellino, W. , title =. 2026 , url =
2026
-
[141]
Suzuki, Yuichi and DeKeyser, Robert , year =
-
[142]
The Psychology of Second Language Acquisition , year =
D. The Psychology of Second Language Acquisition , year =
-
[143]
Proceedings of the Sixth International Conference on Learning Analytics & Knowledge , pages =
Drachsler, Hendrik and Greller, Wolfgang , title =. Proceedings of the Sixth International Conference on Learning Analytics & Knowledge , pages =. 2016 , isbn =. doi:10.1145/2883851.2883893 , abstract =
2016
-
[144]
2026 , eprint=
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback , author=. 2026 , eprint=
2026
-
[145]
The Eaquals Framework for Language Teacher Training and Development , year =
-
[146]
2026 , eprint=
Beyond Accuracy: Towards a Robust Evaluation Methodology for AI Systems for Language Education , author=. 2026 , eprint=
2026
-
[147]
Proceedings of the
Eriksson, Maria and Purificato, Elisa and Noroozian, Amir and Vinagre, Joao and Chaslot, Guillaume and Gomez, Emilia and Fernandez-Llorca, David , title =. Proceedings of the. 2025 , doi =
2025
-
[148]
Educationalization , booktitle =
Fendler, Lynn , editor =. Educationalization , booktitle =. 2018 , publisher =. doi:10.1007/978-3-319-72761-5_81 , url =
2018 doi
-
[149]
, title =
Frick, Theodore W. , title =. TechTrends , volume =. 2024 , doi =
2024
-
[150]
2026 , eprint=
Not Too Short, Not Too Long: How LLM Response Length Shapes People's Critical Thinking in Error Detection , author=. 2026 , eprint=
2026
-
[151]
2026 , eprint=
Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles , author=. 2026 , eprint=
2026
-
[152]
AI, Brain and Child , volume =
Gao, Jie and Cohrssen, Caroline , title =. AI, Brain and Child , volume =. 2026 , doi =
2026
-
[153]
2014 , url=
Automatic Linguistic Annotation ofLarge Scale L2 Databases: The EF-Cambridge Open Language Database(EFCamDat) , author=. 2014 , url=
2014
-
[154]
2026 , month = apr, howpublished =
Ghosh, Avijit and Mai, Yifan and Channing, Georgia and Choshen, Leshem , title =. 2026 , month = apr, howpublished =
2026
-
[155]
Language Learning & Technology , volume =
Godwin-Jones, Robert , title =. Language Learning & Technology , volume =. 2022 , doi =
2022
-
[156]
Guha, Neel and Nyarko, Julian and Ho, Daniel E. and R. LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models , booktitle =. 2023 , doi =
2023
-
[157]
Education Sciences , VOLUME =
Grassini, Simone , TITLE =. Education Sciences , VOLUME =. 2023 , NUMBER =
2023
-
[158]
ChatGPT in education: A blessing or a curse? A qualitative study exploring early adopters’ utilization and perceptions , journal =
Reza. ChatGPT in education: A blessing or a curse? A qualitative study exploring early adopters’ utilization and perceptions , journal =. 2024 , issn =. doi:https://doi.org/10.1016/j.chbah.2023.100027 , url =
2024
-
[159]
2026 , eprint=
Learning Context Matters: Measuring and Diagnosing Personalization Gaps in LLM-Based Instructional Design , author=. 2026 , eprint=
2026
-
[160]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =
He, Chaoqun and Luo, Renjie and Bai, Yuzhuo and Hu, Shengding and Thai, Zhen and Shen, Junhao and Hu, Jinyi and Han, Xu and Huang, Yujie and Zhang, Yuxiang and Liu, Jie and Qi, Lei and Liu, Zhiyuan and Sun, Maosong , title =. Proceedings of the 62nd Annual Meeting of the Assoc...
2024 doi
-
[161]
2021 , eprint=
Measuring Massive Multitask Language Understanding , author=. 2021 , eprint=
2021
-
[162]
Holmes, Wayne and Bialik, Maya and Fadel, Charles , title =
-
[163]
Ethics of AI in Education: Towards a Community-Wide Framework , journal =
Holmes, Wayne and Porayska-Pomsta, Ka\'. Ethics of AI in Education: Towards a Community-Wide Framework , journal =. 2022 , doi =
2022
-
[164]
International Journal of Artificial Intelligence in Education , volume =
Holmes, Wayne , title =. International Journal of Artificial Intelligence in Education , volume =. 2024 , pages =
2024
-
[165]
npj Science of Learning , volume =
Hu, Bihao and Zhu, Jiayi and Pei, Yiying and Gu, Xiaoqing , title =. npj Science of Learning , volume =. 2025 , doi =
2025
-
[166]
Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems , articleno =
Hudig, Anna Ida and Kallina, Emma and Singh, Jatinder , title =. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems , articleno =. 2026 , isbn =. doi:10.1145/3772318.3791968 , abstract =
2026
-
[167]
2026 , eprint=
EduEVAL-DB: A Role-Based Dataset for Pedagogical Risk Evaluation in Educational Explanations , author=. 2026 , eprint=
2026
-
[168]
, title =
Ishikawa, Shunji I. , title =. Learner Corpus Studies in Asia and the World , editor =
-
[169]
Jacobson and Uri Wilensky , title =
Michael J. Jacobson and Uri Wilensky , title =. Journal of the Learning Sciences , volume =. 2006 , publisher =. doi:10.1207/s15327809jls1501\_4 , URL =
2006 doi
-
[170]
2024 , eprint=
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code , author=. 2024 , eprint=
2024
-
[171]
On Assessing the Faithfulness of LLM-generated Feedback on Student Assignments , year =
Qinjin Jia and Jialin Cui and Ruijie Xi and Chengyuan Liu and Parvez Rashid and Ruochi Li and Edward Gehringer , booktitle =. On Assessing the Faithfulness of LLM-generated Feedback on Student Assignments , year =. doi:10.5281/zenodo.12729868 , editor =
-
[172]
and Gillick, Daniel and others , title =
Jurenka, Ivan and Kunesch, Matthias and McKee, Kyle R. and Gillick, Daniel and others , title =. 2024 , note =
2024
-
[173]
Educational Psychology Review , volume =
Kalyuga, Slava , title =. Educational Psychology Review , volume =. 2007 , doi =
2007
-
[174]
ChatGPT for good? On opportunities and challenges of large language models for education , journal =
Enkelejda Kasneci and Kathrin Sessler and Stefan Küchemann and Maria Bannert and Daryna Dementieva and Frank Fischer and Urs Gasser and Georg Groh and Stephan Günnemann and Eyke Hüllermeier and Stephan Krusche and Gitta Kutyniok and Tilman Michaeli and Claudia Nerdel and Jürge...
2023
-
[175]
Matthew and Campos, Daniel Vargas , title =
Kennedy, Wm. Matthew and Campos, Daniel Vargas , title =. Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society , pages =. 2025 , publisher =
2024
-
[176]
Matthew and Vargas Campos, Daniel , title =
Kennedy, Wm. Matthew and Vargas Campos, Daniel , title =. Handbook of Critical Studies in AI for Education , editor =. 2026 , note =
2026
-
[177]
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , year =
Beigman Klebanov, Beata and Madnani, Nitin , title =. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , year =. doi:10.18653/v1/2020.acl-main.697 , url =
2020 doi
-
[178]
Maurya, Kaushal Kumar and Srivatsa, Kv Aditya and Petukhova, Kseniia and Kochmar, Ekaterina , title =. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers...
2025 doi
-
[179]
The NarrativeQA Reading Comprehension Challenge , journal =
Ko. The NarrativeQA Reading Comprehension Challenge , journal =. 2018 , publisher =. doi:10.1162/tacl_a_00023 , url =
2018 doi
-
[180]
and Corbett, Albert T
Koedinger, Kenneth R. and Corbett, Albert T. and Perfetti, Charles , title =. Cognitive Science , volume =. doi:https://doi.org/10.1111/j.1551-6709.2012.01245.x , url =. https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1551-6709.2012.01245.x , abstract =
2012
-
[181]
RELC Journal , volume =
Lucas Kohnke and Benjamin Luke Moorhouse and Di Zou , title =. RELC Journal , volume =. 2023 , doi =. https://doi.org/10.1177/00336882231162868 , abstract =
2023 doi
-
[182]
Krathwohl , title =
David R. Krathwohl , title =. Theory Into Practice , volume =. 2002 , publisher =. doi:10.1207/s15430421tip4104\_2 , URL =
2002 doi
-
[183]
Can ChatGPT support prospective teachers in physics task development? , author =. Phys. Rev. Phys. Educ. Res. , volume =. 2023 , month =. doi:10.1103/PhysRevPhysEducRes.19.020128 , url =
2023 doi
-
[184]
Kulik and J
James A. Kulik and J. D. Fletcher , title =. Review of Educational Research , volume =. 2016 , doi =. https://doi.org/10.3102/0034654315581420 , abstract =
2016 doi
-
[185]
2022 , eprint=
DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation , author=. 2022 , eprint=
2022
-
[186]
, title =
De Costa, Peter I. , title =. Applied Linguistics , volume =. 2007 , month =. doi:10.1093/applin/amm027 , url =
2007 doi
-
[187]
2003 , publisher =
Larsen-Freeman, Diane , title =. 2003 , publisher =
2003
-
[188]
2025 , eprint=
AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms , author=. 2025 , eprint=
2025
-
[189]
Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , year =
Lee, Changyoon and Seonwoo, Yeon and Oh, Alice , title =. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , year =. doi:10.18653/v1/2022.naacl-main.148 , url =
2022 doi
-
[190]
2025 , eprint=
Benchmarking the Pedagogical Knowledge of Large Language Models , author=. 2025 , eprint=
2025
-
[191]
Transactions on Machine Learning Research , issn=
Holistic Evaluation of Language Models , author=. Transactions on Machine Learning Research , issn=. 2023 , url=
2023
-
[192]
Applications of Transformer Neural Networks in Processing Examinee Responses , booktitle =
Lottridge, Susan , editor =. Applications of Transformer Neural Networks in Processing Examinee Responses , booktitle =. 2024 , month =. doi:10.1108/979-8-88730-606-320251003 , url =
2024 doi
-
[193]
Advances in Neural Information Processing Systems , editor=
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering , author=. Advances in Neural Information Processing Systems , editor=. 2022 , url=
2022
-
[194]
, title =
Luckin, Rose and Holmes, Wayne and Griffiths, Mark and Forcier, Louise B. , title =. 2016 , publisher =
2016
-
[195]
Studies in Second Language Acquisition , author=
CORRECTIVE FEEDBACK AND LEARNER UPTAKE: Negotiation of Form inCommunicative Classrooms , volume=. Studies in Second Language Acquisition , author=. 1997 , pages=. doi:10.1017/S0272263197001034 , number=
1997 doi
-
[196]
The 2023 Conference on Empirical Methods in Natural Language Processing , year=
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems , author=. The 2023 Conference on Empirical Methods in Natural Language Processing , year=
2023
-
[197]
and CLÉMENT, RICHARD and DÖRNYEI, ZOLTÁN and NOELS, KIMBERLY A
MACINTYRE, PETER D. and CLÉMENT, RICHARD and DÖRNYEI, ZOLTÁN and NOELS, KIMBERLY A. , title =. The Modern Language Journal , volume =. doi:https://doi.org/10.1111/j.1540-4781.1998.tb05543.x , url =. https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1540-4781.1998.tb05543.x , a...
1998
-
[198]
Artificial Intelligence in Education: 24th International Conference, AIED 2023, Tokyo, Japan, July 3–7, 2023, Proceedings , pages =
McNichols, Hunter and Zhang, Mengxue and Lan, Andrew , title =. Artificial Intelligence in Education: 24th International Conference, AIED 2023, Tokyo, Japan, July 3–7, 2023, Proceedings , pages =. 2023 , isbn =. doi:10.1007/978-3-031-36272-9_30 , abstract =
2023 doi
-
[199]
American Psychologist , volume =
Messick, Samuel , title =. American Psychologist , volume =. 1995 , doi =
1995
-
[200]
Liebenow and Marlene Steinbach and Andrea Horbach and Johanna Fleckenstein , keywords =
Jennifer Meyer and Thorben Jansen and Ronja Schiller and Lucas W. Liebenow and Marlene Steinbach and Andrea Horbach and Johanna Fleckenstein , keywords =. Using LLMs to bring evidence-based feedback into the classroom: AI-generated feedback increases secondary students’ text r...
2024
-
[201]
and Steinberg, Linda S
Mislevy, Robert J. and Steinberg, Linda S. and Almond, Russell G. , title =. Measurement: Interdisciplinary Research and Perspectives , volume =. 2003 , doi =
2003
-
[202]
Exploring the potential of using an AI language model for automated essay scoring , journal =
Atsushi Mizumoto and Masaki Eguchi , keywords =. Exploring the potential of using an AI language model for automated essay scoring , journal =. 2023 , issn =. doi:https://doi.org/10.1016/j.rmal.2023.100050 , url =
2023
-
[203]
Human Feedback Shapes Learner Engagement , author=
Same Feedback, Different Source: How AI vs. Human Feedback Shapes Learner Engagement , author=. 2026 , eprint=
2026
-
[204]
Proceedings of the Eighteenth Conference on Computational Natural Language Learning: Shared Task , year =
Ng, Hwee Tou and Wu, Siew Mei and Briscoe, Ted and Hadiwinoto, Christian and Susanto, Raymond Hendy and Bryant, Christopher , title =. Proceedings of the Eighteenth Conference on Computational Natural Language Learning: Shared Task , year =. doi:10.3115/v1/W14-1701 , url =
-
[205]
1999 , url=
The Cambridge Learner Corpus-Error coding and analysis , author=. 1999 , url=
1999
-
[206]
Proceedings of the Twelfth ACM Conference on Learning @ Scale , pages =
Nie, Allen and Chandak, Yash and Suzara, Miroslav and Malik, Ali and Woodrow, Juliette and Peng, Matt and Sahami, Mehran and Brunskill, Emma and Piech, Chris , title =. Proceedings of the Twelfth ACM Conference on Learning @ Scale , pages =. 2025 , isbn =. doi:10.1145/3698205....
2025
-
[207]
2026 , eprint=
CurricuLLM: Designing Personalized and Workforce-Aligned Cybersecurity Curricula Using Fine-Tuned LLMs , author=. 2026 , eprint=
2026
-
[208]
2013 , publisher =
Norton, Bonny , title =. 2013 , publisher =
2013
-
[209]
2009 , edition =
Ortega, Lourdes , title =. 2009 , edition =
2009
-
[210]
Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , year =
Pang, Richard Yuanzhe and Parrish, Alicia and Joshi, Nitish and Nangia, Nikita and Phang, Jason and Chen, Angelica and Padmakumar, Vishakh and Ma, Johnny and Thompson, Jana and He, He and Bowman, Samuel , title =. Proceedings of the 2022 Conference of the North American Chapte...
2022 doi
-
[211]
Annual Review of Applied Linguistics , volume =
Second Language Anxiety: Construct, Effects, and Sources , author =. Annual Review of Applied Linguistics , volume =. 2023 , doi =
2023
-
[212]
2023 , eprint=
Learning gain differences between ChatGPT and human tutor generated algebra hints , author=. 2023 , eprint=
2023
-
[213]
2025 , month =
Paskov, Patricia and Soder, Lisa and Smith, Everett , title =. 2025 , month =
2025
-
[214]
Education for Life and Work: Developing Transferable Knowledge and Skills in the 21st Century , year =
-
[215]
the state of the art
Perelman, Les , year =. When “the state of the art” is counting words , volume =. Assessing Writing , doi =
-
[216]
and Askell, Amanda and Grosse, Roger and Hernandez, Danny and Ganguli, Deep and Hubinger, Evan and Schiefer, Nicholas and Kaplan, Jared , title =
Perez, Ethan and Ringer, Sam and Lukosiute, Kamile and Nguyen, Karina and Chen, Edwin and Heiner, Scott and Pettit, Craig and Olsson, Catherine and Kundu, Sandipan and Kadavath, Saurav and Jones, Andy and Chen, Anna and Mann, Benjamin and Israel, Brian and Seethor, Bryan and M...
2023
-
[217]
Coursebook Texts as a Helping Hand for Classifying Linguistic Complexity in Language Learners' Writings , booktitle =
Pil. Coursebook Texts as a Helping Hand for Classifying Linguistic Complexity in Language Learners' Writings , booktitle =. 2016 , publisher =
2016
-
[218]
2003 , publisher =
Vygotsky's Educational Theory in Cultural Context , series =. 2003 , publisher =
2003
-
[219]
Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , year =
Rajpurkar, Pranav and Zhang, Jian and Lopyrev, Konstantin and Liang, Percy , title =. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , year =. doi:10.18653/v1/D16-1264 , url =
2016 doi
-
[220]
Computer Assisted Language Learning , volume =
Jim Ranalli , title =. Computer Assisted Language Learning , volume =. 2018 , publisher =. doi:10.1080/09588221.2018.1428994 , URL =
2018
-
[221]
2024 , booktitle=
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices , author=. 2024 , booktitle=
2024
-
[222]
2025 , eprint=
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations , author=. 2025 , eprint=
2025
-
[223]
2025 , eprint=
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks , author=. 2025 , eprint=
2025
-
[224]
Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , articleno =
Sambasivan, Nithya and Kapania, Shivani and Highfill, Hannah and Akrong, Diana and Paritosh, Praveen and Aroyo, Lora M , title =. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , articleno =. 2021 , isbn =. doi:10.1145/3411764.3445518 , abstract =
2021
-
[225]
2025 , eprint=
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects , author=. 2025 , eprint=
2025
-
[226]
Using Interactive Feedback to Improve the Accuracy and Explainability of Question Answering Systems Post-Deployment , url=
Li, Zichao and Sharma, Prakhar and Lu, Xing Han and Cheung, Jackie and Reddy, Siva , year=. Using Interactive Feedback to Improve the Accuracy and Explainability of Question Answering Systems Post-Deployment , url=. doi:10.18653/v1/2022.findings-acl.75 , booktitle=
2022 doi
-
[227]
1996 , address =
Universal Declaration of Linguistic Rights , author =. 1996 , address =
1996
-
[228]
2012 , edition =
Selwyn, Neil , title =. 2012 , edition =
2012
-
[229]
Selwyn, Neil , title =
-
[230]
European Journal of Education , volume =
Selwyn, Neil , title =. European Journal of Education , volume =. 2022 , doi =
2022
-
[231]
Handbook of Automated Essay Evaluation: Current Applications and New Directions , year =
-
[232]
2012 , url=
Contrasting state-of-the-art automated scoring of essays: analysis , author=. 2012 , url=
2012
-
[233]
An Evaluation of Khanmigo, a Generative
Shetye, Shamini , journal =. An Evaluation of Khanmigo, a Generative. 2024 , institution =
2024
-
[234]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =
Shi, Yao and Liang, Rongkeng and Xu, Yong , title =. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =. doi:10.18653/v1/2025.acl-long.1576/ , url =
2025 doi
-
[235]
Shute , title =
Valerie J. Shute , title =. Review of Educational Research , volume =. 2008 , doi =. https://doi.org/10.3102/0034654307313795 , abstract =
2008 doi
-
[236]
Slavin , title =
Robert E. Slavin , title =. Journal of Education for Students Placed at Risk (JESPAR) , volume =. 2017 , publisher =. doi:10.1080/10824669.2017.1334560 , URL =
2017
-
[237]
2024 , eprint=
Socially Responsible Data for Large Multilingual Language Models , author=. 2024 , eprint=
2024
-
[238]
2024 , issue_date =
Stadler, Matthias and Bannert, Maria and Sailer, Michael , title =. 2024 , issue_date =. doi:10.1016/j.chb.2024.108386 , journal =
2024
-
[239]
Automated Feedback and Second Language Writing , booktitle=
Stevenson, Marie and Phakiti, Aek , editor=. Automated Feedback and Second Language Writing , booktitle=. 2019 , pages=
2019
-
[240]
, title =
Stobart, Gordon and Boyd, Elaine and Green, Anthony and Hopfenbeck, Therese N. , title =. 2019 , publisher =
2019
-
[241]
2024 , eprint=
SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research , author=. 2024 , eprint=
2024
-
[242]
Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , year =
Tafjord, Oyvind and Dalvi, Bhavana and Clark, Peter , title =. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , year =. doi:10.18653/v1/2021.findings-acl.317 , url =
2021 doi
-
[243]
and Mueller, Jared and McEachen, William and Mitchell, Will and Carter, Sam and Clark, Jack and Kaplan, Jared and Ganguli, Deep , title =
Tamkin, Alex and McCain, Matthew and Handa, Kanishka and Durmus, Esin and Lovitt, Liane and Rathi, Anushree and Huang, Stephanie and Mountfield, Austin and Hong, Justin and Ritchie, Sam and Stern, Michael and Clarke, Ben and Goldberg, Leo and Sumers, Theodore R. and Mueller, J...
2024
-
[244]
2026 , eprint=
Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education , author=. 2026 , eprint=
2026
-
[245]
2023 , publisher =
Holmes, Wayne and Miao, Fengchun , title =. 2023 , publisher =
2023
-
[246]
2025 , address =
AI and the Future of Education: Disruptions, Dilemmas and Directions , institution =. 2025 , address =
2025
-
[247]
Royal Society Open Science , volume =
Wachter, Sandra and Mittelstadt, Brent and Russell, Chris , title =. Royal Society Open Science , volume =. 2024 , doi =
2024
-
[248]
2024 , url=
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models , author=. 2024 , url=
2024
-
[249]
2024 , url=
Xingyao Wang and Zihan Wang and Jiateng Liu and Yangyi Chen and Lifan Yuan and Hao Peng and Heng Ji , booktitle=. 2024 , url=
2024
-
[250]
and Gardner, Matt , title =
Welbl, Johannes and Liu, Nelson F. and Gardner, Matt , title =. Proceedings of the 3rd Workshop on Noisy User-generated Text , year =. doi:10.18653/v1/W17-4413 , url =
-
[251]
Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , year =
Xie, Qizhe and Lai, Guokun and Dai, Zihang and Hovy, Eduard , title =. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , year =. doi:10.18653/v1/D18-1257 , url =
2018 doi
-
[252]
2026 , eprint=
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios , author=. 2026 , eprint=
2026
-
[253]
2024 , month = aug, day =
Yan, Eugene , title =. 2024 , month = aug, day =
2024
-
[254]
and Laflair, Geoffrey and Verardi, Anthony and Burstein, Jill , title =
Yancey, Kevin P. and Laflair, Geoffrey and Verardi, Anthony and Burstein, Jill , title =. Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023) , year =. doi:10.18653/v1/2023.bea-1.49 , url =
2023 doi
-
[255]
, title =
Yang, Zhilin and Qi, Peng and Zhang, Saizheng and Bengio, Yoshua and Cohen, William and Salakhutdinov, Ruslan and Manning, Christopher D. , title =. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , year =. doi:10.18653/v1/D18-1259 , url =
2018 doi
-
[256]
Systematic review of research on artificial intelligence applications in higher education -- where are the educators? , journal =
Zawacki-Richter, Olaf and Mar. Systematic review of research on artificial intelligence applications in higher education -- where are the educators? , journal =. 2019 , doi =
2019
-
[257]
ETS Research Report Series , volume =
Zechner, Klaus and Chen, Lei and Davis, Larry and Evanini, Keelan and Lee, Chong Min and Leong, Chee Wee and Wang, Xinhao and Yoon, Su-Youn , title =. ETS Research Report Series , volume =. doi:https://doi.org/10.1002/ets2.12080 , url =. https://onlinelibrary.wiley.com/doi/pdf...
-
[258]
2026 , eprint=
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation , author=. 2026 , eprint=
2026
-
[259]
2024 , eprint=
FinBen: A Holistic Financial Benchmark for Large Language Models , author=. 2024 , eprint=
2024
-
[260]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , title =. Proceedings of the 37th International Conf...
2023
-
[261]
Findings of the Association for Computational Linguistics: NAACL 2024 , year =
Zhong, Wanjun and Cui, Ruixiang and Guo, Yiduo and Liang, Yaobo and Lu, Shuai and Wang, Yanlin and Saied, Amin and Chen, Weizhu and Duan, Nan , title =. Findings of the Association for Computational Linguistics: NAACL 2024 , year =. doi:10.18653/v1/2024.findings-naacl.149 , url =
2024 doi
-
[262]
2023 , eprint=
Instruction-Following Evaluation for Large Language Models , author=. 2023 , eprint=
2023
-
[263]
2025 , month = dec, day =
Blanco, Cindy , title =. 2025 , month = dec, day =
2025
-
[264]
2014 , address =
Cambridge. 2014 , address =
2014
-
[265]
2023 , month = oct, type =
Millin, Sandy , title =. 2023 , month = oct, type =
2023
-
[266]
2026 , eprint=
Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials , author=. 2026 , eprint=
2026
-
[267]
and Rodrigo, Ma
Ocumpaugh, Jodi and Baker, Ryan S. and Rodrigo, Ma. Mercedes T. , title =. 2015 , address =
2015
-
[268]
Advances in Neural Information Processing Systems (NeurIPS) , year=
LLM Evaluators Recognize and Favor Their Own Generations , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=
-
[269]
2024 , eprint=
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge , author=. 2024 , eprint=
2024
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.