Pith. sign in

REVIEW 3 major objections 6 minor 52 references

AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics

T0 review · 3 major / 6 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read AI scoring systematically underestimates conceptual understanding when physics explanations are linguistically weak.

desk verdict Consistent underestimation of lower-linguistic-quality physics explanations by 11 AI scorers vs dual-expert ratings; solid design, bounded by single-task N and non-ground-truth experts. read the letter →

arxiv 2607.28210 v1 pith:B35IWALP submitted 2026-07-30 physics.ed-ph

classification physics.ed-ph
keywords AI-basedscoringlanguagebiasconceptualunderstandingphysicseducationautomatedassessmentmultilinguallearnerslargemodelstext-basedexplanations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether AI can score students’ conceptual understanding in physics without being swayed by how well the student writes. Using 116 authentic secondary-school explanations of a sound-transmission task, the authors compared eleven AI scorers—nine machine-learning pipelines and two large language models—against trained experts who had rated conceptual understanding and linguistic quality separately. Agreement with experts was generally good, yet every AI approach was more likely to under-score explanations of lower linguistic quality, while higher linguistic quality rarely produced over-scoring. The pattern matches bias previously seen in physics teachers, which the authors read as a feature of the assessment task itself: understanding is only reachable through language. The practical stakes are highest for multilingual learners and rise as AI moves into higher-stakes grading.

What carries the argument

Scoring-deviation direction (underestimation, agreement, or overestimation of expert conceptual-understanding scores), modelled with multinomial logistic regression on expert linguistic-quality scores while covarying expert conceptual-understanding scores to block scale-induced confounding.

What would settle it

Hold demonstrated conceptual content fixed while systematically degrading only lexical and syntactic quality of the same explanations, and test whether AI underestimation relative to the same experts still rises with lower linguistic quality.

Watch

Extended reading notes

Core claim

Across all eleven AI-based scoring approaches, explanations rated lower in linguistic quality by experts were systematically more likely to receive lower conceptual-understanding scores from the AI than from experts—roughly 1.3 to 3.9 times higher odds of underestimation per one-unit drop in linguistic quality—whereas higher linguistic quality showed no comparable link to overestimation in most approaches.

Load-bearing premise

That expert conceptual-understanding scores are independent enough of language that AI–expert underestimations can be attributed to AI language bias rather than shared entanglement of the two constructs.

Editorial extensions

If this is right

  • Multilingual and international learners risk systematic disadvantage if AI scoring scales into summative decisions.
  • Agreement with expert ratings is necessary but not sufficient; evaluations must also check whether disagreements track construct-irrelevant linguistic quality.
  • Human-in-the-loop alone may not remove the bias where teachers and AI err on the same texts; routing low-linguistic-quality explanations for mindful human review is a proposed safeguard.
  • Language bias should be treated as intrinsic to inferring conceptual understanding from text-based explanations, not as a quirk of one scorer type.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same underestimation pattern is likely in automated scoring of open explanations across other sciences and writing-heavy STEM assessments.
  • Fairness audits for educational AI should treat overall lexical–syntactic sophistication—not only spelling or surface grammar—as a construct-irrelevant risk factor.
  • Decoupling prompts or specialized embeddings may shrink but not erase the bias if conceptual evidence is itself carried by linguistic form.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript tests whether AI-based scorers of physics explanations can separate conceptual understanding (CU) from linguistic quality (LQ). Using 116 authentic Grade-9 German explanations of a spacewalk/sound-transmission task, each dual-rated by trained experts on independent CU (0–2) and LQ (0–4) schemes, the authors compare nine ML pipelines (3 embeddings × 3 classifiers) and two LLM prompters (GPT-4.1, GPT-5-mini) to expert CU. Agreement is moderate-to-good (67–78%). Multinomial logistic regressions of deviation direction (under / agree / over), with expert CU as covariate, show that lower LQ systematically raises the odds of underestimation relative to experts across all eleven approaches (~1.3–3.9× per LQ unit; five coefficients p<0.05, two p≈0.05), while higher LQ shows no comparable link to overestimation except in two science-embedding ML pipelines. The authors interpret this as a language bias resembling that previously reported for physics teachers and discuss fairness stakes for multilingual learners.

Significance. If the pattern holds, the result is consequential for physics education assessment: it shows that both classical ML and modern LLM scorers reproduce a construct-irrelevant LQ–underestimation association even when prompted or trained to score content, and that standard agreement-with-experts metrics miss this. Strengths include dual expert schemes with strong reliability (Krippendorff α≈0.91–0.92; Kendall W≈0.91–0.94), CU controlled as covariate to block scale-induced confounding, explicit handling of LLM stochasticity via 1,000 nested resamples and Rubin pooling, full coefficient tables (S2–S3), and directional consistency across heterogeneous scoring architectures. The link to teacher language bias and the fairness framing for multilingual learners are timely as AI scoring moves toward higher-stakes use. The work is a solid empirical contribution to physics education research and AI-in-assessment, not a methods breakthrough.

major comments (3)
  1. [Methods §4.5; Tables S2–S3; Table 1] Methods §4.5 and Tables S2–S3: several multinomial models show signs of separation or sparse-cell instability (e.g., underestimation intercepts ≈ −20.7 to −20.8 with SE ≈ 68–73 for German_Semantic RVM and paraphrase-multilingual MLR; correspondingly extreme CU ORs). With N=116, a three-category outcome, and CU as covariate, cell counts for underestimation are small (Table 1: often ~6–15%). The directional LQ pattern is still visible, but the claim that bias “emerged across every” approach should be qualified by model diagnostics (e.g., reporting underestimation base rates, checking Firth/penalized multinomial or collapsing rare Δ=±2, and flagging unstable fits). Without that, some of the larger 1/OR values (up to ~3.9) and non-significant but directionally consistent coefficients are hard to interpret at face value.
  2. [Discussion; Methods §4.2] Discussion and Methods §4.2: the central quantity is Δ = AI CU − expert CU, so “underestimation” is defined relative to experts who, by the paper’s own framing, face the same inferential entanglement of language and understanding. High inter-rater α/W and separate schemes make the expert reference the best available dual-trained standard, and the CU covariate correctly blocks the main mechanical confound. Still, residual shared language influence cannot be ruled out, so the load-bearing interpretation should be stated more tightly as: AI–expert negative deviations are systematically associated with lower LQ—not as a pure demonstration that AI alone underestimates true CU. A short explicit statement of what would falsify the AI-specific reading (e.g., experimental LQ manipulation holding CU evidence fixed, as the authors themselves propose) would strengthen the claim without overclaiming
  3. [Abstract; Results; Discussion limitations] Results / generalizability: evidence is from one German open-ended task, one age band, and N=116. Directional consistency across 11 scorers is impressive within this corpus, but the title and abstract’s universal phrasing (“systematically underestimates… across every AI-based scoring approach”) invites over-reading. The limitation paragraph already notes content/age/language/format scope; the Results and abstract should mirror that scope more carefully (e.g., “in this corpus, across all eleven implementations tested”) so the central empirical claim stays proportionate to the design.
minor comments (6)
  1. [Abstract; Methods §4.2] Abstract and §4.2: duplicate “the the extent” in the operationalization of conceptual understanding.
  2. [Discussion] Discussion paragraph on science-embedding overestimation: “Fig. 1, (vii), (vii), (ix)” duplicates (vii); should be (vii)–(ix).
  3. [Fig. 1] Fig. 1: forest plot is central; ensure coefficient order and shaded ranges remain legible in grayscale and that the note’s pipeline key (i)–(xi) matches the plotted order exactly.
  4. [Introduction] Intro body text as provided has multiple missing spaces (“uncertaintyabouttherole”, “agrowingbodyofstudies”, etc.). Please proof the production PDF for tokenization/spacing artifacts.
  5. [Table 1; Methods §4.4–4.5] Table 1 note and §4.4: LLM proportions use N=1,160 run-level observations; a brief reminder in the table note that regression inference uses the Rubin-pooled single-run-per-text design would reduce reader confusion.
  6. [Supplementary Information] Supplementary prompt: valuable for reproducibility; consider also depositing the exact API model snapshots/dates and the ML training code or seed settings if journal policy allows.

Circularity Check

1 steps flagged · score 1.0 of 10

No derivation circularity: AI–expert underestimation is an empirical association, not forced by definition; only minor self-citation continuity on the teacher-bias resemblance claim.

  1. self citation load bearing [Abstract; Discussion (teacher-bias resemblance); Methods 4.1–4.2 (dataset/rubrics)]
    "Notably, this language bias closely resembles that previously reported for physics teachers, suggesting the difficulty lies less in any particular assessor than in the nature of inferring conceptual understanding from text-based explanations itself. ... The present study builds on a dataset that was originally developed as part of a research project examining how physics teachers assess students’ text-based scientific explanations. ... (Feser, 2019; Feser & Höttecke, 2021)."

    The secondary claim that AI language bias matches physics-teacher bias rests on the lead author’s own prior teacher-assessment studies that also supplied the corpus and dual scoring schemes. That is program continuity, not an independent external benchmark for the resemblance claim. It does not force the primary AI–expert underestimation coefficients, which are newly estimated against the expert reference and do not reduce by construction to those citations.

full rationale

The paper’s load-bearing result is that lower expert-rated linguistic quality predicts higher odds that AI conceptual-understanding scores fall below expert CU scores (Δ < 0), consistently across nine ML pipelines and two LLMs. That claim is not self-definitional: Δ is AI minus expert CU, linguistic quality is a separately trained ordinal score, and expert CU is entered as a covariate precisely to block the scale-induced confound. Nothing is fitted to the bias coefficient and then relabeled a prediction; ML scores are out-of-fold, LLM scores are prompt-only with nested resampling. Dataset and dual rubrics reuse the lead author’s prior corpus (Feser 2019; Feser & Höttecke 2021), and the interpretive claim that the pattern “closely resembles” physics-teacher language bias cites that same line plus Tajmel—related-program continuity, not a uniqueness theorem or ansatz that forces the AI result. Expert CU is acknowledged as imperfect reference, not ground truth; that is a validity bound, not circular reduction of Eq. X to Eq. Y. Score 1 for non-load-bearing self-citation on the secondary teacher comparison only.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

This is an empirical assessment study. The load-bearing commitments are domain assumptions about what the expert rubrics measure, the validity of treating expert CU as the reference for ‘under/overestimation,’ and standard statistical modeling choices—not new physical entities or first-principles axioms. Free parameters are ordinary ML/LLM hyperparameters and the ordinal scoring scales inherited from prior instruments.

free parameters (3)
  • ML classifier hyperparameters (PCA dims / QDA covariance shrinker; MLR L1 strength; RVM RBF width) = pipeline-specific best CV configuration (not numerically tabulated per fold)
    Grid-searched inside the same 10-fold CV used to report accuracy; chosen to maximize out-of-fold accuracy rather than fixed a priori.
  • LLM decoding settings (GPT-4.1 temperature=1.0; GPT-5-mini reasoning effort & verbosity = medium) = temperature 1.0; reasoning/verbosity medium
    Set a priori but control stochastic score distribution that enters the bootstrap regression.
  • Ordinal expert scales (CU 0–2; linguistic quality sum 0–4) = CU categories 0/1/2; LQ sum of lexical+syntactic 0–4
    Category cut-points from prior instrument development; they define the discrete Δ and the linguistic-quality predictor units that produce reported ORs.
assumptions (5)
  • domain assumption Conceptual understanding of the spacewalk phenomena can be operationalized as internal and external scientific consistency on a 3-level ordinal scheme independent of linguistic form.
    Methods §4.2; scoring scheme from Feser 2019 used as expert reference for all Δ calculations.
  • domain assumption Lexical and syntactic quality (excluding spelling/punctuation) summed to 0–4 adequately represents ‘linguistic quality’ relevant to assessment bias.
    Methods §4.2 linguistic-quality scheme; this is the sole focal predictor of under/overestimation.
  • domain assumption After covarying expert CU, the partial association of linguistic quality with Δ direction identifies language bias rather than floor/ceiling artifacts of the 0–2 scale.
    Methods §4.5 multinomial logistic regression justification; central to causal interpretation of coefficients.
  • standard math Standard multinomial logistic regression, stratified CV, and Rubin pooling of bootstrap LLM replicates are appropriate for these discrete outcomes and nested stochastic scores.
    Methods §4.3–4.5; conventional stats toolkit, not paper-specific mathematics.
  • ad hoc to paper Pooled out-of-fold ML predictions and ten LLM runs per text are adequate stand-ins for ‘the’ AI-assigned score when studying bias direction.
    Authors explicitly prioritize well-performing objects of bias analysis over unbiased deployment error estimates (§4.3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics." pith.science (2026). https://pith.science/paper/B35IWALP

@misc{pith2026260728210,
  author       = {Pith},
  title        = {Pith review of: AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B35IWALP}},
  note         = {Machine review of arXiv:2607.28210}
}
read the original abstract

Explaining physical phenomena is central to physics learning, because students' explanations provide evidence of their conceptual understanding. Because conceptual understanding can only be inferred through language rather than observed directly, distinguishing conceptual understanding from linguistic quality represents a fundamental challenge for assessment. In this study, we examined whether AI-based scoring approaches can assess students' conceptual understanding independently of the linguistic quality of their text-based explanations in physics. We compared conceptual understanding scores generated by nine machine learning-based scoring approaches and two large language model-based scoring approaches against human-expert-assigned scores for 116 secondary-school students' explanations. Despite generally good agreement with expert-assigned scores, explanations of lower linguistic quality were systematically more likely to be underestimated, i.e. receiving lower AI-generated conceptual understanding scores than experts assigned - a language bias that emerged across every AI-based scoring approach. Higher linguistic quality showed no comparable link to overestimation. Notably, this language bias closely resembles that previously reported for physics teachers, suggesting the difficulty lies less in any particular assessor than in the nature of inferring conceptual understanding from text-based explanations itself. The stakes fall hardest on multilingual learners, whose language proficiency may be misread as weaker understanding, and grow as AI-based scoring takes on higher-stakes decisions.

Figures

Figures reproduced from arXiv: 2607.28210 by the authors.

Figure 2
Figure 2. Distribution and relationship of expert-assigned scores for conceptual understanding and linguistic quality. Note. Individual observations represent the 116 student responses to the spacewalk task and are displayed as jittered points to reduce overplotting within the discrete score categories. The solid line represents a fitted linear regression and the depicted R² describes how much variance in conceptual understan… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 13 canonical work pages

  1. [1]

    Educational Researcher, 33 (1), 4–14

    Abedi,J.(2004).TheNoChildLeftBehindActandEnglishlanguagelearners:Assessment and accountability issues. Educational Researcher, 33 (1), 4–14. https://doi.org/10.3102/0013189X033001004

  2. [2]

    Avenia-Tapper, B., & Llosa, L. (2015). Construct relevant or irrelevant? The role of linguistic complexity in the assessment of English language learners’ science knowledge. Educational Assessment, 20(2), 95–111. https://doi.org/10.1080/10627197.2015.1028622

  3. [3]

    S., & Hawn, A

    Baker, R. S., & Hawn, A. (2022). Algorithmic bias in education. International Journal of Artificial Intelligence in Education, 32(4), 1052–1092. https://doi.org/10.1007/s40593- 021-00285-9

  4. [4]

    Braaten, M., & Windschitl, M. (2011). Working toward a stronger conceptualization of scientific explanation for science education. Science Education, 95(4), 639–669. https://doi.org/10.1002/sce.20449

  5. [5]

    A., Salinas, A., Mahotiere, M., Lee, O., & Secada, W

    Buxton, C. A., Salinas, A., Mahotiere, M., Lee, O., & Secada, W. G. (2013). Leveraging cultural resources through teacher pedagogical reasoning: Elementary grade teachers analyze second language learners’ science problem solving. Teaching and Teacher Education,32(1), 31–42

  6. [6]

    Science, 356(6334), 183–186

    Caliskan,A.,Bryson,J.J.,&Narayanan,A.(2017).Semanticsderivedautomaticallyfrom language corpora contain human-like biases. Science, 356(6334), 183–186. https://doi.org/10.1126/science.aal4230

  7. [7]

    Ferrara, E. (2023). Fairness and bias in artificial intelligence: A brief survey of sources, impacts, and mitigation strategies. Sci, 6(1), 3. https://doi.org/10.3390/sci6010003

  8. [8]

    Journal of Science Teacher Education, 20

    Feser,M.S.,&Höttecke,D.(2021).ExploringtheRoleofLanguageinPhysicsTeachers’ Everyday Assessment Practice. Journal of Science Teacher Education, 20. https://doi.org/10.1080/1046560X.2021.1890926

Show all 52 references
  1. [9]

    O., Rossi, R

    Gallegos, I. O., Rossi, R. A., Barrow, J., Tanjim, M. M., Kim, S., Dernoncourt, F., Yu, T., Zhang, R., & Ahmed, N. K. (2024). Bias and fairness in large language models: A survey. Computational Linguistics, 50(3), 1097–1179. https://doi.org/10.1162/coli_a_00524

  2. [10]

    https://doi.org/10.18608/jla.2023.7793

    Grimm,A.,Steegh,A.,Kubsch,M.,&Neumann,K.(2023).Learninganalyticsinphysics education:Equity-focuseddecision-makinglacksguidance!JournalofLearningAnalytics, 10(1), 71–84. https://doi.org/10.18608/jla.2023.7793

  3. [11]

    Ha, M., & Nehm, R. H. (2016). The impact of misspelled words on automated computer scoring: A case study of scientific explanations. Journal of Science Education and Technology, 25(3), 358–374. https://doi.org/10.1007/s10956-015-9598-9

  4. [12]

    Härtig, H., Bernholt, S., Prechtl, H., & Retelsdorf, J. (2015). Unterrichtssprache im Fachunterricht – Stand der Forschung und Forschungsperspektiven am Beispiel des Textverständnisses.ZeitschriftfürDidaktikderNaturwissenschaften ,21(1), 55–67

  5. [13]

    Horbach, A., Ding, Y., & Zesch, T. (2017). The influence of spelling errors on content scoring performance. Proceedings of the 4th Workshop on Natural Language Processing Techniques for Educational Applications (NLPTEA 2017), 45--53. https://aclanthology.org/W17-5908/

  6. [14]

    Jansen, T., Machts, N., Möller, J., & Keller, S. (2021). Don’t just judge the spelling! The influence of spelling on assessing second-language student essays.Frontline Learning Research,9 (1), 85–99

  7. [15]

    Kang, H., Thompson, J., & Windschitl, M. (2014). Creating opportunities for students to show what they know: The role of scaffolding in assessing scientific explanations.Science Education,98(4), 674–704. https://doi.org/10.1002/sce.21123

  8. [16]

    Kim,C.,Passonneau,R.J.,Lee,E.,SheikhiKarizaki,M.,Gnesdilow,D.,&Puntambekar, S. (2026). NLP ‐enabled automated assessment of scientific explanations: Towards eliminating linguistic discrimination. British Journal of Educational Technology, 57(1), 79–111. https://doi.org/10.1111...

  9. [17]

    Kortemeyer, G., Nöhl, J., & Onishchuk, D. (2024). Grading assistance for a handwritten thermodynamics exam using artificial intelligence: An exploratory study. Physical Review Physics Education Research, 20(2), 020144. https://doi.org/10.1103/PhysRevPhysEducRes.20.020144

  10. [18]

    Lee, G.-G., Latif, E., Wu, X., Liu, N., & Zhai, X. (2024). Applying large language models and chain-of-thought for automatic scoring. Computers and Education: Artificial Intelligence, 6, 100213. https://doi.org/10.1016/j.caeai.2024.100213

  11. [19]

    Li, B., Qunhan, X., & Mao, C. (2026). Differences between human and AI scoring: A meta-analysis of english language assessments. Sc ientific Reports, 16(1), 17412. https://doi.org/10.1038/s41598-026-48053-w

  12. [20]

    Llosa, L. (2017). Assessing Students’ Content Knowledge and Language Proficiency. In: Shohamy, E., Or, I., May, S. (eds) Language Testing and Assessment. Encyclopedia of Language and Education. Springer, Cham. https://doi.org/10.1007/978-3-319-02261- 1_33

  13. [21]

    Luykx, A., Lee, O., Hart, J., & Deaktor, R. (2007). Cultural and Home Language Influences on Children’s Responses to Science Assessments.TeachersCollegeRecord , 109(4), 897–926

  14. [22]

    Lyon, E. G. (2013). Learning to assess science in linguistically diverse classrooms: Tracking growth in secondary science preservice teachers’ assessment expertise.Science Education,97(3), 442–467. https://doi.org/10.1002/sce.21059

  15. [23]

    Madras, D., Pitassi, T., & Zemel, R. (2018). Predict responsibly: Improving fairness and accuracybylearningtodefer.Proceedingsofthe32ndInternationalConferenceonNeural Information Processing System, NIPS’18, 6150–6160

  16. [24]

    Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35. https://doi.org/10.1145/3457607

  17. [25]

    Independently published

    Molnar,C.(2023).InterpretingmachinelearningmodelswithSHAP:AguidewithPython examples and theory on Shapley values (1st ed.). Independently published

  18. [26]

    Resnik, P. (2025). Large language models are biased because they are large language models.ComputationalLinguistics,51(3),885–906.https://doi.org/10.1162/coli_a_00558

  19. [27]

    R., & Lovorn, M

    Rezaei, A. R., & Lovorn, M. (2010). Reliability and validity of rubrics for assessment through writing. Assessing Writing, 15 (1), 18–39. https://doi.org/10.1016/j.asw.2010.01.003

  20. [28]

    Rodrigues, L., Xavier, C., Costa, N., Gasevic, D., & Mello, R. F. (2025). Is GPT-4 fair? An empirical analysis in automatic short answer grading. Computers and Education: Artificial Intelligence, 8, 100428. https://doi.org/10.1016/j.caeai.2025.100428

  21. [29]

    Rosenthal, J. W. (1996).TeachingSciencetoLanguageMinorityStudents:TheoryandPractice . Multilingual Matters Ltd

  22. [30]

    P., & Marshall, J

    Scannell, D. P., & Marshall, J. C. (1966). The effect of selected composition errors on grades assigned to essay examinations.American Educational Research Journal, 3(2), 125–130. https://doi.org/10.3102/00028312003002125

  23. [31]

    Schleppegrell, M. J. (2004). The Language of Schooling. A Functional Linguistics Perspective. Routledge

  24. [32]

    R., Pratt, K

    Schumm, W. R., Pratt, K. K., Hartenstein, J. L., Jenkins, B. A., & Johnson, G. A. (2013). Determining statistical significance (alpha) and reporting statistical trends: Controversies, issues, and facts.ComprehensivePsychology ,2(1), Article 10

  25. [33]

    (2017).NaturwissenschaftlicheBildunginderMigrationsgesellschaft

    Tajmel, T. (2017).NaturwissenschaftlicheBildunginderMigrationsgesellschaft . Springer Fachmedien Wiesbaden.https://doi.org/10.1007/978-3-658-17123-0

  26. [34]

    Wang, M., Chen, Y., Huang, X., & Lai, Y. (2026). Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity. Computers and Education: Artificial Intelligence, 10, 100568. https://doi.org/10.1016/j.caeai.2026.100568

  27. [35]

    M., Xi, X., & Breyer, F

    Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice, 31(1), 2–13. https://doi.org/10.1111/j.1745-3992.2011.00223.x

  28. [36]

    C., Stuhlsatz, M

    Zhai, X., Haudek, K. C., Stuhlsatz, M. A. M., & Wilson, C. (2020b). Evaluation of construct-irrelevant variance yielded by machine and human scoring of a science teacher PCK constructed response assessment. Studies in Educational Evaluation, 67, 100916. https://doi.org/10.1016...

  29. [37]

    W., Haudek, K

    Zhai, X., Yin, Y., Pellegrino, J. W., Haudek, K. C., & Shi, L. (2020a). Applying machine learning in science assessment: A systematic review.StudiesinScienceEducation ,56(1), 111–151. https://doi.org/10.1080/03057267.2020.1735757

  30. [38]

    The present study builds on a dataset that was originally developed as part of a research project examining how physics teachers assess students’ text-based scientific explanations

    Methods 4.1Data To investigate whether AI-based scoring approaches can infer students’ conceptual understanding independently of the linguistic quality of their text-based explanations, we required a dataset of authentic student responses for which conceptual understanding and...

  31. [39]

    Feser, M. S. (2019). Physiklehrkräfte korrigieren Schülertexte. Eine Explorationsstudie zur fachlich-konzeptuellen und sprachlichen Leistungsfeststellung und -beurteilung im Physikunterricht. Logos

  32. [40]

    Enders, C. K. (2010). Applied missing data analysis. Guilford Press

  33. [41]

    Hastie,T.,Tibshirani,R.,&Wainwright,M.(2015).Statisticallearningwithsparsity:The Lassoandgeneralizations(0ed.).ChapmanandHall/CRC.https://doi.org/10.1201/b18401

  34. [42]

    Fluck, H.-R. (1997). Fachdeutsch in Naturwissenschaft und Technik. Einführung in die Fachsprachen und die Didaktik/Methodik des fachorientierten Fremdsprachenunterrichts (Deutsch als Fremdsprache). Julius Groos Verlag

  35. [43]

    Latif, E., Lee, G.-G., Neumann, K., Kastorff, T., & Zhai, X. (2024). G-SciEdBERT: A contextualized LLM for science assessment tasks in German (Version 2). arXiv. https://doi.org/10.48550/ARXIV.2402.06584

  36. [44]

    (2024, December 11)

    Huggingface. (2024, December 11). German_Semantic_V3. Web page. https://huggingface.co/aari1995/German_Semantic_V3

  37. [45]

    OpenAI.(2024).GPT-4systemcard.https://cdn.openai.com/papers/gpt-4-system-card.pdf

  38. [46]

    Lengyel, D., & Roth, H.-J. (2012). Beobachtung der Schreibentwicklung in der Sekundarstufe I. In S. Fürstenau & M. Gomolla (Hrsg.), Migration und schulischer Wandel: Leistungsbeurteilung (S. 123–136). Springer VS

  39. [47]

    arXiv:1908.10084v1

    Reimers,N.,&Gurevych,I.(2019).Sentence-BERT:Sentenceembeddingsusingsiamese BERT-networks. arXiv:1908.10084v1. https://doi.org/10.48550/ARXIV.1908.10084

  40. [48]

    OpenAI. (2025). GPT-5 system card. https://openai.com/de-DE/index/gpt-5-system-card/

  41. [49]

    Tharwat, A. (2016). Linear vs. quadratic discriminant analysis classifier: A tutorial. International Journal of Applied Pattern Recognition, 3(2), 145. https://doi.org/10.1504/IJAPR.2016.079050

  42. [50]

    (2012).QualitativeContentAnalysisinPractice

    Schreier, M. (2012).QualitativeContentAnalysisinPractice . Sage Publications Ltd

  43. [52]

    eingebe, ordne ihn genau einer der drei Kategorien (K1, K2, K3) zu und gib das Ergebnis nur im folgenden JSON-Format aus: {

    Tipping, M. E. (1999). The relevance vector machine. Advances in Neural Information Processing Systems, 12. Supplementary Information Table S1.Performance of the AI-based scoring approaches. ML-basedscoringapproaches Accuracy F1-Score Cohen’sκ Text-embeddingmodels SupervisedML...

  44. [2016]

    produces curved, quadratic boundaries; because it struggles with high-dimensional data, a principal component analysis was placed upstream to reduce the embedding dimension. Regularized multinomial logistic regression(MLR; Hastie, 2015) constructs simple linear decision bounda...

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.