{"id":"ac98df97-3cc7-4a53-a61b-32b0804577d1","arxiv_id":"2607.28210","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Every tested AI scorer systematically underestimates conceptual understanding in linguistically weaker secondary-school physics explanations relative to trained experts.","lead":"AI scorers of physics explanations systematically give lower conceptual-understanding scores when the writing is linguistically weaker, even when human experts rate the science as sound. The same underestimation pattern appears across nine ML pipelines and two LLMs, and it mirrors known teacher bias—raising fairness stakes for multilingual students as automated scoring scales.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection beyond the reader's already-flagged expert-reference limit; the underestimation pattern is internally well-supported.","rationale":"The strongest claim is a consistent correlational pattern (lower expert LQ → higher odds of AI underestimation of CU vs. the same experts) across a fully crossed set of embedding+classifier pipelines and two prompted LLMs. The paper supplies the necessary safeguards for that claim: independent dual schemes, high inter-rater reliability, CU as covariate to block scale-induced association, full coefficient tables, and proper handling of LLM run variance. The single load-bearing vulnerability is precisely the one the reader named—the non-ground-truth status of expert CU—which the authors themselves flag and which already justifies the CONDITIONAL verdict. No stronger internal inconsistency, mis-specified model, or unacknowledged confound appears in the reported analyses. Therefore the reader's verdict, confidence, and weakest-assumption diagnosis require no adjustment; the concrete residualization check is only a useful robustness verification, not a expected falsifier.","tokens_in":18264,"tokens_out":541,"duration_ms":13659,"concrete_test":"Re-fit the eleven multinomial logits after residualizing expert LQ on expert CU (or using only the LQ residual as predictor) and confirm that the underestimation coefficients remain negative and of similar magnitude; if they collapse toward zero, residual shared variance was driving the result and the bias claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption is correctly identified and is already the paper's own stated limitation (Discussion; Methods 4.2): expert CU scores are the best available dual-trained reference (high α/W) but not language-free ground truth, so AI–expert negative deviations could partly reflect residual shared entanglement rather than pure AI language bias. That concern is real and bounds causal attribution and generalizability, yet it does not undermine the reported descriptive claim. The design already blocks the main mechanical confound (CU covariate in the multinomial logits; Tables S2–S3), the negative LQ–underestimation coefficients are directionally consistent across all 11 pipelines (5 at p<0.05, 2 at p<0.10; ORs 1.3–3.9), overestimation shows no symmetric pattern, and LLM stochasticity is handled by nested resampling + Rubin's rules. No hidden statistical or design flaw is load-bearing beyond what the reader already conditioned on. The central empirical pattern therefore stands as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript tests whether AI-based scorers of physics explanations can separate conceptual understanding (CU) from linguistic quality (LQ). Using 116 authentic Grade-9 German explanations of a spacewalk/sound-transmission task, each dual-rated by trained experts on independent CU (0–2) and LQ (0–4) schemes, the authors compare nine ML pipelines (3 embeddings × 3 classifiers) and two LLM prompters (GPT-4.1, GPT-5-mini) to expert CU. Agreement is moderate-to-good (67–78%). Multinomial logistic regressions of deviation direction (under / agree / over), with expert CU as covariate, show that lower LQ systematically raises the odds of underestimation relative to experts across all eleven approaches (~1.3–3.9× per LQ unit; five coefficients p<0.05, two p≈0.05), while higher LQ shows no comparable link to overestimation except in two science-embedding ML pipelines. The authors interpret this as a language bias resembling that previously reported for physics teachers and discuss fairness stakes for multilingual learners.","tokens_in":18519,"tokens_out":1464,"duration_ms":52151,"significance":"If the pattern holds, the result is consequential for physics education assessment: it shows that both classical ML and modern LLM scorers reproduce a construct-irrelevant LQ–underestimation association even when prompted or trained to score content, and that standard agreement-with-experts metrics miss this. Strengths include dual expert schemes with strong reliability (Krippendorff α≈0.91–0.92; Kendall W≈0.91–0.94), CU controlled as covariate to block scale-induced confounding, explicit handling of LLM stochasticity via 1,000 nested resamples and Rubin pooling, full coefficient tables (S2–S3), and directional consistency across heterogeneous scoring architectures. The link to teacher language bias and the fairness framing for multilingual learners are timely as AI scoring moves toward higher-stakes use. The work is a solid empirical contribution to physics education research and AI-in-assessment, not a methods breakthrough.","major_comments":[{"comment":"Methods §4.5 and Tables S2–S3: several multinomial models show signs of separation or sparse-cell instability (e.g., underestimation intercepts ≈ −20.7 to −20.8 with SE ≈ 68–73 for German_Semantic RVM and paraphrase-multilingual MLR; correspondingly extreme CU ORs). With N=116, a three-category outcome, and CU as covariate, cell counts for underestimation are small (Table 1: often ~6–15%). The directional LQ pattern is still visible, but the claim that bias “emerged across every” approach should be qualified by model diagnostics (e.g., reporting underestimation base rates, checking Firth/penalized multinomial or collapsing rare Δ=±2, and flagging unstable fits). Without that, some of the larger 1/OR values (up to ~3.9) and non-significant but directionally consistent coefficients are hard to interpret at face value.","section":"Methods §4.5; Tables S2–S3; Table 1"},{"comment":"Discussion and Methods §4.2: the central quantity is Δ = AI CU − expert CU, so “underestimation” is defined relative to experts who, by the paper’s own framing, face the same inferential entanglement of language and understanding. High inter-rater α/W and separate schemes make the expert reference the best available dual-trained standard, and the CU covariate correctly blocks the main mechanical confound. Still, residual shared language influence cannot be ruled out, so the load-bearing interpretation should be stated more tightly as: AI–expert negative deviations are systematically associated with lower LQ—not as a pure demonstration that AI alone underestimates true CU. A short explicit statement of what would falsify the AI-specific reading (e.g., experimental LQ manipulation holding CU evidence fixed, as the authors themselves propose) would strengthen the claim without overclaiming","section":"Discussion; Methods §4.2"},{"comment":"Results / generalizability: evidence is from one German open-ended task, one age band, and N=116. Directional consistency across 11 scorers is impressive within this corpus, but the title and abstract’s universal phrasing (“systematically underestimates… across every AI-based scoring approach”) invites over-reading. The limitation paragraph already notes content/age/language/format scope; the Results and abstract should mirror that scope more carefully (e.g., “in this corpus, across all eleven implementations tested”) so the central empirical claim stays proportionate to the design.","section":"Abstract; Results; Discussion limitations"}],"minor_comments":[{"comment":"Abstract and §4.2: duplicate “the the extent” in the operationalization of conceptual understanding.","section":"Abstract; Methods §4.2"},{"comment":"Discussion paragraph on science-embedding overestimation: “Fig. 1, (vii), (vii), (ix)” duplicates (vii); should be (vii)–(ix).","section":"Discussion"},{"comment":"Fig. 1: forest plot is central; ensure coefficient order and shaded ranges remain legible in grayscale and that the note’s pipeline key (i)–(xi) matches the plotted order exactly.","section":"Fig. 1"},{"comment":"Intro body text as provided has multiple missing spaces (“uncertaintyabouttherole”, “agrowingbodyofstudies”, etc.). Please proof the production PDF for tokenization/spacing artifacts.","section":"Introduction"},{"comment":"Table 1 note and §4.4: LLM proportions use N=1,160 run-level observations; a brief reminder in the table note that regression inference uses the Rubin-pooled single-run-per-text design would reduce reader confusion.","section":"Table 1; Methods §4.4–4.5"},{"comment":"Supplementary prompt: valuable for reproducibility; consider also depositing the exact API model snapshots/dates and the ML training code or seed settings if journal policy allows.","section":"Supplementary Information"}],"recommendation":"minor_revision","confidential_remarks":"Fit for a physics-education / assessment journal is good. The expert-reference and single-task limits are real but already largely owned by the authors; I would not reject on them. Watch that the title’s causal tone is softened slightly in revision so it matches the relative-to-experts estimand. No concerns about misconduct or citation manipulation; continuity with the lead author’s prior teacher-bias work is disclosed and appropriate."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: across nine ML pipelines and two LLMs, lower expert-rated linguistic quality raised the odds of underestimating conceptual understanding relative to trained experts (roughly 1.3–3.9× per unit drop), while higher linguistic quality did not symmetrically drive overestimation. That pattern is new for overall lexical/syntactic quality in physics explanations, not just spelling, and it holds directionally for every approach they tried.\n\nWhat they did well is the dual-scheme expert reference (α ≈ 0.91–0.92, Kendall W ≈ 0.91–0.94), the CU covariate that blocks the obvious scale confound, full coefficient tables, and proper handling of LLM stochasticity via 1,000 nested resamples plus Rubin pooling. The 3×3 embedding×classifier design plus two GPT variants makes the consistency harder to dismiss as model-specific. Citations to the teacher-bias literature (including their own prior line) are appropriate continuity, not circularity; Δ is AI minus expert CU and LQ is a separate predictor.\n\nSoft spots are real but already flagged by the authors. Experts are not language-free ground truth—conceptual understanding is only inferable from text—so some shared entanglement could remain in the negative deviations. N = 116 on one German Grade-9 spacewalk task, same-CV hyperparameter tuning, and no public code/data limit how far you can push the claim. Overestimation only appears for the science-specific embedding in two cases; that is a minor model-specific note, not a collapse of the main result. No load-bearing statistical flaw beyond those bounds.\n\nThis is for people working on automated scoring of science explanations, equity for multilingual learners, and validity checks that go past raw agreement. It deserves a serious referee. I would engage: cite the underestimation pattern when discussing AI scoring fairness, and bring it to reading group if the group cares about assessment validity.","headline":"Consistent underestimation of lower-linguistic-quality physics explanations by 11 AI scorers vs dual-expert ratings; solid design, bounded by single-task N and non-ground-truth experts.","tokens_in":19167,"tokens_out":485,"would_cite":true,"duration_ms":12025,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"AI scoring systematically underestimates conceptual understanding when physics explanations are linguistically weak.","keywords":["AI-based scoring","language bias","conceptual understanding","physics education","automated assessment","multilingual learners","large language models","text-based explanations"],"falsifier":"Hold demonstrated conceptual content fixed while systematically degrading only lexical and syntactic quality of the same explanations, and test whether AI underestimation relative to the same experts still rises with lower linguistic quality.","tokens_in":19081,"feed_emoji":"📝","tokens_out":768,"duration_ms":38574,"temperature":0.7,"pith_summary":"This paper asks whether AI can score students’ conceptual understanding in physics without being swayed by how well the student writes. Using 116 authentic secondary-school explanations of a sound-transmission task, the authors compared eleven AI scorers—nine machine-learning pipelines and two large language models—against trained experts who had rated conceptual understanding and linguistic quality separately. Agreement with experts was generally good, yet every AI approach was more likely to under-score explanations of lower linguistic quality, while higher linguistic quality rarely produced over-scoring. The pattern matches bias previously seen in physics teachers, which the authors read as a feature of the assessment task itself: understanding is only reachable through language. The practical stakes are highest for multilingual learners and rise as AI moves into higher-stakes grading.","feed_headline":"AI underestimates weak writers' physics understanding","feed_subtitle":"Across eleven scorers, lower linguistic quality raised the odds of under-scoring conceptual grasp","key_machinery":"Scoring-deviation direction (underestimation, agreement, or overestimation of expert conceptual-understanding scores), modelled with multinomial logistic regression on expert linguistic-quality scores while covarying expert conceptual-understanding scores to block scale-induced confounding.","core_discovery":"Across all eleven AI-based scoring approaches, explanations rated lower in linguistic quality by experts were systematically more likely to receive lower conceptual-understanding scores from the AI than from experts—roughly 1.3 to 3.9 times higher odds of underestimation per one-unit drop in linguistic quality—whereas higher linguistic quality showed no comparable link to overestimation in most approaches.","pith_inferences":["The same underestimation pattern is likely in automated scoring of open explanations across other sciences and writing-heavy STEM assessments.","Fairness audits for educational AI should treat overall lexical–syntactic sophistication—not only spelling or surface grammar—as a construct-irrelevant risk factor.","Decoupling prompts or specialized embeddings may shrink but not erase the bias if conceptual evidence is itself carried by linguistic form."],"forward_implications":["Multilingual and international learners risk systematic disadvantage if AI scoring scales into summative decisions.","Agreement with expert ratings is necessary but not sufficient; evaluations must also check whether disagreements track construct-irrelevant linguistic quality.","Human-in-the-loop alone may not remove the bias where teachers and AI err on the same texts; routing low-linguistic-quality explanations for mindful human review is a proposed safeguard.","Language bias should be treated as intrinsic to inferring conceptual understanding from text-based explanations, not as a quirk of one scorer type."],"fun_headline_variants":["AI scorers underrate physics grasp in linguistically weak explanations","All eleven AI approaches underestimate weaker writers' conceptual grasp","Lower language quality raises odds of AI under-scoring physics understanding","AI mirrors teacher bias against linguistically weaker physics explanations","Linguistic weakness linked to systematic AI underestimation of concepts"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That expert conceptual-understanding scores are independent enough of language that AI–expert underestimations can be attributed to AI language bias rather than shared entanglement of the two constructs.","fun_headline_variants_meta":{"raw":{"variants":["AI scorers underrate physics grasp in linguistically weak explanations","All eleven AI approaches underestimate weaker writers' conceptual grasp","Lower language quality raises odds of AI under-scoring physics understanding","AI mirrors teacher bias against linguistically weaker physics explanations","Linguistic weakness linked to systematic AI underestimation of concepts"]},"model":"grok-4.5","effort":"low","cost_usd":0.00488,"raw_usage":{"total_tokens":1369,"prompt_tokens":779,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":48804000,"prompt_tokens_details":{"text_tokens":779,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":525,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":779,"tokens_out":65,"duration_ms":9977,"temperature":1.0,"reasoning_tokens":525,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T14:33:42.957194+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Hold demonstrated conceptual content fixed while systematically degrading only lexical and syntactic quality of the same explanations, and test whether AI underestimation relative to the same experts still rises with lower linguistic quality.","supporting_citations":[],"review_version":1}