{"id":"cd92a920-aa66-46d2-b1a1-2a4ddc35e243","arxiv_id":"2604.00259","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Open-weight LLMs show moderate holistic agreement (QWK≈0.6) but large negative bias on grammar/conventions traits, detectable at sample sizes as small as 5 essays.","lead":"This paper tested five open-weight AI essay graders against human scores on three datasets, comparing holistic and trait-level scoring. It found the models are systematically harsher than humans on grammar-style traits, and that this bias can be detected with surprisingly small human-labeled samples.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Nmin in §4.4 is a descriptive full-data interval, not a prospective power analysis; the 'detectable with N≤10' claim is unestablished.","rationale":"The reader's weakest assumption is the single-reference HCS and inter-rater noise. That is a legitimate limitation and is acknowledged by the authors, but it is less decisive for LOC traits such as Grammar/Conventions, where agreement among human raters is typically higher. The more load-bearing weakness is the Nmin methodology: the paper uses it to support the abstract's 'detectable with very small validation sets' and the conclusion's 'not artifacts of sampling variability.' As written, the procedure is not a valid prospective sample-size estimate, and the reported Nmin values are therefore not evidence for the deployment recommendation. Because this is an analytical flaw rather than a contradiction in the observed bias values, the reader's CONDITIONAL verdict remains appropriate; the paper would need either a corrected power analysis or softer claims about small-sample detectability.","tokens_in":11376,"tokens_out":10116,"duration_ms":101714,"concrete_test":"Recompute Nmin as a power curve using the released model scores: for each N in {5,10,15,...,100}, draw 1,000 independent random subsamples of size N (without replacement) from the relevant dataset split; for each subsample compute a 95% bootstrap percentile CI for mean bias (10,000 resamples); record the proportion of subsamples whose CI excludes zero. Report the smallest N at which the detection proportion reaches 80% (or 95%) for Llama-3.1-70B on ELLIPSE Grammar/Conventions under the Keywords prompt. If the detection proportion at N=5 is near 95%, the paper's small-sample claim survives; if it is far lower (e.g., <50%), the Nmin values should be reinterpreted as descriptive statistics rather than deployment-relevant sample sizes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that LOC bias is 'stable' and 'detectable with very small validation sets' rests on the Nmin bootstrap analysis in §4.4. As described, the procedure computes a 95% bootstrap percentile interval for mean bias and records the smallest N at which that interval excludes zero. But this is not a sample-size or power calculation. The event that the full dataset's bootstrap distribution excludes zero at N=5 does not equal the probability that an independent 5-essay validation set will yield a confidence interval excluding zero. A proper prospective analysis would draw many independent subsamples of size N, compute a bootstrap CI for each, and report the detection rate. The current Table 4/5 values (e.g., median Nmin=5 for ELLIPSE, Grammar/Conventions Nmin=5) conflate a large realized effect size with a high probability of detecting it in a small fresh sample. The conclusion in §6 that 'observed biases are not artifacts of sampling variability but reflect stable model behavior' depends on this step. This is a real statistical gap in the paper's most actionable claim, though it does not necessarily overturn the existence of a negative LOC bias.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates five open-weight instruction-tuned LLMs as zero-shot essay scorers on three open datasets (ASAP 2.0, ELLIPSE, DREsS), comparing holistic and analytic scoring under two prompt strategies (Keywords vs. Guidelines). It reports agreement (QWK, EA), directional bias (mean signed error), score compression, length sensitivity, and a bootstrap-based minimum sample size (Nmin) for detecting non-zero mean bias. The central claims are that (i) strong open-weight models achieve moderate holistic agreement (QWK≈0.6) but this does not transfer to analytic scoring; (ii) there is a large and stable negative bias on Lower-Order Concern (LOC) traits such as Grammar and Conventions; (iii) keyword-based prompts generally outperform rubric-style guidelines for analytic multi-trait scoring; and (iv) the LOC bias is often detectable with very small validation sets (median Nmin=5 on ELLIPSE), supporting a bias-correction-first deployment strategy.","tokens_in":11654,"tokens_out":3621,"duration_ms":37334,"significance":"If the central claims hold, the paper provides a reproducible, multi-dataset characterization of a trait-dependent harshness in LLM essay scoring, with practical implications for using small bias-estimation sets rather than full fine-tuning. Strengths include the use of open datasets and open-weight models, deterministic greedy decoding, a clear separation of agreement and bias metrics, and an explicit acknowledgment of the single-reference-score limitation. However, the most actionable claim — that LOC bias is detectable with N≤10 essays — rests on a descriptive bootstrap statistic that is not a prospective power analysis, and several secondary claims lack uncertainty quantification. The paper is a useful contribution to the LLM-as-judge and AES literature, but the statistical support for the small-sample detectability claim needs substantial additional analysis.","major_comments":[{"comment":"The abstract and §6 claim that LOC bias is 'detectable with very small validation sets' (often N≤10) and that the bootstrap analysis shows biases are 'not artifacts of sampling variability.' The procedure in §4.4, however, computes, for each model–trait–split–strategy combination, the smallest N at which a 95% bootstrap percentile interval for the mean bias excludes zero, using 10,000 bootstrap resamples. As described, this is a descriptive statement about the full dataset's bias relative to its own variability; it is not a prospective sample-size or power calculation. The event that the full-data bootstrap distribution excludes zero at N=5 does not imply that an independent validation set of 5 essays would yield a confidence interval excluding zero with any stated probability. Moreover, the description is ambiguous about the sampling universe: are resamples drawn from the full dataset a","section":"§4.4, Tables 4–5, Abstract, §6"},{"comment":"The footnote states: 'All bias estimates are statistically different from zero (p<0.001) according to a Wilcoxon Signed-Rank test.' No test statistics, exact p-values, sample sizes, or multiple-comparison corrections are provided. Given the large number of model–trait–split–strategy combinations (e.g., 120 on ELLIPSE in Table 4), unadjusted p<0.001 across all comparisons is not credible as stated and is not verifiable from the manuscript. If the Wilcoxon test is used, the report should provide the distribution of p-values or a proper correction (e.g., FDR) and specify which comparisons are being tested. This is relevant to the paper's characterization of bias as 'systematic' rather than noise.","section":"Table 2 footnote"},{"comment":"The claim that 'concise keyword-based prompts generally outperform longer rubric-style prompts in multi-trait analytic scoring' is based on point estimates of QWK and EA without confidence intervals or significance tests. For example, on ELLIPSE Cohesion under Keywords vs. Guidelines, QWK changes from 0.566 to 0.412 with a large shift in bias (−0.12 vs. −0.55), but the paper does not quantify uncertainty around these values. Given sample sizes of thousands of essays, some differences may be statistically meaningful, but the reader cannot assess which prompt effects are robust. A bootstrap or other resampling-based CI for QWK and EA, or at least for the QWK difference, should be reported for the key comparisons underlying RQ3.","section":"§3.4, Table 3, RQ3"},{"comment":"The definition of 'bias' is entirely relative to the dataset-provided Human Consensus Score (HCS), which is a single reference score per essay. The paper acknowledges in the Limitations that 'Without modeling inter-rater variability, it is difficult to disentangle true model error from legitimate disagreement.' But the abstract and §5 make stronger statements, e.g., 'models often score these traits more harshly than human raters' and 'large and stable negative directional bias.' If HCS is noisy or reflects a single rater's leniency, the observed negative bias on LOC traits could be partly an artifact of the reference rather than a stable property of the models. Because this is the paper's central empirical finding, the interpretation should be tempered to 'bias relative to the dataset-provided HCS,' and the potential impact of HCS noise on the LOC-vs-HOC contrast should be discussed more","section":"§3.4, §4.2, §5, Limitations"}],"minor_comments":[{"comment":"The table header contains a typo: '#Combinations of model–trait–split–strategy' should be '#Combinations' or '#Combinations (model–trait–split–strategy)'. Also, the meaning of 'Bias Not Reached' is clear from the text but could be defined in the caption.","section":"§4.4, Table 4"},{"comment":"The notation for bias is introduced twice: as 'Bias (Mean Signed Error)' and then with the formula. The formula itself is correct, but the symbol μ̂_bias is not used consistently in Tables 2–5; the tables just say 'Bias.' Consider defining the abbreviation in a single place and using it throughout.","section":"§3.4"},{"comment":"The correlation comparison (r_Human vs. r_LLM) is reported without confidence intervals or a test of the difference. Given the large sample sizes, the difference between 0.71 and 0.47 on ASAP is likely significant, but the paper does not quantify this. Also, Pearson correlation with word counts may be influenced by outliers; a robust alternative (e.g., Spearman) would strengthen the verbosity discussion.","section":"§4.3"},{"comment":"The table header has 'T raits' with a stray space. Also, the table reports average word counts and standard deviations but not the standard deviations for the full dataset; consider adding the split sizes consistently.","section":"§3.1, Table 1"},{"comment":"The reference list contains several entries with incomplete metadata (e.g., items 1, 2, 3, 5, 8, 11, 13, 14, 15, 19, 21, 23 lack page numbers or full proceedings details in some cases). Please ensure all references are complete and consistently formatted.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and practically relevant question, and the core observation of a negative LOC-trait bias appears robust across models and datasets based on the aggregate tables. However, the headline 'detectable with N≤10' claim is not supported by the current statistical procedure, and the Wilcoxon footnote in Table 2 is unverifiable. Both points are load-bearing for the paper's central argument and deployment recommendation. The revision should add a prospective subsample-based detection analysis and proper uncertainty quantification for the key metric comparisons. I see no evidence of circularity or fabricated data; the concerns are about statistical interpretation, not about the existence of the bias pattern itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: if you work on LLM-as-a-judge or automated essay scoring, this paper is worth reading for the per-trait bias decomposition. But the headline sample-size claims are overinterpreted, and the bias-correction recommendation is never actually tested.\n\nWhat's new: a systematic comparison of holistic vs. analytic scoring for open-weight LLMs across three datasets and two prompt strategies. The trait-level breakdown is the real contribution. Models show a consistent, large negative bias on lower-order traits like grammar and conventions, which is a practically useful result. The prompt effect reversal — keywords beat full rubric guidelines in multi-trait analytic scoring, but full descriptions help on holistic — is a clean, testable finding. The paper is transparent about models, decoding, and data, so the results are reproducible.\n\nSoft spots, in order. First, §4.4's Nmin is a post-hoc descriptive summary of the full dataset, not a sample-size calculation. The bootstrap interval excluding zero at N=5 means the observed sample has a large effect, but it doesn't tell you the probability that a fresh 5-essay validation set will detect it. That distinction matters for the deployment recommendation. Second, the bias-correction-first strategy is advocated but never tested — there is no experiment showing corrected scores actually improve agreement. Third, the Wilcoxon p<0.001 claim is stated without any test details, which matters given the number of comparisons. Fourth, no confidence intervals are reported for QWK or exact agreement, so the prompt effects are point estimates. The single-rater HCS limitation is acknowledged and is a genuine constraint, but it doesn't kill the main finding — the LOC bias is large, consistent across datasets and models, and unlikely to be pure noise.\n\nNet: the empirical matrix is a genuine contribution, and the central bias observation holds as a descriptive result. The overreach is in the inferential framing, not the data. A good referee can push the authors to redo the sample-size analysis prospectively and to test the correction.\n\nI'd send this to peer review. It deserves a serious referee, not a desk reject. The claims need scaling back and a couple of analyses added, but the work is honest and reproducible. I'd bring it to a reading group focused on LLM evaluation and would cite it as evidence of trait-level bias in essay scoring.","headline":"Solid, reproducible empirical core on trait-level LLM scoring bias, but the small-sample detection claims outrun the bootstrap analysis.","tokens_in":12118,"tokens_out":3061,"would_cite":true,"duration_ms":31842,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper aims to establish that zero-shot LLM essay scorers systematically under-score lower-order writing traits such as grammar and conventions, that this harshness is stable rather than random, and that small human-scored validation set","keywords":["automated essay scoring","LLM-as-a-judge","holistic vs analytic rubrics","scoring bias","lower-order concerns","bootstrap minimum sample size","prompt strategy","zero-shot evaluation"],"falsifier":"Take an analytic-scoring corpus where every essay has multiple independent human raters, and recompute the mean signed bias for grammar and conventions against the average of all raters (and against each rater separately). If the negative bias disappears or becomes inconsistent across raters, the claim of stable model harshness fails; if it persists against every rater, the claim is strengthened.","tokens_in":11293,"feed_emoji":"📝","tokens_out":6086,"duration_ms":59720,"temperature":0.7,"pith_summary":"Open-weight, instruction-tuned LLMs, when asked to score essays with no fine-tuning, do not miss evenly across rubric dimensions. The paper finds that they agree reasonably with human raters on overall essay quality (quadratic weighted kappa around 0.6) but consistently under-grade lower-order concern traits such as grammar, conventions, and syntax, often by about one point on a five-point scale. This negative bias is stable: bootstrap analysis detects it with validation sets as small as five essays on the six-trait analytic corpus, while higher-order traits such as content and organization need much larger samples or are not detectable at all. The paper also shows that concise keyword-style prompts beat detailed rubric text for multi-trait analytic scoring, while the reverse holds for holistic scoring. If correct, the practical conclusion is bias-correction-first deployment: estimate the systematic offset from a tiny set of human scores and subtract it, rather than fine-tune.","feed_headline":"LLM essay scorers under-grade grammar traits by about a point","feed_subtitle":"Even 5–10 human-scored essays expose the bias, so correction without fine-tuning is realistic.","key_machinery":"The paper's analytical engine is the HOC/LOC distinction (higher-order concerns = global discourse traits; lower-order concerns = local linguistic form) used to organize trait-level results, paired with a bootstrap minimum-sample-size (Nmin) procedure: at each sample size N, draw 10,000 resamples, build a 95% percentile confidence interval for mean signed bias, and record the smallest N at which the interval excludes zero. This tool turns \"is the model biased?\" into \"how many human-scored essays prove the bias?\" The prompt contrast (keyword labels vs full rubric guidelines) is the other manipulated factor, used to show that instruction detail interacts with scoring granularity.","core_discovery":"The central discovery is trait-dependent directional bias. In zero-shot scoring against dataset-provided human consensus scores, models assigned numerically lower scores than humans on Lower-Order Concern traits—grammar, conventions, syntax, vocabulary, phraseology—while Higher-Order Concern traits (content, organization, cohesion) often showed smaller or less consistent deviation. On the six-trait corpus, the median minimum sample size for detecting non-zero bias was 5 essays; on the holistic corpus, some configurations required 515–1,460 essays, and on one content trait the bias was not detectable. The authors interpret this as evidence that the LOC harshness reflects stable model behavior","pith_inferences":["The paper links the LOC harshness to the absence of learner-specific context; an immediate testable extension is to add writer metadata (e.g., L1 or grade level) to prompts and see whether the negative bias shrinks.","The keywords-over-guidelines reversal suggests that long rubric text may overload the model when several traits are judged at once; a prompt-engineering experiment that presents traits one at a time or in a fixed order could isolate the mechanism.","The bootstrap Nmin is, in effect, a ready-made audit rule: score a handful of essays, resample the bias, and decide whether to deploy or correct a scorer. The authors stop short of stating a formal decision threshold, so a practical next step is to derive one with explicit trade-offs.","The results imply that any future claim of \"LLMs match humans on essay scoring\" should be qualified by rubric granularity and trait type, since aggregate agreement can hide large trait-level offsets."],"forward_implications":["Systematic LOC bias can be corrected with a small bias-estimation set, so a low-cost deployment path exists: measure the offset on 5–20 human-scored essays and apply it to raw zero-shot scores.","Analytic trait scores should not be used as exact diagnostics: even the best model reached only moderate agreement (QWK around 0.32–0.46) on multi-trait rubrics, and near-zero mean bias can coexist with low agreement because of score compression.","Prompt design is not one-size-fits-all: keyword prompts are preferable for multi-trait analytic scoring, while full rubric descriptions help holistic scoring.","Bias-detection sample sizes should be planned per trait, not per model: lower-order traits need very small sets, higher-order traits need hundreds or are undetectable within available data."],"fun_headline_variants":["LLM essay graders dock grammar points—detectable with 5 essays","Grammar bias in LLM scoring shows up in tiny samples","LLM graders punish grammar; 5 essays reveal the bias","Systematic grammar under-scoring by LLMs, fixable without tuning","LLM essay scoring: grammar harshness is stable and detectable"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper treats each dataset's single human consensus score as the true gold standard, so if that reference is noisy or reflects one rater's leniency, the measured model harshness is partly an artifact of the reference rather than a stable property of the model.","fun_headline_variants_meta":{"raw":{"variants":["LLM essay graders dock grammar points—detectable with 5 essays","Grammar bias in LLM scoring shows up in tiny samples","LLM graders punish grammar; 5 essays reveal the bias","Systematic grammar under-scoring by LLMs, fixable without tuning","LLM essay scoring: grammar harshness is stable and detectable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1148,"prompt_tokens":797,"completion_tokens":351,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":261}},"tokens_in":541,"tokens_out":351,"duration_ms":4553,"temperature":1.0,"reasoning_tokens":261,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T16:58:46.586552+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take an analytic-scoring corpus where every essay has multiple independent human raters, and recompute the mean signed bias for grammar and conventions against the average of all raters (and against each rater separately). If the negative bias disappears or becomes inconsistent across raters, the claim of stable model harshness fails; if it persists against every rater, the claim is strengthened.","supporting_citations":[],"review_version":1}