Pith. sign in

REVIEW 3 major objections 3 minor 1 cited by

Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory

T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper uses generalizability theory to show that LLM essay scoring is less reliable than human scoring on AP Chinese writing tasks, but composite human-plus-LLM scoring improves reliability.

desk verdict First G-theory analysis of LLM essay scoring, honest and useful, but the random-facet assumption for seven specific LLMs is the load-bearing weakness. read the letter →

arxiv 2507.19980 v2 pith:RNFEI5F3 submitted 2025-07-26 cs.CL

classification cs.CL
keywords largelanguagemodelsautomatedessayscoringgeneralizabilitytheoryreliabilitywritingassessmenthybridAPChinesehuman-AIcomparison
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models can reliably replace trained human raters in high-stakes second-language writing assessment. Using generalizability theory on 120 AP Chinese essays scored by two humans and seven LLMs, it finds that human scoring is more reliable overall, that LLM scores are reasonably consistent for story-narration tasks but less so for email responses, and that combining human and AI scores raises reliability above AI-only scoring. The practical upshot is that LLM autoscoring is not yet a substitute for human raters, but it can be a useful component in a hybrid scoring design.

What carries the argument

The machinery is multivariate generalizability theory (G theory), which decomposes score variance into facets instead of treating error as one undifferentiated term. The main design is $p^{\bullet} \times t^{\circ} \times r^{\bullet}$: persons and raters crossed with a fixed task-type facet (story narration vs email), with tasks nested within task type. A companion design $p^{\bullet} \times r^{\circ} \times t^{\bullet}$ treats rater type (human vs AI) as the fixed facet for composite-scoring questions. Variance and covariance components estimated in a G study feed D studies that project generalizability coefficients for different numbers of tasks and raters, with 0.8 as the benchmark. This lets the paper say where reliability comes from—person variance, task difficulty, rater stringency, and their interactions—rather than reporting only a single inter-rater correlation.

What would settle it

Take a new set of prompts and a different roster of LLMs, score a comparable batch of essays with the same rubric, and rerun the same G and D studies; if the human-rater advantage, the story-narration advantage, and the composite-scoring benefit do not reappear with similar generalizability coefficients, the random-facet assumption and the hybrid recommendation would not survive.

Watch

Extended reading notes

Core claim

On its own terms, the paper's discovery is a set of variance decompositions: LLM raters show larger person-task and person-rater interaction variances than humans, which lowers generalizability coefficients; with two tasks and two raters, humans reach 0.81 while AI reaches 0.71 on the holistic score. Story narration is scored more consistently than email response by both rater types, and the gap is larger for AI (0.132 vs 0.036 at two raters and two tasks). Composite scoring with one human and one or more AI raters produces higher generalizability than AI-only scoring in most conditions, though adding a second AI rater under proportional weighting can reduce the coefficient for email responses. The inference the authors draw is that a hybrid design, not LLM-only scoring, is the viable path for large-scale writing assessment.

Load-bearing premise

The result stands on the idea that the two prompts used for each task type and the seven LLMs chosen are representative stand-ins for all possible prompts and all possible LLMs, so the reliability numbers are meant to generalize; if that idea fails, the numbers describe only this set of 120 essays and seven models.

Editorial extensions

If this is right

  • Human raters are more reliable than LLM raters at equal numbers of tasks and raters; for two tasks and two raters, the holistic-score generalizability coefficient is 0.81 for humans versus 0.71 for AI.
  • LLM scoring is more dependable for story narration than for email response, so the structure of the task matters and visual-prompt narration tasks are the safer target for autoscoring.
  • Composite scoring that includes at least one human rater beats AI-only scoring at the same total rater count, so hybrid designs are the practical route to reliability.
  • Adding a second AI rater to a human rater can lower composite reliability when proportional weighting overweights AI, meaning rater-count decisions should be made with weights in mind.
  • Across analytic domains, task completion scores are slightly more reliable than delivery or language use for both rater types, and LLMs are especially variable on the more subjective language dimensions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the finding replicates, test administrators could use D-study curves as a calibration table: pick the combination of human and AI raters that reaches a target reliability coefficient at lowest cost, rather than assuming more raters always helps.
  • The random-facet assumption invites a sampling experiment: draw different LLMs and prompts from the same universe and check whether the variance components stabilize; the answer would tell whether these coefficients are design constants or model-specific artifacts.
  • The non-monotone composite result suggests that optimal weighting, not just rater count, is the adjustable screw; alternative weights might make AI-heavy composites more attractive than the proportional-weighting case shown here.
  • The paper does not test fine-tuned models, so a natural extension is to see whether fine-tuning closes the gap to human reliability more than the few-shot prompting used here.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper applies multivariate generalizability theory to compare the scoring reliability of two human raters and seven LLM raters (ChatGPT-3.5, GPT-4, GPT-4o, GPT-o1, Gemini 1.5, Gemini 2.0, Claude 3.5 Sonnet) on 120 AP Chinese exam essays from 30 students across two task types (story narration and email response) and four scoring dimensions (holistic, task completion, delivery, language use). The authors estimate G-study variance components and run D studies to project reliability under different numbers of tasks and raters, including hybrid human-AI composite designs. The main findings reported are that human raters yield higher generalizability coefficients than AI raters, story narration scores are more reliable than email response scores, task completion is the most reliable domain, and composite scoring with at least one human rater can improve reliability, though adding AI raters does not always help.

Significance. If the results are valid, the paper makes a useful empirical contribution to the growing literature on LLM-based automated essay scoring by demonstrating that current LLMs are not yet full substitutes for human raters in a high-stakes language exam, and by quantifying the reliability of hybrid scoring designs. The use of multivariate generalizability theory to model task and rater facets in this context is a methodological strength, as is the transparent reporting of the full AI scoring prompt in Appendix A. The paper also provides variance-component tables that could be re-analyzed by other researchers. However, the generalizability of the findings to 'LLMs' as a class is seriously limited by the treatment of the seven specific, dated proprietary models as a random sample from a well-defined universe of LLMs, and this limitation directly affects the central claims in the abstract and research questions.

major comments (3)
  1. [Generalizability Theory Analysis (Methods)] The treatment of the seven AI raters as a random facet is load-bearing and unsupported. The paper states that 'both t and r are treated as random facets because the specific tasks and raters included in the study are considered representative samples drawn from larger universes,' but also that the seven AI engines were 'intentionally selected to represent a diverse range of architectures and capabilities' and were scored over an 18-month period with changing model versions. These statements are in direct tension: intentional selection for diversity and temporal drift do not define a universe of exchangeable LLM raters. Under G theory, the rater variance component and the resulting generalizability coefficients for AI raters estimate a hypothetical universe of interchangeable raters; with these seven fixed, dated artifacts, the estimates conflate true model differences, version changes, and prompt-protocol effects. The abstract's claim that 'LLMs demonstrated reasonable consistency' is consequently supported only for these seven specific models, not for LLMs as a class. The authors should either provide a principled justification for the random-facet assumption (e.g., by defining the universe and arguing exchangeability) or reframe the analysis with model as a fixed facet and report model-specific and model-average reliability.
  2. [Results: Mixed Use of Human and AI Raters (Research Question 4); Abstract] The abstract's claim that 'composite scoring that incorporates both human and AI raters improved reliability' is an overstatement of the reported results. In Figure 5, for the email response task, the generalizability coefficient for the 1 human + 2 AI condition is lower than for the 1 human + 1 AI condition, which the authors attribute to the differential weighting (0.33/0.67 vs. 0.5/0.5). The paper does not present a comparison of the composite conditions against the best human-only condition (e.g., two human raters), nor does it show a consistent monotonic improvement from adding AI raters. The only unambiguous comparison is between 1 human + 2 AI and 0 human + 3 AI, where the human-inclusive composite is higher. The composite-scoring conclusion should be qualified to state that reliability depends on the weighting scheme and the task type, and the abstract should reflect this conditionality.
  3. [D studies and Tables 1-4] None of the reported variance components or generalizability coefficients are accompanied by standard errors, confidence intervals, or any other measure of sampling variability. With 30 persons, only two prompts per task type, and two (or seven) raters, the G-study estimates are likely to be highly imprecise; indeed, several variance components are negative and set to zero (e.g., Σ_t for the AI rater ER condition and Σ_tr for human SN in Table 1), which indicates substantial sampling noise. The D-study coefficients are therefore point estimates whose differences (e.g., .81 vs. .71 for the n_t=2, n_r=2 condition, or the .036 difference between SN and ER for human raters) may not be statistically meaningful. The authors should report uncertainty intervals (e.g., bootstrap or Bayesian intervals) for the key coefficients, or at minimum discuss the likely magnitude of sampling error, before drawing comparative conclusions such as 'SN scores were more reliable than ER scores' and 'using more than two raters may not be necessary.'
minor comments (3)
  1. [References] The text cites 'Hyland (2003)' in the literature review, but the reference list contains Hyland (2019) 'Second language writing'; the citation and list are inconsistent.
  2. [D studies] The dashed line at 0.8 in Figures 2-5 is used as the target reliability benchmark, but no justification is given for this specific threshold; the paper should either cite a source for the criterion or acknowledge it as an arbitrary design choice.
  3. [Methods] The paper states that 'This paper presents a subset of the results' after describing the multivariate analyses for each domain score, but it does not specify which analyses were omitted; readers may want to know whether the omitted results are available elsewhere.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the reliability coefficients are computed from observed scores via standard G-theory decompositions, and the paper's claims are not defined in terms of fitted parameters.

full rationale

The paper's derivation chain is empirical, not definitional. G-study variance components are estimated from the 120 independently produced human and AI scores ('These data were used to estimate G study variance and covariance components, which served as the basis for various D studies'), and the D-study coefficients are computed from those components by standard generalizability formulas. No target quantity is defined in terms of a fitted parameter, and no parameter is fitted to a subset of the data and then renamed as a prediction. The use of benchmark responses in the AI prompts is a calibration or training protocol; the reported outcome is still the reliability of the resulting scores, so this is a design condition, not a circular step. The only self-citation, Song and Tang (2025), supports the incidental 0-6 score-scale statement ('each ranging from 0 to 6 (Song & Tang, 2025)') and is not load-bearing for the main comparison; the scale values are observable in the data themselves. The composite-weighting results are explicitly tied to the chosen weights ('the weights were 0.33 and 0.67, respectively, giving twice as much weight to AI raters'), so the finding that adding an AI rater can reduce reliability is a transparent mathematical consequence of the weighting scheme, not a hidden fit. The random-facet assumption for the seven specific LLMs is a legitimate generalization or validity limitation, but it does not make the derivation circular because the estimated coefficients describe the data actually analyzed rather than being assumed in advance.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The ledger contains no new theoretical entities. The free parameters are design choices (composite weights and the 0.8 benchmark) rather than fitted theoretical constants. The main load-bearing assumptions are the random-facet assumption and the stability of variance component estimates, both of which the small convenience sample makes fragile.

free parameters (2)
  • Composite scoring weights in D studies = 0.50/0.50 for 1 human + 1 AI; 0.33/0.67 for 1 human + 2 AI
    Weights are set proportional to the number of raters of each type, a hand-chosen convention. This choice directly drives the finding that adding an AI rater can lower composite reliability (Research Question 4, Figure 5).
  • Target reliability benchmark = 0.80
    The 0.8 line in Figures 2 and 4 is chosen by the investigators as the threshold for deciding whether a D study design has acceptable reliability; it is a judgment call, not derived from data.
assumptions (3)
  • domain assumption Tasks and raters are random facets sampled from larger universes of prompts and LLM raters.
    G theory requires this assumption for generalization beyond the observed prompts and models. The paper asserts representativeness but provides no sampling frame (Generalizability Theory Analysis).
  • standard math ANOVA-based variance component estimates are sufficiently stable for D-study extrapolation with n_p=30 and two prompts per task type.
    Standard G-theory formulas are used, but with small samples and negative components replaced by zero the estimated coefficients are approximate. This is implicit throughout Tables 1 to 4.
  • domain assumption Prompt-based calibration with human-scored benchmark responses places LLM scores on the same scale as human scores.
    AI raters were given sample responses benchmarked at score levels from College Board and previously human-scored responses (Data Collection). The human-AI reliability comparison presupposes this calibration is valid.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory." pith.science (2026). https://pith.science/paper/RNFEI5F3

@misc{pith2026250719980,
  author       = {Pith},
  title        = {Pith review of: Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RNFEI5F3}},
  note         = {Machine review of arXiv:2507.19980}
}
read the original abstract

This study investigates the estimation of reliability for large language models (LLMs) in scoring writing tasks from the AP Chinese Language and Culture Exam. Using generalizability theory, the research evaluates and compares score consistency between human and AI raters across two types of AP Chinese free-response writing tasks: story narration and email response. These essays were independently scored by two trained human raters and seven AI raters. Each essay received four scores: one holistic score and three analytic scores corresponding to the domains of task completion, delivery, and language use. Results indicate that although human raters produced more reliable scores overall, LLMs demonstrated reasonable consistency under certain conditions, particularly for story narration tasks. Composite scoring that incorporates both human and AI raters improved reliability, which supports that hybrid scoring models may offer benefits for large-scale writing assessments.

Figures

Figures reproduced from arXiv: 2507.19980 by the authors.

Figure 2
Figure 2. summarizes the generalizability coefficients obtained from the 16 D studies. The X-axis represents the number of tasks (𝑛$ & ) ranging from 1 to 4, while the four lines correspond to the number of raters (𝑛% & ) from 1 to 4. The solid line at 0.8 serves as a benchmark for assessing whether the reliability meets the investigator’s desired threshold [PITH_FULL_IMAGE:figures/full_fig_p017_2.png] view at source ↗
Figure 3
Figure 3. Generalizability coefficients: Comparison of SN and ER task types. Comparison Across Domain-Specific Scores (Research Question 3) As previously mentioned, in addition to the overall holistic scores, three domain-specific scores were generated: task completion (TC), delivery (DL), and language use (LU). A MGT analysis using the 𝑝• × 𝑡∘ × 𝑟• design was conducted for each domain score, and the estimated G study varianc… view at source ↗
Figure 4
Figure 4. Generalizability coefficients: Comparison across domain-specific scores. Mixed Use of Human and AI Raters (Research Question 4) The last research question is concerned about the use of both human and AI raters for scoring. To address this question, a multivariate 𝑝• × 𝑟∘ × 𝑡• design was employed, in which the [PITH_FULL_IMAGE:figures/full_fig_p022_4.png] view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: Generalizability coefficients: Mixed use of human and AI raters. Inter-rater Reliability Often, an inter-rater coefficient is used as a reliability index for ratings. This coefficient is typically computed using the Pearson correlation, based on examinees' scores on a …

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Across three enterprise agent benchmarks, agent main effects are under 3% of total score variance, so leaderboard order reflects task specialization rather than a general capability advantage.

Reference graph

Works this paper leans on

8 extracted references · 7 canonical work pages · cited by 1 Pith paper

  1. [1]

    1 Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory Dan Song, University of Iowa Won-Chan Lee, University of Iowa Hong Jiao, University of Maryland Abstract This study investigates the estimation of reliability for large language models (LLMs) in scoring writing tasks from the AP Chinese Language and Cu...

  2. [3]

    As previously mentioned, in addition to the overall holistic scores, three domain-specific scores were generated: task completion (TC), delivery (DL), and language use (LU). A MGT analysis using the 𝑝•×𝑡∘×𝑟• design was conducted for each domain score, and the estimated G study variance and covariance components are presented in Tables 2 and 3 for human an...

  3. [4]

    Can composite scoring (human + AI raters) improve the overall reliability of writing assessments? Methods To address the research questions, multivariate generalizability theory (MGT) was employed, with each question involving a distinct design structure. The MGT approach offers a flexible framework for quantifying measurement error and reliability, parti...

  4. [5]

    G study variance and covariance components were estimated for the seven score effects associated with the 𝑝•×𝑡∘×𝑟• design

    v p r t p t r v v = Task Type (SN & ER) v = Rater Type (Human & AI) 15 To address this question, analyses focused on the overall holistic scores. G study variance and covariance components were estimated for the seven score effects associated with the 𝑝•×𝑡∘×𝑟• design. Table 1 presents these estimates, providing a comparison between human and AI raters. Ne...

  5. [6]

    These coefficients are notably lower than the traditional inter-rater coefficients, as they account for variability across both raters and tasks. Table 4 Estimated Inter-rater Reliability Coefficients: Overall Scores Human Rater AI Rater 𝑛%&=1 & 𝑛$&=1: 𝑟 random and 𝑡 fixed SN ER SN ER .750 .850 .540 .651 𝑛%&=1 and 𝑛$&=1: Both 𝑟 and 𝑡 random SN ER SN ER .5...

  6. [7]

    https://doi.org/10.1016/j.jsp.2017.12.005 31 Wilson, J., Chen, D., Sandbank, M

    Journal of School Psychology, 68, 19–37. https://doi.org/10.1016/j.jsp.2017.12.005 31 Wilson, J., Chen, D., Sandbank, M. P., & Hebert, M. (2019). Generalizability of automated scores of writing quality in Grades 3–5. Journal of Educational Psychology, 111, 619–640. https://doi.org/10.1037/edu0000311 Xiao, C., Ma, W., Song, Q., Xu, S. X., Zhang, K., Wang, ...

  7. [9]

    Similarly, Yancey et al

    While AES using GPT can achieve a certain level of accuracy, it still falls short of full agreement with human raters and should therefore be used alongside human evaluation. Similarly, Yancey et al. (2023) evaluated GPT-4’s few-shot capabilities in predicting Common European Framework of Reference for Languages levels for short essays written by second-l...

  8. [2023]

    576-584)

    (pp. 576-584). Zhao, R., Zhuang, Y ., Zou, D., Xie, Q., & Yu, P. L. (2023). AI-assisted automated scoring of picture-cued writing tasks for language assessment. Education and Information Technologies, 28(6), 7031-7063. 32 Appendix A. AI Training Protocol for Grading SN2 Student Samples Based on this prompt, grading rubric, and student writing samples, gra...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.