{"id":"7290fa1a-8329-4d80-9764-2772b99728be","arxiv_id":"2507.19980","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"LLM-based raters are less reliable than trained human raters on AP Chinese writing, but hybrid human plus AI composite scoring improves reliability over AI-only scoring.","lead":"This study uses generalizability theory to see how reliably large language models and trained human raters score AP Chinese writing tasks, based on 120 essays scored by 2 humans and 7 AI models. It finds that human scores are more reliable overall, but AI scores can be reasonably consistent for story narration, and hybrid scoring with at least one human rater improves reliability over AI-only scoring.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The random-facet assumption for the seven LLM raters is load-bearing and unsupported: the reported reliability coefficients describe this convenience sample of models, not 'LLMs' in general.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the seven LLMs and four prompts are treated as random samples from larger universes, but they are convenience selections. This matters because the central claim is phrased about 'LLMs' generally, not about the seven specific models studied. If the random-facet assumption fails, the generalizability coefficients do not estimate reliability for the population of LLMs; they describe only the particular models included. The paper's own methods section explicitly asserts the random-facet assumption, while the data collection section describes intentional, non-random selection of models, so the mismatch is internal and not merely a matter of external skepticism. This concern is addressable: the authors could reframe the claim as conditional on the observed models, or provide a defined sampling frame for LLMs and prompts, or report leave-one-model-out sensitivity. Because the within-sample descriptive findings (e.g., these seven models are less reliable than two human raters under these prompts) are still meaningful, the appropriate verdict remains conditional rather than rejection. No additional fatal flaw was identified; the absence of uncertainty intervals is a related secondary issue that strengthens the need for caution but is less fundamental than the unsupported generalization target.","tokens_in":15218,"tokens_out":5308,"duration_ms":70081,"concrete_test":"Re-run the AI-rater D-study as a leave-one-model-out analysis: for each of the seven LLMs, recompute the G-study variance components and the SN and ER generalizability coefficients using the other six models, and also estimate a version with model as a fixed facet. If the leave-one-out SN coefficients or the SN-versus-ER gap vary by more than about 0.05 to 0.10 across subsets, or if the fixed-facet estimates differ materially from the random-facet estimates, the reported 'LLM reliability' is an artifact of the particular seven-model convenience sample and the abstract should be limited to those models rather than to 'LLMs' generally.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim generalizes from seven specific LLMs to 'LLMs' as a class, but the G-theory design treats raters as a random facet. The paper states that 'both t and r are treated as random facets because the specific tasks and raters included in the study are considered representative samples drawn from larger universes.' For the AI raters, this assumption is not credible: the seven models (ChatGPT-3.5, GPT-4, GPT-4o, GPT-o1, Gemini 1.5, Gemini 2.0, Claude 3.5 Sonnet) were 'intentionally selected to represent a diverse range of architectures and capabilities' and were scored over an 18-month period with changing model versions. They are not exchangeable draws from a well-defined universe of LLMs. Under G theory, the rater variance component and the resulting generalizability coefficients estimate a hypothetical universe of interchangeable raters. With these seven fixed, dated artifacts, the variance component conflates true rater differences with model-version and prompt-protocol effects, and the D-study coefficients are conditional on this exact set. The abstract's claim that 'LLMs demonstrated reasonable consistency' is therefore only supported for these seven models unless the random-rater assumption is justified or the analysis is reframed with model as a fixed facet. The small numbers of persons (30) and prompts (2 per task type) further mean the coefficients have no uncertainty intervals, but the primary load-bearing problem is the generalization target itself.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies multivariate generalizability theory to compare the scoring reliability of two human raters and seven LLM raters (ChatGPT-3.5, GPT-4, GPT-4o, GPT-o1, Gemini 1.5, Gemini 2.0, Claude 3.5 Sonnet) on 120 AP Chinese exam essays from 30 students across two task types (story narration and email response) and four scoring dimensions (holistic, task completion, delivery, language use). The authors estimate G-study variance components and run D studies to project reliability under different numbers of tasks and raters, including hybrid human-AI composite designs. The main findings reported are that human raters yield higher generalizability coefficients than AI raters, story narration scores are more reliable than email response scores, task completion is the most reliable domain, and composite scoring with at least one human rater can improve reliability, though adding AI raters does not always help.","tokens_in":15453,"tokens_out":3956,"duration_ms":45765,"significance":"If the results are valid, the paper makes a useful empirical contribution to the growing literature on LLM-based automated essay scoring by demonstrating that current LLMs are not yet full substitutes for human raters in a high-stakes language exam, and by quantifying the reliability of hybrid scoring designs. The use of multivariate generalizability theory to model task and rater facets in this context is a methodological strength, as is the transparent reporting of the full AI scoring prompt in Appendix A. The paper also provides variance-component tables that could be re-analyzed by other researchers. However, the generalizability of the findings to 'LLMs' as a class is seriously limited by the treatment of the seven specific, dated proprietary models as a random sample from a well-defined universe of LLMs, and this limitation directly affects the central claims in the abstract and research questions.","major_comments":[{"comment":"The treatment of the seven AI raters as a random facet is load-bearing and unsupported. The paper states that 'both t and r are treated as random facets because the specific tasks and raters included in the study are considered representative samples drawn from larger universes,' but also that the seven AI engines were 'intentionally selected to represent a diverse range of architectures and capabilities' and were scored over an 18-month period with changing model versions. These statements are in direct tension: intentional selection for diversity and temporal drift do not define a universe of exchangeable LLM raters. Under G theory, the rater variance component and the resulting generalizability coefficients for AI raters estimate a hypothetical universe of interchangeable raters; with these seven fixed, dated artifacts, the estimates conflate true model differences, version changes, and prompt-protocol effects. The abstract's claim that 'LLMs demonstrated reasonable consistency' is consequently supported only for these seven specific models, not for LLMs as a class. The authors should either provide a principled justification for the random-facet assumption (e.g., by defining the universe and arguing exchangeability) or reframe the analysis with model as a fixed facet and report model-specific and model-average reliability.","section":"Generalizability Theory Analysis (Methods)"},{"comment":"The abstract's claim that 'composite scoring that incorporates both human and AI raters improved reliability' is an overstatement of the reported results. In Figure 5, for the email response task, the generalizability coefficient for the 1 human + 2 AI condition is lower than for the 1 human + 1 AI condition, which the authors attribute to the differential weighting (0.33/0.67 vs. 0.5/0.5). The paper does not present a comparison of the composite conditions against the best human-only condition (e.g., two human raters), nor does it show a consistent monotonic improvement from adding AI raters. The only unambiguous comparison is between 1 human + 2 AI and 0 human + 3 AI, where the human-inclusive composite is higher. The composite-scoring conclusion should be qualified to state that reliability depends on the weighting scheme and the task type, and the abstract should reflect this conditionality.","section":"Results: Mixed Use of Human and AI Raters (Research Question 4); Abstract"},{"comment":"None of the reported variance components or generalizability coefficients are accompanied by standard errors, confidence intervals, or any other measure of sampling variability. With 30 persons, only two prompts per task type, and two (or seven) raters, the G-study estimates are likely to be highly imprecise; indeed, several variance components are negative and set to zero (e.g., Σ_t for the AI rater ER condition and Σ_tr for human SN in Table 1), which indicates substantial sampling noise. The D-study coefficients are therefore point estimates whose differences (e.g., .81 vs. .71 for the n_t=2, n_r=2 condition, or the .036 difference between SN and ER for human raters) may not be statistically meaningful. The authors should report uncertainty intervals (e.g., bootstrap or Bayesian intervals) for the key coefficients, or at minimum discuss the likely magnitude of sampling error, before drawing comparative conclusions such as 'SN scores were more reliable than ER scores' and 'using more than two raters may not be necessary.'","section":"D studies and Tables 1-4"}],"minor_comments":[{"comment":"The text cites 'Hyland (2003)' in the literature review, but the reference list contains Hyland (2019) 'Second language writing'; the citation and list are inconsistent.","section":"References"},{"comment":"The dashed line at 0.8 in Figures 2-5 is used as the target reliability benchmark, but no justification is given for this specific threshold; the paper should either cite a source for the criterion or acknowledge it as an arbitrary design choice.","section":"D studies"},{"comment":"The paper states that 'This paper presents a subset of the results' after describing the multivariate analyses for each domain score, but it does not specify which analyses were omitted; readers may want to know whether the omitted results are available elsewhere.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":"The core technical content (G-study variance component estimation, D-study design) appears sound, but the paper's central generalization to 'LLMs' as a class rests on a random-facet assumption for seven convenience-sampled proprietary models. If the authors reframe the claims to the specific models or justify the universe of generalization, the paper could become acceptable. I would not recommend rejection on the random-facet issue alone, but the current formulation is not publishable without this fix. The lack of uncertainty intervals is an additional barrier to publishing the comparative reliability claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a competent, first-of-its-kind application of multivariate generalizability theory to LLM-based scoring of AP Chinese writing, and the authors do not oversell their findings. The main result—that LLMs score less reliably than trained humans, do better on story narration than email tasks, and that hybrid human-AI composites can beat AI-only designs—is backed by standard G-theory machinery and reported with unusual transparency. It deserves a serious referee, though the generalization from seven named LLMs to \"LLMs\" as a class is the weakest joint in the argument.\n\nWhat is actually new: nobody has run a multivariate G-study with LLMs as raters on a high-stakes language exam. The extension from Wilson et al. and Chen et al. on traditional AES to this rater class is real, and the D-study comparisons across task types and scoring domains are useful for assessment practice. The appendix with the full AI training protocol is a welcome piece of reproducibility. The G-study variance components are standard, and the paper is candid about its limitations.\n\nWhere it gets soft. The load-bearing assumption is that the seven AI models—ChatGPT-3.5, GPT-4, 4o, o1, Gemini 1.5/2.0, Claude 3.5—are a random sample from a universe of \"LLMs.\" They are not. They are a convenience sample selected \"to represent a diverse range,\" measured over 18 months as model versions changed. The rater variance component therefore conflates model differences with version drift, and the generalizability coefficients are conditional on this exact set. The paper states the assumption but gives no defense for it. Reframing model as a fixed facet, or at least acknowledging that the target universe is \"the seven selected models\" rather than \"LLMs,\" would be more honest. The small sample (30 students, 120 essays) and the zeroing of negative variance components also mean the coefficients carry no uncertainty intervals; this is a minor issue on its own but compounds the generalization problem.\n\nNone of this breaks the paper. The central comparison between human and AI raters on this dataset is sound, and the hybrid-scoring conclusion is, if anything, understated. The paper would get value from a revision that either justifies the random-rater assumption or narrows the claim. I would send it out for review.","headline":"First G-theory analysis of LLM essay scoring, honest and useful, but the random-facet assumption for seven specific LLMs is the load-bearing weakness.","tokens_in":16009,"tokens_out":1603,"would_cite":true,"duration_ms":20124,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper uses generalizability theory to show that LLM essay scoring is less reliable than human scoring on AP Chinese writing tasks, but composite human-plus-LLM scoring improves reliability.","keywords":["large language models","automated essay scoring","generalizability theory","reliability","writing assessment","hybrid scoring","AP Chinese","human-AI comparison"],"falsifier":"Take a new set of prompts and a different roster of LLMs, score a comparable batch of essays with the same rubric, and rerun the same G and D studies; if the human-rater advantage, the story-narration advantage, and the composite-scoring benefit do not reappear with similar generalizability coefficients, the random-facet assumption and the hybrid recommendation would not survive.","tokens_in":14985,"feed_emoji":"🤖","tokens_out":7097,"duration_ms":81287,"temperature":0.7,"pith_summary":"This paper asks whether large language models can reliably replace trained human raters in high-stakes second-language writing assessment. Using generalizability theory on 120 AP Chinese essays scored by two humans and seven LLMs, it finds that human scoring is more reliable overall, that LLM scores are reasonably consistent for story-narration tasks but less so for email responses, and that combining human and AI scores raises reliability above AI-only scoring. The practical upshot is that LLM autoscoring is not yet a substitute for human raters, but it can be a useful component in a hybrid scoring design.","feed_headline":"Hybrid human-AI scoring beats AI-only on writing tests","feed_subtitle":"Generalizability study of AP Chinese essays finds LLM ratings lag humans but a composite design restores reliability.","key_machinery":"The machinery is multivariate generalizability theory (G theory), which decomposes score variance into facets instead of treating error as one undifferentiated term. The main design is $p^{\\bullet} \\times t^{\\circ} \\times r^{\\bullet}$: persons and raters crossed with a fixed task-type facet (story narration vs email), with tasks nested within task type. A companion design $p^{\\bullet} \\times r^{\\circ} \\times t^{\\bullet}$ treats rater type (human vs AI) as the fixed facet for composite-scoring questions. Variance and covariance components estimated in a G study feed D studies that project generalizability coefficients for different numbers of tasks and raters, with 0.8 as the benchmark. This lets the paper say where reliability comes from—person variance, task difficulty, rater stringency, and their interactions—rather than reporting only a single inter-rater correlation.","core_discovery":"On its own terms, the paper's discovery is a set of variance decompositions: LLM raters show larger person-task and person-rater interaction variances than humans, which lowers generalizability coefficients; with two tasks and two raters, humans reach 0.81 while AI reaches 0.71 on the holistic score. Story narration is scored more consistently than email response by both rater types, and the gap is larger for AI (0.132 vs 0.036 at two raters and two tasks). Composite scoring with one human and one or more AI raters produces higher generalizability than AI-only scoring in most conditions, though adding a second AI rater under proportional weighting can reduce the coefficient for email responses. The inference the authors draw is that a hybrid design, not LLM-only scoring, is the viable path for large-scale writing assessment.","pith_inferences":["If the finding replicates, test administrators could use D-study curves as a calibration table: pick the combination of human and AI raters that reaches a target reliability coefficient at lowest cost, rather than assuming more raters always helps.","The random-facet assumption invites a sampling experiment: draw different LLMs and prompts from the same universe and check whether the variance components stabilize; the answer would tell whether these coefficients are design constants or model-specific artifacts.","The non-monotone composite result suggests that optimal weighting, not just rater count, is the adjustable screw; alternative weights might make AI-heavy composites more attractive than the proportional-weighting case shown here.","The paper does not test fine-tuned models, so a natural extension is to see whether fine-tuning closes the gap to human reliability more than the few-shot prompting used here."],"forward_implications":["Human raters are more reliable than LLM raters at equal numbers of tasks and raters; for two tasks and two raters, the holistic-score generalizability coefficient is 0.81 for humans versus 0.71 for AI.","LLM scoring is more dependable for story narration than for email response, so the structure of the task matters and visual-prompt narration tasks are the safer target for autoscoring.","Composite scoring that includes at least one human rater beats AI-only scoring at the same total rater count, so hybrid designs are the practical route to reliability.","Adding a second AI rater to a human rater can lower composite reliability when proportional weighting overweights AI, meaning rater-count decisions should be made with weights in mind.","Across analytic domains, task completion scores are slightly more reliable than delivery or language use for both rater types, and LLMs are especially variable on the more subjective language dimensions."],"supporting_citations":[{"why":"Supplies the generalizability theory estimation framework and D-study logic on which all coefficients rest.","marker":"Brennan (2001)"},{"why":"Foundational source of G theory's variance-component decomposition used throughout the analysis.","marker":"Cronbach et al. (1972)"},{"why":"Earlier G-theory analysis of automated essay scoring that the study extends to LLM raters and high-stakes second-language writing.","marker":"Wilson et al. (2019)"},{"why":"Multivariate G-theory application to human and automated ratings; precedent for composite reliability questions.","marker":"Chen, Hebert, and Wilson (2022)"},{"why":"Prior evidence that ChatGPT-era LLM scoring falls short of full human agreement, motivating the reliability comparison.","marker":"Mizumoto and Eguchi (2023)"},{"why":"Source of the prompt-calibration recommendation and evidence that GPT-4 lags traditional models, informing the AI training protocol.","marker":"Yancey et al. (2023)"},{"why":"Shows few-shot prompting improves LLM scoring accuracy, justifying the task-specific training documents used before AI scoring.","marker":"Lee et al. (2024)"},{"why":"Supplies the four-scores-per-essay scoring framework (holistic plus task completion, delivery, and language use) applied to the AP Chinese data.","marker":"Song and Tang (2025)"}],"fun_headline_variants":["Hybrid scoring boosts reliability over AI-only on essays","AI raters lag humans, but hybrid scoring recovers reliability","LLM scoring less reliable, hybrid with humans wins","For writing tests, mix human and AI raters for reliability","Composite human-AI scoring bests AI-only on AP essays"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result stands on the idea that the two prompts used for each task type and the seven LLMs chosen are representative stand-ins for all possible prompts and all possible LLMs, so the reliability numbers are meant to generalize; if that idea fails, the numbers describe only this set of 120 essays and seven models.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid scoring boosts reliability over AI-only on essays","AI raters lag humans, but hybrid scoring recovers reliability","LLM scoring less reliable, hybrid with humans wins","For writing tests, mix human and AI raters for reliability","Composite human-AI scoring bests AI-only on AP essays"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000165,"raw_usage":{"total_tokens":1204,"prompt_tokens":851,"completion_tokens":353,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":467,"completion_tokens_details":{"reasoning_tokens":271}},"tokens_in":467,"tokens_out":353,"duration_ms":4251,"temperature":1.0,"reasoning_tokens":271,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T13:50:51.184039+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a new set of prompts and a different roster of LLMs, score a comparable batch of essays with the same rubric, and rerun the same G and D studies; if the human-rater advantage, the story-narration advantage, and the composite-scoring benefit do not reappear with similar generalizability coefficients, the random-facet assumption and the hybrid recommendation would not survive.","supporting_citations":[],"review_version":1}