{"id":"093b8db2-505d-4a6f-b3af-c531af8e5a23","arxiv_id":"2605.25652","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM judges on Thai bar exam essays converge on the majority human reading and none reproduce the minority human reading on rubric-ambiguous cells.","lead":"This paper compares LLM judges and human bar examiners scoring the same Thai bar exam essays under identical conditions. It finds LLMs align with the majority human view on ambiguous rubric items but fail to reproduce the minority human interpretation.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Validity of the 2-1 human split as rubric-permitted ambiguity (vs. examiner idiosyncrasy) rests only on rubric silence and N=3 raters","rationale":"The reader's weakest_assumption exactly isolates the condition required for the asymmetry claim to be meaningful. The small human panel is the point at which that condition is least secure; confirming or refuting it via additional raters would directly test whether the observed LLM convergence is an artifact of an idiosyncratic human reading.","tokens_in":1889,"tokens_out":339,"duration_ms":17451,"concrete_test":"Have 4-6 additional Bar Council-trained examiners independently score the same 5 contested cells under identical inputs; if none assign scores in the 1-2 band while the 6-8 band is reproduced, the minority reading lacks confirmation as a valid alternative interpretation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The claim that zero of 26 LLMs reproduce the minority reading (examiner A's 1-2 band on the 5 contested cells) requires that A's lower-band interpretation is a legitimate, rubric-permitted alternative rather than an outlier. The paper grounds this in the rubric's silence on how to penalize omission of a decisive statutory citation when the final answer is otherwise correct. With only three examiners and a 2-1 split, however, the data cannot distinguish genuine ambiguity from one examiner applying a stricter personal threshold. No additional evidence (e.g., explicit rubric clauses, inter-examiner rationale transcripts, or larger rater pool) is described that would confirm both bands are defensible under the full Bar Council regulation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports results from an identical-inputs protocol in which three Bar Council examiners and a 26-LLM panel independently score the same 15 Thai bar-exam free-form answers against the official grading regulation and gold answer. It finds near-universal convergence on 10 cells where the rubric is prescriptive, but a 2-1 human split on the remaining 5 cells (where the rubric is silent on penalizing omission of a decisive statutory citation); 22 LLMs cluster with the majority human band, three occupy the middle gap, and none reproduce the minority (examiner A) reading. An instrumented three-LLM anchor sub-panel yields higher internal consistency (α=0.77) than the human panel (α=0.36), which the authors attribute to systematic convergence on the majority reading.","tokens_in":2029,"tokens_out":482,"duration_ms":26308,"significance":"If the central empirical pattern holds, the work demonstrates that LLM-judge panels can systematically favor one of two rubric-permitted human interpretations when the grading regulation is under-specified, implying that any benchmark that selects judges by maximizing agreement with a human reference panel will inherit this asymmetry by construction. The identical-inputs design, determinism probes, input ablations, and bootstrap CIs on the anchor panel constitute concrete methodological strengths that allow direct comparison of observed score distributions.","major_comments":[{"comment":"The claim that zero of the 26 LLMs reproduce the minority human reading on the 5 contested cells presupposes that examiner A's lower-band scores (1-2) constitute a legitimate, rubric-permitted alternative rather than an idiosyncratic threshold. This assumption rests on the rubric's silence regarding omission of a decisive statutory citation when the final answer is otherwise correct, together with the observed 2-1 split. With only three human raters, however, the data cannot distinguish genuine ambiguity from examiner error or personal judgment; no additional validation (expanded rater pool, inter-examiner rationale transcripts, or explicit rubric clauses) is described that would confirm both bands are defensible under the full Bar Council regulation. This is load-bearing for the headline asymmetry result.","section":"discussion of the 5 contested cells and the 2-1 human split"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful reading and for identifying a key interpretive assumption in our analysis of the contested cells. We address the major comment below.","responses":[{"response":"We agree that n=3 prevents us from definitively separating rubric ambiguity from individual examiner judgment. The manuscript's interpretation rests on (i) the official regulation's explicit silence on penalizing omission of a decisive statutory citation when the final answer is otherwise correct and (ii) the fact that the 2-1 split occurred among three Bar Council-trained examiners rather than random disagreement. We do not claim the split proves equal validity of both readings; we claim it demonstrates observable human variation on an under-specified point. The central empirical pattern—that 22 LLMs cluster with the majority band, three occupy the middle gap, and none consistently reproduce examiner A's band—remains informative even under a conservative reading that treats A's scores as idiosyncratic: it shows the LLM population does not sample the full range of observed human scores on these cells. We will add a limitations paragraph explicitly noting that the small human panel precludes strong claims about the source of the split.","revision_made":"partial","referee_comment":"The claim that zero of the 26 LLMs reproduce the minority human reading on the 5 contested cells presupposes that examiner A's lower-band scores (1-2) constitute a legitimate, rubric-permitted alternative rather than an idiosyncratic threshold. This assumption rests on the rubric's silence regarding omission of a decisive statutory citation when the final answer is otherwise correct, together with the observed 2-1 split. With only three human raters, however, the data cannot distinguish genuine ambiguity from examiner error or personal judgment; no additional validation (expanded rater pool, inter-examiner rationale transcripts, or explicit rubric clauses) is described that would confirm both bands are defensible under the full Bar Council regulation. This is load-bearing for the headline asymmetry result."}],"tokens_in":1652,"tokens_out":459,"duration_ms":31239,"standing_objections":["We cannot supply an expanded rater pool, inter-examiner rationale transcripts, or new explicit clauses in the Bar Council regulation, as these lie outside the original study design and the fixed external materials provided by the Bar Council."]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper shows zero of 26 LLMs reproduce the minority human reading on the five contested cells, while most align with the B/C majority band. The identical-inputs protocol makes that asymmetry easy to see.\n\nThe setup is straightforward and useful. Three Bar Council examiners and the LLM panel score the same 15 answers against the same question, regulation, gold answer, and candidate response. On the 10 cells where the rubric is clear, everyone converges. On the five where it is silent about penalizing a missing decisive citation, humans split 2-1 and LLMs do not. The anchor panel alphas (0.77 LLM vs 0.36 human) and the fact that the majority cluster crosses model sizes and vendors give the convergence claim some weight. This is the kind of direct test that matters for LLM-as-judge work in legal grading.\n\nThe soft spot is the human side. The claim that both bands are legitimate interpretations of the rubric depends on the 2-1 split reflecting real ambiguity rather than one examiner applying a stricter personal rule. With only three raters and no additional rationale transcripts or larger pool, the data cannot separate those possibilities. The stress-test note is correct on this point.\n\nThe paper is for people building or auditing LLM judges for subjective legal or essay tasks. A reader working on agreement metrics or bias in automated evaluation will get something concrete from the protocol and the specific result.\n\nIt deserves peer review. The design is clean enough and the asymmetry finding has practical implications even if the human ambiguity part needs more support.","headline":"LLMs cluster with the majority human band on ambiguous cells but the 2-1 split rests on thin evidence from three examiners.","tokens_in":2542,"tokens_out":397,"would_cite":false,"duration_ms":21315,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLM judges converge on majority human reading and ignore minority interpretation on ambiguous rubric cells","keywords":["LLM judges","inter-rater stability","Thai bar exam","rubric ambiguity","human-AI comparison","legal essay evaluation","asymmetric agreement","free-form scoring"],"falsifier":"Finding even one LLM that consistently assigns scores in examiner A's 1-2 band on all five contested cells would falsify the claim that zero LLMs reproduce the minority human reading.","tokens_in":2798,"feed_emoji":"⚖️","tokens_out":796,"duration_ms":26116,"temperature":0.7,"pith_summary":"The paper tests both the assumption that expert inter-rater stability forms a single ceiling and that LLM agreement with that ceiling proves judge stability. It runs an identical-inputs protocol on 15 Thai bar-exam answers scored by three Bar Council examiners and 26 LLMs. On ten cells with clear rubric rules all raters converge, but on five cells left open by the rubric about missing statutory citations the humans split into two coherent bands. LLMs follow the upper band in 22 cases, sit in the gap in three cases, and only one approaches the lower band without matching it consistently. Zero LLMs reproduce the minority human reading.","feed_headline":"Zero LLMs match minority human scores on contested bar-exam cells","feed_subtitle":"When Thai bar examiners split on omitted citations, 26 LLMs converge on the majority reading and none reproduce the minority view.","key_machinery":"The identical-inputs protocol that supplies the exact same question, official grading regulation, gold answer, and candidate answer to both the three human examiners and the 26-LLM panel for direct comparison of score distributions on contested cells.","core_discovery":"When the rubric does not prescribe how to grade a correct final answer that omits a decisive statutory citation, the human panel splits between a B/C majority at scores 6-8 and examiner A at scores 1-2. Of 26 LLMs, 22 score in or near the B/C band, three occupy the silent middle gap, and only GPT-5.4 Nano approaches A's band without consistently scoring within it. No LLM reproduces the minority reading on the contested cells. The B/C cluster spans every model size, vendor, and price tier tested, while an instrumented three-LLM anchor sub-panel reaches alpha 0.77 against the human panel's alpha of 0.36.","pith_inferences":["LLM judges may systematically narrow the range of acceptable legal interpretations even when the rubric permits multiple coherent readings.","The protocol could be extended to other domains with ambiguous scoring rules to test whether one-sided LLM convergence is specific to legal essays.","Methods that deliberately encourage LLMs to sample from minority human bands on contested cells could be developed and measured against this baseline.","Stability metrics alone do not guarantee that an LLM judge captures the full distribution of valid human judgments."],"forward_implications":["A benchmark that selects its LLM judge by maximizing agreement with a human reference panel will inherit the asymmetry toward the majority reading by construction.","The high LLM-panel alpha of 0.77 versus human-panel alpha of 0.36 reflects systematic convergence on one reading rather than balanced reproduction of both.","The B/C-direction cluster includes models from every size, vendor, and price tier tested.","An instrumented three-LLM anchor sub-panel carries determinism probes and input ablations yet still inherits the majority bias."],"fun_headline_variants":["LLMs skip minority examiner band on Thai bar citation splits","Zero LLMs match A's low scores in contested bar-exam cells","LLM judges cluster with B/C majority, none hit human minority","22 LLMs join B/C band on rubric-silent bar essay cells","No model reproduces examiner A's reading on omitted citations"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The three human examiners represent the full spectrum of valid interpretations permitted by the rubric, and the observed split reflects genuine rubric ambiguity rather than examiner error or idiosyncratic judgment.","fun_headline_variants_meta":{"raw":{"variants":["LLMs skip minority examiner band on Thai bar citation splits","Zero LLMs match A's low scores in contested bar-exam cells","LLM judges cluster with B/C majority, none hit human minority","22 LLMs join B/C band on rubric-silent bar essay cells","No model reproduces examiner A's reading on omitted citations"]},"model":"grok-4.3","cost_usd":0.003755,"raw_usage":{"total_tokens":1988,"prompt_tokens":918,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":37553000,"prompt_tokens_details":{"text_tokens":918,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":987,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":918,"tokens_out":83,"duration_ms":9387,"temperature":1.0,"reasoning_tokens":987,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T22:09:11.784862+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Finding even one LLM that consistently assigns scores in examiner A's 1-2 band on all five contested cells would falsify the claim that zero LLMs reproduce the minority human reading.","supporting_citations":[],"review_version":1}