{"id":"8ffd5465-7d6c-4587-8407-d0af35bc3cc1","arxiv_id":"2608.09289","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using two AI rewrites of 3,830 student texts, the authors separate grammar errors from unnatural phrasing and map 40 grammar structures into four instructional zones.","lead":"This paper tests a new way to separate grammar mistakes from unnatural-sounding English in Japanese students' writing using an AI that creates two versions of each text. It finds that structures like articles and modals cause accuracy errors while forms like -ing are avoided, and maps these onto a chart for teachers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The idiomatic revision layer violates its own \"WITHOUT adding or removing content\" constraint (Table 1: +18% words, -31% sentences), so the idiomatic gap conflates LLM elaboration with native-like idiomaticity, making the typology's vertical axis an artifact of prompt non-compliance.","rationale":"The reader correctly identifies the idiomatic revision layer as the weakest premise, but the paper's own Table 1 makes the problem sharper than \"unvalidated LLM.\" The revision prompt explicitly forbids adding or removing content, yet the output layer is 18% longer and has 31% fewer sentences. That observable discrepancy means the idiomatic gap is not even an internally consistent operationalization of \"natural revision of the same content.\" This directly threatens the central claim that accuracy and idiomatic gaps are distinct, measurable learner diagnostics: the vertical axis of the typology could be measuring the model's propensity to elaborate and subordinate rather than learners' underuse or overuse of specific grammar structures. The paper deserves credit for a transparent pipeline, deterministic CEFR-J extraction, and sensible statistical reporting, and the limitation is partially acknowledged in Section 5.1. However, the specific prompt-violation mechanism is not acknowledged and is testable with a relatively simple alignment analysis and a constrained re-run. The reader's CONDITIONAL verdict remains appropriate: the framework is a plausible proof-of-concept if reinterpreted as \"LLM revision preferences,\" but full acceptance requires demonstrating that the idiomatic gap survives a content-preserving revision constraint and, ideally, comparison with human native-speaker revisions. I therefore keep the verdict unchanged rather than escalating to rejection, because the concern, while central, can be settled empirically and the authors have already flagged human validation as necessary next work.","tokens_in":14629,"tokens_out":4337,"duration_ms":45109,"concrete_test":"Sample 100 literal-idiomatic pairs from the corpus and use word alignment or a semantic similarity measure to compute the proportion of lexical content tokens in the idiomatic revision that have no aligned counterpart in the literal correction. If more than 5% of content tokens are unaligned, the prompt's no-content-change constraint was violated. Then re-run the full pipeline on the same sample with a stricter revision prompt that preserves sentence boundaries and limits length change to within 5%; if the idiomatic-gap rankings in Figure 3 and the quadrant assignments in Figure 4 shift materially, the reported idiomatic gaps are artifacts of LLM elaboration rather than of idiomatic preferences.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is not merely that gemma4 is unvalidated; it is that the idiomatic revision layer visibly violates the input constraint that defines it. Section 3.3 instructs the model to rewrite \"WITHOUT adding or removing content,\" yet Table 1 shows the Idiomatic Revision layer is 18.3% longer than Literal Correction (182,671 vs. 154,478 words) and contains 31% fewer sentences (12,998 vs. 18,871), raising mean sentence length from 8.19 to 14.05. A revision that preserves content can legitimately change length, but an 18% word increase combined with sentence merging indicates the model added cohesive devices, elaborations, or new content. Consequently Freq_idiomatic - Freq_literal (Section 3.4) is not a clean \"idiomatic gap\": a large share of the reported shifts, including V-ING +42.21, MD.would.AFF +28.67, CL_after.etc +26.20, and TA.PRPF.AFF +13.71, can be driven by the LLM writing more and with more subordination rather than by native-like preference given the same content. Since the vertical axis of the RQ3 typology (Figure 4) is built entirely on this difference, the main empirical conclusion is compromised by prompt non-compliance. This is an internal inconsistency, not just a missing native-speaker validation, and it is directly testable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a layered LLM-correction pipeline to separate grammatical accuracy from native-like idiomaticity in Japanese junior high school EFL writing. For 3,830 writing samples from 120 students, the authors generate a literal error-correction layer and an idiomatic native-revision layer, apply the CEFR-J regex grammar extractor, and define an accuracy gap (Freq_literal - Freq_raw) and an idiomatic gap (Freq_idiomatic - Freq_literal). Results are presented as frequency shifts across three layers, leading to a two-dimensional instructional typology that maps error rates against idiomatic gaps to assign distinct pedagogical priorities to grammar structures such as articles, third-person -s, modals, -ing forms, and SVO patterns.","tokens_in":14905,"tokens_out":3650,"duration_ms":34710,"significance":"If the gaps are validly measured, the framework is a useful and scalable diagnostic tool: it gives teachers a transparent, reproducible way to distinguish structures learners attempt but execute incorrectly from structures they avoid or overuse, and the CEFR-J alignment and deterministic regex extractor are strengths for classroom deployment. The paper also makes a genuine contribution by attempting to operationalize the accuracy/idiomaticity distinction at feature level rather than as a global error count. However, the significance is conditional: the idiomatic gap rests entirely on an unvalidated LLM proxy for native norms, and the reported revision layer violates the very constraint that defines what the idiomatic gap is supposed to measure.","major_comments":[{"comment":"The idiomatic revision instruction requires the model to rewrite \"WITHOUT adding or removing content,\" yet Table 1 shows the Idiomatic Revision layer contains 182,671 words versus 154,478 in Literal Correction (an 18.3% increase) and 12,998 sentences versus 18,871 (a 31% decrease), raising mean sentence length from 8.19 to 14.05. This substantial expansion and sentence merging indicate that the model added cohesive devices, elaboration, or subordination, so Freq_idiomatic - Freq_literal is not a clean measure of native-like preference for the same content. A large share of the reported idiomatic shifts (e.g., V-ING +42.21, MD.would.AFF +28.67, CL_after.etc +26.20, TA.PRPF.AFF +13.71) may be driven by the model writing more and with more complex syntax rather than by native-like idiomaticity. Since the vertical axis of Figure 4 is built on this difference, the central typology is compromised by prompt non-compliance. This is directly testable by comparing revisions that honor the length constraint against the current output.","section":"§3.3, Table 1, Appendix"},{"comment":"The idiomatic gap is operationalized entirely through gemma4:31b's revisions, with no native-speaker corpus, human ratings, or native judgments to anchor the \"native norms\" against which underuse and overuse are defined. The paper itself acknowledges in §5.1 that \"manual validation was ad hoc and restricted to a small subset of outputs\" and lists \"native speaker correction and revisions\" as future work. As a result, the finding that -ing forms and hypothetical would are underused while simple present and can are overused is, by construction, a statement about what gemma4 adds or removes during revision, not an independently validated claim about Japanese learners' distance from native writing. If gemma4's revision preferences differ from real native norms (for example, if it over-generates nonrestrictive relatives or present perfect), the idiomatic gap values and the quadrant assignments in Figure 4 are artifacts of model style. A concrete remedy is to collect native-speaker revisions or native-corpus frequency benchmarks on a held-out sample and compare the resulting gap values.","section":"§4.2, §5.1"},{"comment":"The accuracy gap is computed as the difference in regex-extractor frequencies between raw and literal-corrected text. This presumes that gemma4's literal correction layer preserves each learner's underlying syntactic structure while repairing only surface errors, and that the CEFR-J regex extractor recognizes the repaired forms as instances of the intended structures. No reliability evidence is provided for the literal corrections: there is no comparison with expert human corrections on a sample, and the extractor's precision and recall on raw versus corrected text are not reported. For example, the +17.27 Δ for DT.the could overstate article omission if the literal correction inserts articles in contexts where the learner's raw form would not have been tagged at all. The paper should include a validation sample with human-corrected texts and report extractor agreement on both layers to support the accuracy-gap interpretation.","section":"§3.4, Table 2"}],"minor_comments":[{"comment":"The text contains a typo: \"TO.to_d o\" should be \"TO.to_do\".","section":"§4.3"},{"comment":"The Error Rate formula is garbled in the typeset text (\"Correctedraw\" and \"Uncorrectedraw\" appear without proper subscripts); it should be presented as (Freq_literal - Freq_raw) / Freq_literal × 100 for readability.","section":"§4.3"},{"comment":"The sentence describing log-ratio thresholds (\"≥ 1 and ≥ 2 indicate doubling/halving and quadrupling/quartering, respectively\") is ambiguous; log ratio 1 indicates doubling and -1 indicates halving, so the thresholds should be described separately for positive and negative values.","section":"§3.4"},{"comment":"The limitation that manual validation was ad hoc appears only in the final Discussion; it would be more transparent to state this limitation in the Methodology section (§3.3) when the LLM processing is introduced.","section":"§5.1"},{"comment":"The column header \"Words / Sentence\" is ungrammatical; it should be \"Words per Sentence.\"","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper's core idea is timely and the pipeline is reproducible in principle, but the load-bearing idiomatic-gap measure is not yet anchored to any external native norm, and the reported revision layer visibly violates its own content-preservation instruction. These issues are fixable within the manuscript's scope (e.g., re-running with stricter constraints, adding native-speaker validation on a subsample, and tempering the claims), so I do not recommend rejection, but the current evidence does not support the strong typological conclusions. The fit with the journal's applied NLP/education scope is good."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a useful proof-of-concept with a load-bearing flaw in the idiomatic half of its diagnostic. Worth engaging, but it needs major revision, not light editing.\n\nWhat's new and good: the layered pipeline (literal correction vs. idiomatic revision) applied to 3,830 junior-high samples through a deterministic CEFR-J regex extractor is a genuine new application. The two-dimensional typology (error rate vs. idiomatic gap) is a practical way to present feedback options to teachers. The accuracy-gap results largely replicate known Japanese EFL patterns—articles, third-person -s, modals—which is reassuring. Using log-likelihood with log ratio and an open-source model are sensible choices for reproducibility.\n\nNow the soft spots, in proportion. The idiomatic gap is defined as Freq_idiomatic minus Freq_literal, where both layers are outputs of the same LLM. No native-speaker corpus or human ratings anchor it. The paper admits this, which is honest. But the stress-test note points to something worse, and it holds up on reading: the idiomatic revision prompt says \"WITHOUT adding or removing content,\" while Table 1 shows the revision layer is 18% longer and has 31% fewer sentences, and the paper's own text says the LLM \"introduced more explicit connectors, elaborated ideas, and added cohesive devices.\" That directly contradicts the constraint. So the idiomatic gap conflates elaboration with native-like idiomaticity. The big positive shifts—V-ING +42, would +29, subordinate clauses +26—could simply reflect the model writing more and merging sentences, not learner underuse relative to native norms. The vertical axis of the typology is therefore compromised by an internal inconsistency, not just a missing validation.\n\nThe accuracy gap (literal vs. raw) is more defensible, though it still depends on the regex extractor recognizing repaired forms. Minor concern. Code and data are absent; the prompt list is \"available on request,\" which limits reproducibility.\n\nThe paper is clearly written and does not overstate its readiness—the authors flag manual validation as ad hoc and call for native correction as future work. But the abstract still frames the results as \"relative to native norms,\" which is not supported.\n\nWho is this for? Applied linguists and educational NLP researchers building automated feedback tools. Teachers might find the typology intuitive, but they should not act on the idiomatic axis until it is re-anchored. This deserves a serious referee: the framework is novel and the main flaw is fixable with native validation, a content-preservation check on the revisions, and a revised abstract. Send it to review, but expect major revision.","headline":"Promising two-axis diagnostic for EFL writing, but the idiomatic half rests on a revision layer that visibly violates its own content-preservation constraint.","tokens_in":15480,"tokens_out":1938,"would_cite":false,"duration_ms":20173,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that grammatical accuracy and native-like idiomaticity are distinct, measurable gaps in learner writing, and that mapping both on the same texts assigns each grammar structure a different instructional priority.","keywords":["automated writing evaluation","learner corpus","EFL","grammatical accuracy","idiomaticity","LLM","CEFR-J","Japanese EFL writing"],"falsifier":"Ask independent native-speaker raters to revise a random sample of the same 3,830 texts with the same two instructions, and compare the idiomatic gap rankings from human revisions with those from the model; if the ordering of underused and overused structures changes materially, the typology is an artifact of the revision model rather than a property of learner writing.","tokens_in":14381,"feed_emoji":"📝","tokens_out":12955,"duration_ms":102280,"temperature":0.7,"pith_summary":"This paper sets out to show that grammatical accuracy and native-like idiomaticity are not the same thing in learner writing, and that both can be measured separately on the same texts. The authors take 3,830 diary-style submissions from 120 Japanese junior high school students, produce two LLM-generated variants of each text—a literal grammatical correction and a freely idiomatic revision—and count how often CEFR-J grammar patterns appear in each version. The frequency shifts define an accuracy gap and an idiomatic gap, and the two gaps place every grammar item in one of four instructional quadrants. If the argument holds, teachers can decide from corpus evidence whether a structure is misused, avoided, or overused, instead of treating every deviation as a single undifferentiated error.","feed_headline":"Two gaps, not one, separate grammar from naturalness in EFL writing","feed_subtitle":"A layered LLM pipeline maps 3,830 Japanese students' texts into a four-quadrant teaching typology.","key_machinery":"The load-bearing mechanism is the layered LLM-correction pipeline: a raw learner text is first rewritten with a prompt that fixes only spelling, punctuation, and local grammar while explicitly keeping the student's sentence structure, then rewritten again with a prompt that freely naturalises the phrasing. Counting the same grammar patterns in all three layers with the deterministic regex-based CEFR-J grammar extractor turns each prompt into a measurable frequency shift. The log-likelihood statistic $G^2$ decides which shifts are significant, and the log-ratio measure with a smoothing constant of $0.1$ controls for the baseline-frequency trap, so rare structures are not swamped by high-volume defaults.","core_discovery":"The paper's central claim is that when the same learner text is passed through a literal error-correction layer and then an idiomatic-revision layer, the normalized frequency differences between consecutive layers are diagnostic measures with distinct pedagogical meanings. The accuracy gap, $\\mathrm{Freq}_{\\mathrm{literal}}-\\mathrm{Freq}_{\\mathrm{raw}}$, captures structures attempted but produced with errors; the idiomatic gap, $\\mathrm{Freq}_{\\mathrm{idiomatic}}-\\mathrm{Freq}_{\\mathrm{literal}}$, captures structures that are grammatically correct but underused or overused relative to native-like revisions. On the accuracy side, definite articles ($+17.27$ per 10,000 words), third-person singular $-s$, and the modals would and could show the largest gains after correction. On the idiomatic side, $-ing$ forms and would are the most underused, while simple present verbs, subject-verb-object patterns, and can are overused. Mapping the top 40 structures by error rate against idiomatic gap produces four quadrants—high-error/more-natural, low-error/avoided, accurate/overused, and high-error/overused—each with its own instructional target.","pith_inferences":["A testable extension would be to run the same two prompts on argumentative or academic essays: diary prompts likely inflate first-person pronouns, simple present, and can, so the quadrant assignments may shift with genre.","The idiomatic gap should be read as relative to one revision model's style until it is calibrated against native-speaker revisions or a native corpus; the vertical axis is the least externally anchored part of the method.","The framework would be strengthened by per-learner and per-prompt stability checks, since the current analysis aggregates 3,830 texts and does not yet directly measure individual avoidance or overuse patterns.","Applying the same layered-correction procedure to learners from other first-language backgrounds would test whether the overuse and avoidance profiles are transfer-specific or developmental."],"forward_implications":["Articles and third-person singular $-s$ sit in the high-error, more-natural quadrant, so the paper implies they need both form-focused instruction and production practice rather than exposure alone.","Simple present, can, and SVO patterns land in the accurate/overused quadrant, implying instruction should offer lexical and syntactic alternatives instead of more grammar correction.","Low-error, avoided forms such as non-finite verb forms, first-person pronouns, and determiners like some and any require awareness-raising and input, not further form-focused teaching.","Because the pipeline produces a per-text profile cheaply, it could generate individual diagnostic reports rather than only aggregate class-level patterns."],"supporting_citations":[{"why":"Supplies the CEFR-J grammar profile and the regex extractor used to count the grammar structures in all three text layers.","marker":"Ishii & Tono, 2016"},{"why":"Provides the prior corpus profiling of Japanese EFL overuse and underuse that the accuracy and idiomatic gap results extend.","marker":"Ishii & Tono, 2018"},{"why":"Gives the $G^2$ log-likelihood test used to determine which frequency shifts are statistically significant.","marker":"Dunning, 1993"},{"why":"Supplies the log-ratio measure used to interpret the size of shifts relative to baseline frequency.","marker":"Hardie, 2014"},{"why":"Accounts for why third-person singular $-s$ is disproportionately difficult for Japanese learners, a result the accuracy gap corroborates.","marker":"Muroya, 2018"},{"why":"Documents Japanese learners' article omission patterns that the article accuracy gap builds on.","marker":"Hinenoya & Lyster, 2015"},{"why":"Establishes the method of comparing learner output against corrected parallel texts that grounds the layered-corpus approach.","marker":"Granger & Rayson, 1998"},{"why":"Explains L1-driven avoidance of phrasal verbs and post-nominal modification that the paper invokes for relative-clause underuse.","marker":"Strong et al., 2026"},{"why":"Supports the claim that modern LLMs reliably repair surface grammar in learner texts, justifying the literal correction layer.","marker":"Coyne et al., 2023"},{"why":"Supplies the human-annotation evidence the paper cites for treating LLM idiomatic revisions as fluent or native-like.","marker":"Labib et al., 2026"}],"fun_headline_variants":["Two gaps, not one: grammar vs naturalness in EFL writing","Accuracy gaps and idiomatic gaps: two separate problems in EFL","LLM layers split grammar errors from unnaturalness in EFL","Four quadrants map grammar and idiomatic gaps in EFL"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument stands or falls on whether the idiomatic revisions produced by the local open-source language model are a valid stand-in for native-like English usage, since the paper's manual checks covered only a small subset of outputs and no native-speaker corpus anchors the idiomatic gap.","fun_headline_variants_meta":{"raw":{"variants":["Two gaps, not one: grammar vs naturalness in EFL writing","Accuracy gaps and idiomatic gaps: two separate problems in EFL","LLM layers split grammar errors from unnaturalness in EFL","Four quadrants map grammar and idiomatic gaps in EFL"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000564,"raw_usage":{"total_tokens":2714,"prompt_tokens":1020,"completion_tokens":1694,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":636,"completion_tokens_details":{"reasoning_tokens":1624}},"tokens_in":636,"tokens_out":1694,"duration_ms":11291,"temperature":1.0,"reasoning_tokens":1624,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:04:00.722215+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask independent native-speaker raters to revise a random sample of the same 3,830 texts with the same two instructions, and compare the idiomatic gap rankings from human revisions with those from the model; if the ordering of underused and overused structures changes materially, the typology is an artifact of the revision model rather than a property of learner writing.","supporting_citations":[],"review_version":1}