{"id":"3d2e45c7-79f6-427b-a5ef-02233169ff4c","arxiv_id":"2412.12117","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review and position paper arguing that generative AI should be used as a supplement to teacher guidance in secondary writing instruction, with controls to prevent student dependency.","lead":"This position paper reviews how generative AI like ChatGPT might help or hurt students learning to write in secondary school, and argues that AI should support, not replace, teachers. It offers practical recommendations such as process-based assessment and controlled testing, but it presents no new experimental data of its own.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The core recommendation rests on an untested premise that students must generate drafts themselves; the paper's own cited evidence partly contradicts this, leaving the policy argument empirically underdetermined.","rationale":"The reader correctly identifies the active-production premise as the weakest assumption. The paper is a narrative review and position piece, so it does not make a novel falsifiable empirical claim; its central claim is a recommendation. As such, the verdict UNVERDICTED is appropriate. My concern reinforces rather than alters that verdict: the recommendation is grounded in a causal assumption about writing development that the paper itself admits is not yet supported by solid studies. The paper does cite relevant literature that could be interpreted as challenging the assumption, which makes the concern concrete rather than hypothetical. A randomized experiment would settle whether AI-assisted drafting harms, helps, or is neutral for writing development; until then, the recommendation remains an opinion informed by logic and prior writing research, not an evidence-based conclusion. The paper should receive credit for being transparent about the lack of systematic measurements and for presenting a balanced set of uses for AI (feedback, brainstorming, counterargument generation). However, the specific 'own words' requirement in the Writing Phase remains an empirical claim about an AI context, and the paper's own cited studies show the direction of evidence is not uniform. Hence the verdict should remain UNVERDICTED, and the proposed experiment would provide the missing evidence.","tokens_in":15848,"tokens_out":3131,"duration_ms":25760,"concrete_test":"Run a pre-registered randomized experiment with secondary students over one semester, comparing three conditions: (A) unassisted drafting, (B) AI-produced draft that students then revise and expand under teacher guidance, and (C) AI feedback on student-written drafts. Blind-rated pre/post writing competence (e.g., essays scored on content, structure, and voice) across conditions. If condition B shows non-inferior gains relative to A (within a pre-specified equivalence margin), the 'own words' premise in the Writing Phase section is falsified and the recommendation's foundation weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In the Writing Phase section, the paper asserts: 'it is essential for students to use their own words and create a draft' and 'It is not easy to see how AI can contribute sensibly in this phase without interfering with or even replacing the necessary process of the student formulating the text in their own language.' This is a causal claim about learning: active text generation is necessary for skill development, and delegating generation impairs it. The paper provides no empirical test of this in an AI context; it explicitly says 'we must largely rely on conclusions based on logic and reason.' Its own cited evidence cuts against a hard version of the claim. Marzuki et al. (2023) report unanimous positive impacts of AI writing tools on content and structure, Levine et al. (2024) find students can use ChatGPT without bypassing planning, drafting, and revising, and Doshi and Hauser (2024) show AI can boost creativity for less-creative individuals. If AI-assisted drafting can scaffold or model text production, the 'own words' requirement is not obviously necessary, and the paper's emphasis on controls and restrictions is underdetermined. This is load-bearing because the entire recommendation against AI-enabled shortcuts hinges on this premise, yet the paper itself acknowledges the absence of solid studies.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a conceptual review and position paper on the potential effects of generative AI (such as ChatGPT) on secondary-school students' writing competence and personal voice. It argues for a balanced approach in which AI supplements rather than replaces teacher guidance, and it recommends process-based assessments, authentic and personally meaningful assignments, and instruction in critical thinking as safeguards. The paper is organized around the stages of the writing process (pre-writing, writing, and revision/feedback), with additional sections on self- and peer assessment, AI-text detection, characteristics of AI-generated language, and implications for teachers' work.","tokens_in":16047,"tokens_out":4853,"duration_ms":43260,"significance":"The paper offers a timely and readable synthesis of recent literature and provides concrete, usable examples of prompts and assignment types, which is a genuine practical strength. It is also honest about the state of evidence, explicitly stating that 'we must largely rely on conclusions based on logic and reason' because 'there are yet no solid studies' (Introduction and Writing and AI sections). This transparency is commendable. However, the paper is a narrative synthesis rather than a systematic review, and its central recommendation rests on an empirical premise about student drafting that is asserted rather than tested. These limitations reduce its evidentiary weight but do not eliminate its value as a considered position statement for educators and policymakers.","major_comments":[{"comment":"The claim that 'it is essential for students to use their own words and create a draft' and that 'it is not easy to see how AI can contribute sensibly in this phase without interfering with or even replacing the necessary process of the student formulating the text in their own language' is presented as self-evident, but it is actually a substantive empirical premise about how writing skill develops. The manuscript itself cites evidence that is in tension with a hard version of this claim: Levine et al. (2024) found that upper secondary students can use ChatGPT as a constructive writing asset 'without bypassing the essential stages of planning, drafting, and revising'; Marzuki et al. (2023) report unanimous teacher-perceived positive impacts on content and structure; and Doshi and Hauser (2024) show creativity benefits for less-creative individuals. The authors should either qualify the claim, specify conditions under which AI-assisted drafting is or is not harmful, or present a theoretical argument for why these findings do not transfer to secondary school writing development. As written, the central policy recommendation against AI-enabled shortcuts depends on this undefended premise.","section":"Writing Phase"},{"comment":"The paper acknowledges that 'there are yet no solid studies clarifying this issue' and that it must rely on logic and reason. That admission is honest, but the paper does not then adopt a transparent argumentative structure that would let the reader evaluate its reasoning. In particular, it does not specify how the literature was selected, what counts as relevant evidence, or how the authors move from cognitive writing research (e.g., Bereiter and Scardamalia; Graham and Perin) to prescriptions about AI use. Several recommendations appear to be based on professional judgment rather than on cited studies, yet they are stated in the same assertive tone as evidence-based claims. The authors should clarify the nature of the contribution (e.g., position paper, research synthesis) and, if the argument is primarily from first principles, make the chain of reasoning and its premises explicit so that readers can assess the inference.","section":"Writing and AI / Introduction"},{"comment":"The paper's target population is secondary education, but many of the empirical studies it invokes involve university students or adult writers. Doshi and Hauser (2024) studied adult participants recruited online; Marzuki et al. (2023) surveyed EFL educators in Indonesian universities; Steiss et al. (2024) evaluated ChatGPT feedback on writing more generally, not specifically in secondary school; and Jeon and Lee (2024) review a body of work that is largely outside the K–12 context. The authors should address the extent to which findings from higher education and adult populations transfer to secondary school students, especially given the paper's concern for low-performing readers and 'functional illiterates.' This is load-bearing because the recommendations are specifically scoped to secondary education.","section":"Generalizability of cited evidence"}],"minor_comments":[{"comment":"The text cites 'Dosher et al. (2024)', but the reference list contains only 'Doshi, A. R. & Hauser, O. P. (2024)'; please correct the in-text citation to Doshi and Hauser (2024).","section":"Pre-writing Phase"},{"comment":"The text cites '(Steere, 2024)', but no such reference appears in the reference list; the nearby reference 'Shere, E. (2024)' about the anatomy of an AI essay is presumably intended, so please correct either the citation or the reference list.","section":"Characteristics of ChatGPT Language"},{"comment":"The discussion of Scarfe et al. (2024) refers to 'these exam sensors' and says 'These sensors were experts'; this wording is confusing. If the authors mean that human expert examiners were unable to detect the AI-generated submissions, they should write 'expert examiners' or 'experienced markers' rather than 'sensors.'","section":"Cheating and AI-generated Texts"},{"comment":"The in-text citation 'Fleckensten et al. (2024)' does not match the reference list entry 'Fleckenstein, J., Meyer, J., Jansen, T., Keller, S. D., Köller, O., & Möller, J. (2024)'; please align the spelling.","section":"Cheating and AI-generated Texts"},{"comment":"The citation 'Hogdes & Kirschner, 2024' is a typo; it should be 'Hodges & Kirschner, 2024' to match the reference list.","section":"Writing and AI"},{"comment":"The final sentence of the abstract and the final sentence of the introduction are nearly identical ('We argue for a balanced approach...'). This repetition should be removed or substantially rewritten to avoid redundancy.","section":"Abstract / Introduction"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is more of an informed essay than a research article, and it would benefit from being explicitly framed as a position paper with a clearly stated argumentative method. The self-citation (Eriksen, 2018) is modest and appropriate. If the editors see a strong fit with the journal's scope, revision along the lines of the major comments could make the contribution publishable; otherwise the paper might be better suited to a practitioner-oriented venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a position piece, not a research preprint, so judge it as a narrative review. The useful thing is the practical synthesis: it walks through pre-writing, writing, revision, and feedback stages, gives concrete prompts and a worked example of ChatGPT feedback, and lands on a sensible balanced recommendation. It is honest that solid studies are missing, and it flags the AI-detection problem (Scarfe et al.) and teacher detectability (Fleckenstein et al.) clearly.\n\nWhat is new is not the argument—the balanced approach already appears in Hodges & Kirschner and others—but the detailed Nordic classroom context and the worked examples. That gives it real value for teachers and teacher educators.\n\nThe soft spot is the load-bearing premise: the paper asserts that students must formulate drafts in their own words for writing competence to develop. It says this is essential, yet its own cited evidence complicates it. Levine et al. show students using ChatGPT without bypassing planning, drafting, and revising; Marzuki et al. report unanimous positive impacts on content and structure; Doshi and Hauser find AI boosts creativity for less creative writers. None of that refutes the premise, but it is enough to demand a softer formulation or an explicit engagement with the counter-evidence. The paper admits it relies on 'logic and reason,' which is thin for a policy argument that includes technological controls and assessment changes. There are also minor reference inconsistencies (Dosher/Doshi, Hogdes/Hodges, Gollins/Collins, Shere/Steere) that an editor should clean up.\n\nWho is this for? Teachers, teacher educators, and education policymakers in secondary writing instruction. Researchers looking for new empirical findings or a novel framework will not find it. It is a well-written, coherent review that deserves a serious referee in a venue that publishes position papers or teaching-oriented research. I would send it to peer review and accept it after revision, with a request to temper the strength of the 'own words' claim and deal more fairly with the studies that point the other way.","headline":"A practical, honest position review with a load-bearing but untested assumption about students needing to compose in their own words.","tokens_in":16570,"tokens_out":2348,"would_cite":false,"duration_ms":22134,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues AI should support, not replace, the student's own drafting in secondary writing instruction.","keywords":["AI in education","writing competence","writing voice","process-based assessment","ChatGPT","secondary education","argumentative writing","feedback"],"falsifier":"A controlled longitudinal study would settle it: randomly assign secondary students writing the same argumentative tasks to either draft entirely in their own words or freely draft with ChatGPT and then revise, keep assignments and feedback constant, and compare blind-scored writing competence at the end of the year. If the AI-drafting group matches or exceeds the own-words group, the paper's central premise is wrong.","tokens_in":15632,"feed_emoji":"✍️","tokens_out":5895,"duration_ms":51578,"temperature":0.7,"pith_summary":"This paper argues that generative AI can strengthen secondary-school writing instruction only if it is used as a supplement to, not a substitute for, the student's own writing. The authors claim AI belongs in the pre-writing and revision phases—as an idea generator, a devil's advocate for counterarguments, and a source of instant, criteria-linked feedback—but not in the drafting phase, where students must formulate text in their own words. They recommend process-based assessments, assignments that require personal reflection, and technological controls so that AI use does not erode writing competence or personal voice. The stakes are practical: schools must decide how to respond to ChatGPT's arrival without either banning a useful tool or letting it hollow out the writing instruction that builds literacy.","feed_headline":"Use AI for feedback, not for drafting student essays","feed_subtitle":"A review of secondary-school writing instruction makes the case for process-based assessment and personal tasks.","key_machinery":"The load-bearing mechanism is the staged model of the writing process—pre-writing, writing, revision, and publishing—combined with the requirement that the student's own words carry the writing phase. Within that model, AI's legitimate roles are defined by phase: in pre-writing it can activate ideas and supply counterarguments; in revision it can give concrete, criteria-based feedback that would otherwise exceed a teacher's time; in the writing phase it should not intervene. The paper also leans on the knowledge-transforming model of writing, in which converting ideas into text itself refines understanding, to explain why delegating drafting is costly, and on process-based assessment to make the student's journey visible.","core_discovery":"The paper's central claim is that the writing process itself is where writing competence develops, and that AI should be inserted at its edges, not its center. It distinguishes writing competence, the technical ability to compose a well-structured text, from writing voice, the personal expression that makes a text distinctive, and argues that both are put at risk when students delegate text generation to AI. On this view, AI is useful as a brainstorming partner, as a generator of counterarguments that strengthen argumentative texts, and as an instant feedback provider during revision, but the drafting phase must remain the student's own work. The paper therefore advocates a balanced approach: keep AI out of high-stakes, unmonitored writing, redesign assignments around authentic purposes and personal experience, and make assessment process-based so that learning-inhibiting shortcuts are not rewarded.","pith_inferences":["The authors do not spell this out, but the staged-process argument implies that acceptable AI use should be defined by phase: allowed for idea generation and feedback, prohibited for drafting in graded work, and made explicit in school regulations.","A testable extension would compare two-year writing growth in classrooms using AI only in pre-writing and revision against classrooms where AI is also allowed in drafting, holding assignments and assessment constant.","The logic implies that AI's role should shrink as the high-stakes nature of writing grows, suggesting different rules for low-stakes exploratory writing versus final graded products."],"forward_implications":["Teachers should redesign assignments so that success depends on personal experience, local knowledge, or individual judgment, making AI-generated responses hard to pass off as authentic.","Schools should weight process-based evidence—drafts, discussions, revisions—alongside final products when grading, weakening the incentive to copy AI text.","AI feedback can be scaled to provide immediate, criteria-focused comments in revision, especially for large classes, as long as prompts are short and concrete.","Exams and controlled writing conditions will remain necessary, since detection tools and even experienced teachers cannot reliably identify AI-generated texts.","Students can use AI as a tactical resource—finding counterarguments, testing rhetorical moves—without bypassing the planning and drafting stages."],"supporting_citations":[{"why":"Supplies the knowledge-transforming model of writing that explains why formulating text in one's own words refines understanding.","marker":"Bereiter & Scardamalia, 2013"},{"why":"Provides the evidence that students who work through writing stages perform better than those who do not.","marker":"Graham & Perin, 2007"},{"why":"Compares AI feedback with feedback from trained educators and finds human feedback superior, motivating continued teacher oversight.","marker":"Steiss et al., 2024"},{"why":"Offers the collaboration lens for human-AI writing and the idea of hybridized text origins.","marker":"Baron, 2023"},{"why":"Frames the risk of learning-inhibiting shortcuts and the case for process-based assessment.","marker":"Hodges & Kirschner, 2024"},{"why":"Shows that AI-generated exam submissions largely evade detection and receive higher grades, supporting the need for controlled writing conditions.","marker":"Scarfe et al., 2024"},{"why":"Demonstrates that teachers cannot reliably distinguish AI-generated from student-written texts, reinforcing reliance on process evidence.","marker":"Fleckenstein et al., 2024"},{"why":"Shows students using ChatGPT as a writing support without bypassing planning, drafting, and revising stages.","marker":"Levine et al., 2024"}],"fun_headline_variants":["AI as writing aid, not author: keep student voice","Let AI brainstorm, but students must draft","AI for feedback, not drafting: education review","Process-based writing with AI: feedback over generation","Balance AI in writing: supplement, don't replace teacher"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument stands on the premise that writing competence grows mainly through the student's own active text production, so letting AI generate drafts necessarily bypasses that growth; if AI-assisted drafting or feedback could transfer writing skill without full student authorship, the case against AI shortcuts would weaken.","fun_headline_variants_meta":{"raw":{"variants":["AI as writing aid, not author: keep student voice","Let AI brainstorm, but students must draft","AI for feedback, not drafting: education review","Process-based writing with AI: feedback over generation","Balance AI in writing: supplement, don't replace teacher"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00015,"raw_usage":{"total_tokens":1125,"prompt_tokens":802,"completion_tokens":323,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":418,"completion_tokens_details":{"reasoning_tokens":249}},"tokens_in":418,"tokens_out":323,"duration_ms":3273,"temperature":1.0,"reasoning_tokens":249,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:16:40.855005+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled longitudinal study would settle it: randomly assign secondary students writing the same argumentative tasks to either draft entirely in their own words or freely draft with ChatGPT and then revise, keep assignments and feedback constant, and compare blind-scored writing competence at the end of the year. If the AI-drafting group matches or exceeds the own-words group, the paper's central premise is wrong.","supporting_citations":[{"cited_title":"Turing Test","cited_arxiv_id":null,"evidence_quote":"Shows that AI-generated exam submissions largely evade detection and receive higher grades, supporting the need for controlled writing conditions."}],"review_version":1}