{"id":"b9455daa-d3cd-4686-bd21-322e0ece11a9","arxiv_id":"2412.07200","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Modifying AI-generated text while writing improves essay quality on vocabulary, sentence complexity, and cohesion, while accepting AI text unchanged lowers quality.","lead":"This study analyzed logs of 1,445 AI-assisted writing sessions to see how students' interactions with a text generator affect essay quality. Writers who modified AI suggestions produced more sophisticated, complex, and cohesive essays, while those who accepted AI text unchanged produced lower-quality essays.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No confidence intervals or hypothesis tests are reported for the ATEs in Table 2, so the core claim that T3 'consistently improves' essay quality is not statistically supported.","rationale":"The reader's verdict is CONDITIONAL, centered on unobserved confounding. I agree that confounding is a threat, but I think the more immediate, load-bearing gap is the absence of any uncertainty quantification on the ATEs. A causal estimate with no confidence interval cannot support 'consistently improved' or 'saw a decrease'—those are inferential statements. The paper's refutation tests are frequently misinterpreted: a high p-value in Placebo/RCC/DSR indicates the refuted estimate is not significantly different from the original, not that the original differs from zero. In fact, if the original ATE is zero, the refutation p-values would also tend to be high, so these tests cannot rescue the claim. The dataset's nested structure (1,445 sessions, 63 writers) further means that naive standard errors would be too small; a bootstrap stratified by writer is a minimal check. This does not invalidate the paper, but it makes the central claim conditional on an analysis the manuscript does not report. I therefore keep the reader's CONDITIONAL verdict (UNCHANGED), adding this specific test as a precondition.","tokens_in":15960,"tokens_out":4934,"duration_ms":79150,"concrete_test":"Bootstrap the X-learner with 1,000 resamples of writing sessions stratified by writer, and compute percentile-based 95% confidence intervals for each ATE in Table 2. Also report a cluster-robust or mixed-effects variant. If any 95% CI for T3 on Y1, Y2, or Y3 includes zero, the claim of consistent improvement is unsupported. If all CIs exclude zero, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central causal claim is that T3 (modify GAI text) improves essay quality across Y1, Y2, and Y3, while T2 (accept without revision) degrades them. Table 2 presents only point ATE estimates, with no standard errors, confidence intervals, or hypothesis tests. The X-learner returns a point estimate; the reported refutation p-values (RCC, Placebo, DSR) only check whether a perturbed estimate differs from the original estimate—they do not test whether the original ATE differs from zero. Thus ATE_T3,Y1 = 0.102, ATE_T3,Y2 = 0.963, and ATE_T3,Y3 = 0.008 (and the negative T2 estimates) may all be sampling noise. This concern is prior to unobserved confounding: even with perfect adjustment, a null ATE would falsify the headline. Additionally, 1,445 sessions are nested in only 63 writers; any uncertainty quantification must account for writer-level clustering, or CIs will be overconfident. The authors' own Limitations section acknowledges missing writer-level confounders, but it does not address the absence of inferential statistics.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how different patterns of interaction with generative AI (GAI) during writing affect the quality of the resulting essay. Using the CoAuthor dataset (1,445 writing sessions from 63 writers), the authors define three binary treatments: T1 (seeking suggestions but not accepting them), T2 (accepting GAI suggestions without revision), and T3 (accepting and then modifying GAI suggestions). Outcomes are four text-based quality measures: lexical sophistication, syntactic complexity, text cohesion, and gender bias. The authors construct a DAG, apply the X-learner with back-door adjustment on five confounders, and report average treatment effects (ATEs). They claim that T3 consistently improves lexical sophistication, syntactic complexity, and cohesion, while T2 reduces them, and that all three behaviors reduce gender bias. The paper includes refutation checks (Random Common Cause, Placebo Treatment, Data Subset Refuter) and SHAP-based subgroup analyses.","tokens_in":16161,"tokens_out":5330,"duration_ms":55472,"significance":"If the causal claims were adequately supported, the paper would make a useful contribution to GAI-assisted writing research and educational assessment. It addresses a timely question, uses a publicly available dataset, and is among the first to attempt causal effect estimation for writing behaviors in this setting. The authors provide a clearly stated DAG, a reproducible code repository, and several refutation checks, which are strengths. However, the central quantitative claims are currently not backed by inferential statistics: reported ATEs come without standard errors, confidence intervals, or tests against zero, and the paper does not account for the nested structure of the data. These gaps are load-bearing because the abstract and discussion state that effects are 'significant' and 'consistent.' The manuscript therefore needs substantial revision before its conclusions can be accepted.","major_comments":[{"comment":"The ATE estimates are reported without standard errors, confidence intervals, or p-values testing the null hypothesis ATE = 0. The p-values in the RCC, Placebo, and DSR columns test whether the refutation estimate differs from the original point estimate, not whether the original ATE differs from zero. For example, ATE_T3,Y2 = 0.963 may be well within sampling noise; the reported refutation p-values do not rule out a true effect of zero. Consequently, the abstract and Discussion statements that T3 'consistently and significantly improves' all three quality measures are not supported by the evidence as presented. Please report inferential statistics for each ATE, and correct for multiple comparisons across the twelve treatment-outcome pairs.","section":"Section 4.1, Table 2"},{"comment":"The 1,445 writing sessions are produced by only 63 writers, so the observations are not independent. The paper does not account for writer-level clustering in the X-learner estimation or in the refutation procedures. Any confidence intervals or significance tests added in response to the previous comment must be cluster-robust at the writer level (or use a multilevel model); otherwise, uncertainty is understated and the refutation p-values are overconfident. The present analysis gives no indication of how much of the variation in outcomes is between-writer versus within-writer, which is essential for interpreting effects of behaviors that are inherently writer-level tendencies.","section":"Sections 3.1 and 4.1"},{"comment":"Treatments T1, T2, and T3 are defined by dichotomizing each behavioral frequency at the sample median (Section 3.2, third paragraph). No sensitivity analysis is reported for this choice of threshold, and the median split is effectively a free parameter. It is plausible that the estimated ATEs, especially the smaller effects on Y3 (text cohesion) and Y4 (gender bias), are sensitive to the cutoff. Please report results at alternative thresholds (e.g., quartiles) or use a continuous treatment representation (e.g., dose-response or ordinal treatment) to assess the robustness of the qualitative conclusions.","section":"Section 3.2, treatment binarization"},{"comment":"The back-door adjustment set (C1-C5) is assumed sufficient for identifiability, but the Limitations section acknowledges that writer-level variables such as GAI literacy are missing. If writing skill, motivation, or GAI literacy affects both treatment assignment (behavioral pattern) and essay quality, the ATEs are biased. Because the paper's headline claims are causal, this is a load-bearing threat, not a routine caveat. A concrete sensitivity analysis (e.g., E-values or a negative-control outcome) would help bound the potential bias from unmeasured confounding and is necessary to support the strength of the current causal language.","section":"Section 3.2 and Limitations"}],"minor_comments":[{"comment":"The numbering of confounders is inconsistent: the bullet list in Section 3.2 calls language background C2 and writing topic C3, while Table 1 and Table 3 use the reverse labeling. Please align the numbering throughout the text, tables, and figures.","section":"Section 3.2 and Table 1"},{"comment":"The description of T3 contains a typo: 'Ratio between teh accepted GAI suggestions' should be 'Ratio between the accepted GAI suggestions.'","section":"Table 1"},{"comment":"The phrase 'T3 (seek suggestions -> first accept and the revise)' should read 'T3 (seek suggestions -> accept and then revise).'","section":"Section 4.1, fourth paragraph"},{"comment":"The beeswarm plots are referenced extensively in Section 4.2, but the captions in the manuscript do not describe the axis labels or the meaning of dot colors, and the axes are not legible in the provided version. Please add explicit labels and a legend.","section":"Figures 2-4"},{"comment":"The sentence 'the imbalance in data distribution (e.g., no-native writers vs. native writers)' contains a typo: 'no-native' should be 'non-native.'","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and important question, and the public dataset and code link are assets. However, the complete absence of uncertainty quantification for the central ATE estimates is a serious gap that currently prevents the headline claims from being supported. The issues are addressable within the scope of the paper (adding standard errors, cluster-robust inference, sensitivity analyses), so I recommend major revision rather than rejection. Please also ask the authors to clarify the confounder numbering inconsistency and to ensure the DAG and beeswarm plots are legible in the final version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is the first to run causal inference (X-Learner with back-door adjustment) on the CoAuthor data, and it frames a genuinely useful pedagogical question: whether accepting AI text without revision, or revising it, causally changes essay quality. Credit where due: the dataset is public, the DAG is explicit, the five confounders are named, code is released, and the outcomes are external text metrics rather than self-reports. The refutation tests (RCC, Placebo, DSR) are a reasonable robustness habit. So the paper is not a throwaway.\n\nThe load-bearing problem: Table 2 reports only point ATEs. No standard errors, no confidence intervals, no hypothesis tests. The refutation p-values >0.05 tell you the estimate is stable under perturbations, not that it differs from zero. With 1,445 sessions nested in 63 writers, the effective sample size is far smaller than the session count, and ignoring writer-level clustering overstates precision. The Discussion calls results 'significant' and 'consistently improve,' but no test in the paper supports those words. That is not a minor omission; it is the difference between an established effect and a suggestive pattern.\n\nSome softer spots: treatments are binarized at the sample median with no sensitivity analysis around that threshold, and the subgroup trends in Table 3 are selected post hoc from beeswarm plots by eye. Those are hypothesis-generating, not confirmatory. The Limitations honestly flag missing writer-level confounders like GAI literacy, which is good, but they do not mention the absent inferential statistics.\n\nNone of this means the effect is fake. The directions are plausible and line up with earlier correlational work. But as written, the paper overstates what it has shown. A careful reader should treat the ATEs as descriptive point estimates, not established causal effects.\n\nWho gets value: researchers in AI-assisted writing and learning analytics will want to cite this as the first causal attempt on CoAuthor, provided they cite it as an approach rather than as evidence. The paper deserves a serious referee: the question is relevant, the method choice is defensible, and the gaps are fixable. I would send it out with a request for major revision—add uncertainty quantification with writer-level clustering, test ATEs against zero, run sensitivity analyses on the median split and unobserved confounding, and tone down 'significant' and 'consistently' unless the tests support them.","headline":"First causal-inference attempt on CoAuthor is worth engaging, but the headline ATE claims lack the standard errors and tests needed to support 'significant' and 'consistently improves.'","tokens_in":16718,"tokens_out":1948,"would_cite":true,"duration_ms":24719,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Frequent modification of AI-generated text improves lexical sophistication, syntactic complexity, and text cohesion, while verbatim acceptance degrades them.","keywords":["generative AI-assisted writing","causal inference","X-learner","essay quality","lexical sophistication","syntactic complexity","text cohesion","gender bias"],"falsifier":"A reanalysis that adds writer-level covariates, such as independent writing ability measured before AI exposure, GAI literacy, motivation, or prior topic knowledge, to the adjustment set would falsify the causal claim if the T3 average treatment effects on lexical sophistication, syntactic complexity, or cohesion collapse toward zero. A more direct falsifier is a randomized experiment in which writers are instructed either to revise or to accept AI suggestions verbatim; observing no difference in the three text-quality metrics would refute the paper's causal interpretation.","tokens_in":15773,"feed_emoji":"✍️","tokens_out":6109,"duration_ms":61223,"temperature":0.7,"pith_summary":"This paper tries to establish that the way a student engages with AI writing suggestions changes the quality of the final essay, not just whether AI is used. On 1,445 logged GPT-3-assisted writing sessions from the CoAuthor dataset, it compares three behaviors: asking for suggestions and not accepting them (T1), accepting them unchanged (T2), and accepting then revising them (T3). Its central result is causal: frequent T3 raises lexical sophistication, syntactic complexity, and text cohesion, while frequent T2 lowers all three. It also finds that all three behaviors reduce gender bias in essays, with the largest reduction for T1. If these estimates hold, teachers can use observable writing-process logs to distinguish meaningful engagement from passive delegation.","feed_headline":"Revising AI text raises essay quality; accepting it drops","feed_subtitle":"Writers who modify AI suggestions improve vocabulary, syntax, and cohesion; passive acceptance degrades them.","key_machinery":"The argument is carried by a causal graph plus the X-learner meta-algorithm, a machine-learning approach that combines treated and untreated observations to estimate average and individual treatment effects. Behaviors are encoded as binary treatments (above or below the median frequency of the pattern), the four essay metrics as outcomes, and five confounders—writing genre, writing topic, language background, GPT temperature, and frequency penalty—as the back-door adjustment set. The back-door criterion identifies the causal estimands from observational data, and X-learner estimates both average and individual treatment effects. The paper uses random common cause, placebo treatment, and data subset refutations to check that the estimates are not artifacts of the model.","core_discovery":"On the paper's own terms, the central discovery is that accepting AI-generated text and revising it (T3) is the only one of the three GAI-assisted writing behaviors that consistently improves the three text-quality measures; the estimated average treatment effects are positive for lexical sophistication, syntactic complexity, and text cohesion (0.102, 0.963, and 0.008 in the paper's reported ATE table). Accepting AI suggestions verbatim (T2) is associated with the largest negative effects on all three, and seeking suggestions without accepting them (T1) lowers lexical sophistication and syntactic complexity while slightly raising cohesion. The same causal model yields a separate finding: all three behaviors reduce gender bias in the essays, with T1 producing the largest reduction, which the paper interprets as evidence that human independent writing itself introduces measurable linguistic bias. Because the effects pass the paper's refutation checks, the authors present these as credible causal relationships rather than correlations.","pith_inferences":["Going beyond the paper: the estimated average treatment effects come from an adult crowd-sourced writing population, so a natural extension is to test whether the T3 advantage is larger for lower-skill writers, who may have more room to learn from revision, or for higher-skill writers, who may revise more effectively.","An editorial inference: if verbatim acceptance genuinely lowers quality, current AI autocomplete interfaces may be nudging users toward worse text by making acceptance the default action; redesigning them to require a deliberate edit could be a low-cost experiment.","The gender-bias result for T1, where seeking suggestions without accepting them reduces bias the most, suggests that merely reading AI suggestions may change a writer's attention; this could be tested directly by comparing essays written after exposure to de-biased brainstorming prompts against a no-prompt control.","The authors list missing GAI literacy as a limitation; a follow-up could also record revision quality, such as the number of semantic changes made to AI text, which would make the hypothesized learning mechanism testable rather than inferred from behavior alone."],"forward_implications":["Instructors should treat final essay quality alone as insufficient evidence of learning in GAI-assisted writing; process logs showing whether suggestions were revised carry assessment-relevant information.","Pedagogical guidance should push students toward critically editing AI suggestions rather than accepting them verbatim, since verbatim acceptance is estimated to reduce all three text-quality measures.","Non-native English writers may gain lexical sophistication and syntactic complexity from GAI suggestions, but they also show higher gender bias when writing independently or when revising AI text, suggesting targeted support is needed.","The benefit of revising AI suggestions depends on context: it improves cohesion more in creative writing than in argumentative writing, and the effects shift with GPT temperature and frequency penalty settings.","All three GAI behaviors reduce gender bias relative to independent writing, so AI-assisted writing can serve as a partial bias-mitigation strategy even when suggestions are accepted unchanged."],"supporting_citations":[{"why":"Supplies the 1,445 GPT-3-assisted writing sessions and log traces that all analyses use.","marker":"[44]"},{"why":"Provides the X-learner algorithm used to estimate average and individual treatment effects.","marker":"[39]"},{"why":"States the back-door criterion used to identify causal effects from the causal DAG.","marker":"[48]"},{"why":"Defines lexical sophistication, syntactic complexity, and text cohesion as the outcome measures.","marker":"[18]"},{"why":"Motivates the gender-bias outcome and supplies the downstream gender-bias measure used for Y4.","marker":"[61]"},{"why":"Defines the Genbit gender-bias metric used to score the Y4 outcome.","marker":"[55]"},{"why":"Provides the refutation methods (random common cause, placebo, data subset) used to validate the estimates.","marker":"[56]"},{"why":"Supplies the planning-translating-reviewing writing model used to interpret the three behaviors as levels of cognitive engagement.","marker":"[26]"}],"fun_headline_variants":["Edit AI drafts for better essays; accept at your own risk","Modifying AI suggestions boosts essay quality, study finds","Active AI text editing lifts writing quality; passive use hurts","Tweak AI text to improve essays; don't just accept it","Writers who revise AI text score higher on essay quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the five logged confounders—genre, topic, language background, GPT temperature, and frequency penalty—are enough to block all back-door paths, so unmeasured writer traits like skill, motivation, or GAI literacy do not distort the estimated effects.","fun_headline_variants_meta":{"raw":{"variants":["Edit AI drafts for better essays; accept at your own risk","Modifying AI suggestions boosts essay quality, study finds","Active AI text editing lifts writing quality; passive use hurts","Tweak AI text to improve essays; don't just accept it","Writers who revise AI text score higher on essay quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000667,"raw_usage":{"total_tokens":3088,"prompt_tokens":1037,"completion_tokens":2051,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":1968}},"tokens_in":653,"tokens_out":2051,"duration_ms":16283,"temperature":1.0,"reasoning_tokens":1968,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:00:57.153147+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reanalysis that adds writer-level covariates, such as independent writing ability measured before AI exposure, GAI literacy, motivation, or prior topic knowledge, to the adjustment set would falsify the causal claim if the T3 average treatment effects on lexical sophistication, syntactic complexity, or cohesion collapse toward zero. A more direct falsifier is a randomized experiment in which writers are instructed either to revise or to accept AI suggestions verbatim; observing no difference in the three text-quality metrics would refute the paper's causal interpretation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 1,445 GPT-3-assisted writing sessions and log traces that all analyses use."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the X-learner algorithm used to estimate average and individual treatment effects."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"States the back-door criterion used to identify causal effects from the causal DAG."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines lexical sophistication, syntactic complexity, and text cohesion as the outcome measures."},{"cited_title":"Sengupta, R","cited_arxiv_id":null,"evidence_quote":"Defines the Genbit gender-bias metric used to score the Y4 outcome."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the planning-translating-reviewing writing model used to interpret the three behaviors as levels of cognitive engagement."}],"review_version":1}