{"id":"a7321c1d-f072-456d-b33f-190737c4a413","arxiv_id":"2607.15006","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"Frequent AI users show more variable and slightly more overestimating authorship calibration scores than light users, though both group means are near zero.","lead":"This paper introduces a concept called authorship calibration, defined as the difference between how much users think they wrote in an AI-assisted document and how much they actually wrote. Analyzing the CoAuthor writing dataset, it reports that users who call on AI more often show wider, less consistent calibration scores than lighter AI users.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The only significance test compares signed calibration scores, not the absolute miscalibration the abstract claims; both group means are ~0, so 'more accurate' is unsupported.","rationale":"The reader's weakest assumption was the statistical independence of sessions, which is a valid methodological concern that undermines the reported p-value. I agree that clustering by participant is a serious issue. However, the more load-bearing problem is that the paper's only statistical test does not actually test the construct it claims to support. The abstract and discussion frame the result in terms of accuracy ('misjudge,' 'more accurate'), but the Mann-Whitney U test is applied to signed calibration scores, which measures a directional shift (under- vs overestimation), not the magnitude of miscalibration. Both group means are essentially zero, so even a significant difference in distributions would imply a negligible practical difference in signed bias. The 'accuracy' conclusion would require a comparison of absolute errors, which is absent. This is a conceptual gap that invalidates the central claim regardless of the independence issue. I therefore partially agree with the reader: the independence problem is real and should be fixed, but the construct mismatch is the decisive flaw. Since the reader's verdict of REJECT remains appropriate, I do not propose changing it.","tokens_in":8342,"tokens_out":4667,"duration_ms":52953,"concrete_test":"Re-run the analysis using the absolute calibration error |Declared - Actual| as the outcome. Compare Low- vs High-AI groups with a user-level analysis (e.g., compute each participant's mean absolute error across sessions, then compare groups via Mann-Whitney or a two-sample t-test) or a mixed-effects model with random intercepts for participants to account for session nesting. Report the effect size (e.g., Cohen's d) and a confidence interval. If the difference is not statistically significant or the effect size is trivially small (|d| < 0.2), the abstract's 'more accurate' claim for low-AI users is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that high-AI users 'tend to misjudge their authorship' while low-AI users 'exhibit more accurate authorship calibration.' Accuracy of calibration is formally defined in Section 3 as the absolute value |Declared - Actual|, with larger values indicating poorer calibration. However, the only inferential statistic reported in Section 5.2 is a Mann–Whitney U test on the signed calibration scores (Declared - Actual). That test compares whether one group's distribution tends to produce larger signed values (negative for low-AI, positive for high-AI) — it does not compare the degree of miscalibration, i.e., the absolute deviation from zero. The reported means are -0.004 and +0.003, both effectively zero, with standard deviations 0.112 and 0.146. Any conclusion that one group is 'more accurate' would require comparing absolute errors (or another accuracy metric) between groups; no such test is reported. The discussion's claim of 'more accurate authorship calibration' for low-AI users appears to rest on visual inspection of heatmaps and a subjective reading of SD differences, not on a formal statistical comparison of the construct the paper defines. Even if the Mann–Whitney test were valid and significant, it would not support the accuracy claim. Thus the central empirical assertion is not established by the evidence provided.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 'authorship calibration,' defined as the signed difference between a user's declared proportion of authorship and their actual proportion in AI-assisted writing, using the CoAuthor dataset. It categorizes 1,252 sessions into Low- and High-AI usage groups by median number of AI calls, reports descriptive statistics and a Mann–Whitney U test, and concludes that high-AI users tend to misjudge their authorship while low-AI users are more accurately calibrated. The paper frames the contribution as relevant to metacognition and learning analytics.","tokens_in":8602,"tokens_out":2180,"duration_ms":26294,"significance":"If the central claim were empirically established, the concept of authorship calibration could be a useful lens for studying human-AI collaboration and its educational implications. The paper has strengths: it builds on a public, rich dataset; the operationalization of calibration is simple and computable; and the authors provide code for reproducibility. However, the significance hinges on the empirical finding that AI usage frequency affects calibration accuracy, and that finding is not supported by the reported statistical analysis.","major_comments":[{"comment":"The abstract claims that heavy AI users 'tend to misjudge their authorship' whereas low-AI users 'exhibit more accurate authorship calibration.' Accuracy of calibration is defined in Eq. (1) and the text as the absolute magnitude |Declared − Actual|. However, the only inferential test reported in §5.2 is a Mann–Whitney U test on the signed calibration scores (Declared − Actual). With group means of −0.004 and +0.003, both effectively zero, this test does not show that one group misjudges more. It at most shows a difference in the location/shape of the signed distributions. No test compares absolute miscalibration between groups. The discussion's claim of 'more accurate authorship calibration' for low-AI users is therefore unsupported by the evidence presented.","section":"Abstract / §5.2"},{"comment":"The statistical comparison treats each writing session as an independent observation, but the 1,252 filtered sessions come from only 60 authors, with multiple sessions per author (Section 4.1). Sessions from the same author are likely correlated, violating the independence assumption of the Mann–Whitney U test. The reported p < 0.05 is therefore suspect. The analysis should account for clustering, for example by using a mixed-effects model or by aggregating scores per author before comparing groups.","section":"§5.2 / §4.1"}],"minor_comments":[{"comment":"The text refers to 'heatmap confirms this pattern ... (See Figure 3a)' and later 'heatmap distribution ... (See Figure 5b)'. Figure 3 is a calibration curve, not a heatmap; the internal cross-references should be corrected.","section":"§5.2"},{"comment":"There are typographical and grammatical errors, e.g., 'users from the the Low-AI usage group' and 'our results confirms that the accuracy of authorship calibration vary.' A careful proofread is needed.","section":"§5.2"},{"comment":"The operationalization of 'declared authorship' relies on a single survey item ('—% of the essay/story is written by me...'). The paper does not discuss the validity or reliability of this item as a measure of perceived authorship, especially when users may interpret 'written by me' differently in AI-assisted contexts.","section":"§4.3"},{"comment":"Several references are incomplete or informal (e.g., [4] is a Medium post, [24] is titled 'Genai et al. cocreation, authorship, ownership...'). The authors should ensure all citations meet the venue's standards.","section":"References"}],"recommendation":"reject","confidential_remarks":"The paper's central empirical assertion is not supported by the reported statistics: the Mann–Whitney test is run on signed scores while the accuracy claim concerns absolute deviations, and the group means are both essentially zero. Combined with the non-independence of sessions within authors, the main finding cannot be considered reliable. A revised version that tests absolute miscalibration with appropriate clustering might be reconsidered, but the current manuscript does not establish its primary conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the construct is worth discussing, but the headline claim doesn't follow from the analysis reported. The paper introduces 'authorship calibration' as declared minus actual authorship, applies it to the CoAuthor dataset, and compares Low- vs High-AI usage groups. That comparison is new relative to the cited literature, and the paper is clearly written, with a sensible operationalization for a first pass.\n\nThe soft spots are load-bearing. The abstract claims high-AI users 'tend to misjudge their authorship' and low-AI users show 'more accurate authorship calibration.' But the only inferential test is a Mann-Whitney U on signed calibration scores. That test tells you whether one group's scores tend to be larger (i.e., more positive), not whether one group is further from zero. The reported means are -0.004 and +0.003 — both effectively perfect calibration on average. The larger SD in the high group (0.146 vs 0.112) is suggestive, but no test on absolute miscalibration is reported. So the central claim is unestablished.\n\nThe independence assumption is also violated: 1,252 sessions from 60 authors are treated as independent observations. Multiple sessions per author are almost certainly correlated; a mixed model with random intercepts for authors, or a user-level analysis, is needed. Without it, the p-value means little. The median split into Low/High is data-dependent, and no robustness checks are given. Code is promised but not actually available ('will be publicly available upon acceptance'), which makes the analysis hard to verify. And the measurement of 'actual authorship' comes from CoAuthor metadata, with a single survey item for declared authorship, and no independent validation of the construct.\n\nThat said, I don't think this is a bad paper at heart. The concept of authorship calibration is plausible and could be useful in learning analytics. The dataset is public and rich. A reanalysis with absolute scores, mixed models, and effect sizes could well change the picture. As it stands, the paper overclaims.\n\nI'd send it to peer review rather than desk-reject, because the construct and research question deserve a careful referee — but I'd expect the authors to do the reanalysis before publication. The reader's reject verdict is fair for the current version.","headline":"Construct is plausible, but the paper's central accuracy claim is not tested by its own statistics.","tokens_in":9108,"tokens_out":2859,"would_cite":false,"duration_ms":30534,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"People who frequently use AI for writing tend to misjudge how much of the final text is their own work, while occasional users gauge it more accurately.","keywords":["generative AI","authorship calibration","AI-assisted writing","metacognitive calibration","human-AI collaboration","learning analytics","self-regulated learning","overestimation of authorship"],"falsifier":"Reanalyze the session-level data with a multilevel model that includes writer as a random effect and compute the contrast between low- and high-AI-usage sessions within writers; if the contrast is not robustly negative (or the confidence interval includes zero), the frequency conclusion fails. An independent replication with a per-session authorship question asked immediately after writing would also test whether the effect is a survey-recall artifact.","tokens_in":8183,"feed_emoji":"✍️","tokens_out":4789,"duration_ms":51744,"temperature":0.7,"pith_summary":"This paper introduces a measurable construct it calls authorship calibration: how closely a writer's declared percentage of personal contribution in an AI-assisted text matches the percentage they actually wrote. Analyzing session logs and self-reports from a public dataset of AI-assisted writing, the authors find that calibration varies widely across users, with many sessions showing little error but some deviating by more than 30 percentage points in either direction. The central result is that the frequency of AI use tracks calibration: people who rely heavily on AI tend to overestimate their own authorship, while people who use it less are better calibrated and if anything slightly underestimate their contribution. If true, this suggests that heavy AI assistance can blur a writer's sense of where their work ends and the machine's begins, which matters for learning because accurate self-assessment is a precondition for effective metacognitive monitoring and self-regulated learning.","feed_headline":"Heavy AI help makes writers overestimate their own contribution","feed_subtitle":"An analysis of over 1,200 AI-assisted writing sessions ties frequent AI use to inflated authorship estimates, a risk for learning.","key_machinery":"The load-bearing instrument is the authorship calibration score, defined as declared authorship minus actual authorship, where actual authorship is computed from text-provenance metadata in the interaction logs. This score operationalizes the otherwise abstract idea of awareness: zero means perfect calibration, positive means overestimation, negative means underestimation. The second piece is the median split on number of AI calls to form low- and high-usage groups, compared with a rank-based test for distributional difference.","core_discovery":"The paper's central claim is that authorship, a construct traditionally assumed to be transparent to the writer, becomes opaque in AI-assisted writing, and the degree of opacity is systematically related to how much AI help a user requests. Concretely, the authors define authorship calibration as the difference between a user's declared authorship (the percentage they say they wrote) and their actual authorship (the percentage of final sentences attributable to their own keystrokes). Across 1,252 sessions they find a near-zero average calibration score but high variability, and when they split sessions at the median number of AI calls, the heavy-AI group shows significantly lower calibration","pith_inferences":["The effort-blend explanation suggests a testable sharpening: if asked about the process (edits, prompts, selections) rather than the final product, heavy AI users may claim what they 'worked on' rather than what they 'wrote'; an experiment that varies how the authorship question is framed could show calibration depends on frame.","The same subtraction-of-declared-and-actual metric could be applied to code generation, image editing, or spreadsheet formulas, where 'authorship' is harder to define but equally relevant to learning analytics.","Since the dataset pools multiple sessions per writer, a per-writer analysis (e.g., mixed-effects model) would reveal whether the heavy AI effect is driven by a few habitually overestimating users or is robust across individuals.","The near-zero mean may hide a dissociation: light users are slightly underconfident and heavy users overconfident, meaning that the 'calibration problem' is really two distinct errors with possibly different mechanisms (conservatism vs. effort blend)."],"forward_implications":["In educational settings, miscalibrated authorship can distort metacognitive monitoring, potentially leading to overconfidence, reduced effort, or unrealistic self-assessment of learning.","Authorship calibration offers a concrete, computable outcome measure for evaluating human-AI writing tools and interaction designs.","Interventions that make text provenance visible (e.g., highlighting which sentences came from AI suggestions) could be evaluated by whether they move calibration scores toward zero.","The near-zero group mean shows most writers are reasonably calibrated, so the problem is not universal but concentrated in high-use sessions, suggesting targeted interventions for that group.","If heavy AI use degrades calibration, then pedagogical guidance on 'responsible AI use' may need to address self-awareness, not just product quality or plagiarism."],"fun_headline_variants":["AI power users misjudge their writing credit","Frequent AI users overrate their authorship","AI heavy use skews writers' sense of authorship","Heavy AI reliance distorts perceived authorship","AI heavy users misjudge how much they wrote"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The statistical comparison treats each writing session as an independent observation, but the data contain many sessions from the same writer, and sessions from one writer are likely correlated; if that correlation is strong, the reported group difference between light and heavy AI users could be an artifact rather than a real effect.","fun_headline_variants_meta":{"raw":{"variants":["AI power users misjudge their writing credit","Frequent AI users overrate their authorship","AI heavy use skews writers' sense of authorship","Heavy AI reliance distorts perceived authorship","AI heavy users misjudge how much they wrote"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1113,"prompt_tokens":657,"completion_tokens":456,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":401,"completion_tokens_details":{"reasoning_tokens":385}},"tokens_in":401,"tokens_out":456,"duration_ms":5493,"temperature":1.0,"reasoning_tokens":385,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:24:56.845286+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reanalyze the session-level data with a multilevel model that includes writer as a random effect and compute the contrast between low- and high-AI-usage sessions within writers; if the contrast is not robustly negative (or the confidence interval includes zero), the frequency conclusion fails. An independent replication with a per-session authorship question asked immediately after writing would also test whether the effect is a survey-recall artifact.","supporting_citations":[],"review_version":1}