{"id":"35cfa8fc-5b41-41f2-b568-ab73873660c8","arxiv_id":"2504.13629","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Using 627,000 arXiv abstracts, this paper detects ChatGPT-style revisions and reports heterogeneous adoption rates plus a convergence of detected users' writing toward senior-author style.","lead":"A detector trained on ChatGPT-style revisions finds that around 22 percent of computer-science abstracts on arXiv were AI-polished by the end of 2023, with adoption highest among non-native speakers and junior researchers. The same data links detected AI use to a convergence of junior, male, and non-native authors' writing toward a senior-like style, raising both quality and homogenization concerns.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The convergence DiD is circular: GPT adoption is assigned by a classifier trained on GPT-3.5 revisions, while the outcome measures similarity to that same revision style; the effect may reflect the detector's own signal.","rationale":"The paper's most defensible contribution is the descriptive adoption measurement: a large-scale, reproducible pipeline with reported accuracy on a synthetic test set, and the pre-period normalization is a sensible correction for baseline classifier error. However, the headline causal claim depends on the DiD in Section 4.4, and that DiD cannot separate the effect of GPT use from the detector's construction. The reader's rationale already names this circularity; our emphasis differs from the formal 'weakest_assumption' (detector generalization and ground-truth contamination), hence partial agreement. If external validation of treatment labels were supplied, or if the DiD were re-run with an instrument for ChatGPT access (for example, staggered country-level availability or organizational API access), the convergence result could be rescued. As it stands, the classifier's 96-99% test accuracy is not evidence against this concern because the test set is generated by the same revision process that defines the outcome. The release-date error in Section 2.1 is a separate validity threat to the human ground truth, and the reported '740-fold' growth is arithmetically inconsistent with the stated 0.47% to 13.8% change (about 29-fold); neither is the central structural flaw, but both reinforce the need for caution. The reader's REJECT verdict remains appropriate.","tokens_in":30856,"tokens_out":4621,"duration_ms":46766,"concrete_test":"Draw a stratified random sample of roughly 2,000 arXiv abstracts from January through December 2023 that the classifier labels as GPT-adopted, plus an equal number labeled non-adopted; obtain external ground truth by surveying corresponding authors or using explicit acknowledgments or declarations of ChatGPT use, and re-estimate the Figure 10-11 difference-in-differences with this externally validated treatment indicator, keeping the senior-similarity outcome computed from raw text. If the DiD coefficient attenuates to zero, the circularity concern is confirmed; if it survives, the convergence claim is independent of the detector.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is circularity in the convergence difference-in-differences. Treatment ('GPT adoption') is assigned by classifiers (Section 2.1) fine-tuned to separate arXiv abstracts from their own GPT-3.5-Turbo revisions under six fixed prompts, so a text is labeled GPT-revised precisely when it is stylistically close to those synthetic revisions. The outcome in Section 4.3 is cosine similarity to the same GPT-revised template; in Section 4.4 it is similarity to senior writing, which Table 6 shows is itself mimicked by the GPT revisions. Thus Figures 10 and 11 may be measuring the classifier's own signal rather than an independent behavioral change: post-November-2022 abstracts labeled 'adopters' are, by construction, closer to the style that defines convergence. High F1 on the synthetic test set (Table 2) does not resolve this; it only confirms separation between original abstracts and their own GPT-3.5 rewrites. The release-date error in Section 2.1 (GPT-3's API launched in June 2020, not only after November 2021) compounds the training-label problem, but the structural issue is that treatment and outcome share the same stylistic measurement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a detector for GPT-revised arXiv abstracts by fine-tuning sentence-transformer models on pairs of human-written abstracts and GPT-3.5-Turbo revisions produced under six prompts, across eight arXiv disciplines (48 binary and 8 multiclass models). It applies the detector to 627,384 arXiv abstracts from 2021-2023, reports aggregate adoption rates (e.g., 22% of Computer Science abstracts by the end of 2023), heterogeneous adoption by field, nativeness, ethnicity, gender, and seniority, analyzes differences in ten writing rules between original and revised texts, and uses a difference-in-differences-style comparison of cosine similarities to claim that GPT use drives junior, male, and non-native researchers' writing toward senior styles. The central claims are: (i) reliable detection of GPT revisions at scale, (ii) large and heterogeneous adoption, and (iii) causal style convergence from GPT adoption.","tokens_in":31040,"tokens_out":8866,"duration_ms":78605,"significance":"If the detection and difference-in-differences results were valid, this would be a valuable large-scale measurement of LLM adoption in scientific writing, with clear implications for science policy and the study of AI-generated content. The strengths are the large corpus, the explicit separation of six revision prompts, and the internal consistency of the synthetic test-set metrics, with out-of-sample F1 scores above 0.95 for most prompts. However, the evaluation is entirely internal: the test set is generated by the same model and prompts used for training, and no external validation against real ChatGPT-edited abstracts is provided. The convergence analysis is subject to a direct circularity concern, the stated ground-truth period for 'human-written' text is factually contaminated by GPT-3's June 2020 API release, and the reported difference-in-differences is not actually estimated with any formal specification. Unless these issues are resolved, the headline adoption and convergence numbers are not supported.","major_comments":[{"comment":"The training ground truth is not reliably human-written. The paper states that texts updated before November 30, 2021 are used because 'GPT-3 was publicly released only after this period,' but the GPT-3 API was publicly released in June 2020, before the training window (articles updated before October 1, 2021) and the test window (October 1-November 30, 2021). Abstracts in the 'human' class may therefore already include GPT-3-revised text. This contaminates the training labels for all 48 classifiers and invalidates the reported F1 scores as evidence of detection quality; it also breaks the assumption that pre-November-2022 ground-truth abstracts are uncontaminated.","section":"Section 2.1"},{"comment":"The detector's external validity is unestablished. All training and test revisions were created by GPT-3.5-Turbo under six fixed prompts written by the authors, and the held-out test set in Table 2 is from the same pipeline. High precision and recall on this test set only show that the model can separate original abstracts from their own GPT-3.5 rewrites. Real-world ChatGPT use involves different prompts, different model versions, and possibly human post-editing. No validation against externally labeled real-world GPT-revised abstracts is reported, so the absolute adoption levels in Section 3.2 and the treatment assignments in Sections 4.3 and 4.4 are not calibrated for the population to which they are applied.","section":"Section 3.2, Table 2"},{"comment":"The convergence difference-in-differences is circular. GPT adoption is assigned by classifiers fine-tuned to recognize the exact GPT-3.5 revision style, while the thought-experiment outcome is cosine similarity to the version-6 GPT revision and the real-world outcome is similarity to senior writing, which Table 6 shows is itself mimicked by the GPT revisions. An abstract labeled as an adopter is therefore closer to the convergence target by construction once ChatGPT is used, irrespective of any independent behavioral mechanism. Figures 8-11 may thus be measuring the detector's stylistic signal rather than a genuine convergence of researchers' writing practices.","section":"Sections 4.3-4.4, Figures 10-11"},{"comment":"No formal difference-in-differences estimates are reported. The text claims that a difference-in-differences analysis shows GPT-driven convergence, but the figures present only monthly averages of cosine similarity, with no regression coefficients, standard errors, or pre-trend tests. It is therefore impossible to assess the statistical significance or magnitude of the convergence, or to control for group-specific time trends and composition changes. The absence of a formal DiD specification is a load-bearing gap for the paper's central causal claim.","section":"Sections 4.3-4.4, Figures 10-11"},{"comment":"The pre-period adjustment assumes a constant false-positive rate. The paper subtracts the average pre-November-2022 adoption rate from all subsequent months to remove 'misclassification.' This is valid only if the classifier's false-positive rate is constant over time. If it drifts because human writing style evolves, because the arXiv population changes, or because the model's operating point interacts with text length or field composition, the adjusted adoption levels are biased. No evidence is provided for a constant false-positive rate, and the per-discipline pre-period levels in Figure 2 vary substantially, which suggests the adjustment is at best approximate.","section":"Section 3.2"}],"minor_comments":[{"comment":"The heading 'Mothods' should be 'Methods.'","section":"Section 2.1"},{"comment":"The text says the increase from 0.47% to 13.8% is a '740-fold increase'; the correct ratio is approximately 29-fold (13.8 divided by 0.47).","section":"Section 3.2"},{"comment":"'Table ?? presents the main results' contains an unresolved cross-reference; the intended table is presumably Table 5.","section":"Section 3.4"},{"comment":"The confusion matrix shows that revisions 4 and 5 are classified correctly only 23.76% and 18.46% of the time, respectively; the prompt-level adoption and revision-mix results in Section 3.3 should explicitly discuss this near-chance performance.","section":"Table 4"},{"comment":"The note says 'arVix' instead of 'arXiv.'","section":"Table 9"}],"recommendation":"reject","confidential_remarks":"To the editor: the gap between the paper's claims and its evidence is unusually large. The ground-truth contamination (the GPT-3 API was released in June 2020) is a factual, easily checked error, and the convergence analysis is structurally circular. Even with a corrected release date, the paper would need an externally validated detector and a formal DiD regression to support its headline numbers. I recommend rejection. The topic is timely and the large-scale corpus is a useful starting point, so a future version with external validation and non-circular outcome measures could be worth reconsidering."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a big-data measurement paper with a genuinely interesting descriptive core—discipline- and prompt-specific detection of GPT-revised arXiv abstracts—and a central causal claim that doesn't survive contact with its own design. The convergence difference-in-differences is circular, and the paper has two avoidable numerical/factual mistakes that make me trust the authors' attention to detail less than I'd like.\n\nWhat's new and good: they build 48 fine-tuned classifiers (8 fields × 6 prompts) on 343k arXiv abstracts, with synthetic GPT-3.5 revisions. They show high out-of-sample F1 (~0.95–0.99). The adoption curves—CS 22%, EE&SS 21%, Math 3.5% by end of 2023—are striking if the detector generalizes. Heterogeneity by nativeness, ethnicity, and seniority is a useful descriptive map. The writing-rule analysis (word count, hedge words, tense) is a reasonable thought experiment.\n\nThe main problem is Sections 4.3–4.4. Treatment ('GPT adopters') is defined by the detector, which is trained to separate original abstracts from their own GPT-3.5 revisions. So a text is labeled GPT-revised exactly when it is stylistically close to those revisions. Then the outcome in 4.3 is cosine similarity to the GPT-revised version, and in 4.4 similarity to senior style, which Table 6 shows GPT mimics. Of course adopters converge to GPT style—the label and the outcome share the same measurement. The high F1 on the synthetic test set doesn't break the circularity; it just confirms the classifier can tell originals from their own rewrites.\n\nOther issues: pre-period subtraction assumes a constant false-positive rate, which is plausible but not defended. The '740-fold increase' is arithmetically wrong: 0.47% to 13.8% is about 29-fold. And Section 2.1 states GPT-3 was released after Nov 2021, but the GPT-3 API launched in June 2020. That is factually wrong and matters for the training-label assumption.\n\nWho is this for? People studying AI-mediated academic writing or science-of-science. They would cite the adoption numbers, but should treat the convergence results with caution. The paper deserves peer review in the sense that a serious referee should see it—the data collection is substantial and the descriptive statistics are potentially valuable. But the convergence analysis as written is not interpretable as a causal effect. I'd recommend sending it out, with the expectation that major revision is needed, or that the authors reposition it as a descriptive measurement paper and drop or rewrite the DiD.","headline":"Useful large-scale adoption measurements undermined by a circular convergence design and a few embarrassing errors.","tokens_in":31589,"tokens_out":2632,"would_cite":false,"duration_ms":24207,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By the end of 2023, about one in five computer-science abstracts on arXiv had been revised with ChatGPT, and GPT-assisted revision is measurably pulling junior, male, and non-native researchers' writing closer to the style of senior…","keywords":["ChatGPT detection","arXiv abstracts","scientific writing style","generative AI adoption","difference-in-differences","LLM text classification","writing convergence","heterogeneous adoption"],"falsifier":"Take a hand-labeled corpus of abstracts whose authors disclose or can independently verify their ChatGPT use with unrestricted prompts, and test whether the paper's classifiers reproduce the reported adoption rates; a mismatch would show the detector does not generalize beyond its six fixed prompts, and checking pre-2021 abstracts against known GPT-3-era output patterns would expose whether the supposed human ground truth is contaminated.","tokens_in":30597,"feed_emoji":"🤖","tokens_out":10589,"duration_ms":85884,"temperature":0.7,"pith_summary":"This paper claims that ChatGPT-assisted revision of arXiv abstracts became widespread within a year of ChatGPT's release and that its adoption was uneven across disciplines and researcher groups. Using more than 627,000 papers posted between January 2021 and December 2023, the authors fine-tune discipline- and prompt-specific classifiers to detect GPT-style revisions, and report that by the end of 2023 about 22% of Computer Science abstracts had been revised with GPT, against 3.5% in Mathematics. They also find that GPT revisions shorten abstracts, push writing toward present tense and formal neutrality, and, according to a difference-in-differences analysis, drive junior, male, and non-native researchers' writing closer to the style of senior researchers. The paper matters because it turns speculation about AI's footprint in scientific writing into quantitative, group-level estimates, and it frames the main risk as homogenization of academic expression rather than simple quality improvement.","feed_headline":"ChatGPT revised 22% of computer science abstracts by end of 2023","feed_subtitle":"Adoption ranges from 3.5% in math to 22% in CS; GPT revision pulls junior writers toward senior style","key_machinery":"The machinery is a family of fine-tuned sentence-transformer classifiers, one per field and revision prompt, trained to separate human-written arXiv abstracts from GPT-3.5-Turbo versions generated under six fixed zero-shot prompts (clarity, formality, objectivity, readability, grammar, and comprehensive revision). A second component is the writing-rule regression: ten quantitative measures of style such as word count, sentence count, present-tense ratio, hedge-word frequency, superlatives, and evocative words, compared within the same article using article fixed effects between the original and GPT-revised text. The convergence analysis uses bag-of-words cosine similarity between original and GPT-revised abstracts, and between junior and senior, male and female, and native and non-native author groups, in a difference-in-differences design that compares adopters with non-adopters before and after ChatGPT's release.","core_discovery":"The central discovery is that ChatGPT-revised scientific abstracts are detectable at scale and that the detectable signal is strong enough to map adoption and style convergence across groups. The paper constructs a ground-truth corpus from 343,461 abstracts last updated before ChatGPT's release, generates revised versions with GPT-3.5-Turbo along six prompt dimensions (clarity, formality, objectivity, readability, grammar, and comprehensive revision), and fine-tunes 48 field- and prompt-specific binary classifiers plus 8 multiclass classifiers on those pairs, with out-of-sample F1 scores above 0.95. Applying these classifiers to 627,384 arXiv abstracts, the paper reports adoption rising from 0.47% in January 2023 to 13.8% in December 2023, with end-2023 rates ranging from 3.5% in Mathematics to 22% in Computer Science. In the style analysis, GPT revisions reduce word count by more than 25% for targeted prompts, cut hedge-word usage by about a third, and in 8 of 11 writing-rule comparisons move text toward the profile of senior authors. The difference-in-differences results then show real-world convergence: junior, male, and non-native researchers who adopt GPT shift toward senior and native writing patterns, while non-adopters and female authors show little or no convergence.","pith_inferences":["The measurement of convergence may be partly circular: the classifier that labels an abstract as GPT-revised was trained on the same GPT style features used to define convergence, so part of the reported junior-to-senior convergence could reflect the detector's template rather than an independent stylistic shift.","Because the synthetic revisions come from GPT-3.5-Turbo under six fixed prompts, real-world use of newer models or different instructions is likely under-detected; the true late-2023 adoption rates could be higher than 22%, and the revision-type shares may not describe actual prompt choices.","The 'democratizing' reading, that GPT raises the writing of junior and non-native researchers, has a plausible selection alternative: less fluent writers choose GPT, so part of the observed improvement reflects who adopts, not what GPT does; the within-paper fixed-effects comparisons mitigate this but do not fully rule it out.","A direct extension would apply the same detector to other genres, such as peer reviews, grant proposals, or non-English abstracts, where similar convergence could either speed knowledge exchange or erase recognizable author voice; the paper's framework is ready-made for that test."],"forward_implications":["By the end of 2023, roughly 22% of Computer Science abstracts and 21% of EE&SS abstracts on arXiv had been revised with GPT, while Mathematics stood at 3.5%.","GPT revision makes abstracts shorter and more uniform: targeted prompts cut word count by more than 25%, and all six revisions reduce hedge words by roughly one-third.","Junior researchers adopt GPT at about three times the rate of senior researchers, and non-native and East Asian researchers adopt more heavily, so GPT is functioning as an equalizer of writing proficiency.","Real-world convergence toward senior-author style appears only among researchers who actually adopt GPT; non-adopters' writing similarity stays flat after ChatGPT's release.","Disciplines differ in convergence: juniors in EE&CS, Biology, Economics & Finance, and Statistics move toward senior style, while juniors in Mathematics show little to no convergence."],"supporting_citations":[{"why":"Provides the fine-tuned detection approach the paper adapts into discipline- and prompt-specific classifiers.","marker":"Chen et al. [2023]"},{"why":"Supplies the monitoring-at-scale framing for AI-modified content that motivates detecting GPT revisions on arXiv.","marker":"Liang et al. [2024a]"},{"why":"Establishes that ChatGPT-written abstracts are hard for humans to distinguish from real ones, motivating the need for trained classifiers.","marker":"Gao et al. [2023]"},{"why":"Supplies the ten empirical writing rules the paper uses to measure GPT's effect on clarity, conciseness, tense, hedging, and formality.","marker":"Weinberger et al. [2015]"},{"why":"Zero-shot prompting is the method used to instruct GPT-3.5-Turbo to produce the revised abstracts that form the training corpus.","marker":"Brown [2020]"},{"why":"Sentence-transformer embeddings provide the base representation on which the fine-tuned classifiers are built.","marker":"Reimers [2019]"}],"fun_headline_variants":["LLM adoption in arXiv abstracts hits 13.8% by Dec 2023","GPT revisions align junior authors with senior writing style","ChatGPT detectable in 22% of CS, 3.5% of math abstracts","AI edits cut hedge words by a third in scientific abstracts","Study maps LLM adoption and style convergence across 627k papers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The adoption estimates all rest on the assumption that a classifier trained on GPT-3.5-Turbo's responses to six fixed prompts recognizes real-world ChatGPT revisions, and that abstracts last updated before November 2021 are uncontaminated human text; if real-world prompts differ, or earlier abstracts already contained LLM writing, every adoption number shifts.","fun_headline_variants_meta":{"raw":{"variants":["LLM adoption in arXiv abstracts hits 13.8% by Dec 2023","GPT revisions align junior authors with senior writing style","ChatGPT detectable in 22% of CS, 3.5% of math abstracts","AI edits cut hedge words by a third in scientific abstracts","Study maps LLM adoption and style convergence across 627k papers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001382,"raw_usage":{"total_tokens":5621,"prompt_tokens":995,"completion_tokens":4626,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":4532}},"tokens_in":611,"tokens_out":4626,"duration_ms":28619,"temperature":1.0,"reasoning_tokens":4532,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:02:59.186289+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a hand-labeled corpus of abstracts whose authors disclose or can independently verify their ChatGPT use with unrestricted prompts, and test whether the paper's classifiers reproduce the reported adoption rates; a mismatch would show the detector does not generalize beyond its six fixed prompts, and checking pre-2021 abstracts against known GPT-3-era output patterns would expose whether the supposed human ground truth is contaminated.","supporting_citations":[],"review_version":1}