Pith. sign in

REVIEW 5 major objections 5 minor 21 references

Divergent LLM Adoption and Heterogeneous Convergence Paths in Research Writing

T0 review · 5 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read By the end of 2023, about one in five computer-science abstracts on arXiv had been revised with ChatGPT, and GPT-assisted revision is measurably pulling junior, male, and non-native researchers' writing closer to the style of senior…

desk verdict Useful large-scale adoption measurements undermined by a circular convergence design and a few embarrassing errors. read the letter →

arxiv 2504.13629 v1 pith:A2GJSI2P submitted 2025-04-18 cs.CL cs.AIecon.GNq-fin.EC

classification cs.CLcs.AIecon.GNq-fin.EC
keywords ChatGPTdetectionarXivabstractsscientificwritingstylegenerativeAIadoptiondifference-in-differencesLLMtextclassificationconvergenceheterogeneous
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that ChatGPT-assisted revision of arXiv abstracts became widespread within a year of ChatGPT's release and that its adoption was uneven across disciplines and researcher groups. Using more than 627,000 papers posted between January 2021 and December 2023, the authors fine-tune discipline- and prompt-specific classifiers to detect GPT-style revisions, and report that by the end of 2023 about 22% of Computer Science abstracts had been revised with GPT, against 3.5% in Mathematics. They also find that GPT revisions shorten abstracts, push writing toward present tense and formal neutrality, and, according to a difference-in-differences analysis, drive junior, male, and non-native researchers' writing closer to the style of senior researchers. The paper matters because it turns speculation about AI's footprint in scientific writing into quantitative, group-level estimates, and it frames the main risk as homogenization of academic expression rather than simple quality improvement.

What carries the argument

The machinery is a family of fine-tuned sentence-transformer classifiers, one per field and revision prompt, trained to separate human-written arXiv abstracts from GPT-3.5-Turbo versions generated under six fixed zero-shot prompts (clarity, formality, objectivity, readability, grammar, and comprehensive revision). A second component is the writing-rule regression: ten quantitative measures of style such as word count, sentence count, present-tense ratio, hedge-word frequency, superlatives, and evocative words, compared within the same article using article fixed effects between the original and GPT-revised text. The convergence analysis uses bag-of-words cosine similarity between original and GPT-revised abstracts, and between junior and senior, male and female, and native and non-native author groups, in a difference-in-differences design that compares adopters with non-adopters before and after ChatGPT's release.

What would settle it

Take a hand-labeled corpus of abstracts whose authors disclose or can independently verify their ChatGPT use with unrestricted prompts, and test whether the paper's classifiers reproduce the reported adoption rates; a mismatch would show the detector does not generalize beyond its six fixed prompts, and checking pre-2021 abstracts against known GPT-3-era output patterns would expose whether the supposed human ground truth is contaminated.

Watch

Extended reading notes

Core claim

The central discovery is that ChatGPT-revised scientific abstracts are detectable at scale and that the detectable signal is strong enough to map adoption and style convergence across groups. The paper constructs a ground-truth corpus from 343,461 abstracts last updated before ChatGPT's release, generates revised versions with GPT-3.5-Turbo along six prompt dimensions (clarity, formality, objectivity, readability, grammar, and comprehensive revision), and fine-tunes 48 field- and prompt-specific binary classifiers plus 8 multiclass classifiers on those pairs, with out-of-sample F1 scores above 0.95. Applying these classifiers to 627,384 arXiv abstracts, the paper reports adoption rising from 0.47% in January 2023 to 13.8% in December 2023, with end-2023 rates ranging from 3.5% in Mathematics to 22% in Computer Science. In the style analysis, GPT revisions reduce word count by more than 25% for targeted prompts, cut hedge-word usage by about a third, and in 8 of 11 writing-rule comparisons move text toward the profile of senior authors. The difference-in-differences results then show real-world convergence: junior, male, and non-native researchers who adopt GPT shift toward senior and native writing patterns, while non-adopters and female authors show little or no convergence.

Load-bearing premise

The adoption estimates all rest on the assumption that a classifier trained on GPT-3.5-Turbo's responses to six fixed prompts recognizes real-world ChatGPT revisions, and that abstracts last updated before November 2021 are uncontaminated human text; if real-world prompts differ, or earlier abstracts already contained LLM writing, every adoption number shifts.

Editorial extensions

If this is right

  • By the end of 2023, roughly 22% of Computer Science abstracts and 21% of EE&SS abstracts on arXiv had been revised with GPT, while Mathematics stood at 3.5%.
  • GPT revision makes abstracts shorter and more uniform: targeted prompts cut word count by more than 25%, and all six revisions reduce hedge words by roughly one-third.
  • Junior researchers adopt GPT at about three times the rate of senior researchers, and non-native and East Asian researchers adopt more heavily, so GPT is functioning as an equalizer of writing proficiency.
  • Real-world convergence toward senior-author style appears only among researchers who actually adopt GPT; non-adopters' writing similarity stays flat after ChatGPT's release.
  • Disciplines differ in convergence: juniors in EE&CS, Biology, Economics & Finance, and Statistics move toward senior style, while juniors in Mathematics show little to no convergence.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The measurement of convergence may be partly circular: the classifier that labels an abstract as GPT-revised was trained on the same GPT style features used to define convergence, so part of the reported junior-to-senior convergence could reflect the detector's template rather than an independent stylistic shift.
  • Because the synthetic revisions come from GPT-3.5-Turbo under six fixed prompts, real-world use of newer models or different instructions is likely under-detected; the true late-2023 adoption rates could be higher than 22%, and the revision-type shares may not describe actual prompt choices.
  • The 'democratizing' reading, that GPT raises the writing of junior and non-native researchers, has a plausible selection alternative: less fluent writers choose GPT, so part of the observed improvement reflects who adopts, not what GPT does; the within-paper fixed-effects comparisons mitigate this but do not fully rule it out.
  • A direct extension would apply the same detector to other genres, such as peer reviews, grant proposals, or non-English abstracts, where similar convergence could either speed knowledge exchange or erase recognizable author voice; the paper's framework is ready-made for that test.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper develops a detector for GPT-revised arXiv abstracts by fine-tuning sentence-transformer models on pairs of human-written abstracts and GPT-3.5-Turbo revisions produced under six prompts, across eight arXiv disciplines (48 binary and 8 multiclass models). It applies the detector to 627,384 arXiv abstracts from 2021-2023, reports aggregate adoption rates (e.g., 22% of Computer Science abstracts by the end of 2023), heterogeneous adoption by field, nativeness, ethnicity, gender, and seniority, analyzes differences in ten writing rules between original and revised texts, and uses a difference-in-differences-style comparison of cosine similarities to claim that GPT use drives junior, male, and non-native researchers' writing toward senior styles. The central claims are: (i) reliable detection of GPT revisions at scale, (ii) large and heterogeneous adoption, and (iii) causal style convergence from GPT adoption.

Significance. If the detection and difference-in-differences results were valid, this would be a valuable large-scale measurement of LLM adoption in scientific writing, with clear implications for science policy and the study of AI-generated content. The strengths are the large corpus, the explicit separation of six revision prompts, and the internal consistency of the synthetic test-set metrics, with out-of-sample F1 scores above 0.95 for most prompts. However, the evaluation is entirely internal: the test set is generated by the same model and prompts used for training, and no external validation against real ChatGPT-edited abstracts is provided. The convergence analysis is subject to a direct circularity concern, the stated ground-truth period for 'human-written' text is factually contaminated by GPT-3's June 2020 API release, and the reported difference-in-differences is not actually estimated with any formal specification. Unless these issues are resolved, the headline adoption and convergence numbers are not supported.

major comments (5)
  1. [Section 2.1] The training ground truth is not reliably human-written. The paper states that texts updated before November 30, 2021 are used because 'GPT-3 was publicly released only after this period,' but the GPT-3 API was publicly released in June 2020, before the training window (articles updated before October 1, 2021) and the test window (October 1-November 30, 2021). Abstracts in the 'human' class may therefore already include GPT-3-revised text. This contaminates the training labels for all 48 classifiers and invalidates the reported F1 scores as evidence of detection quality; it also breaks the assumption that pre-November-2022 ground-truth abstracts are uncontaminated.
  2. [Section 3.2, Table 2] The detector's external validity is unestablished. All training and test revisions were created by GPT-3.5-Turbo under six fixed prompts written by the authors, and the held-out test set in Table 2 is from the same pipeline. High precision and recall on this test set only show that the model can separate original abstracts from their own GPT-3.5 rewrites. Real-world ChatGPT use involves different prompts, different model versions, and possibly human post-editing. No validation against externally labeled real-world GPT-revised abstracts is reported, so the absolute adoption levels in Section 3.2 and the treatment assignments in Sections 4.3 and 4.4 are not calibrated for the population to which they are applied.
  3. [Sections 4.3-4.4, Figures 10-11] The convergence difference-in-differences is circular. GPT adoption is assigned by classifiers fine-tuned to recognize the exact GPT-3.5 revision style, while the thought-experiment outcome is cosine similarity to the version-6 GPT revision and the real-world outcome is similarity to senior writing, which Table 6 shows is itself mimicked by the GPT revisions. An abstract labeled as an adopter is therefore closer to the convergence target by construction once ChatGPT is used, irrespective of any independent behavioral mechanism. Figures 8-11 may thus be measuring the detector's stylistic signal rather than a genuine convergence of researchers' writing practices.
  4. [Sections 4.3-4.4, Figures 10-11] No formal difference-in-differences estimates are reported. The text claims that a difference-in-differences analysis shows GPT-driven convergence, but the figures present only monthly averages of cosine similarity, with no regression coefficients, standard errors, or pre-trend tests. It is therefore impossible to assess the statistical significance or magnitude of the convergence, or to control for group-specific time trends and composition changes. The absence of a formal DiD specification is a load-bearing gap for the paper's central causal claim.
  5. [Section 3.2] The pre-period adjustment assumes a constant false-positive rate. The paper subtracts the average pre-November-2022 adoption rate from all subsequent months to remove 'misclassification.' This is valid only if the classifier's false-positive rate is constant over time. If it drifts because human writing style evolves, because the arXiv population changes, or because the model's operating point interacts with text length or field composition, the adjusted adoption levels are biased. No evidence is provided for a constant false-positive rate, and the per-discipline pre-period levels in Figure 2 vary substantially, which suggests the adjustment is at best approximate.
minor comments (5)
  1. [Section 2.1] The heading 'Mothods' should be 'Methods.'
  2. [Section 3.2] The text says the increase from 0.47% to 13.8% is a '740-fold increase'; the correct ratio is approximately 29-fold (13.8 divided by 0.47).
  3. [Section 3.4] 'Table ?? presents the main results' contains an unresolved cross-reference; the intended table is presumably Table 5.
  4. [Table 4] The confusion matrix shows that revisions 4 and 5 are classified correctly only 23.76% and 18.46% of the time, respectively; the prompt-level adoption and revision-mix results in Section 3.3 should explicitly discuss this near-chance performance.
  5. [Table 9] The note says 'arVix' instead of 'arXiv.'

Circularity Check

2 steps flagged · score 6.0 of 10

Convergence DiD is partly circular: 'GPT adoption' is assigned by a classifier trained on GPT-3.5 revisions, while the convergence outcome measures similarity to that same revision style (or to senior style that the revisions mimic).

  1. self definitional [Section 2.1 (classifier training) and Section 4.3 (Figures 8-9)]
    "For each article, we compare the abstract with its GPT-revised version (version 6) by constructing a bag-of-words representation and calculating the cosine similarity between the two versions. ... Second, following the release of ChatGPT, the writing style increasingly converges to that of the GPT-revised abstracts, irrespective of author seniority. However, this convergence is limited to junior and senior authors who actively use GPT for article revisions."

    Adoption is not observed; it is imputed by classifiers (Section 2.1) fine-tuned to separate original abstracts from their own GPT-3.5-Turbo revisions, so a text is labeled 'GPT-revised' precisely when it is stylistically close to those synthetic revisions. The outcome in Section 4.3/Figure 8 is the cosine similarity between the same abstract and the version-6 GPT revision. Hence the 'convergence to GPT-style writing' among adopters is, to a large extent, the classifier's own decision rule reappearing as a time trend: adopters are defined as texts near the GPT-revision distribution, and convergence is measured as movement toward that same distribution.

  2. fitted input called prediction [Section 4.4, building on Section 4.1 and Table 6]
    "A plausible explanation is that GPT-revised articles resemble those written by more experienced authors, as shown in Table 6. Thus, leveraging GPT for revisions may push juniors' writing styles closer to those of senior authors."

    The real-world DiD labels juniors as 'adopters' with the same Section 2.1 detector, which keys on resemblance to GPT-3.5 revisions. Section 4.1/Table 6 establishes that those GPT revisions shift text toward senior writing in 8 of 11 measured rules. The outcome in Section 4.4 is similarity to senior writing. Therefore, the detector selects juniors who already write in a senior-like (GPT-like) style, and the outcome measures the same stylistic distance; the convergence estimate is inflated, if not largely produced, by the shared stylistic measurement. The paper's own explanation makes the bridge explicit, showing that the treatment assignment and the outcome are linked through the same GPT-revision style rather than through an independent behavioral measure.

full rationale

The prevalence estimates in Section 3 and the controlled within-article comparison in Tables 6-7 are not circular per se: the classifier application is a standard detection task, and the GPT-revision experiment independently documents style shifts. The circularity lies in the convergence analysis. In Section 4.3 the treatment indicator ('uses GPT') comes from classifiers trained to recognize GPT-3.5-Turbo revisions of the same abstracts, and the outcome is cosine similarity to the version-6 GPT revision; an abstract is classified as an adopter precisely when it is close to the revision distribution, and convergence is measured as increasing closeness to that same distribution. In Section 4.4 the outcome is similarity to senior writing, but Table 6 shows GPT revisions move text toward senior style, so the detector's positive class is enriched for senior-like style; the observed convergence of 'adopters' to senior writing is thus inflated by the shared measurement. Separately, Section 2.1's claim that GPT-3 was publicly released only after November 2021 is factually incorrect (the GPT-3 API launched in June 2020), which undermines the purity of the pre-2022 'human-written' ground truth; this is a correctness risk rather than circularity, but it compounds the label problem. Because the headline adoption and quality results retain independent content, the paper is partially circular rather than wholly so.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claims rest on a multi-stage pipeline whose key assumptions are the purity of the pre-ChatGPT ground truth, the representativeness of the synthetic revisions, and the time-invariance of detector error. None of these are independently validated, and one is factually wrong as stated.

free parameters (3)
  • Pre-period false-positive baseline = Per-discipline average GPT-detection rate before November 2022
    Raw detection rates are adjusted by subtracting the average pre-ChatGPT adoption rate, a data-fitted offset that assumes the misclassification rate is constant over time.
  • Seniority thresholds = 10 published papers or 10 years of experience
    Authors are classified as senior or junior using ad hoc cutoffs chosen without sensitivity analysis.
  • Six revision prompts = Six prompt templates covering clarity, formality, objectivity, readability, grammar, and comprehensive revision
    The detector is built around this specific prompt set, and real-world ChatGPT usage is assumed to match these categories.
assumptions (5)
  • domain assumption Pre-November-2021 arXiv abstracts are human-written and GPT-free
    Section 2.1 states GPT-3 was publicly released only after November 2021, but GPT-3's API launched in June 2020, so this grounding assumption is false.
  • domain assumption False-positive detection rate is time-invariant
    Section 3.2 subtracts a fixed pre-period baseline from all months, assuming misclassification does not drift as human writing evolves.
  • domain assumption GPT-3.5-Turbo revisions under the six prompts represent real ChatGPT usage
    Section 2 trains and evaluates detectors only on these synthetic revisions; real usage with different prompts, models, and editing behaviors is assumed to be captured.
  • domain assumption Name-based gender and ethnicity inference is accurate
    Figures 5 and 6 infer ethnicity and gender via an unnamed online package; measurement error in these variables is not modeled.
  • domain assumption Bag-of-words cosine similarity captures writing-style convergence
    Section 4.3 treats cosine similarity over term vectors as a measure of style convergence, ignoring word-choice diversity and semantic content.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Divergent LLM Adoption and Heterogeneous Convergence Paths in Research Writing." pith.science (2026). https://pith.science/paper/A2GJSI2P

@misc{pith2026250413629,
  author       = {Pith},
  title        = {Pith review of: Divergent LLM Adoption and Heterogeneous Convergence Paths in Research Writing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A2GJSI2P}},
  note         = {Machine review of arXiv:2504.13629}
}
read the original abstract

Large Language Models (LLMs), such as ChatGPT, are reshaping content creation and academic writing. This study investigates the impact of AI-assisted generative revisions on research manuscripts, focusing on heterogeneous adoption patterns and their influence on writing convergence. Leveraging a dataset of over 627,000 academic papers from arXiv, we develop a novel classification framework by fine-tuning prompt- and discipline-specific large language models to detect the style of ChatGPT-revised texts. Our findings reveal substantial disparities in LLM adoption across academic disciplines, gender, native language status, and career stage, alongside a rapid evolution in scholarly writing styles. Moreover, LLM usage enhances clarity, conciseness, and adherence to formal writing conventions, with improvements varying by revision type. Finally, a difference-in-differences analysis shows that while LLMs drive convergence in academic writing, early adopters, male researchers, non-native speakers, and junior scholars exhibit the most pronounced stylistic shifts, aligning their writing more closely with that of established researchers.

Figures

Figures reproduced from arXiv: 2504.13629 by the authors.

Figure 1
Figure 1. Vector Representation of Various Versions [PITH_FULL_IMAGE:figures/full_fig_p040_1.png] view at source ↗
Figure 2
Figure 2. Percentage of Texts Identified as being Revised by ChatGPT [PITH_FULL_IMAGE:figures/full_fig_p041_2.png] view at source ↗
Figure 3
Figure 3. Percentage of Texts Identified being Revised by ChatGPT in various Di [PITH_FULL_IMAGE:figures/full_fig_p042_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Native v.s. Non-Native in Using GPT to Revise Abstracts [PITH_FULL_IMAGE:figures/full_fig_p043_4.png]
Figure 5
Figure 5. Figure 5: Enthinicity Difference in Using GPT to Revise Articles [PITH_FULL_IMAGE:figures/full_fig_p044_5.png]
Figure 6
Figure 6. Figure 6: Gender Difference in ChatGPT Adoption Note: This figure examines gender differences in adopting GPT for revising articles. We first use a machine learning algorithm to estimate the gender of each author based on their name 4 . Each month, we divide the sample into subg…
Figure 7
Figure 7. Figure 7: Seniorioty Difference in GPT Adoption (a) Panel A (b) Panel B Note: This figure examines the differences in seniority in adopting GPT for revising articles. We use two measures as proxies for academic research experience: the number of academic papers written and years…
Figure 8
Figure 8. Figure 8: Heterogeneous Impact of GPT Adoption (Seniority) [PITH_FULL_IMAGE:figures/full_fig_p047_8.png]
Figure 9
Figure 9. Figure 9: Heterogeneous Impact of GPT-Adoption (Gender/Native) [PITH_FULL_IMAGE:figures/full_fig_p048_9.png]
Figure 10
Figure 10. Figure 10: Heterogeneity in difference-in-difference of Textual Similarity (Seniority) [PITH_FULL_IMAGE:figures/full_fig_p049_10.png]
Figure 11
Figure 11. Figure 11: Heterogeneity in difference-in-difference of Textual Similarity [PITH_FULL_IMAGE:figures/full_fig_p050_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

21 extracted references · 6 canonical work pages

  1. [4]

    On the possibilities of ai-generated text detection

    Souradip Chakraborty, Amrit Singh Bedi, Sicheng Zhu, Bang An, Dinesh Manocha, and Furong Huang. On the possibilities of ai-generated text detection. arXiv preprint arXiv:2304.04736,

  2. [6]

    Note: This figure displays the percentage of abstracts identified as having been written by ChatGPT, based on predictions from our trained multi-label large language model (LLM)

    41 Figure 3: Percentage of Texts Identified being Revised by ChatGPT in various Di- mensions. Note: This figure displays the percentage of abstracts identified as having been written by ChatGPT, based on predictions from our trained multi-label large language model (LLM). Specifically, the model classifies the original abstract, where revision 0 represent...

  3. [7]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 ,

  4. [8]

    Three bricks to consolidate watermarks for large language models

    Pierre Fernandez, Antoine Chaffin, Karim Tit, Vivien Chappelier, and Teddy Furon. Three bricks to consolidate watermarks for large language models. In 2023 IEEE International Workshop on Information Forensics and Security (WIFS), pages 1–6. IEEE,

  5. [9]

    Gltr: Statistical detection and visualization of generated text

    Sebastian Gehrmann, Hendrik Strobelt, and Alexander M Rush. Gltr: Statistical detection and visualization of generated text. arXiv preprint arXiv:1906.04043 ,

  6. [10]

    Monitoring ai- modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews

    Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, et al. Monitoring ai- modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews. arXiv preprint arXiv:2403.07183 , 2024a. Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Yi Din...

  7. [12]

    Release strategies and the social impacts of language models

    Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. Release strategies and the social impacts of language models. arXiv preprint arXiv:1908.09203,

  8. [13]

    Large language models in medicine

    Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. Large language models in medicine. Nature medicine, 29(8):1930–1940,

Show all 21 references
  1. [14]

    Authorship attribution for neural text generation

    Adaku Uchendu, Thai Le, Kai Shu, and Dongwon Lee. Authorship attribution for neural text generation. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) , pages 8384–8395,

  2. [15]

    Yihan Wu, Zhengmian Hu, Hongyang Zhang, and Heng Huang

    https://mit-genai.pubpub.org/pub/24gsgdjx. Yihan Wu, Zhengmian Hu, Hongyang Zhang, and Heng Huang. Dipmark: A stealthy, efficient and resilient watermark for large language models. arXiv preprint arXiv:2310.07710,

  3. [16]

    Whose chatgpt? unveiling real-world educational inequalities introduced by large language models

    Renzhe Yu, Zhen Xu, Sky CH-Wang, and Richard Arum. Whose chatgpt? unveiling real-world educational inequalities introduced by large language models. arXiv preprint arXiv:2410.22282,

  4. [17]

    Provable robust watermarking for ai-generated text

    Xuandong Zhao, Prabhanjan Ananth, Lei Li, and Yu-Xiang Wang. Provable robust watermarking for ai-generated text. arXiv preprint arXiv:2306.17439 ,

  5. [18]

    Maths”, Physics as “Phys

    30 6 Figures and Tables Table 1: The Number of Articles by Field and Year in arVix This table shows the number of articles by field and year. We abbreviate Mathematics as “Maths”, Physics as “Phys”, Computer Science as “CS”, Electrical Engineering and System Science as “EE&SS”...

  6. [20]

    The figure shows no significant gender difference in GPT usage, as both male and female authors follow nearly the same trend in adopting GPT for revisions

    Each month, we divide the sample into subgroups based on gender and calculate the percentage of ab- stracts revised by GPT. The figure shows no significant gender difference in GPT usage, as both male and female authors follow nearly the same trend in adopting GPT for revision...

  7. [21]

    We compute the average cosine similarity for articles authored by senior and junior researchers separately each month

    by constructing a bag-of- words representation and calculating the cosine similarity between the two versions. We compute the average cosine similarity for articles authored by senior and junior researchers separately each month. Authors are considered senior they have at leas...

  8. [2004]

    All that’s’ human’is not gold: Evaluating human evaluation of generated text

    Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. All that’s’ human’is not gold: Evaluating human evaluation of generated text. arXiv preprint arXiv:2107.00061 ,

  9. [2019]

    A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity

    Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023 ,

  10. [2020]

    Sentence-bert: Sentence embeddings using siamese bert-networks

    N Reimers. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084,

  11. [2021]

    Language models are few-shot learners

    Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165,

  12. [2023]

    Gpt-sentinel: Distinguishing human and chatgpt generated content

    Yutian Chen, Hao Kang, Vivian Zhai, Liangze Li, Rita Singh, and Bhiksha Raj. Gpt-sentinel: Distinguishing human and chatgpt generated content. arXiv preprint 26 arXiv:2305.07969,

  13. [2024]

    Real or fake? learning to discriminate machine from human gener- ated text

    Anton Bakhtin, Sam Gross, Myle Ott, Yuntian Deng, Marc’Aurelio Ranzato, and Arthur Szlam. Real or fake? learning to discriminate machine from human gener- ated text. arXiv preprint arXiv:1906.03351 ,

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.