Pith. sign in

REVIEW 5 cited by

ChatGPT "contamination": estimating the prevalence of LLMs in the scholarly literature

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.16887 v1 pith:MXJXPLGL submitted 2024-03-25 cs.DL

classification cs.DL
keywords keywordsprevalencescholarlychatgptliteraturellm-assistedpublishingacademic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The use of ChatGPT and similar Large Language Model (LLM) tools in scholarly communication and academic publishing has been widely discussed since they became easily accessible to a general audience in late 2022. This study uses keywords known to be disproportionately present in LLM-generated text to provide an overall estimate for the prevalence of LLM-assisted writing in the scholarly literature. For the publishing year 2023, it is found that several of those keywords show a distinctive and disproportionate increase in their prevalence, individually and in combination. It is estimated that at least 60,000 papers (slightly over 1% of all articles) were LLM-assisted, though this number could be extended and refined by analysis of other characteristics of the papers or by identification of further indicative keywords.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback

    cs.CL 2025-08 conditional novelty 6.0 of 10

    People prefer text containing the words that an instruction-tuned model uses far more than its base version, linking human feedback training to LLM word overuse.

  2. Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English

    cs.CL 2025-08 conditional novelty 6.0 of 10

    After ChatGPT's release, science and tech podcast speakers used AI-associated words like 'surpass' and 'align' more often, while control synonyms showed no average shift.

  3. Exploring the Structure of AI-Induced Language Change in Scientific English

    cs.CL 2025-06 conditional novelty 6.0 of 10

    In PubMed abstracts, AI-associated 'spiking' words rise together with their synonyms rather than replacing them, and declining words show less systematic, more organic patterns.

  4. Human-LLM Coevolution: Evidence from Academic Writing

    cs.CL 2025-02 conditional novelty 6.0 of 10

    After ChatGPT-style words were publicly flagged in early 2024, their frequency in arXiv abstracts dropped, while other common LLM-favored words kept rising, suggesting authors are adapting their writing to avoid detection.

  5. GPT Editors, Not Authors: The Stylistic Footprint of LLMs in Academic Preprints

    cs.CL 2025-05 reject novelty 5.0 of 10

    Across 2,408 arXiv preprints, LLM-typical word usage does not cluster in any section, indicating that AI assistance, when used, is uniform rather than limited to specific parts of a paper.

Pith tools