REVIEW 3 major objections 6 minor 2 cited by
The Widespread Adoption of Large Language Model-Assisted Writing Across Society
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper estimates that LLMs assisted with roughly 18% of consumer complaint text, 24% of corporate press releases, 10% of small-firm job postings, and 14% of UN press releases by late 2024.
desk verdict Solid cross-domain measurement of the post-ChatGPT surge, but the absolute adoption percentages need debiasing and external validation before I'd quote them. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine of the analysis is a domain-adapted text-mixture estimator. For each corpus, word-frequency distributions are built from two reference pools: human writing collected before ChatGPT's release and synthetic text produced by prompting current GPT models. Fitting a mixture model to observed monthly text yields an estimate of $\alpha$, the proportion of sentences substantially modified by LLMs. The same estimator is fitted separately for each press-release platform, job category, and complaint/UN corpus, and its bias is checked by mixing pre-ChatGPT human text with GPT output at known concentrations. That calibration, with prediction error below 3.3 percentage points in the validation tables, is what turns raw detection scores into population-level adoption numbers.
What would settle it
Take a corpus of documents whose true AI-assistance status is known independently—for example, job postings or press releases drafted with an LLM and then edited by humans, with editing logs intact—and run this estimator on it. If the estimated $\alpha$ falls well below the true fraction of LLM-influenced text, the reported population figures are best read as lower bounds rather than point estimates.
Extended reading notes
Core claim
The central claim is that LLM-assisted writing has become a measurable, large-scale feature of public and commercial text, not an edge case. Using a statistical estimator of the fraction $\alpha$ of sentences that were generated or substantially modified by an LLM, the paper reports adoption levels of about 18% for financial consumer complaints, 23–24% for at least one corporate press-release platform, up to 15% for job postings from young small firms, and 14% for UN press releases by the end of the study period. The paper further claims that these adoption curves share a common shape: a lag of several months after ChatGPT's debut, a steep rise through 2023, and a plateau by 2024, which it reads as either saturation or the growing indistinguishability of AI output.
Load-bearing premise
The method assumes that text generated by current GPT models in the validation setup is a faithful stand-in for all real-world LLM-assisted writing, and that the pre-ChatGPT false-positive rate stays constant through 2024.
Editorial extensions
If this is right
- Adoption of LLM-assisted writing is now a majority-adjacent phenomenon in corporate press releases, with the top platform reaching about 24% of text by late 2024.
- The consistent plateau across domains suggests the first wave of adoption had largely run its course by 2024, whether through saturation or through models becoming harder to detect.
- Smaller and younger firms lead in job-posting adoption, with post-2015 firms reaching 10–15% in some roles, pointing to an organizational-age gradient in AI uptake.
- Geographic and demographic heterogeneity is modest but real: more urbanized areas show higher complaint adoption, while lower-education areas show slightly higher rates.
- International organizations, exemplified by UN press releases, reach roughly 14% LLM-modified content, indicating institutional adoption in high-stakes communication.
Reading between the lines
- Because the method misses heavily edited or highly human-like LLM text, the paper's own numbers are lower bounds; the true prevalence in late 2024 could be materially higher than 18–24%.
- If the plateau reflects detector blindness rather than saturation, apparent stabilization may mask continued growth—a testable prediction when future detectors calibrated on newer models are applied to the same 2024 corpora.
- A direct extension would compare these population estimates with self-reported usage from surveys or platform telemetry, which could cross-validate the framework without relying on synthetic ground truth.
- The finding that consumer complaints in lower-education areas show higher LLM adoption suggests these tools may be functioning as an equalizer in consumer advocacy, but the paper does not test whether AI-assisted complaints are more likely to receive redress.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper applies a word-frequency mixture estimator, originally developed by the authors to measure AI-modified text in peer reviews, to four large English-language text corpora: 687,241 consumer complaints to the CFPB, 537,413 corporate press releases from Newswire/PRNewswire/PRWeb, 304.3 million LinkedIn job postings, and 15,919 UN press releases. The central empirical claim is that LLM-assisted writing rose sharply starting 3–5 months after ChatGPT's November 2022 release and then plateaued, reaching roughly 18% of consumer complaint text, up to 24% of corporate press release text, just below 10% of small-firm job postings, and nearly 14% of UN press releases by late 2024. The paper also reports heterogeneity in adoption by geography, urbanization, education, and firm age/size. The detection method's ground truth is built entirely from a synthetic pipeline in which GPT-3.5-turbo compresses pre-ChatGPT human text into bullet skeletons and re-expands it, and the validation sets mix only this synthetic positive class with pre-ChatGPT human text.
Significance. If the headline percentages are unbiased, the paper provides the first population-level, cross-domain measurement of LLM-assisted writing, with immediate relevance to policy discussions about AI adoption, labor markets, and institutional communication. The study's strengths are its scale, the consistency of the temporal pattern across four independent domains, the use of an open and transparent estimator rather than commercial black-box detectors, and the robustness check across GPT model versions in Supplementary Figure 4. The authors also honestly acknowledge that heavily edited or human-like LLM text escapes detection and frame their numbers as lower bounds. However, the central estimates inherit a load-bearing external-validity assumption: that synthetic GPT-generated outlines-and-expansions are representative of real-world human-in-the-loop LLM-assisted writing. That assumption is not tested against any independently labeled real documents, and the validation tables in the supplement show a systematic positive bias at every ground-truth level, including at zero.
major comments (3)
- [Supplementary Information, Model Fitting; Supp. Figs. 5–6; Supp. Tables 1–5] The ground truth for the 'LLM-assisted' class is entirely synthetic: pre-ChatGPT human text is compressed into bullet-point skeletons and then re-expanded by GPT-3.5-turbo (Supp. Figs. 5–6), and the validation corpora in Supp. Tables 1–5 mix only this synthetic positive class with pre-ChatGPT human text. No independently labeled set of real documents produced by humans using LLMs in the wild is used for validation. The reported prediction error of less than 3.3 percentage points therefore measures how well the estimator recovers the fraction of text that resembles this specific two-prompt pipeline, not the fraction of text actually written or substantially modified by LLMs. Direct prompting, human editing of model drafts, and other model families can produce different lexical signatures. The paper's 'lower bound' framing covers heavily edited or very human-like LLM output being missed, but it does not cover the opposite risk: human text whose style drifts toward GPT-like patterns, from either temporal style drift or humans imitating LLM output. This assumption is load-bearing for every headline percentage.
- [Supp. Tables 1–5] The validation tables show a systematic positive bias of roughly 2–3 percentage points at every ground-truth level, including at α=0. For example, Supp. Table 1 reports an estimate of 1.8% at ground-truth 0.0%; Supp. Table 3 reports 2.9%, 2.1%, and 2.3% for PRNewswire, PRWeb, and Newswire at α=0; and Supp. Table 5 reports 2.0% for Scientist at α=0. This offset is not subtracted or otherwise corrected in the headline estimates. For the smallest headline claim ('just below 10%' in small-firm job postings), this is a substantial relative correction, and even for the largest estimates (18–24%) it is a non-negligible absolute correction. The manuscript should either explicitly debias the estimates using the measured false-positive rates at α=0 and at the relevant mixing levels, or demonstrate that the qualitative conclusions are unchanged after applying the corresponding correction.
- [Fig. 1; Results; Discussion] The stabilization/plateau pattern is presented as a main finding, but the interpretation is confounded with possible changes in detector sensitivity over time. Because the synthetic positive class is generated by specific GPT models (GPT-3.5-turbo and GPT-4 variants), improvements in LLM indistinguishability—or in human writers' tendency to imitate LLM style—could cause the estimated fraction to flatten or decline even if true adoption continues to rise. The manuscript acknowledges this in the Discussion and in footnote 2, but the abstract and results still present stabilization as an empirical regularity. The validation in Supp. Tables 1–5 only assesses calibration on pre-ChatGPT human text mixed with synthetic LLM text; it does not test whether the detector's sensitivity is stationary across 2023–2024. A sensitivity analysis that varies the assumed sensitivity trajectory over time (e.g., by re-generating the synthetic positive class with successive model versions and showing the time-series conclusions are robust) is needed to support the plateau claim.
minor comments (6)
- [Results, LLM Adoption in LinkedIn Job Postings; Fig. 1] The sentence 'Using the sample of small companies based on the number of vacancies posted, our findings reveal... (Fig. 1d, Fig. 4)' refers to Figure 1d, which displays UN press releases; the correct panel for job postings appears to be Fig. 1c.
- [Supplementary Figure 4] The caption states that GPT-3.5-turbo, 'used in main analysis,' was 'released January 25, 2024'; GPT-3.5-turbo was released earlier, so this date likely refers to a specific model snapshot rather than the model family, and the caption should be corrected or clarified.
- [Results, LLM Adoption in Corporate Press Releases] The Introduction says adoption surged '3-4 months' after ChatGPT's release, while this section says 'about 2 quarters post rollout'; please reconcile the timing statements.
- [Results, Geographic and Demographic Disparities] The p-values ('less than 0.001') and 'highly statistically significant' claims are reported without specifying the statistical test, the unit of analysis, or whether any multiple-comparison correction was applied; please add these details.
- [Abstract] The consumer-complaint data end in August 2024, so 'By late 2024' is slightly imprecise; please adjust to 'by August 2024' or 'by late summer 2024.'
- [Supplementary Information, LinkedIn Job Posting Data] The definition of small firms as 'companies with either 10 or fewer registered employees in 2021 or companies posting less than or equal to about 2 postings per year' is ambiguous in light of the later statement that the median number of postings is 3; please clarify whether the threshold is 2 or 3.
Circularity Check
No formal circularity: synthetic ground truth is a validity limitation, not a self-referential reduction; the post-ChatGPT surge is external.
full rationale
The paper's derivation chain is not circular in the formal sense. The population estimates are produced by a word-frequency mixture estimator whose positive class is constructed by prompting GPT-3.5-turbo to compress a pre-ChatGPT human text into a bullet skeleton and then expand it (Supp. Figs. 5-6). This means the validation error (<3.3 percentage points) is calibrated on synthetic 'LLM-modified' text rather than on independently labeled real documents, and the headline percentages inherit that construct as a validity limitation; the paper explicitly acknowledges that heavily edited or human-like LLM output is missed and offers the numbers as a lower bound. But a calibration gap is not a circular reduction: the fitted model does not take the headline values as inputs, no equation in the paper is defined in terms of its own outputs, and the post-ChatGPT temporal surge in estimated alpha is an external, non-encoded pattern. The self-citations to the authors' prior method [12] are supporting references, and the paper re-describes the fitting procedure and points to code, so the citations do not substitute for the derivation. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (1)
- Alpha estimator word-frequency model parameters =
not disclosed
assumptions (4)
- domain assumption Pre-ChatGPT corpora (2021, or 2019 for UN) are representative of human writing without LLM assistance.
- ad hoc to paper Synthetic LLM-modified texts generated by GPT models are representative of real-world LLM-assisted writing in 2023-2024.
- domain assumption The false-positive rate measured on pre-ChatGPT text remains stable over the study window.
- domain assumption English-language text is sufficient for the claimed adoption patterns.
Cite this review
Pith. "Pith review of The Widespread Adoption of Large Language Model-Assisted Writing Across Society." pith.science (2026). https://pith.science/paper/TU35BESO
@misc{pith2026250209747,
author = {Pith},
title = {Pith review of: The Widespread Adoption of Large Language Model-Assisted Writing Across Society},
year = {2026},
howpublished = {\url{https://pith.science/paper/TU35BESO}},
note = {Machine review of arXiv:2502.09747}
}
read the original abstract
The recent advances in large language models (LLMs) attracted significant public and policymaker interest in its adoption patterns. In this paper, we systematically analyze LLM-assisted writing across four domains-consumer complaints, corporate communications, job postings, and international organization press releases-from January 2022 to September 2024. Our dataset includes 687,241 consumer complaints, 537,413 corporate press releases, 304.3 million job postings, and 15,919 United Nations (UN) press releases. Using a robust population-level statistical framework, we find that LLM usage surged following the release of ChatGPT in November 2022. By late 2024, roughly 18% of financial consumer complaint text appears to be LLM-assisted, with adoption patterns spread broadly across regions and slightly higher in urban areas. For corporate press releases, up to 24% of the text is attributable to LLMs. In job postings, LLM-assisted writing accounts for just below 10% in small firms, and is even more common among younger firms. UN press releases also reflect this trend, with nearly 14% of content being generated or modified by LLMs. Although adoption climbed rapidly post-ChatGPT, growth appears to have stabilized by 2024, reflecting either saturation in LLM adoption or increasing subtlety of more advanced models. Our study shows the emergence of a new reality in which firms, consumers and even international organizations substantially rely on generative AI for communications.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 2 Pith papers
-
Political Ideology Shifts in Large Language Models
LLMs shift their Political Compass answers when adopting synthetic personas, with shifts growing with scale, asymmetric between right- and left-leaning cues, and tracking persona themes.
-
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
Seven prototypical collaboration behaviors, such as asking for more outputs, asking questions, and adding content, explain most variation in how users follow up with writing assistants in the wild.
Reference graph
Works this paper leans on
-
[1]
Blueprint for an AI bill of rights (2022)
The White House Office of Science and Technology Policy. Blueprint for an AI bill of rights (2022). Accessed: 2023-09-08
work page 2022
-
[2]
Bommasani, R. et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2022)
arXiv 2022
-
[3]
Artificial intelligence index report 2024
Stanford Institute for Human-Centered Artificial Intelligence. Artificial intelligence index report 2024. Tech. Rep., Stanford University (2024)
work page 2024
-
[4]
Dwivedi, Y . K.et al. Opinion paper:“so what if chatgpt wrote it?” multidisciplinary perspectives on opportunities, challenges and implications of generative conversational ai for research, practice and policy. Int. J. Inf. Manag. 71, 102642, DOI: 10.1016/j.ijinfomgt.2023.102642 (2023)
arXiv 2023
-
[5]
Kasneci, E. et al. Chatgpt for good? on opportunities and challenges of large language models for education. Learn. Individ. Differ. 103, 102274, DOI: 10.1016/j.lindif.2023.102274 (2023)
arXiv 2023
-
[6]
M., Gebru, T., McMillan-Major, A
Bender, E. M., Gebru, T., McMillan-Major, A. & Smitchell, M. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, 610–623, DOI: 10.1145/3442188.3445922 (ACM, 2021)
arXiv 2021
-
[7]
R., John Mellor & other authors
Laura Weidinger, M. R., John Mellor & other authors. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359 (2021)
arXiv 2021
-
[8]
Humlum, A. & Vestergaard, E. The adoption of chatgpt. IZA Discuss. Pap. No. 16992 DOI: 10.2139/ssrn. 4827166 (2024)
doi:10.2139/ssrn 2024
Show all 30 references
-
[9]
& Deming, D
Bick, A., Blandin, A. & Deming, D. J. The rapid adoption of generative ai. Working Paper 32966, National Bureau of Economic Research (2024). DOI: 10.3386/w32966. Revised February 2025
2024 doi
-
[10]
& Peskoff, D
Brooks, C., Eggert, S. & Peskoff, D. The rise of ai-generated content in wikipedia (2024). 2410.08044
2024 arXiv
-
[11]
& Shin, J
Shin, M., Kim, J. & Shin, J. The adoption and efficacy of large language models: Evidence from consumer complaints in the financial industry. Available at SSRN 5004194 (2024)
2024
-
[12]
Liang, W. et al. Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews. arXiv preprint arXiv:2403.07183 (2024)
2024 arXiv
-
[13]
Liang, W. et al. Mapping the increasing use of LLMs in scientific papers. In First Conference on Language Modeling (2024)
2024
-
[14]
& Zou, J
Liang, W., Yuksekgonul, M., Mao, Y ., Wu, E. & Zou, J. Y . Gpt detectors are biased against non-native english writers. arXiv (2023). 2304.02819
2023 arXiv
-
[15]
& Raymond, L
Brynjolfsson, E., Li, D. & Raymond, L. Generative ai at work. Natl. Bureau Econ. Res. DOI: 10.3386/w31161 (2023)
2023 doi
-
[16]
M., Singhal, A
Rogers, E. M., Singhal, A. & Quinlan, M. M. Diffusion of innovations. In Salwen, M. B. & Stacks, D. W. (eds.) An Integrated Approach to Communication Theory and Research , 432–448 (Routledge, 2008), 2 edn
2008
-
[17]
Kalyani, A. et al. The diffusion of new technologies. Working Paper 28999, National Bureau of Economic Research (2021). DOI: 10.3386/w28999
2021 doi
-
[18]
Kalyani, A. et al. The diffusion of new technologies*. The Q. J. Econ. qjaf002, DOI: 10.1093/qje/qjaf002 (2025). https://academic.oup.com/qje/advance-article-pdf/doi/10.1093/qje/qjaf002/61489798/qjaf002.pdf
2025 doi
-
[19]
Davis, F. D. Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Q. 13, 319–340, DOI: 10.2307/249008 (1989)
1989 doi
-
[20]
G., Davis, G
Venkatesh, V ., Morris, M. G., Davis, G. B. & Davis, F. D. User acceptance of information technology: Toward a unified view. MIS quarterly 425–478 (2003). 8/23
2003
-
[21]
I., Parasuraman, A
Rojas-Mendez, J. I., Parasuraman, A. & Papadopoulos, N. Demographics, attitudes, and technology readiness: A cross-cultural analysis and model validation. Mark. Intell. & Plan. 35, 18–39 (2017)
2017
-
[22]
Foster, A. D. & Rosenzweig, M. R. Microeconomics of technology adoption. Annu. Rev. Econ. 2, 395–424 (2010)
2010
-
[23]
Jakesch, M., French, M., Ma, X., Hancock, J. T. & Naaman, M. Ai-mediated communication: How the perception that profile text was written by ai affects trustworthiness. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , 1–13 (2019)
2019
-
[24]
Bias in perception of art produced by artificial intelligence
Hong, J.-W. Bias in perception of art produced by artificial intelligence. In Human-Computer Interaction. Interaction in Context: 20th International Conference, HCI International 2018, Las V egas, NV , USA, July 15–20, 2018, Proceedings, Part II 20, 290–303 (Springer, 2018)
2018
-
[25]
& Naaman, M
Kadoma, K., Metaxa, D. & Naaman, M. Generative ai and perceptual harms: Who’s suspected of using llms? arXiv preprint arXiv:2410.00906 (2024)
2024 arXiv
-
[26]
& Hodson, J
Babina, T., Fedyk, A., He, A. & Hodson, J. Artificial intelligence, firm growth, and product innovation. J. Financial Econ. 151, 103745, DOI: 10.1016/j.jfineco.2023.103745 (2024)
2024
-
[27]
& Horton, J
Wiles, E. & Horton, J. J. More, but worse: The impact of ai writing assistance on the supply and quality of job posts. Mass. Inst. Technol. (MIT) Sloan (2024)
2024
-
[28]
Generative AI Top 150: The World’s Most Used AI Tools
Van Rossum, D. Generative AI Top 150: The World’s Most Used AI Tools. https://www.flexos.work/learn/ generative-ai-top-150 (2024)
2024
-
[29]
The Global AI Talent Tracker (2024)
MacroPolo. The Global AI Talent Tracker (2024). 9/23 a d Consumer Complaint UN Press Release bPress Release cLinkedIn Job Posting Figure 1. Temporal dynamics of large language model (LLM) adoption across diverse writing domains. Analysis of LLM-generated or substantially modif...
2024
-
[2024]
Our analysis primarily focused on the full body text. Due to the limited number of articles post-ChatGPT introduction available from Newswire, we conducted detailed robustness checks only on PR Newswire and PRWeb data, which provided sufficient volume for heterogeneity analysi...
1945
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.