REVIEW 3 major objections 5 minor 105 references
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Human unscripted speech showed a significant post-2022 increase in LLM-associated words, while baseline synonyms did not shift, evidence that AI language may be entering the human language system.
desk verdict Novel, honestly-hedged evidence of AI-adjacent words rising in unscripted speech, but the target-word selection from prior written trends makes the convergence claim provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing tool is the weighted log frequency ratio: for each lemma, the log of post-2022 occurrences per million divided by pre-2022 occurrences per million, smoothed with Laplace's +0.5 and weighted by inverse variance (capped at 20), tested with a z-test. The probe is a list of 34 LLM-associated words established in prior work, of which 20 occur in the corpus; the control is a list of 117 semantically related synonyms, of which 87 occur. The comparison between the target words and the synonym baseline is what separates a genuine convergence with LLM patterns from general lexical drift.
What would settle it
A direct falsifier would be to fit each target word's pre-2022 frequency trend in the same podcast corpus and check whether the post-2022 values lie within the extrapolated pre-2022 range; if they do, the claimed post-2022 convergence disappears. Alternatively, a control corpus of unscripted speech from communities with minimal LLM exposure that shows no rise in these words would falsify the claim of a general seep-in effect.
Extended reading notes
Core claim
In unscripted spoken English drawn from 1,326 podcast episodes, the frequency of a curated set of 20 words associated with large language model output increased significantly after ChatGPT's release (weighted mean log frequency ratio = 0.210, z = 3.725, p < 0.001). The 87 synonym baselines showed no significant directional change (0.033, z = 1.277). The paper takes this contrast as evidence that the post-2022 lexical shift is not confined to AI-generated text or to heavily scripted academic writing; it argues that some LLM-favored words are beginning to appear in spontaneous human speech, which it frames as a possible seep-in effect where AI linguistic patterns enter the human language system. The authors are explicit that this does not prove causation: the same words were already trending upward before 2022, and the shift could be natural language change accelerated by AI exposure rather than a direct influence.
Load-bearing premise
The target word list is treated as an unbiased probe of LLM influence, but it was constructed from words that had already risen in written usage after ChatGPT, so the spoken increase could reflect the same pre-existing upward trajectory rather than a new convergence with AI patterns.
Editorial extensions
If this is right
- If human spontaneous speech is measurably adopting LLM-favored words, then the post-2022 lexical shift cannot be fully attributed to direct text generation; the human language system itself is changing.
- The seep-in effect described in the paper implies that repeated exposure to model output can alter speakers' default vocabulary, which parallels the paper's concern that value misalignment could spread through the same mechanism.
- The divergence between model language and human norms — the paper's notion of misalignment — becomes a feedback loop: human speech feeds training data, model output feeds back into speech, potentially homogenizing language over time.
- The study's reliance on unscripted speech illustrates the broader problem of human-authorship indeterminacy, making unscripted speech a necessary data source for linguistic research that aims to track human language.
Reading between the lines
- A testable extension would be to track the same 20 words in a cohort whose LLM exposure is measured directly (e.g., through device logs or self-report), and compare their spoken usage before and after exposure; the paper's causal question would then be answerable.
- The baseline synonym set is not matched on pre-existing trend, so part of the target effect could be regression to the mean; an editor-level check would fit pre-2022 trajectories for each word and ask whether post-2022 changes exceed the extrapolated trend.
- The finding that delve did not rise significantly while words like surpass did suggests the phenomenon is word-specific; isolating which lexical properties (formality, salience, semantic niche) predict adoption could sharpen the mechanism beyond the paper's general convergence claim.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether unscripted spoken English, drawn from 22.1 million words of science and technology podcasts, shows increased use of lexical items associated with large language models after ChatGPT's release in 2022. The authors analyze 20 target words previously identified in their own prior work as increasing in written text, compare pre-2022 and post-2022 frequencies using a weighted mean log frequency ratio, and report a significant increase (0.210, z = 3.725, p < 0.001) against a non-significant baseline of 87 synonyms (0.033, z = 1.277). They interpret this as moderate convergence of human lexical choices with LLM-associated patterns while explicitly leaving causation open. The paper includes a new spoken corpus, openly available code and data, and a candid discussion of limitations, including the absence of a counterfactual.
Significance. If the reported effect were robust, the paper would be a substantive contribution to the emerging literature on LLM influence on human language, with implications for alignment research and for linguistic methodology in an era of AI-generated text. The choice of unscripted podcast speech is a real strength, since it avoids the human-authorship indeterminacy that contaminates written corpora, and the authors are unusually transparent about their data, code, and analytical choices. The main quantitative result is, however, not yet sufficient to support the paper's central inference, because the target list is selected on the very outcome being measured and the baseline is not matched on pre-existing trends. The finding is better characterized, at present, as a descriptive observation about words previously found to increase in written post-ChatGPT text also increasing in one spoken register, rather than as evidence of convergence of the human language system with LLM patterns.
major comments (3)
- [Methods, Target words; Results] The central inference is circular in a way the manuscript does not fully address. The 20 target words are taken from Galpin, Anderson, and Juzek (2025), the authors' own prior study, which selected the original 34 words precisely because their written usage increased after ChatGPT's release. Testing the same words on a new spoken corpus therefore does not independently probe 'LLM-associated vocabulary'; it re-tests a list selected on the outcome of interest. The baseline synonym set is semantically related but is not matched on pre-2022 frequency, register, or pre-existing trajectory. The Discussion itself concedes that many of these words were already trending upward before 2022 (citing Matsui 2024; Juzek and Ward 2025). Because the pre-2022 spoken trajectories are never reported, the contrast between 0.210 and 0.033 may simply reflect a continuation of a pre-existing upward drift of formal academic vocabulary in this tech/science speaker population. The authors should report per-word pre-2022 trends and add control words matched on frequency, register, and prior slope, or explicitly reframe the claim as descriptive rather than evidence of convergence.
- [Results; Data processing and analysis] The main group-level result rests on a single weighted-mean z-test with no confidence interval and no measure of between-word variability beyond sampling weights. The weights are inverse-variance estimates based on frequencies, capped at 20, and a Laplace smoothing constant of 0.5 is applied; these are free parameters whose influence is not examined. The reported z = 3.725 may therefore overstate precision, especially because 6 of the 20 target words decline, two significantly ('crucial' and 'realm'), and five of the significant increases include 'align,' which the authors themselves flag as potentially topic-driven. I request confidence intervals for the weighted mean, a leave-one-out analysis, and a random-effects or bootstrap analysis that accounts for heterogeneity across words. Without these, the claim of a 'robust group-level inference' is not established.
- [Discussion; Results, Table 1] The significant increase for 'align' is explicitly acknowledged in the Discussion to 'need further analysis, as its rise may be linked to a broader shift in discourse, viz. the growing conversation around aligning LLMs with human values.' Because 'align' is one of only five significant increases among the target words, the group-level effect may be inflated by a topical confound rather than by lexical convergence. A sensitivity analysis excluding 'align,' or controlling for the frequency of the discourse topic (e.g., occurrences of 'alignment' or 'AI alignment' in the same episodes), is needed before the group-level result can be interpreted as a general lexical shift.
minor comments (5)
- [Appendix B, Tables 3 and 4] The table captions say 'pre-2022 vs. post-2024' but the text and analysis use 'pre-2022 vs. post-2022' (with 2022 excluded); these captions should be corrected.
- [Discussion, General discussion] The sentence 'Extensive qualitative analyses are projects in their own right' appears twice in the same paragraph; one instance should be deleted.
- [Throughout] There are several typographical and encoding issues, including 'V ariation' (Biber reference), 'F orms' (Goffman reference), and 'T¨urkiye' in the Introduction; these should be cleaned up.
- [Results, Figure 2] The two panels of Figure 2 are not visually labeled in the text as left (target) and right (baseline); adding explicit panel labels or a caption note would improve clarity.
- [Results, Table 1 and Appendix B] The column heading 'p ≤ 0.05' is ambiguous because it reports whether the chi-square test is significant at an uncorrected alpha of 0.05; the heading should say 'Significant at uncorrected α = 0.05' to avoid implying a corrected threshold.
Circularity Check
The target and baseline word lists are imported from the authors' own prior paper that selected words for post-2022 written increases; the spoken increase is then presented as convergence, making the probe selected on outcome, though the new spoken corpus and open causal framing keep the central claim partly independent.
-
fitted input called prediction
[Methods, 'Target words' subsection; Results, weighted-mean log-ratio analysis; Discussion, 'General discussion']
"Target words: For the AI-related words under investigation, we used the 34 words identified in Galpin, Anderson, and Juzek (2025), which were drawn from prior literature and their own analyses. For any of the 34 words, it was established through a formal analysis that there has been a considerable increase in the words' usage pre- and post-ChatGPT's release in 2022."
The probe set is not an independent measure of LLM-associated vocabulary: it was constructed in the authors' own prior paper by selecting words that rose in written usage after ChatGPT's release. The same post-2022 increase is then re-tested in spoken data and interpreted as convergence with LLM patterns. Because the target words were selected on the post-2022 written increase, the spoken increase is a selected-on-outcome result rather than an independent confirmation; the baseline synonyms also come from the same self-cited paper and are not matched on pre-2022 trend or register. The Discussion's admission that many of these words were already trending upward before 2022 (citing Matsui 2024 and Juzek and Ward 2025) reinforces that the list is not an unbiased probe of LLM influence.
full rationale
The paper's statistical derivation itself—computing weighted-mean log frequency ratios for target and baseline words from the podcast corpus—is self-contained and not definitionally forced: no parameter is fitted to the spoken outcomes, and the spoken frequency changes are new data. The circularity concern sits upstream of that calculation. Both the target word set and the synonym baseline are taken from Galpin, Anderson, and Juzek (2025), a prior paper by the same authors, and that prior paper selected the words because they increased in written usage after ChatGPT's release. Re-using that list to test for a post-2022 increase in spoken usage is a selected-on-outcome design: if written and spoken frequencies share any common drift, the target words are expected to rise again regardless of LLM influence. The baseline synonyms, also from the self-cited paper, are semantically related but not matched on pre-2022 frequency, register, or trajectory; the Discussion itself concedes that many target words were already trending upward before 2022 (Matsui 2024; Juzek and Ward 2025). These issues undermine the interpretation of the 0.210 vs 0.033 contrast as LLM convergence, but they do not make the spoken measurement equivalent to the input by definition. Hence score 4 rather than 6 or higher.
Assumptions & free parameters
free parameters (2)
- Laplace smoothing constant =
0.5
- Weight cap for inverse-variance weights =
20
assumptions (4)
- domain assumption The target word list from Galpin, Anderson, and Juzek (2025) is a valid operationalization of LLM-associated vocabulary.
- domain assumption Conversational podcast speech is sufficiently unscripted and free of AI-generated content to serve as genuine human language production.
- domain assumption The 2022 ChatGPT release can be treated as a break point for a pre/post comparison without an explicit counterfactual.
- standard math The inverse-variance weighted z-test with Laplace smoothing is a valid group-level test for corpus frequency changes.
invented entities (1)
-
'Seep-in' effect
Cite this review
Pith. "Pith review of Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English." pith.science (2026). https://pith.science/paper/SXX23ZX4
@misc{pith2026250800238,
author = {Pith},
title = {Pith review of: Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English},
year = {2026},
howpublished = {\url{https://pith.science/paper/SXX23ZX4}},
note = {Machine review of arXiv:2508.00238}
}
read the original abstract
In recent years, written language, particularly in science and education, has undergone remarkable shifts in word usage. These changes are widely attributed to the growing influence of Large Language Models (LLMs), which frequently rely on a distinct lexical style. Divergences between model output and target audience norms can be viewed as a form of misalignment. While these shifts are often linked to using Artificial Intelligence (AI) directly as a tool to generate text, it remains unclear whether the changes reflect broader changes in the human language system itself. To explore this question, we constructed a dataset of 22.1 million words from unscripted spoken language drawn from conversational science and technology podcasts. We analyzed lexical trends before and after ChatGPT's release in 2022, focusing on commonly LLM-associated words. Our results show a moderate yet significant increase in the usage of these words post-2022, suggesting a convergence between human word choices and LLM-associated patterns. In contrast, baseline synonym words exhibit no significant directional shift. Given the short time frame and the number of words affected, this may indicate the onset of a remarkable shift in language use. Whether this represents natural language change or a novel shift driven by AI exposure remains an open question. Similarly, although the shifts may stem from broader adoption patterns, it may also be that upstream training misalignments ultimately contribute to changes in human language use. These findings parallel ethical concerns that misaligned models may shape social and moral beliefs.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Aitchison, J. 2005. Language Change. In Cobley, P., ed., The Routledge Companion to Semiotics and Linguistics, chapter 7. London: Routledge
2005
-
[4]
Akintande, O. J. 2023. Algorithmic Bias: When stigmatization becomes a perception: The stigmatized become endangered. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, 966--971
2023
-
[5]
Al-Sharqi, L.; and Abbasi, I. S. 2020. The Influence of Technology on English Language and Literature. English Language Teaching, 13(7): 1
2020
-
[6]
I.; Babaei, H.; LeJeune, D.; Siahkoohi, A.; and Baraniuk, R
Alemohammad, S.; Casco-Rodriguez, J.; Luzi, L.; Humayun, A. I.; Babaei, H.; LeJeune, D.; Siahkoohi, A.; and Baraniuk, R. G. 2023. Self-consuming generative models go mad. arXiv preprint arXiv:2307.01850
arXiv 2023
-
[7]
Ali, J.; Kleindessner, M.; Wenzel, F.; Budhathoki, K.; Cevher, V.; and Russell, C. 2023. Evaluating the fairness of discriminative foundation models in computer vision. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, 809--833
2023
-
[8]
Amato, R.; Lacasa, L.; Díaz-Guilera, A.; and Baronchelli, A. 2018. The dynamics of norm change in the cultural evolution of language. Proceedings of the National Academy of Sciences, 115(33): 8260–8265
2018
Show all 105 references
-
[9]
Bao, T.; Zhao, Y.; Mao, J.; and Zhang, C. 2025. Examining linguistic shifts in academic writing before and after the launch of ChatGPT: a study on preprint papers. Scientometrics, 1--31
2025
-
[10]
M.; Gebru, T.; McMillan-Major, A.; and Shmitchell, S
Bender, E. M.; Gebru, T.; McMillan-Major, A.; and Shmitchell, S. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, 610--623
2021
-
[11]
Biber, D. 1991. Variation across speech and writing. Cambridge university press
1991
-
[12]
Bland, M. 2015. An introduction to medical statistics. Oxford university press
2015
-
[13]
L.; Barocas, S.; Daum \'e III, H.; and Wallach, H
Blodgett, S. L.; Barocas, S.; Daum \'e III, H.; and Wallach, H. 2020. Language (technology) is power: A critical survey of" bias" in nlp. arXiv preprint arXiv:2005.14050
2020 arXiv
-
[14]
Bochkarev, V.; Solovyev, V.; and Wichmann, S. 2014. Universals versus historical contingencies in lexical evolution. Journal of the Royal Society Interface, 11(101): 20140841
2014
-
[15]
Y.; Saligrama, V.; and Kalai, A
Bolukbasi, T.; Chang, K.-W.; Zou, J. Y.; Saligrama, V.; and Kalai, A. T. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. Advances in neural information processing systems, 29
2016
-
[16]
Briesch, M.; Sobania, D.; and Rothlauf, F. 2023. Large language models suffer from their own output: An analysis of the self-consuming training loop. arXiv preprint arXiv:2311.16822
2023 arXiv
-
[17]
Bybee, J. L. 2006. From usage to grammar: The mind's response to repetition. Language, 82(4): 711--733
2006
-
[18]
Chen, Y.; Zhao, W.; Breitbarth, A.; Stoeckel, M.; Mehler, A.; and Eger, S. 2024. Syntactic language change in english and german: Metrics, parsers, and convergences. arXiv preprint arXiv:2402.11549
2024 arXiv
-
[19]
F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30
2017
-
[20]
Clark, H. H. 1996. Using Language. Cambridge University Press
1996
-
[21]
H.; Poliak, M.; Regev, T.; Haskins, A
Clark, T. H.; Poliak, M.; Regev, T.; Haskins, A. J.; Gibson, E.; and Robertson, C. 2025. The relationship between surprisal, prosody, and backchannels in conversation reflects intelligibility-oriented pressures. PsyArXiv preprints
2025
-
[22]
Comas-Forgas, R.; Koulouris, A.; and Kouis, D. 2025. ‘AI-navigating’or ‘AI-sinking’? An analysis of verbs in research articles titles suspicious of containing AI-generated/assisted content. Learned Publishing, 38(1): e1647
2025
-
[23]
Crystal, D. 2009. Txtng: The Gr8 Db8
2009
-
[24]
Djalolovna, M. S. 2025. The Influence of Technology on Spoken Language: Texting, Voice Assistants, and Speech Patterns. Web of Technology: Multidimensional Research Journal, 3(1): 34--38
2025
-
[25]
Eisenstein, E. 1980. The Printing Press as an Agent of Change: Communications and Cultural Transformations in Early-Modern Europe. Cambridge University Press
1980
-
[26]
T.; and Gibson, E
Fedorenko, E.; Piantadosi, S. T.; and Gibson, E. A. F. 2024. Language is primarily a tool for communication rather than thought. Nature, 630(8017): 575--586
2024
-
[27]
Fillmore, C. J. 2006. Frame Semantics. In Geeraerts, D.; Dirven, R.; and Taylor, J. R., eds., Cognitive Linguistics: Basic Readings, volume 34 of Cognitive Linguistics Research, 373--400. Berlin: Mouton de Gruyter
2006
-
[28]
Floridi, L. 2025. Distant Writing: Literary Production in the Age of Artificial Intelligence. Minds and Machines, 35(3): 1--26
2025
-
[29]
Floridi, L.; Cowls, J.; Beltrametti, M.; Chatila, R.; Chazerand, P.; Dignum, V.; Luetge, C.; Madelin, R.; Pagallo, U.; Rossi, F.; et al. 2018. AI4People—an ethical framework for a good AI society: opportunities, risks, principles, and recommendations. Minds and machines, 28: 689--707
2018
-
[30]
Gabriel, I. 2020. Artificial intelligence, values, and alignment. Minds and machines, 30(3): 411--437
2020
-
[31]
Galpin, R.; Anderson, B.; and Juzek, T. S. 2025. Lexical underrepresentation matters as much as lexical overrepresentation. arXiv preprint arXiv:2501.20200
2025
-
[32]
Gee, J. P. 2014. An introduction to discourse analysis: Theory and method. routledge
2014
-
[33]
Geng, M.; Chen, C.; Wu, Y.; Chen, D.; Wan, Y.; and Zhou, P. 2024. The Impact of Large Language Models in Academia: from Writing to Speaking. arXiv preprint arXiv:2409.13686
2024 arXiv
-
[34]
Ghosh, S.; and Caliskan, A. 2023. Chatgpt perpetuates gender bias in machine translation and ignores non-gendered pronouns: Findings across bengali and five other low-resource languages. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, 901--912
2023
-
[35]
Goffman, E. 1981. Forms of talk. University of Pennsylvania Press
1981
-
[36]
Grabowicz, P.; Perello, N.; and Takatsu, K. 2023. Learning from Discriminatory Training Data. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, 752--763
2023
-
[37]
contamination
Gray, A. 2024. ChatGPT" contamination": estimating the prevalence of LLMs in the scholarly literature. arXiv preprint arXiv:2403.16887
2024 arXiv
-
[38]
Hagen, L.; Huo, J.; and Nguyen, A. 2025. Elon Musk's AI Chatbot, Grok, Started Calling Itself 'MechaHitler'. NPR
2025
-
[39]
Hartvigsen, T.; Gabriel, S.; Palangi, H.; Sap, M.; Ray, D.; and Kamar, E. 2022. Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. arXiv preprint arXiv:2203.09509
2022 arXiv
-
[40]
Haslwanter, T. 2016. An introduction to statistics with python. With applications in the life sciences. Switzerland: Springer International Publishing
2016
-
[41]
Hataya, R.; Bao, H.; and Arai, H. 2023. Will large-scale generative models corrupt future datasets? In Proceedings of the IEEE/CVF International Conference on Computer Vision, 20555--20565
2023
-
[42]
Helen, J. W. 2010. The Effect of Internet Devices on Children's Language Development. Contemporary Issues in Communication Science Disorders, 37: 141--148
2010
-
[43]
Hendrycks, D.; Burns, C.; Basart, S.; Critch, A.; Li, J.; Song, D.; and Steinhardt, J. 2020. Aligning ai with shared human values. arXiv preprint arXiv:2008.02275
2020 arXiv
-
[44]
Huang, B.; Chen, C.; and Shu, K. 2025. Authorship attribution in the era of llms: Problems, methodologies, and challenges. ACM SIGKDD Explorations Newsletter, 26(2): 21--43
2025
-
[45]
Jaiswal, S. 2024. Uncovering Gender Biases in Human-AI Platforms. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7(2), 21--22
2024
-
[46]
Jin, H.; Sung, S.; Park, S.; Baik, S.; and Han, Y.-S. 2025. TRAPDOC: Deceiving LLM Users by Injecting Imperceptible Phantom Tokens into Documents. arXiv preprint arXiv:2506.00089
2025
-
[47]
Jobin, A.; Ienca, M.; and Vayena, E. 2019. The global landscape of AI ethics guidelines. Nature machine intelligence, 1(9): 389--399
2019
-
[48]
Jumelet, J.; Weissweiler, L.; and Bisazza, A. 2025. MultiBLiMP 1.0: A massively multilingual benchmark of linguistic minimal pairs. arXiv preprint arXiv:2504.02768
2025 arXiv
-
[49]
Jurafsky, D.; and Martin, J. H. 2024. Speech and Language Processing. Online draft. 3rd ed. draft, Feb 3, 2024 release
2024
-
[50]
S.; and Ward, Z
Juzek, T. S.; and Ward, Z. B. 2025. Why Does ChatGPT" Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models. In Proceedings of the 31st International Conference on Computational Linguistics (COLING 2025)
2025
-
[51]
S.; and Sinnott-Armstrong, W
Keswani, V.; Conitzer, V.; Heidari, H.; Borg, J. S.; and Sinnott-Armstrong, W. 2024. On the Pros and Cons of Active Learning for Moral Preference Elicitation. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, 711--723
2024
-
[52]
G.; Horv \'a t, E.- \'A .; and Lause, J
Kobak, D.; M \'a rquez, R. G.; Horv \'a t, E.- \'A .; and Lause, J. 2024. Delving into ChatGPT usage in academic writing through excess vocabulary. arXiv preprint arXiv:2406.07016
2024 arXiv
-
[53]
Koppenburg, P. 2024. Tweet on 01 April 2024. https://x.com/PKoppenburg/status/1774757167045788010. Accessed: 2025-01-23
2024
-
[54]
Kotek, H.; Dockum, R.; and Sun, D. 2023. Gender bias and stereotypes in large language models. In Proceedings of the ACM collective intelligence conference, 12--24
2023
-
[55]
Koziev, I. 2025. Detecting Spelling and Grammatical Anomalies in Russian Poetry Texts. arXiv preprint arXiv:2505.04507
2025 arXiv
-
[56]
Krielke, M.-P. 2021. Relativizers as markers of grammatical complexity: A diachronic, cross-register study of English and German. Bergen Language and Linguistics Studies, 11(1): 91--120
2021
-
[57]
a ndische Universit \
Krielke, M.-P. 2023. Optimizing scientific communication: the role of relative clauses as markers of complexity in English and German scientific writing between 1650 and 1900. Saarl \"a ndische Universit \"a ts-und Landesbibliothek
2023
-
[58]
Krielke, M.-P. 2024. Cross-linguistic Dependency Length Minimization in scientific language: Syntactic complexity reduction in English and German in the Late Modern period. Languages in Contrast, 24(1): 133--163
2024
-
[59]
Krielke, M.-P.; Talamo, L.; Fawzi, M.; and Knappen, J. 2022. Tracing syntactic change in the scientific genre: Two Universal Dependency-parsed diachronic corpora of scientific English and German. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, 4808--4816
2022
-
[60]
Labov, W. 2011. Principles of linguistic change, volume 3: Cognitive and cultural factors, volume 3. John Wiley & Sons
2011
-
[61]
Leidinger, A.; and Rogers, R. 2024. How are LLMs mitigating stereotyping harms? Learning from search engine studies. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, 839--854
2024
-
[62]
Li, Y.; Breithaupt, F.; Hills, T.; Lin, Z.; Chen, Y.; Siew, C. S. W.; and Hertwig, R. 2023. How Cognitive Selection Affects Language Change. Proceedings of the National Academy of Sciences, 121(1). PNAS Direct Submission. Published December 27, 2023
2023
-
[63]
Liang, W.; Zhang, Y.; Wu, Z.; Lepp, H.; Ji, W.; Zhao, X.; Cao, H.; Liu, S.; He, S.; Huang, Z.; et al. 2024. Mapping the increasing use of llms in scientific papers. arXiv preprint arXiv:2404.01268
2024 arXiv
-
[64]
Liu, J.; and Bu, Y. 2024. Towards the relationship between AIGC in manuscript writing and author profiles: evidence from preprints in LLMs. arXiv preprint arXiv:2404.15799
2024 arXiv
-
[65]
L \"u cking, A.; Abrami, G.; Hammerla, L.; Rahn, M.; Baumartz, D.; Eger, S.; and Mehler, A. 2024. Dependencies over times and tools (DoTT). In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 20...
2024
-
[66]
Matsui, K. 2024. Delving into PubMed Records: Some Terms in Medical Writing Have Drastically Changed after the Arrival of ChatGPT. medRxiv, 2024--05
2024
-
[67]
E.; and Weinert, R
Miller, J. E.; and Weinert, R. 1998. Spontaneous spoken language: Syntax and discourse. Oxford University Press
1998
-
[68]
D.; Allo, P.; Taddeo, M.; Wachter, S.; and Floridi, L
Mittelstadt, B. D.; Allo, P.; Taddeo, M.; Wachter, S.; and Floridi, L. 2016. The ethics of algorithms: Mapping the debate. Big Data & Society, 3(2): 2053951716679679
2016
-
[69]
V.; and Peters, H
Montani, I.; Honnibal, M.; Boyd, A.; Landeghem, S. V.; and Peters, H. 2023. explosion/spaCy: v3.7.2: Fixes for APIs and requirements. Version v3.7.2, Zenodo
2023
-
[70]
Naik, R.; and Nushi, B. 2023. Social biases through the text-to-image generation lens. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, 786--808
2023
-
[71]
Narayanan Venkit, P.; Gautam, S.; Panchanadikar, R.; Huang, T.-H.; and Wilson, S. 2023. Unmasking nationality bias: A study of human perception of nationalities in ai-generated articles. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, 554--565
2023
-
[72]
Neff, G.; and Nagy, P. 2016. Talking to bots: Symbiotic agency and the case of Tay. International Journal of Communication, 10: 4915--4931
2016
-
[73]
Norhashim, H.; and Hahn, J. 2024. Measuring Human-AI Value Alignment in Large Language Models. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, 1063--1073
2024
-
[74]
A.; Lester, J
Omiye, J. A.; Lester, J. C.; Spichak, S.; Rotemberg, V.; and Daneshjou, R. 2023. Large language models propagate race-based medicine. NPJ Digital Medicine, 6(1): 195
2023
-
[75]
Omrani Sabbaghi, S.; Wolfe, R.; and Caliskan, A. 2023. Evaluating biased attitude associations of language models in an intersectional context. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, 542--553
2023
-
[76]
OpenAI. 2024. Tweet on 08 April 2024. https://x.com/ChatGPTapp/status/1777221658807521695. Accessed: 2024-08-12
2024
-
[77]
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 27730--27744
2022
-
[78]
Pack, A.; and Maloney, J. 2023. Potential Affordances of Generative AI in Language Education: Demonstrations and an Evaluative Framework. Teaching English with Technology, 23(2)
2023
-
[79]
Perez, E.; Ringer, S.; Lukosiute, K.; Nguyen, K.; Chen, E.; Heiner, S.; Pettit, C.; Olsson, C.; Kundu, S.; Kadavath, S.; et al. 2023. Discovering language model behaviors with model-written evaluations. In Findings of the Association for Computational Linguistics: ACL 2023, 13...
2023
-
[80]
Python Software Foundation . 2024. Python 3
2024
-
[81]
W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I
Radford, A.; Kim, J. W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I. 2022. Whisper. https://github.com/openai/whisper
2022
-
[82]
D.; Ermon, S.; and Finn, C
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36
2024
-
[83]
Russell, S.; Dewey, D.; and Tegmark, M. 2015. Research priorities for robust and beneficial artificial intelligence. AI magazine, 36(4): 105--114
2015
-
[84]
S.; Kumar, A.; Balasubramanian, S.; Wang, W.; and Feizi, S
Sadasivan, V. S.; Kumar, A.; Balasubramanian, S.; Wang, W.; and Feizi, S. 2023. Can AI-generated text be reliably detected? arXiv preprint arXiv:2303.11156
2023 arXiv
-
[85]
Sandra, G. 2011. 7. Greiffenstern, Sandra. 2010. The influence of computers, the internet and computer-mediated communication on everyday English. English and American Studies in German, 2011(1): 16--18
2011
-
[86]
Schmalz, V.; and Tack, A. 2025. Can GPTZ ero ' s AI Vocabulary Distinguish Between LLM -Generated and Student-Written Essays? In Kochmar, E.; Alhafni, B.; Bexte, M.; Burstein, J.; Horbach, A.; Laarmann-Quante, R.; Tack, A.; Yaneva, V.; and Yuan, Z., eds., Proceedings of the 20...
2025
-
[87]
Shumailov, I.; Shumaylov, Z.; Zhao, Y.; Gal, Y.; Papernot, N.; and Anderson, R. 2023. The curse of recursion: Training on generated data makes models forget. arXiv preprint arXiv:2305.17493
2023 arXiv
-
[88]
A.; Abid, A.; Fisch, A.; Brown, A
Srivastava, A.; Rastogi, A.; Rao, A.; Shoeb, A. A.; Abid, A.; Fisch, A.; Brown, A. R.; Santoro, A.; Gupta, A.; Garriga-Alonso, A.; et al. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on machine learning research
2023
-
[89]
Tomasello, M. 2005. Constructing a language: A usage-based theory of language acquisition. Harvard university press
2005
-
[90]
Traugott, E. C. 1985. On regularity in semantic change. Journal of Literary Semantics, 14(3): 155--173
1985
-
[91]
P.; and Sri Nidhya, G
Vanisree, M.; Sharma, S.; Singh, L.; Kantharaja, K. P.; and Sri Nidhya, G. 2024. Effects of Technology and Social Media on English Language Use and Acquisition: A Study on Language Variation and Change. Nanotechnology Perceptions, 20(S7): 490--497
2024
-
[92]
Wang, G.; Wang, H.; Sun, X.; Wang, N.; and Wang, L. 2023. Linguistic complexity in scientific writing: A large-scale diachronic study from 1821 to 1920. Scientometrics, 128(1): 441--460
2023
-
[93]
Warstadt, A.; Parrish, A.; Liu, H.; Mohananey, A.; Peng, W.; Wang, S.-F.; and Bowman, S. R. 2020. BLiMP: The benchmark of linguistic minimal pairs for English. Transactions of the Association for Computational Linguistics, 8: 377--392
2020
-
[94]
Weber-Wulff, D.; Anohina-Naumeca, A.; Bjelobaba, S.; Folt \`y nek, T.; Guerrero-Dib, J.; Popoola, O.; S igut, P.; and Waddington, L. 2023. Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19(1): 1--39
2023
-
[95]
Weissweiler, L.; Mahowald, K.; and Goldberg, A. 2025. Linguistic generalizations are not rules: Impacts on evaluation of LMs. arXiv preprint arXiv:2502.13195
2025
-
[96]
A.; Anderson, K.; Kohli, P.; Coppin, B.; and Huang, P.-S
Welbl, J.; Glaese, A.; Uesato, J.; Dathathri, S.; Mellor, J.; Hendricks, L. A.; Anderson, K.; Kohli, P.; Coppin, B.; and Huang, P.-S. 2021. Challenges in detoxifying language models. arXiv preprint arXiv:2109.07445
2021 arXiv
-
[97]
Wilson, K.; and Caliskan, A. 2024. Gender, race, and intersectional bias in resume screening via language model retrieval. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, 1578--1590
2024
-
[98]
Wu, M.; and Aji, A. F. 2025. Style Over Substance: Evaluation Biases for Large Language Models. In Rambow, O.; Wanner, L.; Apidianaki, M.; Al-Khalifa, H.; Eugenio, B. D.; and Schockaert, S., eds., Proceedings of the 31st International Conference on Computational Linguistics, 2...
2025
-
[99]
Wu, Y.; Bizzoni, Y.; Moreira, P.; and Nielbo, K. 2024. Perplexing canon: A study on GPT-based perplexity of canonical and non-canonical literary works. In Proceedings of the 8th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humaniti...
2024
-
[100]
Yakura, H.; Lopez-Lopez, E.; Brinkmann, L.; Serna, I.; Gupta, P.; and Rahwan, I. 2024. Empirical evidence of Large Language Model's influence on human spoken communication. arXiv preprint arXiv:2409.01754
2024 arXiv
-
[101]
Zamaraeva, O.; Flickinger, D.; Bond, F.; and G \'o mez-Rodr \' guez, C. 2025. Comparing LLM-generated and human-authored news text using formal syntactic theory. arXiv preprint arXiv:2506.01407
2025 arXiv
-
[102]
Zhang, R.; Ouni, J.; and Eger, S. 2024. Cross-lingual cross-temporal summarization: Dataset, models, evaluation. Computational Linguistics, 50(3): 1001--1047
2024
-
[103]
Zhang, X.; Xiong, W.; Chen, L.; Zhou, T.; Huang, H.; and Zhang, T. 2024. From lists to emojis: How format bias affects model alignment. arXiv preprint arXiv:2409.11704
2024 arXiv
-
[104]
Zhang, Y. 2014. A Review of Principles of Linguistic Change: Cognitive and Cultural Factors. Open Journal of Modern Linguistics, 4: 65--68
2014
-
[105]
M.; Stiennon, N.; Wu, J.; Brown, T
Ziegler, D. M.; Stiennon, N.; Wu, J.; Brown, T. B.; Radford, A.; Amodei, D.; Christiano, P.; and Irving, G. 2019. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593
2019 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.