REVIEW 3 major objections 3 minor 36 references
Politicians vs ChatGPT. A study of presuppositions in French and Italian political communication
T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Compared with the politicians' speeches they imitate, ChatGPT 3.5 texts carry more potentially manipulative presuppositions, concentrated in change-of-state slogans like 'we must build our future'.
desk verdict A genuinely new comparison of presupposition use in human vs ChatGPT political speech, but the central frequency claim depends on a non-blind annotation step that needs a reliability check. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The unit that carries the argument is the potentially manipulative presupposition (PMP): a presupposition triggered by a dedicated linguistic form, such as a definite description or a change-of-state verb, whose presupposed content is not bona fide true but tendentious, subjective, or false. The authors identify PMPs in the corpus, annotate their presupposition triggers and their discourse functions (criticism, self-praise, praise-of-others, stance-taking), and then compare the frequency, form, and function of PMPs between politicians' speeches and ChatGPT's imitations; the statistical comparisons rest on the resulting annotation.
What would settle it
Annotate the same sixteen texts again with coders blind to whether each text is a politician's speech or a ChatGPT output, and compute PMP rates; if the ChatGPT advantage disappears or reverses under blind coding, the central claim is not supported.
Extended reading notes
Core claim
The paper's central claim is that ChatGPT-generated political texts are not just superficial imitations of politicians' speeches: they differ systematically in their use of implicit content. Across the corpus, normalized frequencies of potentially manipulative presuppositions are higher in the ChatGPT texts than in the politicians' speeches, and the difference is statistically significant though small in effect size. The trigger types also differ: definite descriptions are the most common PMP trigger in politicians' speeches, whereas change-of-state verbs are the most common in ChatGPT texts. Discourse functions diverge too, with stance-taking massively overrepresented in the AI texts and criticism positively associated with the politicians' texts; Macron's data are the one case where the distribution of functions does not differ. The authors interpret the ChatGPT pattern as repetition and vagueness arising from next-word prediction, which selects high-frequency slogan-like constructions such as 'we must build our future'.
Load-bearing premise
The entire frequency comparison rests on the two annotators' judgment of which presuppositions count as 'questionable', made without blind conditions and without a formal agreement measure at that identification step.
Editorial extensions
If this is right
- If the central claim is correct, AI-generated political speech is not merely vaguer than human speech; it packages contestable positions as taken-for-granted facts more densely, which makes them harder to challenge.
- The predominance of change-of-state verbs in ChatGPT texts means the AI's manipulative profile is especially concentrated in slogans like 'we must build our future', where the presupposed need or goal is smuggled in as shared.
- PMP frequency offers a measurable proxy for the manipulative potential of AI-written political texts, and the trigger-type distribution gives an observable signature that could be tracked automatically.
- The Macron exception indicates that the content of the prompt can shape the discourse functions of PMPs in the output, so the manipulative profile of generated texts is partly under the control of whoever writes the prompt.
- Because the same slogans recur across politicians, languages, and political orientations, the pattern points to a model-level tendency rather than to faithful imitation of any individual speaker's style.
Reading between the lines
- One extension this comparison does not run is a perceptual test: the paper shows the AI texts contain more PMPs, but not that readers are more persuaded or less likely to notice them; connecting the frequency difference to actual persuasion would complete the argument.
- Because change-of-state triggers are lexically identifiable, the finding suggests a cheap automatic screen for manipulative presuppositions in AI-generated texts, even though the PMP judgment itself requires human annotation.
- The paper notes that most of ChatGPT's training data is in English; a natural test is whether the same PMP gap appears in English-language political texts or in languages with even less training data, which could reveal whether the effect grows with linguistic distance from pretraining data.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper compares the use of 'potentially manipulative presuppositions' (PMPs) in authentic French and Italian political speeches with texts generated by ChatGPT 3.5 that mimic those politicians. Based on a manually annotated corpus of eight human and eight chatbot texts, the authors report (Q1) that ChatGPT texts contain more PMPs per 1,000 words, (Q2) that change-of-state verbs are more frequent triggers in ChatGPT texts while definite descriptions dominate in politicians' texts, and (Q3) that stance-taking is proportionally more common in ChatGPT texts while criticism dominates in politicians' texts. The statistical analysis includes chi-square tests, Fisher's exact tests, and a loglinear model, with effect sizes. The authors frame the work as an exploratory contribution to the pragmatics of LLM-generated political discourse and openly list limitations concerning corpus size and stochastic generation.
Significance. If the frequency and distribution differences are robust, the paper provides a concrete, falsifiable empirical observation about a widely feared property of LLM-generated political speech: the heavier reliance on implicit, potentially manipulative content, especially via repetitive change-of-state slogans. The authors use an established theoretical framework for presupposition and manipulative discourse, and they make their corpus and R code available on OSF, which is a strength for reproducibility. The statistical toolkit is appropriate for categorical corpus data, and the transparency about limitations in Section 7 is commendable. However, the central claim's validity hinges entirely on the reliability of the first annotation step, PMP identification, for which no interrater measure is reported.
major comments (3)
- [Section 5.1] The central frequency claim (Q1, abstract, Section 6) rests on the identification of PMPs, a context-sensitive judgment of whether presupposed content is 'non-bona fide true' or tendentious. The manuscript states that the two annotators (also the first and second authors) identified PMPs independently and then resolved disagreements through consensus negotiation, discarding doubtful items, and that 'no formal analysis based on interrater agreement indices was carried out at this stage.' The agreement indices reported in Tables 2 and 3 cover only Presupposition Trigger and Discourse Function after the PMP set was already fixed, so they do not validate the frequency comparison. Because the annotators knew which texts were produced by politicians and which by ChatGPT, expectation bias could systematically shift PMP counts, and the reported chi-square difference (χ2 = 10.28, df = 3, p < .05; Cramér's V = .16) is small enough that a modest annotation shift could alter or erase it. A blinded reannotation of PMP identification by independent annotators, with a PMP/non-PMP agreement measure (e.g., Cohen's kappa on a sample), is needed before the frequency claim can be treated as robust.
- [Section 5.2 and Table 1] The normalized frequency comparison in Figure 2 is based on per-1,000-word rates, but the ChatGPT texts are roughly 35–40% shorter than the politicians' speeches (Table 1). More importantly, the chi-square test in Section 5.2 treats each PMP occurrence as an independent observation, ignoring text-level clustering. With only eight texts per group, the effective sample size for the frequency comparison is very small, and the p-value would not account for potential between-text variability. Please report a permutation or mixed-effects test that treats text as a random effect, or otherwise justify the independence assumption for PMP occurrences.
- [Section 4 and Section 6 (Macron exception)] The comparison between human and ChatGPT texts is partly confounded by the prompt design: the prompt in (4) includes an excerpt from the real speech, a requested length of about 5,500 characters (which ChatGPT did not reach), and style guidelines. The authors discuss the Macron self-praise exception as potentially due to the prompt excerpt, but the more general risk is that the discourse-function differences (Figure 4) reflect differences in text length, genre, or the specific seed excerpt rather than a general property of ChatGPT. Please explain how these confounds were controlled, or temper the general claim that ChatGPT texts are intrinsically more stance-taking and less critical.
minor comments (3)
- [Section 5.2] In the sentence reporting the Fisher's exact test results for Macron, 'Fisher's Exact Test = .93' presumably reports the p-value, not the test statistic; please clarify the notation for consistency.
- [References] There are minor reference inconsistencies: 'Altmann' in the epigraph should be 'Altman'; the in-text citation 'Cai, Duang and Haslett 2023' appears to refer to 'Cai, Duan, and Haslett' in the reference list; and 'Lowen and Plonsky 2016' in Section 5.1 is spelled 'Loewen and Plonsky 2016' in the reference list.
- [Section 4] The paper states that the prompt was written in English and the output requested in French or Italian, but example (4) shows only the English prompt; providing the full prompt for both languages in an appendix or OSF would aid replicability.
Circularity Check
No circularity: the frequency comparison is an empirical annotation result, not a derivation that reduces to its inputs.
full rationale
The paper makes an empirical claim about the frequency of potentially manipulative presuppositions (PMPs) in politician speeches versus ChatGPT-generated texts. The counts come from manual annotation (Section 5.1), not from a fitted parameter, an equation, or a theorem whose conclusion is assumed in its premises. The PMP notion and the four discourse functions are taken from prior work, including the authors' own Garassino, Brocca and Masia (2022) and Garassino, Masia and Brocca (2019), but this is a normal use of an established coding framework rather than a load-bearing self-citation that forces the result: the same categories are applied to new texts, and the observed differences (chi-squared = 10.28 for frequency, chi-squared = 16.13 for triggers, chi-squared = 49.19 for functions) are contingent on the annotation. The acknowledged limitations, such as the small corpus, the non-blind PMP identification, the absence of an interrater agreement index at the PMP-identification stage, and the stochastic nature of ChatGPT, are validity and robustness concerns, not circularity. No step in the paper defines its outcome into existence or renames a fitted input as a prediction. The central comparison therefore has independent empirical content.
Assumptions & free parameters
assumptions (3)
- domain assumption Presuppositions can be classified as bona fide true or potentially manipulative (PMP) based on whether the content is objectively verifiable and shared.
- domain assumption The four discourse functions (criticism, self-praise, praise of others, stance-taking) are exhaustive and mutually exclusive for the purposes of annotation.
- domain assumption The selected speeches and generated texts are comparable enough to compare presupposition rates.
Cite this review
Pith. "Pith review of Politicians vs ChatGPT. A study of presuppositions in French and Italian political communication." pith.science (2026). https://pith.science/paper/CJXM2NMU
@misc{pith2026241118403,
author = {Pith},
title = {Pith review of: Politicians vs ChatGPT. A study of presuppositions in French and Italian political communication},
year = {2026},
howpublished = {\url{https://pith.science/paper/CJXM2NMU}},
note = {Machine review of arXiv:2411.18403}
}
read the original abstract
This paper aims to provide a comparison between texts produced by French and Italian politicians on polarizing issues, such as immigration and the European Union, and their chatbot counterparts created with ChatGPT 3.5. In this study, we focus on implicit communication, in particular on presuppositions and their functions in discourse, which have been considered in the literature as a potential linguistic feature of manipulation. This study also aims to contribute to the emerging literature on the pragmatic competences of Large Language Models.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[3]
Artificial intelligence can persuade humans on political issues. OSF Preprints . https://doi.org/10.31219/osf.io/stakv Bender, Emily M. & Gebru, Timnit & McMillan -Major, Angelina & Shmitchell, Shmargaret
-
[6]
Implicitness impact: measuring texts
https://doi.org/10.3389/fcomm.2021.610807 Lombardi Vallauri, Edoardo & Masia, Viviana. Implicitness impact: measuring texts. Journal of Pragmatics 61, 161–184. https://doi.org/10.1016/j.pragma.2013.09.010 Loewen, Shawn & Plonsky, Luke
-
[7]
How Language Models Could Change Disinformation
Truth, Lies, and Automation. How Language Models Could Change Disinformation. Georgetown: Center for Security and Emerging Technology. https://cset.georgetown.edu/publication/truth-lies-and-automation (last accessed on 26.05.2024) Burtell, Matthew & Woodside, Thomas
work page 2024
-
[8]
Artificial influence: An Analysis of AI-Driven persuasion. ArXiv. https://doi.org/10.48550/arXiv.2303.08721 Cai, Zhenguang G. & Duan, Xufeng & Haslett, David A. & Wang, Shuqi & Pickering, Martin J
-
[9]
https://doi.org/10.48550/arXiv.2303.08014 Chaka, Chaka
Does ChatGPT resemble humans in language use? ArXiv. https://doi.org/10.48550/arXiv.2303.08014 Chaka, Chaka. (2023). Detecting AI content in responses generated by ChatGPT, YouChat, and Chatsonic: The case of five AI content detection tools. Journal of Applied Learning and Teaching 6(2). https://doi.org/10.37074/jalt.2023.6.2.12 Cominetti, Federica & Greg...
-
[12]
https://doi.org/10.36227/techrxiv.22683919.v2 Field, Andy & Miles, Jeremy & Field, Zoë
-
[13]
http://pragmatics.gr.jp/content/files/SIP_013/SIP_13_Reboul.pdf (last accessed on 26.05.2024)
1−19. http://pragmatics.gr.jp/content/files/SIP_013/SIP_13_Reboul.pdf (last accessed on 26.05.2024). Reboul, Anne
work page 2024
-
[15]
https://doi.org/10.31235/osf.io/fp87b Goldstein, Josh A
Can AI write persuasive propaganda? SoCArXiv. https://doi.org/10.31235/osf.io/fp87b Goldstein, Josh A. & Sastry, Girish & Muss er, Micah & Gentzel, Matthew, & Sedova, Katerina
Show all 36 references
-
[16]
Stanford Internet Observatory, OpenAI, and Georgetown University’s Center for Security and Emerging Technology
Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations . Stanford Internet Observatory, OpenAI, and Georgetown University’s Center for Security and Emerging Technology. https://fsi.stanford.edu/publication/generative-language-...
2024
-
[23]
Evidence from the online processing of presuppositions in political tweets
Recalling presupposed information. Evidence from the online processing of presuppositions in political tweets. Pragmatics & Cognition 30:1. 92 –119. https://doi.org/10.1075/pc.22011.mas Mercier, Hugo & Sperber, Dan
-
[24]
https://openai.com/ (last accessed on 06.05.2023)
ChatGPT (Version GPT -3.5) [Computer software]. https://openai.com/ (last accessed on 06.05.2023). Pratim Ray, Partha
2023
-
[25]
https://doi.org/10.1016/j.iotcps.2023.04.003 Qiu, Zhuang & Duan, Xufeng & Cai, Zhenguang G
121−154. https://doi.org/10.1016/j.iotcps.2023.04.003 Qiu, Zhuang & Duan, Xufeng & Cai, Zhenguang G
2023 doi
-
[26]
PsyArXiv
Pragmatic Implicature Processing in ChatGPT. PsyArXiv. https://osf.io/preprints/psyarxiv/qtbh9 (last accessed on 24.05.2024). R Core Team
2024
-
[27]
R Foundation for Statistical Computing, Vienna (Version 4.2.2) [Computer software]
R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna (Version 4.2.2) [Computer software]. https://www.R-project.org/ (last accessed on 06.06.2024) Reboul, Anne
2024
-
[31]
Corpus Linguistics and Linguistic Theory 6(2)
Coding coherence relations: Reliability and validity. Corpus Linguistics and Linguistic Theory 6(2). 241−266. https://doi.org/10.1515/cllt.2010.009 Stalnaker, Robert C
2010 doi
-
[32]
https://doi.org/10.1023/A:1020867916902 Vach, Werner & Gerke, Oke
701−721. https://doi.org/10.1023/A:1020867916902 Vach, Werner & Gerke, Oke
-
[33]
A comparison of basic properties
Gwet’s AC1 is not a substitute for Cohen’s kappa. A comparison of basic properties. MethodsX 10, 102212. https://doi.org/10.1016/j.mex.2023.102212 Venables, William N. & Ripley, Brian D. 2002 . Modern Applied Statistics with S (3rd edition). New York: Springer. White, Jules & ...
2023
- [34]
- [35]
-
[42]
https://doi.org/10.1016/j.jneuroling.2016.11.005 Masia, Viviana & Garassino, Davide & Brocca, Nicola & de Saussure, Louis
31–48. https://doi.org/10.1016/j.jneuroling.2016.11.005 Masia, Viviana & Garassino, Davide & Brocca, Nicola & de Saussure, Louis
2016 doi
-
[45]
https://doi.org/10.1016/j.jneuroling.2017.08.002 Downs, Anthony
13–35. https://doi.org/10.1016/j.jneuroling.2017.08.002 Downs, Anthony
2017 doi
-
[111]
https://doi.org/10.1016/j.pragma.2017.06.020 Hoek, Jet & Scholman , Merel C
245−262. https://doi.org/10.1016/j.pragma.2017.06.020 Hoek, Jet & Scholman , Merel C. J
2017 doi
-
[194]
https://doi.org/10.1016/j.pragma.2022.03.024 Garassino, Davide & Masia, Viviana & Brocca, Nicola
9–22. https://doi.org/10.1016/j.pragma.2022.03.024 Garassino, Davide & Masia, Viviana & Brocca, Nicola
2022 doi
-
[210]
https://revistas.uam.es/chimera/article/view/17979 (last access ed on 26.05.2024) Domaneschi, Filippo & Canal, Paolo & Masia, Viviana & Lombardi Vallauri, Edoardo & Bambini, Valentina
2024
-
[1979]
Cognition 7(4)
Syntactic presupposition in sentence comprehension. Cognition 7(4). 363–383. https://doi.org/10.1016/0010-0277(79)90022-2 Politicians vs ChatGPT AI-Linguistica 25 Lombardi Vallauri, Edoardo
-
[2007]
Journal of Computational and Graphical Statistics 16 (3)
Residual-based shadings for visualizing ( conditional) independence. Journal of Computational and Graphical Statistics 16 (3). 507−525. https://doi.org/10.1198/106186007X237856
-
[2008]
Computational Linguistics 34(4)
Inter-Coder agreement for Computational Linguistics. Computational Linguistics 34(4). 555−596. http://doi.org/10.1162/coli.07-034-R2 Barattieri di San Pietro, Chiara & Frau, Federico & Mangiaterra, Veronica & Bambini, Valentina
-
[2010]
Mind & Language 25(4)
Epistemic vigilance. Mind & Language 25(4). 359–393. https://doi.org/10.1111/j.1468-0017.2010.01394.x Spooren, Wilbert & Degand, Liesbeth
2010
-
[2015]
asserted content in online processing
Presuppositions vs. asserted content in online processing. In Schwarz, Florian (ed.), Experimental Perspectives on Presupposition: Davide Garassino, Viviana Masia, Nicola Brocca & Alice Delorme Benites 26 Studies in Theoretical Psycholinguistics, 89−108. Dordrecht: Springer. h...
-
[2016]
Philologie im Netz, PhiN -Beiheft 11, 66−79
Politici in rete o nella rete dei politici? L’implicito nella comunicazione politica italiana in Politicians vs ChatGPT AI-Linguistica 23 Twitter. Philologie im Netz, PhiN -Beiheft 11, 66−79. http://web.fu- berlin.de/phin/beiheft11/b11t06.pdf (last accessed on 26.05.2024) Brow...
2024
-
[2017]
In Proceedings of the 13th Joint ISO-ACL Workshop on Interoperable Semantic Annotation (Isa -13), 1−13
Evaluating discourse annotation: some recent insights and new approaches. In Proceedings of the 13th Joint ISO-ACL Workshop on Interoperable Semantic Annotation (Isa -13), 1−13. https://aclanthology.org/W17-7401/ (last accessed on 18.01.2024) Hoek, Jet & Scholman, Merel C.J. &...
2024
-
[2019]
0.84.1) [Computer software]
Irr: Various Coefficients of Interrater Reliability and Agreement (v. 0.84.1) [Computer software]. https://CRAN.R- project.org/package=irr (last accessed on 06.06.2024) Davide Garassino, Viviana Masia, Nicola Brocca & Alice Delorme Benites 24 Garassino, Davide, & Brocca, Nicol...
2024
- [2020]
-
[2021]
Association for Computing Machinery, 610−623
On the dangers of stochastic parrots: Can Language Models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ‘21) . Association for Computing Machinery, 610−623. https://doi.org/10.1145/3442188.3445922 Brocca, Nicola & Ga...
2021
-
[2022]
Journal of Experimental Political Science 9(1)
All the news that’s fit to fabricate: AI-Generated text as a tool of media misinformation. Journal of Experimental Political Science 9(1). 104−117. https://doi.org/10.1017/XPS.2020.37 Landis, J. Richard & Koch, Gary G
2020 doi
-
[2023]
The pragmatic profile of ChatGPT: Assessing the communicative skills of a conversational agent. In Bambini, Valentina & Barattieri di San Pietro, Chiara (eds ), Multidisciplinary perspectives on ChatGPT and other Artificial Intelligence Models / Prospettive multidisciplinari s...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.