{"id":"845d01d4-9f06-4428-b1f3-9e7ce47c507c","arxiv_id":"2411.18403","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ChatGPT-generated political speeches contain more potentially manipulative presuppositions per 1,000 words than the human politicians' speeches, with a different distribution of triggers and discourse functions.","lead":"This paper compares how often French and Italian politicians use presuppositions (implied background claims) in their speeches versus texts generated by ChatGPT impersonating those politicians. It finds ChatGPT texts contain more potentially manipulative presuppositions, mostly vague slogans, while real politicians use a wider mix of criticism and other functions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The frequency comparison rests on a non-blind PMP annotation without a reliability score for PMP identification; a blinded reannotation is needed before the central claim can be treated as robust.","rationale":"The reader's weakest-assumption diagnosis correctly identifies the load-bearing point: the PMP frequency comparison rests on a non-blind annotation step for which no interrater reliability is reported. My independent reading confirms this is the most serious threat to the central claim, because the label PMP is inherently evaluative and the annotators knew the source of each text. The paper is otherwise transparent: the corpus is small, the authors acknowledge in Section 7 that the limited size 'could lead to biased results,' and the statistical tests are used cautiously with effect sizes reported. However, no amount of downstream statistical rigor can repair an unreliable first coding step. A blinded reannotation with a proper agreement metric is the minimal check that would settle whether the 'more PMPs in ChatGPT' finding is real or an artifact of expectation. If the finding survives, the paper's contribution stands as an exploratory but valuable study; if it does not, the central claim is unsupported. The reader's CONDITIONAL verdict is appropriate; my stress-test does not move it, so the verdict is UNCHANGED.","tokens_in":17079,"tokens_out":2070,"duration_ms":22920,"concrete_test":"Recruit two independent annotators who are not authors and blind them to the politician/ChatGPT origin of every text (e.g., by removing prompts, headers, and any metadata, and by randomizing presentation order). Have them annotate all 16 texts for PMPs using the same codebook from the OSF repository. Compute Cohen's kappa and Gwet's AC1 for PMP identification. Then compare the normalized PMP counts per 1,000 words between the two groups using the original texts' word counts. If the blinded counts no longer show a significantly higher PMP rate for ChatGPT texts, or if the effect size drops below the reported Cramér's V = .16, the central frequency claim is not robust; if the blinded replication reproduces the direction and approximate magnitude, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that ChatGPT texts contain more potentially manipulative presuppositions (PMPs) than politicians' texts—depends entirely on the first annotation step, PMP identification, described in Section 5.1. The authors explicitly state that 'no formal analysis based on interrater agreement indices was carried out at this stage,' and that the two annotators (also the authors) first identified PMPs independently, then resolved disagreements through consensus negotiation, discarding doubtful items. Because both annotators knew which texts were produced by politicians and which by ChatGPT, their judgments of what counts as a 'questionable' presupposition could be systematically shifted by expectation. This matters especially because the label PMP is not an objectively detectable property; it requires evaluating whether presupposed content is 'non-bona fide true' or 'tendentious,' a judgment sensitive to context and framing. The reliability scores reported in Tables 2 and 3 cover only trigger and discourse-function categories after the PMP set was already fixed, so they do not validate the frequency comparison. The authors' own Section 7 limitation that the small corpus 'could lead to biased results' compounds this: with only 8 human and 8 ChatGPT texts, a modest shift in PMP identification could alter or erase the reported χ2 difference. This is not a claim of bad faith; it is a structural validity threat. The strongest quantitative claim is therefore conditional on a blind, independently measured annotation step that the current design lacks.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares the use of 'potentially manipulative presuppositions' (PMPs) in authentic French and Italian political speeches with texts generated by ChatGPT 3.5 that mimic those politicians. Based on a manually annotated corpus of eight human and eight chatbot texts, the authors report (Q1) that ChatGPT texts contain more PMPs per 1,000 words, (Q2) that change-of-state verbs are more frequent triggers in ChatGPT texts while definite descriptions dominate in politicians' texts, and (Q3) that stance-taking is proportionally more common in ChatGPT texts while criticism dominates in politicians' texts. The statistical analysis includes chi-square tests, Fisher's exact tests, and a loglinear model, with effect sizes. The authors frame the work as an exploratory contribution to the pragmatics of LLM-generated political discourse and openly list limitations concerning corpus size and stochastic generation.","tokens_in":17270,"tokens_out":2400,"duration_ms":23866,"significance":"If the frequency and distribution differences are robust, the paper provides a concrete, falsifiable empirical observation about a widely feared property of LLM-generated political speech: the heavier reliance on implicit, potentially manipulative content, especially via repetitive change-of-state slogans. The authors use an established theoretical framework for presupposition and manipulative discourse, and they make their corpus and R code available on OSF, which is a strength for reproducibility. The statistical toolkit is appropriate for categorical corpus data, and the transparency about limitations in Section 7 is commendable. However, the central claim's validity hinges entirely on the reliability of the first annotation step, PMP identification, for which no interrater measure is reported.","major_comments":[{"comment":"The central frequency claim (Q1, abstract, Section 6) rests on the identification of PMPs, a context-sensitive judgment of whether presupposed content is 'non-bona fide true' or tendentious. The manuscript states that the two annotators (also the first and second authors) identified PMPs independently and then resolved disagreements through consensus negotiation, discarding doubtful items, and that 'no formal analysis based on interrater agreement indices was carried out at this stage.' The agreement indices reported in Tables 2 and 3 cover only Presupposition Trigger and Discourse Function after the PMP set was already fixed, so they do not validate the frequency comparison. Because the annotators knew which texts were produced by politicians and which by ChatGPT, expectation bias could systematically shift PMP counts, and the reported chi-square difference (χ2 = 10.28, df = 3, p < .05; Cramér's V = .16) is small enough that a modest annotation shift could alter or erase it. A blinded reannotation of PMP identification by independent annotators, with a PMP/non-PMP agreement measure (e.g., Cohen's kappa on a sample), is needed before the frequency claim can be treated as robust.","section":"Section 5.1"},{"comment":"The normalized frequency comparison in Figure 2 is based on per-1,000-word rates, but the ChatGPT texts are roughly 35–40% shorter than the politicians' speeches (Table 1). More importantly, the chi-square test in Section 5.2 treats each PMP occurrence as an independent observation, ignoring text-level clustering. With only eight texts per group, the effective sample size for the frequency comparison is very small, and the p-value would not account for potential between-text variability. Please report a permutation or mixed-effects test that treats text as a random effect, or otherwise justify the independence assumption for PMP occurrences.","section":"Section 5.2 and Table 1"},{"comment":"The comparison between human and ChatGPT texts is partly confounded by the prompt design: the prompt in (4) includes an excerpt from the real speech, a requested length of about 5,500 characters (which ChatGPT did not reach), and style guidelines. The authors discuss the Macron self-praise exception as potentially due to the prompt excerpt, but the more general risk is that the discourse-function differences (Figure 4) reflect differences in text length, genre, or the specific seed excerpt rather than a general property of ChatGPT. Please explain how these confounds were controlled, or temper the general claim that ChatGPT texts are intrinsically more stance-taking and less critical.","section":"Section 4 and Section 6 (Macron exception)"}],"minor_comments":[{"comment":"In the sentence reporting the Fisher's exact test results for Macron, 'Fisher's Exact Test = .93' presumably reports the p-value, not the test statistic; please clarify the notation for consistency.","section":"Section 5.2"},{"comment":"There are minor reference inconsistencies: 'Altmann' in the epigraph should be 'Altman'; the in-text citation 'Cai, Duang and Haslett 2023' appears to refer to 'Cai, Duan, and Haslett' in the reference list; and 'Lowen and Plonsky 2016' in Section 5.1 is spelled 'Loewen and Plonsky 2016' in the reference list.","section":"References"},{"comment":"The paper states that the prompt was written in English and the output requested in French or Italian, but example (4) shows only the English prompt; providing the full prompt for both languages in an appendix or OSF would aid replicability.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a well-written exploratory study with a clear research question and appropriate statistical handling of the categorical data it analyzes. The main issue is the absence of a reliability check for the PMP identification step, which is the foundation of the frequency claim. This is fixable within the scope of the paper by adding a blinded reannotation or at least a sensitivity analysis. The reader's conditional verdict aligns with my own assessment; I would not reject the paper, but the central claim should not be published without addressing this validity threat."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The direct paired comparison is new: presuppositions in real French and Italian political speeches versus ChatGPT 3.5 impersonations of the same politicians. Prior work looked at presuppositions in human political discourse and at ChatGPT's pragmatic abilities separately, not this specific contrast. The authors use an established framework, report effect sizes, do a loglinear analysis, and put data and code on OSF. They also state limitations plainly. That is real credit.\n\nThe soft spot is load-bearing, and the reader's stress-test note gets it right. In Section 5.1, the two annotators (also the first and second authors) first identify potentially manipulative presuppositions independently, then merge lists through consensus negotiation, discarding doubtful items. They explicitly say no interrater agreement index was computed for this step. The kappa and AC1 values in Tables 2 and 3 cover only trigger and discourse-function categories after the PMP set was already fixed. So those numbers validate the classification of items already chosen, not the choice of which items count as PMPs in the first place. Since the PMP label requires judging whether presupposed content is 'non-bona fide true' or tendentious, and the annotators knew which texts came from ChatGPT and which from politicians, expectation could shift the identification. With eight texts per group, a small shift could change the chi-square result. The authors acknowledge the small corpus, but the non-blind step is the more serious threat.\n\nThat said, the direction of the effect is consistent across several tests, and the descriptive pattern is plausible: ChatGPT texts lean on change-of-state verbs and stance-taking, while politicians use more definite descriptions and criticism. The Macron exception is handled honestly.\n\nThis is an exploratory corpus study for people working on pragmatics and AI-generated text. It deserves a serious referee, not a desk reject. A referee should ask for blind reannotation, an agreement measure on the PMP identification itself, and ideally a larger corpus before the quantitative claims are treated as robust.","headline":"A genuinely new comparison of presupposition use in human vs ChatGPT political speech, but the central frequency claim depends on a non-blind annotation step that needs a reliability check.","tokens_in":17842,"tokens_out":2347,"would_cite":true,"duration_ms":22363,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Compared with the politicians' speeches they imitate, ChatGPT 3.5 texts carry more potentially manipulative presuppositions, concentrated in change-of-state slogans like 'we must build our future'.","keywords":["presupposition","potentially manipulative presupposition","ChatGPT","political communication","implicit communication","French","Italian","discourse functions"],"falsifier":"Annotate the same sixteen texts again with coders blind to whether each text is a politician's speech or a ChatGPT output, and compute PMP rates; if the ChatGPT advantage disappears or reverses under blind coding, the central claim is not supported.","tokens_in":16834,"feed_emoji":"🤖","tokens_out":4492,"duration_ms":41045,"temperature":0.7,"pith_summary":"This paper compares rally speeches by four French and Italian politicians with ChatGPT 3.5 texts that imitate those same politicians, focusing on presuppositions as a vehicle of implicit, possibly manipulative meaning. Its central claim is that, on average, the ChatGPT-generated texts contain more potentially manipulative presuppositions per 1,000 words than the politicians' own speeches, and that the AI texts use them differently: change-of-state verbs dominate, and stance-taking replaces the more varied functions, including criticism, found in human speeches. If true, the result gives a concrete, measurable linguistic signature of how LLM-generated political communication can manipulate by taking contestable positions for granted.","feed_headline":"ChatGPT leans on hidden claims more than politicians do","feed_subtitle":"A French-Italian corpus finds AI speeches dense with change-of-state slogans such as 'we must build our future'.","key_machinery":"The unit that carries the argument is the potentially manipulative presupposition (PMP): a presupposition triggered by a dedicated linguistic form, such as a definite description or a change-of-state verb, whose presupposed content is not bona fide true but tendentious, subjective, or false. The authors identify PMPs in the corpus, annotate their presupposition triggers and their discourse functions (criticism, self-praise, praise-of-others, stance-taking), and then compare the frequency, form, and function of PMPs between politicians' speeches and ChatGPT's imitations; the statistical comparisons rest on the resulting annotation.","core_discovery":"The paper's central claim is that ChatGPT-generated political texts are not just superficial imitations of politicians' speeches: they differ systematically in their use of implicit content. Across the corpus, normalized frequencies of potentially manipulative presuppositions are higher in the ChatGPT texts than in the politicians' speeches, and the difference is statistically significant though small in effect size. The trigger types also differ: definite descriptions are the most common PMP trigger in politicians' speeches, whereas change-of-state verbs are the most common in ChatGPT texts. Discourse functions diverge too, with stance-taking massively overrepresented in the AI texts and criticism positively associated with the politicians' texts; Macron's data are the one case where the distribution of functions does not differ. The authors interpret the ChatGPT pattern as repetition and vagueness arising from next-word prediction, which selects high-frequency slogan-like constructions such as 'we must build our future'.","pith_inferences":["One extension this comparison does not run is a perceptual test: the paper shows the AI texts contain more PMPs, but not that readers are more persuaded or less likely to notice them; connecting the frequency difference to actual persuasion would complete the argument.","Because change-of-state triggers are lexically identifiable, the finding suggests a cheap automatic screen for manipulative presuppositions in AI-generated texts, even though the PMP judgment itself requires human annotation.","The paper notes that most of ChatGPT's training data is in English; a natural test is whether the same PMP gap appears in English-language political texts or in languages with even less training data, which could reveal whether the effect grows with linguistic distance from pretraining data."],"forward_implications":["If the central claim is correct, AI-generated political speech is not merely vaguer than human speech; it packages contestable positions as taken-for-granted facts more densely, which makes them harder to challenge.","The predominance of change-of-state verbs in ChatGPT texts means the AI's manipulative profile is especially concentrated in slogans like 'we must build our future', where the presupposed need or goal is smuggled in as shared.","PMP frequency offers a measurable proxy for the manipulative potential of AI-written political texts, and the trigger-type distribution gives an observable signature that could be tracked automatically.","The Macron exception indicates that the content of the prompt can shape the discourse functions of PMPs in the output, so the manipulative profile of generated texts is partly under the control of whoever writes the prompt.","Because the same slogans recur across politicians, languages, and political orientations, the pattern points to a model-level tendency rather than to faithful imitation of any individual speaker's style."],"supporting_citations":[{"why":"Defines the notion of non-bona fide true, potentially manipulative presupposition that the study counts as a PMP.","marker":"Lombardi Vallauri (2019)"},{"why":"Supplies the four discourse functions and the prior corpus method for studying presuppositions in political communication that the annotation codebook builds on.","marker":"Garassino, Brocca and Masia (2022)"},{"why":"Proposes the working definitions of criticism, self-praise, praise-of-others, and stance-taking used for the function annotation.","marker":"Garassino, Masia and Brocca (2019)"},{"why":"Grounds the theoretical notion of presupposition as common-ground, taken-for-granted information.","marker":"Stalnaker (2002)"},{"why":"Provides the classic list of presupposition triggers, including factives and change-of-state verbs, used to identify PMPs.","marker":"Kiparsky and Kiparsky (1971)"},{"why":"Supplies experimental evidence that presuppositions are processed shallowly, motivating why PMPs are a plausible mechanism of manipulation.","marker":"Lombardi Vallauri (2021)"}],"fun_headline_variants":["ChatGPT packs more hidden assumptions than politicians do","AI political texts: more presuppositions, fewer criticisms","Study: ChatGPT leans on unspoken claims more than politicians","French-Italian corpus: AI uses more implicit assertions than humans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire frequency comparison rests on the two annotators' judgment of which presuppositions count as 'questionable', made without blind conditions and without a formal agreement measure at that identification step.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT packs more hidden assumptions than politicians do","AI political texts: more presuppositions, fewer criticisms","Study: ChatGPT leans on unspoken claims more than politicians","French-Italian corpus: AI uses more implicit assertions than humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1385,"prompt_tokens":792,"completion_tokens":593,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":408,"completion_tokens_details":{"reasoning_tokens":527}},"tokens_in":408,"tokens_out":593,"duration_ms":6630,"temperature":1.0,"reasoning_tokens":527,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:13:10.394594+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Annotate the same sixteen texts again with coders blind to whether each text is a politician's speech or a ChatGPT output, and compute PMP rates; if the ChatGPT advantage disappears or reverses under blind coding, the central claim is not supported.","supporting_citations":[],"review_version":1}