REVIEW 2 major objections 5 minor 24 references
Reducing belief in conspiracy theories as they unfold using large language models
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Short, targeted conversations with a large language model can reduce belief in conspiracy theories as they unfold, and can reduce belief in later conspiracies about subsequent events.
desk verdict LLM dialogues reduce belief in fresh conspiracies, but the 'no facts needed' framing overreaches because the model was always given a curated fact sheet. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the debunking dialogue itself: a minimum five-round conversation in which a large language model (here Gemini Pro 1.5 and 2.5) is instructed to persuade the participant to stop believing the conspiracy "via a thoughtful, evidence-based dialogue," and is grounded in a curated fact base of about 15 verified facts compiled from contemporaneous news reports within days of the event. Because the model's training data predates the event, the fact base is the model's only source of event knowledge in Experiment 1, and web search restricted to factual verification is added in Experiment 2. The persuasive strategies used by the model were then coded sentence-by-sentence, using the six strategy dimensions validated in prior work plus a new "epistemic humility" dimension that the authors added to capture appeals to caution given how little is known. This coding lets the paper trace why the intervention works when the factual record is thin.
What would settle it
Run the same design on the next major unfolding event with four arms: debunking dialogue with the curated fact base, debunking dialogue without any fact base (the model is told only to use general knowledge and ask questions), the static fact list, and an irrelevant dialogue. If the no-fact-base dialogue fails to reduce belief while the fact-based dialogue succeeds, the central claim that dialogue carries the effect would be overturned. A pre-registered replication where the strategy coding shows no increase in epistemic humility would also pressure the proposed mechanism.
Extended reading notes
Core claim
The authors set out to test whether AI-facilitated debunking, which works on well-established conspiracy theories, could also work on conspiracies that emerge in the immediate aftermath of an unfolding event, where little is known and little corrective evidence exists. In two experiments run within days of the events, US adults who expressed conspiratorial views were assigned to an evidence-based debunking dialogue with Gemini Pro 1.5 (Trump) or 2.5 (Kirk), a static information list, or an irrelevant dialogue. In both studies, the debunking dialogue significantly reduced belief in the participant's own articulated conspiracy, and reduced agreement that there was a cover-up or hidden factors, relative to both controls; effect sizes were modest (Cohen's d around .3-.4) on a 0-100 scale. The paper also reports that the debunking turned into a form of prebunking: participants debunked about the first Trump attempt were less likely to believe conspiratorial claims about a second attempt two months later, and participants debunked about the Kirk assassination showed lower endorsement of generic conspiracy beliefs one month later. The authors interpret this as showing that nascent conspiracy beliefs, though confidently held, are not fully formed and can be swayed by appeals to critical thinking and epistemic humility in addition to facts.
Load-bearing premise
The load-bearing premise is that the conversation itself, rather than the curated fact sheet hidden inside it, is what reduces belief: the dialogue condition never runs without the fact base, and the static fact list alone produced a much smaller effect.
Editorial extensions
If this is right
- A short LLM conversation (averaging about 6.9 minutes) can reduce belief in a conspiracy that is still emerging, so real-time deployment in the days after a crisis is feasible in principle.
- The treatment outperformed a static list of the same facts, suggesting that the interactive dialogue, not merely the information content, adds persuasive value.
- Debunking one unfolding conspiracy reduced belief in different conspiracies one to two months later, meaning debunking can double as prebunking.
- The model shifted rhetorical strategy when facts were scarce, using more epistemic humility, Socratic questioning, and source credibility and less rational persuasion, implying the technique can adapt to low-information environments.
- The intervention did not increase trust in the official explanation in the Trump case, where no clear official explanation existed, so it reduces conspiracy belief without necessarily rehabilitating official accounts.
Reading between the lines
- A direct test that removes the curated fact base from the dialogue would separate the conversation effect from the facts-inside-the-conversation effect; the paper's design never runs that condition, so this is the cleanest open question.
- If the strategy coding is right, then injecting epistemic humility prompts into generic chatbots could reproduce part of the effect without event-specific fact sheets, which would make the intervention cheaper to deploy on novel events.
- The same technique that reduces belief in false conspiracies could presumably amplify belief in emerging conspiracies if the model is prompted adversarially; the authors note this risk in the discussion, and quantifying it would be a natural next step.
- The fact that effects were larger for high-belief participants in the Kirk study but not the Trump study hints that the persuasion target may need to be a belief the person has already articulated; future work could test whether the pre-survey articulation itself is part of the mechanism.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports two experiments plus delayed follow-ups testing whether short conversational dialogues with a large language model (LLM) reduce belief in conspiracy theories that emerge in the days after a major event. In Experiment 1 (July 2024 Trump assassination attempt, N = 472 believers) and Experiment 2 (September 2025 Kirk assassination, N = 1035 believers), U.S. adults holding conspiratorial views were randomized to a debunking dialogue with an LLM, a static fact sheet, or an irrelevant dialogue with an LLM. Across both experiments, the debunking dialogue significantly reduced participants' confidence in their own stated conspiracy belief, perceived cover-up, and perceived hidden factors, relative to both control conditions, with small-to-moderate effect sizes. The static information list produced smaller or null effects. Follow-ups found some evidence of reduced belief in later conspiracies: Trump follow-up effects on two measures, and Kirk follow-up effects on a secondary measure and via a persistence analysis. The paper also presents an LLM-based strategy coding, validated against human raters, showing that the model used epistemic humility and Socratic questioning more when little information was available.
Significance. The work addresses an important and topical question: whether AI dialogues can reduce belief in newly emerging conspiracies, for which the usual debunking evidence is scarce. Strengths include a pre-registered Experiment 2, robustness checks with stricter believer criteria (SM4), validated LLM-based strategy coding, open data/code, and consistent main effects across two natural-event case studies. The delayed follow-up design is valuable. If the central claim is accepted, the intervention has practical potential for real-time misinformation response. However, the central theoretical claim that the LLM succeeds 'even in the absence of extensive corrective facts and evidence' is not directly tested, because the dialogue model received a curated fact base in both experiments. The mixed follow-up results for Experiment 2 also warrant more cautious interpretation. The study is a solid demonstration of a fact-grounded conversational debunking effect, but additional design changes or reframing are needed to support the 'unfolding event' claim as stated.
major comments (2)
- [SM1 and Discussion] The paper's central theoretical claim—that the LLM succeeded 'even in the absence of extensive corrective facts and evidence' (Discussion)—is not actually tested, because both experiments embedded a curated fact base into the dialogue model's system instructions (SM1). This fact base included confirmed facts, a debunked-claims inventory, and open-questions markers; in Experiment 1 it was the model's sole source of information about the event. The Debunking Dialogue condition thus always received the persuasion prompt plus this fact sheet, whereas the Irrelevant Dialogue control received neither, and the Information List control received a static version of similar facts. The design cannot separate the contribution of the model's interactive reasoning from the informational content of the fact sheet, and the much smaller effect of the Information List could reflect a lack of interactivity or engagement rather than the model's ability to reason productively with little evidence. The rhetorical-strategy analysis (Results, 'LLM rhetorical strategies') shows the model used epistemic humility and Socratic questioning more when facts were sparse, but those strategies were always deployed alongside the curated fact base. A no-fact dialogue condition (or a dialogue condition with only open-questions markers) would be needed to support the 'unfolding events' claim as stated.
- [Results, 'Effects on subsequent conspiratorial beliefs'; Figure 3C] The pre-registered primary analysis for the Kirk follow-up found no significant effect of the debunking treatment on LDS conspiracy belief (b = -0.115, 95% CI [-2.73, 2.50], p = .93). The evidence for downstream effects in Experiment 2 therefore rests on the secondary measure of popular conspiracy beliefs and on the Lin et al. (2025) persistence analysis, which estimates that 12% of the Kirk debunking effect was still observable in LDS beliefs (p = .02). The abstract's claim that the authors 'observed reduced belief in different conspiracies one to two months later' may overstate this mixed result; the text should present the null primary analysis more prominently and qualify the downstream-effect claim accordingly.
minor comments (5)
- [SM3] The word 'assanation' appears twice in the GPT-4o classification prompts; please correct to 'assassination'.
- [Methods, 'Study 1 conspiracy believers'] The phrase 'A quarter of participants (n=199) answered close to the scale midpoint' is followed by '(between 40 and 60; 50 = Uncertain)', which is clear, but the later sentence 'Examining mean self-ratings by LLM-generated groupings' reports M_Maybe = 48.5; it would be helpful to state explicitly in the main text that the 'maybe' group's mean belief was below the scale midpoint, despite being included in the primary analysis, because this bears on how 'believers' are characterized.
- [Results, 'LLM rhetorical strategies'] The strategy coding validation is reported as 91.75% overall agreement and κ = .62, with per-category metrics deferred to SM8; please add a sentence in the main text pointing to the per-category values, as κ = .62 is moderate and the pattern may vary by strategy.
- [Discussion] The statement that participants' beliefs 'were not fully formed' is presented as an explanation for the intervention's success, but no direct measure of belief crystallization is provided; please label this as a post-hoc hypothesis rather than an empirical conclusion.
- [Methods, 'Study 1 design'] The sentence 'Participants in the conspiracy beliefs and irrelevant dialogue conditions read brief instructions and then engaged in a minimum five-round exchange' could be clarified to indicate whether the five-round minimum applied to both conditions and how the dialogue was terminated; this detail is important for interpreting engagement differences.
Circularity Check
No circularity: the treatment effect is measured on new data, and the fact-sheet confound is a validity caveat rather than a circular derivation.
full rationale
This is an empirical intervention study, not a derivation, and its central claims are not constructed from their inputs. The Debunking Dialogue effect is estimated by OLS regression of post-treatment conspiracy ratings on condition, controlling for pre-treatment ratings (Results, Main effects), using data collected in newly fielded experiments. The outcome measures are participant self-reports and LLM classifications validated against human raters (SM3, SM8; 91.75% agreement, κ=.62), so measurement does not reduce to the model's own outputs. The paper does reuse the authors' prior dialogue paradigm and strategy-coding categories from Costello et al. (refs. 2 and 4), but these citations support instrument continuity, not the empirical conclusion; the effect sizes are computed from new data. The most salient concern—that both experiments embedded a curated fact base of 15 verified facts into the model instructions (SM1), undermining the 'absence of extensive corrective facts' phrasing in the Discussion—is a construct-validity or scope limitation, not a circular step: the treatment effect is still an empirical contrast, and the paper transparently reports the fact base. The static Information List condition provides a partial control for fact provision, showing a smaller effect, but the absence of a no-fact dialogue condition does not make any outcome equal to any input by construction. No fitted parameter is renamed a prediction, no uniqueness theorem is imported, and no central claim reduces to a self-citation. Therefore no significant circularity is present.
Assumptions & free parameters
assumptions (4)
- standard math Linear regression with robust standard errors assumes a linear conditional expectation and valid inference.
- domain assumption Self-reported belief scales measure actual conspiratorial beliefs.
- domain assumption Random assignment to conditions creates exchangeable groups.
- domain assumption The LLM classification of text responses into conspiracy believers is valid.
Cite this review
Pith. "Pith review of Reducing belief in conspiracy theories as they unfold using large language models." pith.science (2026). https://pith.science/paper/J7IFCK6S
@misc{pith2026260806151,
author = {Pith},
title = {Pith review of: Reducing belief in conspiracy theories as they unfold using large language models},
year = {2026},
howpublished = {\url{https://pith.science/paper/J7IFCK6S}},
note = {Machine review of arXiv:2608.06151}
}
read the original abstract
The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational dialogues with a large language model (LLM) can reduce belief in immediately unfolding conspiracies. In experiments conducted in the days following the July 2024 assassination attempt on Donald Trump and the September 2025 assassination of Charlie Kirk, U.S. adults (Experiment 1: N = 472; Experiment 2: N = 1035) holding conspiratorial views about the crisis event engaged in a multi-turn conversation with an LLM prompted to reduce their conspiracy belief. Compared to control participants who either discussed an irrelevant topic with an LLM or viewed a static fact sheet, participants in the LLM treatment showed significantly reduced conspiracy beliefs in both experiments. We also found evidence of downstream effects of the LLM treatment, observing reduced belief in different conspiracies one to two months later in the wake of subsequent crisis events. These results shed light on the psychology of emerging conspiracies and highlight the potential for scalable, cognitively-focused interventions to counteract misinformation in the immediate aftermath of high-profile societal events.
Figures
Reference graph
Works this paper leans on
-
[1]
Farrell, H., Gopnik, A., Shalizi, C. & Evans, J. Large AI models are cultural and social technologies.Science387, 1153–1156 (2025)
work page 2025
-
[2]
Costello, T. H., Pennycook, G. & Rand, D. G. Durably reducing conspiracy beliefs through dialogues with AI. Science385, eadq1814 (2024)
work page 2024
-
[3]
O’Mahony, C., Brassil, M., Murphy, G. & Linehan, C. The efficacy of interventions in reducing belief in conspiracy theories: A systematic review.PLOS ONE18, e0280902 (2023)
work page 2023
-
[4]
Costello, T. H., Pennycook, G. & Rand, D. Just the facts: How dialogues with AI reduce conspiracy beliefs. Preprint athttps://doi.org/10.31234/osf.io/h7n8u_v1(2025)
-
[5]
H., Spinoza-Martín, D., Rand, D
Boissin, E., Costello, T. H., Spinoza-Martín, D., Rand, D. G. & Pennycook, G. Dialogues with large lan- guage models reduce conspiracy beliefs even when the AI is perceived as human.PNAS Nexuspgaf325 (2025). doi:10.1093/pnasnexus/pgaf325. 14
-
[6]
Preprint at https://doi.org/10.31219/osf.io/hfbeu (2024)
Ognyanova, K.et al.Social media spread conspiracy theories after Trump assassination attempt, but believing them was linked to interpersonal discussions. Preprint at https://doi.org/10.31219/osf.io/hfbeu (2024)
-
[7]
Thompson, S. A. With Few Facts About Kirk Shooting, Wild Speculation Abounds.The New York Times(2025)
work page 2025
-
[8]
Power, J. ‘A script’: Texts of alleged Charlie Kirk killer fuel conspiracy theories.Al Jazeera https://www.alja zeera.com/news/2025/9/18/a-script-texts-of-alleged-charlie-kirk-killer-fuel-conspirac y-theories(2025)
work page 2025
Show all 24 references
-
[9]
SPLC. Antisemitic conspiracy theories claim Israel, Mossad to blame for Kirk killing.Southern Poverty Law Center https://www.splcenter.org/resources/hatewatch/charlie-kirk-antisemitic-conspirac y-theories/(2025)
2025
-
[10]
ISD. Some social media users declared Charlie Kirk’s death a ‘false flag’ operation.Institute for Strategic Dialogue https://www.isdglobal.org/media-mention/some-social-media-users-declared-charlie-kir ks-death-a-false-flag-operation/(2025)
2025
-
[11]
Lin, H.et al.Persuading voters using human–artificial intelligence dialogues.Nature648, 394–401 (2025)
2025
-
[12]
A., Krishna, A
Amazeen, M. A., Krishna, A. & Eschmann, R. Cutting the Bunk: Comparing the Solo and Aggregate Effects of Prebunking and Debunking Covid-19 Vaccine Misinformation.Sci. Commun.44, 387–417 (2022)
2022
-
[13]
Q., Hurlstone, M
Tay, L. Q., Hurlstone, M. J., Kurz, T. & Ecker, U. K. H. A comparison of prebunking and debunking interventions for implied versus explicit misinformation.Br. J. Psychol.113, 591–607 (2022)
2022
-
[14]
& van der Linden, S
Lewandowsky, S. & van der Linden, S. Countering Misinformation and Fake News Through Inoculation and Prebunking.Eur. Rev. Soc. Psychol.32, 348–384 (2021)
2021
-
[15]
& Hobbs, T
Eder, S. & Hobbs, T. D. The Quiet Unraveling of the Man Who Almost Killed Trump.The New York Times(2025)
2025
-
[16]
& Peterson, E
Mummolo, J. & Peterson, E. Demand Effects in Survey Experiments: An Empirical Assessment.Am. Polit. Sci. Rev.113, 517–529 (2019)
2019
-
[17]
& Cushman, F
Woodley, L., Roberts-Gaal, X., Calcott, R. & Cushman, F. A. No Evidence of Experimenter Demand Effects in Three Online Psychology Experiments.Open Mind10, 998–1016 (2025)
2025
- [18]
-
[19]
Miah, M. S. U.et al.A multimodal approach to cross-lingual sentiment analysis with ensemble of transformer and LLM.Sci. Rep.14, 9603 (2024)
2024
-
[20]
Rathje, S.et al.GPT is an effective tool for multilingual psychological text analysis.Proc. Natl. Acad. Sci.121, e2308950121 (2024)
2024
-
[21]
S., Pink, S
Mernyk, J. S., Pink, S. L., Druckman, J. N. & Willer, R. Correcting inaccurate metaperceptions reduces Americans’ support for partisan violence.Proc. Natl. Acad. Sci.119, e2116851119 (2022)
2022
-
[22]
Brotherton, R., French, C. C. & Pickering, A. D. Measuring Belief in Conspiracy Theories: The Generic Conspir- acist Beliefs Scale.Front. Psychol.4(2013)
2013
-
[23]
Your goal is to very effectively persuade users to stop believing in conspiracy theories and misinformation surrounding [the event] via a thoughtful, evidence-based dialogue
Stavropoulos, A., Crone, D. L. & Grossmann, I. Shadows of wisdom: Classifying meta-cognitive and morally grounded narrative content via large language models.Behav. Res. Methods56, 7632–7646 (2024). 15 Supplementary Materials SM1. Model instructions for the conspiracy beliefs ...
2024
-
[24]
American Comeback Tour
https://nypost.com/2024/07/14/us-news/grateful-defiant-trump-recounts-surreal-assassination-attempt-at-rally-im- supposed-to-be-dead/ [4] https://www.cbsnews.com/live-updates/trump-rally-shooting-investigation/ [5] https://www. nbcnews.com/nightly-news/video/new-details-on-sho...
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.