Pith. sign in

REVIEW 4 major objections 6 minor 27 references

Effect of Static vs. Conversational AI-Generated Messages on Colorectal Cancer Screening Intent: a Randomized Controlled Trial

T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A single tailored AI message matches a longer chatbot for colorectal cancer screening intent.

desk verdict A well-run pre-registered RCT showing concise AI messages match a longer MI chatbot for CRC screening intent, but the population claim rests on an uncited intention-to-behavior conversion assumption. read the letter →

arxiv 2507.08211 v1 pith:XG2HLDEW submitted 2025-07-10 cs.CY

classification cs.CY
keywords colorectalcancerscreeningAI-generatedmessagesmotivationalinterviewingchatbotlargelanguagemodelpersuasionrandomizedcontrolledtrialintentionpersonalizedhealthmessagingmodalitydifferences
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

In a pre-registered randomized trial of 915 US adults aged 45–75 who had never been screened for colorectal cancer, the authors set out to test how much AI interaction is needed to shift screening intentions. Participants were randomized to a no-message control, expert-written patient education, a single AI-crafted pamphlet message tailored to age, gender, education, and other demographics, or an AI chatbot explicitly prompted to use motivational interviewing. Both AI arms raised stool-test screening intent by roughly 13 points on a 0–100 scale, significantly more than the 7.5-point gain from expert materials; for colonoscopy intent, the AI arms beat the no-message control but not the expert materials. The central finding is that the chatbot did not outperform the single AI message on either outcome, even though participants spent about 3.5 minutes longer conversing with it. The authors conclude that concise, demographically tailored AI messaging may offer a more scalable clinical route than complex conversational agents, particularly for less familiar screening modalities such as stool tests.

What carries the argument

The experiment's load-bearing mechanism is the controlled comparison of two AI formats built from the same model and the same participant inputs: a single static message versus a multi-turn chatbot conversation, both generated by the same large language model using identical demographic tailoring (age, gender, education, political leaning, urbanicity, self-reported health, and time since last primary care visit). A three-minute minimum engagement floor and matched word count isolate the incremental effect of conversational back-and-forth. A supporting analysis uses a published computational framework that tags motivational-interviewing behaviors in the chatbot dialogue, showing the chatbot relied mainly on concrete elaboration and teaching (roughly 80% and 65% of turns) rather than empathy or values exploration, which the authors suggest may explain the absence of a conversational advantage.

What would settle it

A field experiment that randomizes unscreened patients to receive the single AI message versus expert materials and then measures confirmed screening completion (for example, returned fecal immunochemical tests or scheduled colonoscopies) within 12 months; if the AI message does not increase completed screenings despite reproducing the intent gains, the claim that these messages are a scalable path to behavior change would be falsified. A second check: a replication of the four-arm design using a chatbot that demonstrably performs more empathic motivational interviewing (as measured by a validated coding framework) that outperforms the single message would falsify the claim that conversational depth provides no added benefit.

Watch

Extended reading notes

Core claim

The paper's central claim is that, when personalization and minimum exposure time are held constant, conversational depth adds no measurable persuasive benefit over a static AI message for colorectal cancer screening intent. A single large-language-model-generated message and a chatbot guided by a ten-step motivational-interviewing roadmap produced statistically indistinguishable gains in 12-month intent for both stool testing (Cohen's d ≈ 0.60–0.64 vs control) and colonoscopy (d ≈ 0.27–0.34). The AI formats outperformed the expert-written patient education page only for stool-test intent and not for colonoscopy. The authors interpret this as evidence that concise, demographically tailored AI messages can capture most of the persuasive benefit of longer conversations, and that AI persuasiveness is stronger for less established, lower-burden screening options.

Load-bearing premise

The central practical claim rests on the assumption that self-reported 12-month screening intention tracks actual screening completion, with the paper's population projection further assuming that one in three adults who cross the 50-point threshold will complete screening.

Editorial extensions

If this is right

  • Health systems could deploy short, pre-generated personalized AI messages through patient portals, text reminders, or mailed screening kits, capturing most of the persuasive benefit without building or running interactive chatbots.
  • The gains concentrate in stool-test intention, so AI-tailored outreach may be particularly effective when promoting newer, less invasive, or lesser-known screening options.
  • Among participants with low baseline intent (≤50 on the 0–100 scale), the single AI message raised the odds of crossing the 50-point threshold roughly 4.4-fold for colonoscopy and 5.3-fold for stool testing versus control, implying a potentially substantial population effect if the intent-to-behavior conversion holds.
  • The chatbot arm's extra engagement time did not translate into higher intent, suggesting that simply adding interaction time to AI-generated health outreach does not improve persuasion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The equivalence result hints at a minimum effective dose of AI persuasion; a follow-up that systematically shortens the single message or varies which demographic attributes are used for tailoring could locate the point of diminishing returns.
  • Because the paper measures intention rather than completed screening, the practical payoff depends on an unvalidated assumption that one in three adults who cross the 50-point threshold will complete screening; a field experiment tracking actual stool-test returns or colonoscopy claims would test whether these intent gains translate into completed screenings.
  • The stool-test versus colonoscopy asymmetry suggests a generalizable pattern: AI-tailored messaging may be more persuasive for adopting new behaviors than for changing entrenched preferences, and testing this on other preventive behaviors would be a natural extension.
  • Because the chatbot's dialogue was largely informational rather than empathic, the null result may be specific to this prompt design; a chatbot that more faithfully executes reflective listening and values exploration could still outperform a static message.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This manuscript reports a pre-registered four-arm randomized controlled trial of 915 unscreened U.S. adults aged 45–75, randomized to a no-message control (a neutral story), an expert-written JAMA patient page, a single GPT-4.1-generated message tailored to demographics, or a GPT-4.1 chatbot instructed to use motivational interviewing. The primary outcomes are self-reported 12-month intentions to complete colonoscopy and stool-based colorectal cancer screening, measured immediately after exposure. The authors find that both AI arms significantly increase stool-test intent by roughly 13 points on a 0–100 scale versus control, and significantly outperform the expert material on stool-test intent but not on colonoscopy intent. They also find that the chatbot does not significantly outperform the single AI message on either outcome, despite participants spending about 3.5 minutes more with the chatbot. The authors conclude that concise, demographically tailored AI messages may offer a more scalable path to CRC screening than longer conversational agents and generic expert-written materials.

Significance. If the intent effects replicate and if they translate into completed screenings, the result would be practically useful for health-system outreach. The methodological care is a genuine strength: the study is pre-registered, uses baseline-adjusted OLS with HC2 robust standard errors, includes attention checks and reCAPTCHA bot filtering, has clinician review of generated messages, and discloses full prompts in the appendix. The BOLT-based process analysis of chatbot MI behaviors and the GPT-4.1-assisted persuasion-strategy content analysis are useful descriptive contributions. However, the central practical claim relies on an uncited and unsupported assumption linking intention thresholds to completed screenings, and the headline chatbot-versus-message equivalence is a null result with wide confidence intervals that is presented as evidence of no difference. The manuscript is publishable in principle, but the Discussion needs substantial tempering and, where possible, additional analysis.

major comments (4)
  1. [Discussion] The population projection in the Discussion ('Among participants with low baseline intent... over 1,000 additional completed screenings per 10,000 unscreened, low-intent adults') rests on two unverified quantities: a baseline probability of 20% for crossing the 50-point threshold and an assumption that 'one in three individuals who cross the threshold ultimately follow through' with screening. The one-in-three conversion rate is not derived from reference 21 (Power et al.) or any other cited source, and no sensitivity analysis is provided. Because the outcome is self-reported intention measured immediately after a single exposure, even modest demand or social-desirability effects would materially change the projected number. I request that the authors either remove this population projection or replace it with a sensitivity analysis over a realistic range of conversion rates and with a citation-based justification for any central value.
  2. [Results / Appendix Table S3] The repeated framing that the chatbot 'did not outperform' or 'had no added benefit' over the single AI message is used as evidence of equivalence, but the confidence intervals for these contrasts are wide. From Table S3, the chatbot-versus-single-message difference in Cohen's d is 0.04 (SE 0.11) for stool-test intent and -0.08 (SE 0.11) for colonoscopy intent, giving approximate 95% intervals of [-0.18, 0.26] and [-0.30, 0.13]. These intervals are not tight enough to establish that the two formats are equivalent, and no equivalence margin was pre-registered. Please either rephrase these conclusions as a null result with explicitly reported confidence intervals, or conduct an equivalence test with a pre-specified margin.
  3. [Introduction] The Introduction states that the design 'isolates the incremental persuasive value of conversational back-and-forth' by 'equating content length, tailoring inputs, and minimum exposure time.' This is contradicted by the reported word counts: Table S1 shows the chatbot produced an average of 1,226 words (SD 782) per conversation, whereas the single AI message was fixed at 642 words. The comparison is therefore format-plus-content, not format alone. Please correct the description of the design and discuss the implication that the null difference was observed despite the chatbot delivering substantially more content.
  4. [Abstract / Discussion] The claim that 'LLMs appear more persuasive for lesser-known and less-invasive screening approaches like stool testing, but may be less effective for entrenched preferences like colonoscopy' is inferred from separate tests of AI-versus-expert contrasts on two different outcomes. The manuscript does not report a formal arm-by-outcome interaction test, even though both outcomes are measured in the same participants. A repeated-measures or multivariate model (or at least an explicit test of whether the stool-versus-colonoscopy treatment-effect difference is itself significant) is needed to support this differential-modality conclusion. Please add such a test or substantially soften the claim.
minor comments (6)
  1. [Figure 1 caption] The caption begins 'Change change in participant's 12-month CRC screening intention'; the duplicated 'change' should be removed.
  2. [Materials and Methods] The sentence 'The two primary outcome were' should read 'The two primary outcomes were.'
  3. [Materials and Methods] The text says full prompt texts and example outputs are available in Appendix Tables S1-S2, but the appendix lists prompts in Table S10 and outputs in Table S11; the cross-reference should be corrected.
  4. [Appendix Tables S5-S6] The header states that p-values are Benjamini-Hochberg corrected, but the table notes still report conventional significance stars (*p<0.1, **p<0.05, ***p<0.01). Please clarify whether the displayed stars already reflect the BH correction or are raw p-values.
  5. [Methods / Intervention] The 'no message control' condition actually presents a neutral fictional story rather than no message; the label is understandable but could be more accurately stated as a 'neutral-story control' to avoid confusion.
  6. [Appendix Table S3] The sign convention in Table S3 is potentially confusing: the text reports a positive d=0.64 for chatbot versus no-message on stool-test intent, while the table lists -0.64 for the contrast 'No Message - Chatbot.' A note explaining the signed direction would help readers.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the central comparisons are measured human-intention outcomes, and no prediction reduces to a fitted parameter or self-citation.

full rationale

The paper's central claim is an empirical comparison of four randomized arms on self-reported 12-month screening intentions measured on a 0-100 scale. The treatment effects (e.g., stool-test intent gains of 12.87-13.75 points) are estimated from participant responses, not computed from the AI model's own outputs, so there is no self-definitional or fitted-input circularity. The only internal-model step is the exploratory content analysis in which GPT-4.1 labels persuasive strategies in GPT-4.1-generated messages (Appendix Tables S9-S10); this descriptive annotation is not part of the treatment-effect estimation and does not underwrite the main conclusions. The cited preprint by overlapping authors (ref 7) is used only as interpretive context for modality-specific effects, not as evidence for the central chatbot-vs-message comparison, so it is not load-bearing. The Discussion's population projection does rely on an explicit but uncited one-in-three intention-to-behavior conversion assumption, but that is an external validity assumption, not a derivation from the paper's own fitted values; it does not make the measured intent outcomes circular. Overall, the derivation chain from random assignment to measured intent is self-contained.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No fitted constants or invented entities; the ledger records domain assumptions about self-report, control neutrality, and LLM output safety.

assumptions (4)
  • domain assumption Self-reported 12-month intent correlates with actual CRC screening behavior.
    Central outcome is intent; paper cites Power et al. 2008 but does not measure behavior; the claim's practical value depends on this link.
  • domain assumption Eligibility self-report of never having been screened is accurate.
    No verification against medical records; misreporting would dilute or bias arm comparisons.
  • domain assumption The no-message control story does not itself influence screening intent.
    A neutral 642-word fictional story is used as the control anchor for all treatment contrasts.
  • domain assumption The clinicians' review of 100 of the thousands of generated messages generalizes to all LLM outputs.
    Only 25 chatbot + 25 single messages per clinician were checked; the remainder are assumed factually safe and clinically appropriate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Effect of Static vs. Conversational AI-Generated Messages on Colorectal Cancer Screening Intent: a Randomized Controlled Trial." pith.science (2026). https://pith.science/paper/XG2HLDEW

@misc{pith2026250708211,
  author       = {Pith},
  title        = {Pith review of: Effect of Static vs. Conversational AI-Generated Messages on Colorectal Cancer Screening Intent: a Randomized Controlled Trial},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XG2HLDEW}},
  note         = {Machine review of arXiv:2507.08211}
}
read the original abstract

Large language model (LLM) chatbots show increasing promise in persuasive communication. Yet their real-world utility remains uncertain, particularly in clinical settings where sustained conversations are difficult to scale. In a pre-registered randomized controlled trial, we enrolled 915 U.S. adults (ages 45-75) who had never completed colorectal cancer (CRC) screening. Participants were randomized to: (1) no message control, (2) expert-written patient materials, (3) single AI-generated message, or (4) a motivational interviewing chatbot. All participants were required to remain in their assigned condition for at least three minutes. Both AI arms tailored content using participant's self-reported demographics including age and gender. Both AI interventions significantly increased stool test intentions by over 12 points (12.9-13.8/100), compared to a 7.5 gain for expert materials (p<.001 for all comparisons). While the AI arms outperformed the no message control for colonoscopy intent, neither showed improvement xover expert materials. Notably, for both outcomes, the chatbot did not outperform the single AI message in boosting intent despite participants spending ~3.5 minutes more on average engaging with it. These findings suggest concise, demographically tailored AI messages may offer a more scalable and clinically viable path to health behavior change than more complex conversational agents and generic time intensive expert-written materials. Moreover, LLMs appear more persuasive for lesser-known and less-invasive screening approaches like stool testing, but may be less effective for entrenched preferences like colonoscopy. Future work should examine which facets of personalization drive behavior change, whether integrating structural supports can translate these modest intent gains into completed screenings, and which health behaviors are most responsive to AI-supported guidance.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

27 extracted references · 23 canonical work pages

  1. [2]

    JAMA 325 , 1965–1977 (2021)

    US Preventive Services Task Force, Screening for Colorectal Cancer: US Preventive Services Task Force Recommendation Statement. JAMA 325 , 1965–1977 (2021)

  2. [3]

    MMWR Morb Mortal Wkly Rep 72 (2023)

    CDCMMWR, QuickStats: Age-Adjusted Percentage of Adults Aged 50–75 Years Who Received the Recommended Colorectal Cancer Screening, by Sex and Family Income — National Health Interview Survey, United States, 2021. MMWR Morb Mortal Wkly Rep 72 (2023)

  3. [4]

    Star, et al

    J. Star, et al. , Colorectal cancer screening test exposure patterns in US adults 45 to 49 years of age, 2019-2021. J Natl Cancer Inst 116 , 613–617 (2024)

  4. [5]

    Salvi, M

    F. Salvi, M. Horta Ribeiro, R. Gallotti, R. West, On the conversational persuasiveness of GPT-4. Nat Hum Behav 1–9 (2025). https://doi.org/10.1038/s41562-025-02194-6

  5. [6]

    Available at: https://dl.acm.org/doi/epdf/10.1145/3579592 [Accessed 10 July 2025]

    Working With AI to Persuade: Examining a Large Language Model’s Ability to Generate Pro-Vaccination Messages. Available at: https://dl.acm.org/doi/epdf/10.1145/3579592 [Accessed 10 July 2025]

  6. [7]

    N. K. R. Sehgal, et al. , Conversations with AI Chatbots Increase Short-Term Vaccine Intentions But Do Not Outperform Standard Public Health Messaging. [Preprint] (2025). Available at: http://arxiv.org/abs/2504.20519 [Accessed 10 July 2025]

  7. [8]

    Hou, et al

    Z. Hou, et al. , A vaccine chatbot intervention for parents to improve HPV vaccination uptake among middle school girls: a cluster randomized trial. Nat Med 1–8 (2025). https://doi.org/10.1038/s41591-025-03618-6

  8. [9]

    N. N. Long, et al. , Motivational Interviewing to Improve the Uptake of Colorectal Cancer Screening: A Systematic Review and Meta-Analysis. Front. Med. 9 (2022)

Show all 27 references
  1. [10]

    Oster, et al

    C. Oster, et al. , Can Motivational Interviewing Be Delivered Using Artificial Intelligence Chatbots? Evaluating the Capability of Gpt-4o. [Preprint] (2025). Available at: https://papers.ssrn.com/abstract=5281021 [Accessed 10 July 2025]

  2. [11]

    Y. Y. Chiu, A. Sharma, I. W. Lin, T. Althoff, A Computational Framework for Behavioral Assessment of LLM Therapists. [Preprint] (2024). Available at: http://arxiv.org/abs/2401.00820 [Accessed 10 July 2025]

  3. [12]

    J. A. Shapiro, et al. , Screening for Colorectal Cancer in the United States: Correlates and Time Trends by Type of Test. Cancer Epidemiol Biomarkers Prev 30 , 1554–1565 (2021)

  4. [13]

    T. H. Costello, G. Pennycook, D. G. Rand, Durably reducing conspiracy beliefs through dialogues with AI. Science 385 , eadq1814 (2024)

  5. [14]

    Hackenburg, H

    K. Hackenburg, H. Margetts, Evaluating the persuasive influence of political microtargeting with large language models. Proceedings of the National Academy of Sciences 121 , e2403116121 (2024)

  6. [15]

    Hackenburg, et al

    K. Hackenburg, et al. , Scaling language model size yields diminishing returns for single-message political persuasion. Proceedings of the National Academy of Sciences 122 , e2413443122 (2025)

  7. [16]

    J. A. Goldstein, J. Chao, S. Grossman, A. Stamos, M. Tomz, How persuasive is AI-generated propaganda? PNAS Nexus 3 , pgae034 (2024)

  8. [17]

    H. Bai, J. G. Voelkel, S. Muldowney, J. C. Eichstaedt, R. Willer, LLM-generated messages can persuade humans on policy issues. Nat Commun 16 , 6037 (2025)

  9. [18]

    Huang, et al

    Y. Huang, et al. , AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing. [Preprint] (2025). Available at: http://arxiv.org/abs/2505.17380 [Accessed 10 July 2025]

  10. [19]

    Weiss, Health Literacy: A Manual for Clinicians (American Medical Association)

    B. Weiss, Health Literacy: A Manual for Clinicians (American Medical Association)

  11. [20]

    Zhu, et al

    X. Zhu, et al. , National Survey of Patient Factors Associated with Colorectal Cancer Screening Preferences. Cancer Prevention Research 14 , 603–614 (2021)

  12. [21]

    Power, et al

    E. Power, et al. , Understanding Intentions and Action in Colorectal Cancer Screening. Annals of Behavioral Medicine 35 , 285–294 (2008)

  13. [22]

    Stagnaro, et al

    M. Stagnaro, et al. , Representativeness versus Response Quality: Assessing Nine Opt-In Online Survey Samples. [Preprint] (2024). Available at: https://osf.io/h9j2d_v1 [Accessed 10 July 2025]

  14. [23]

    R. M. Jones, K. J. Devers, A. J. Kuzel, S. H. Woolf, Patient-Reported Barriers to Colorectal Cancer Screening. Am J Prev Med 38 , 508–516 (2010)

  15. [24]

    C. N. Klabunde, et al. , Barriers to Colorectal Cancer Screening: A Comparison of Reports From Primary Care Physicians and Average-Risk Adults. Medical Care 43 , 939 (2005)

  16. [25]

    Jin, Screening for Colorectal Cancer

    J. Jin, Screening for Colorectal Cancer. JAMA 325 , 2026 (2021)

  17. [26]

    Wahab, U

    S. Wahab, U. Menon, L. Szalacha, Motivational Interviewing and Colorectal Cancer Screening. Patient Educ Couns 72 , 210–217 (2008)

  18. [27]

    Opel, Identifying, understanding and talking with vaccine-hesitant parents

    D. Opel, Identifying, understanding and talking with vaccine-hesitant parents. University of Washington School of Medicine (2014). Available at: https://www.fondation-merieux.org/wp-content/uploads/2017/03/from-package-to-protection-how-do-we-close-global-coverage-gaps-to-opti...

  19. [28]

    Hello [Patient Name], it’s good to see you today. How are you feeling?

    reCAPTCHA v3. Google for Developers . Available at: https://developers.google.com/recaptcha/docs/v3 [Accessed 10 July 2025]. FIGURES AND TABLES A) B) C) Figure 1. Change change in participant’s 12 ‑ month CRC screening intention (N = 915) Panel A displays raw mean intention (±...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.