Pith. sign in

REVIEW 3 major objections 5 minor 2 references

Prosocial Persuasion at Scale? Large Language Models Outperform Humans in Donation Appeals Across Levels of Personalization

T0 review · 3 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Two preregistered experiments find that donation appeals written by a large language model raise more money than appeals written by humans — even when human writers were paid and targeting their own demographic group.

desk verdict First direct human-vs-LLM benchmark on costly charitable giving, with two preregistered studies and a behavioral outcome; solid and worth citing despite small effects and a missing suspicion check. read the letter →

arxiv 2604.03202 v2 pith:NJBT33RY submitted 2026-04-03 cs.CY

classification cs.CY
keywords prosocialbehaviorlargelanguagemodelspersuasioncharitablegivingpersonalizationdonationappealshuman–AIcomparisonrandomizedexperiments
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

LLM-generated donation appeals outperform human-authored ones at eliciting real charitable giving, not just stated attitudes. In two preregistered studies, participants read six social-media-style posts and split a real $0.10 bonus among charities; machine-written posts received a significantly larger share of the bonus in both studies (14.2% vs 12.7% and 13.3% vs 10.5%). LLM posts also earned more likes and higher persuasiveness ratings. The personalization results were mixed: accurate personalization helped in one study, while falsely personalized posts consistently underperformed — a backfire effect strongest for LLM content. The paper positions LLMs as a viable, scalable tool for prosocial persuasion.

What carries the argument

The core instrument is a 2×3 within-subjects factorial design: six donation appeals per participant, crossing content source (human vs LLM) with personalization level (generic, personalized, falsely personalized), presented as realistic social media posts with a real monetary donation allocation as the primary outcome. The design controls for charity identity by fully crossing charities with conditions, uses random intercepts for participants and charities, and compares human writers of different skill levels against a frontier language model under identical instructions and character limits. The crucial move is measuring behavior — how much of a real $0.10 bonus participants gave — rather t

What would settle it

A replication that discloses AI authorship to half the participants, or asks them to guess which posts were LLM-written, would settle the issue: if the LLM advantage shrinks or reverses among participants who know or correctly guess, the claim that LLM content is inherently more persuasive is undercut.

Watch

Extended reading notes

Core claim

The central claim is that LLM-generated donation appeals are more effective than human-authored appeals at motivating costly prosocial behavior. The evidence comes from two preregistered within-subjects experiments in which each participant saw six appeals crossing human vs machine authorship with generic, personalized, and falsely personalized targeting, and then allocated bonus money to charities. The machine advantage appeared on all three outcomes — donation amount, engagement, and perceived persuasiveness — and held in Study 2 even though the human writers were US laypeople writing for their own demographic group with financial incentives to produce effective posts. The authors interpre

Load-bearing premise

The central comparison assumes participants remained unaware that some posts were machine-written, so the donation gap reflects content quality rather than reactions to perceived AI authorship.

Editorial extensions

If this is right

  • Fundraisers could produce many effective, tailored donation appeals at almost no marginal cost if the content advantage persists outside the lab.
  • Accurate personalization adds modest value, but bad targeting costs donations; generic messages are safer than falsely personalized ones.
  • The LLM persuasion advantage extends to costly prosocial behavior, not just opinions or attitudes, making the result relevant beyond misinformation debates.
  • The advantage is measured under undisclosed authorship; disclosure of AI authorship is a boundary condition flagged in the paper.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the advantage is driven by fluency and rhetorical polish, then future models trained on even more persuasive text may widen the gap, and human writers may need distinct authenticity cues to compete.
  • A natural field test would send LLM- and human-written fundraising messages to real donors with donation amounts tracked; the $0.10 lab allocation is a strong proxy but not real-world giving.
  • The false-personalization backfire suggests a testable principle: personalization is persuasive only when recipients can verify or at least plausibly accept the match; perceived eeriness or manipulation may drive the penalty.
  • One could attempt to decompose the mechanism by engineering human posts to match LLM text on length, readability, emotional language, and call-to-action structure; if the gap disappears, the advantage is stylistic rather than deeper reasoning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports two preregistered online experiments (Study 1: N=658; Study 2: N=642) using a within-subjects design that crosses Content Source (human vs. LLM) with Personalization (generic, personalized, falsely personalized). Participants read six short social-media-style donation appeals, allocated a $0.10 bonus across six charities, rated perceived persuasiveness, and indicated engagement (Like/Neutral/Dislike). In both studies, LLM-generated appeals produced significantly higher donation allocations, engagement, and persuasiveness than human-authored appeals, with small effect sizes (d ≈ 0.11–0.23). Personalized appeals outperformed generic appeals only in Study 2, and falsely personalized appeals underperformed, particularly for LLM content. The authors conclude that LLMs can be effective tools for prosocial persuasion at scale, while cautioning that false personalization backfires.

Significance. If the results hold, this is the first direct benchmark of LLM versus human authors on a costly prosocial behavioral outcome (real donation allocation). The study has notable strengths: preregistration, a priori power analysis, a behavioral dependent variable, multiple complementary outcomes, two studies with different human-writer populations (academically trained students vs. incentivized laypeople), and open data and materials. The main weakness is the absence of any authorship-suspicion probe, which leaves the central content-quality attribution open to a perceived-source confound. Despite this, the consistency of the LLM advantage across outcomes and studies gives reasonable confidence in the descriptive result, though the interpretive claim about LLM persuasive skill needs qualification.

major comments (3)
  1. [Limitations and Future Directions] The undisclosed-authorship design lacks any manipulation check or suspicion probe. The central claim that 'LLM-generated content' is more effective than human content presumes that participants responded to the text rather than to perceived AI authorship. Because the paper cites evidence that an explicit AI label can reverse the LLM advantage (Yin et al., 2024; Osborne & Bailey, 2025), the direction of any perceived-source confound is unclear; it could inflate or deflate the observed effect. Without data on whether participants guessed AI authorship, the content-quality attribution is unsupported. Please provide a suspicion probe if any data exist, or explicitly reframe the conclusion as 'under undisclosed-authorship conditions' and temper the general claim about LLM persuasive skill.
  2. [Study 2, Stimuli Development] LLM posts from Study 1 were reused in Study 2, while human posts were newly written by a different, incentivized writer pool. This confounds the content-source comparison in Study 2 with stimulus vintage and writer-pool differences in the LLM condition. The Study 2 LLM advantage could reflect the particular generated outputs from Study 1 rather than the LLM's capability under Study 2's conditions. A cleaner test would generate new LLM posts using the same instructions as the human writers in Study 2. At minimum, the paper should acknowledge this asymmetry and discuss its implications for the generalizability of the Study 2 effect.
  3. [Measures and Results] The false-personalization condition presupposes participants noticed the mismatch. No manipulation check is reported for whether they perceived the false demographic cues; if they did not, 'falsely personalized' may be functionally equivalent to 'generic.' This bears on the interpretation of the backfire effect and the practical recommendation to avoid false personalization. Please add a manipulation check or temper the conclusions accordingly.
minor comments (5)
  1. [Abstract] The abstract says LLM content 'yielded more donations' without reporting effect sizes or confidence intervals. Consider including the key effect sizes (e.g., d = 0.11 and d = 0.13) to give readers a sense of the magnitude.
  2. [References] The lme4 reference has a malformed author field ('simulate.formula', etc.). Please correct the citation to the standard format (Bates, Mächler, Bolker, & Walker).
  3. [General] The paper states that an overview of preregistration deviations is available at OSF, but it does not summarize these deviations in the text. Briefly listing the deviations would improve transparency and help readers assess the confirmatory status of each result.
  4. [Table 1] The variable 'Familiarity with charity' appears in the table and note but is not described in the Measures section. Please clarify how this variable was measured and why it was included.
  5. [Figures] The figure captions could include the number of observations per condition and specify whether error bars are 95% confidence intervals from the mixed models. This would aid visual interpretation.

Circularity Check

0 steps flagged · score 0.0 of 10

Empirical benchmark with no fitted prediction; no circular derivation chain.

full rationale

This paper reports two preregistered experiments that directly measure donation allocations, engagement, and persuasiveness after exposing participants to donation appeals. The central comparison—LLM-generated versus human-authored content—is an empirical outcome contrast, not a derived quantity. No model parameter is fitted to the donation data and then renamed as a prediction; the preregistered linear mixed-effects analyses treat content source and personalization as manipulated factors and donation amount as an observed dependent variable. No equation defines the target result in terms of its inputs, and no 'prediction' reduces to a fitted value by construction. The self-citations present (e.g., Kleinberg et al., 2024; Festor et al., 2026; Puklavec et al., 2024) are background support for claims about LLM empathy or moral language and are not load-bearing for the empirical result. The paper's own flagged limitation—that authorship was undisclosed—is a potential confound regarding whether participants inferred AI authorship, but this is a validity threat, not circularity: it does not make the outcome equivalent to the stimulus-generation procedure. Thus, no significant circularity is present.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No fitted theoretical constants or free parameters are used. The central claim rests on domain assumptions about the validity of the donation-allocation task, exclusion rules, stimulus representativeness, and source-blindness; none of these are circular, but several are untested assumptions that could affect generalizability.

assumptions (5)
  • domain assumption The six-condition within-subjects allocation of a fixed $0.10 bonus is a valid operationalization of donation amount per appeal.
    Underlies the central dependent variable; allocation is zero-sum across the six conditions (Study 1 and Study 2, Measures and Procedure).
  • domain assumption Preregistered careless-response exclusion removes invalid responders without biasing content-source effects.
    9.2% and 5.9% of participants were excluded; the rule is preregistered but could interact with response styles in unknown ways.
  • domain assumption Posts shown in a social-media wrapper on a single page approximate real-world exposure to donation appeals.
    Ecological-validity claim in Procedure; no field validation is provided.
  • domain assumption gpt-4.5-preview output is representative of LLM-generated appeals, and the selected human writers are representative of human-authored appeals.
    Stimuli use one LLM and two human writer pools; results may not generalize to other models or writer populations.
  • domain assumption Undisclosed authorship means participants did not systematically infer AI authorship.
    The authors acknowledge this in Limitations; there is no manipulation check, and cited work shows that an AI label can reverse the effect.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Prosocial Persuasion at Scale? Large Language Models Outperform Humans in Donation Appeals Across Levels of Personalization." pith.science (2026). https://pith.science/paper/NJBT33RY

@misc{pith2026260403202,
  author       = {Pith},
  title        = {Pith review of: Prosocial Persuasion at Scale? Large Language Models Outperform Humans in Donation Appeals Across Levels of Personalization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NJBT33RY}},
  note         = {Machine review of arXiv:2604.03202}
}
read the original abstract

Large Language Models (LLMs) are increasingly regarded as having the potential to generate persuasive content at scale. While previous studies have focused on the risks associated with LLM-generated misinformation, the role of LLMs in enabling prosocial persuasion is still underexplored. We investigate whether donation appeals authored by LLMs are as effective as those written by humans across degrees of personalization. Two preregistered online experiments (Study 1: N = 658; Study 2: N = 642) manipulated Personalization (generic vs. personalized vs. falsely personalized) and Content source (human vs. LLM) and presented participants with donation appeals for charities. We assessed how participants distributed their bonus money across the charities, how they engaged with the donation appeals, and how persuasive they found them. In both experiments, LLM-generated content yielded statistically significantly higher donation amounts than human-authored content, although the absolute differences were small; it also resulted in higher engagement and was rated as more persuasive. There was a gain associated with personalization (Study 2) and a penalty for false personalization (Study 1). Our results suggest that LLMs may be a suitable technology for generating content that can encourage prosocial behavior.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

2 extracted references · 1 linked inside Pith

  1. [1091]

    https://doi.org/10.1177/0093650220961965 Zettler, I., & Strandsbjerg, C. F. (2025). Personalized interventions. Current Opinion in Psychology, 66, 102147. https://doi.org/10.1016/j.copsyc.2025.102147 Zhu, Q., Chong, L., Yang, M., & Luo, J. (2024). Reading users’ minds from what they say: An investigation into llm-based empathic mental inference. Internati...

  2. [6037]

    when” not “if

    https://doi.org/10.1038/s41467-025-61345-5 Bajpai, S., Sameer, A., & Fatima, R. (2025). Insights into Moral Reasoning of AI: A Comparative Study Between Humans and Large Language Models. Journal of Media Ethics, 1–15. Barclay, P., & Barker, J. L. (2020). Greener than thou: People who protect the environment are more cooperative, compete to be environmenta...

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.