Pith. sign in

REVIEW 3 major objections 4 minor 2 cited by

Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns

T0 review · 3 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read Large language models systematically vary persuasive style with recipient gender, producing communal language for female targets and agentic language for male targets across all 13 models and 16 languages.

desk verdict Solid, well-verified audit of gender-stereotypical persuasive style in 13 LLMs; direction is credible, but per-category significance needs multiple-comparison correction and judge-noise propagation. read the letter →

arxiv 2601.05751 v2 pith:PO2AWWPJ submitted 2026-01-09 cs.CL cs.AI

classification cs.CLcs.AI
keywords persuasivelanguagegenderbiasLLM-as-judgestereotypesagenticandcommunaltraitsCialdiniprinciplesrhetoricalappealsmultilingualevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to establish whether large language models change the style of their persuasive language when the recipient's gender changes, and to measure how big that shift is. It builds a controlled evaluation framework in which each model writes to paired prompts that differ only in one attribute—recipient gender, sender intent, or output language—and an LLM judge then scores the pairs on 19 established categories of persuasion (rhetorical appeals, six classic persuasion principles, agentic/communal traits, interaction goals, and tone). Across all 13 models tested, responses addressed to female recipients were significantly more affectionate, polite, relational, communal, and emotional, while responses to male recipients were more direct, instrumental, agentic, and logical. The same gendered pattern appeared in a single-model study across 16 languages, and the framework also revealed shifts when intent was framed as noble versus ignoble. If the pattern is real, everyday uses of LLMs for drafting messages will systematically reproduce and potentially amplify gender-stereotypical styles at scale.

What carries the argument

The load-bearing mechanism is a pairwise-prompt evaluation framework: each test prompt is duplicated with a single swapped attribute (e.g., 'female coworker' vs 'male coworker'), the model generates both responses, and an LLM judge scores which text shows more of each of 19 categories on a symmetric -3 to +3 scale; positional bias is mitigated by scoring both orders and symmetrising. Aggregate per-category scores are tested with a signed-rank test, and a Treatment Gap (the sum of absolute mean differences) quantifies how strongly a model differentiates between treatments. The framework's credibility rests on verification steps: a second judge model correlates strongly with the first, replaci

What would settle it

Re-run the core gender experiment on a random subset of the 13 models with multiple-comparison correction and a judge whose outputs are verified against a larger, carefully recruited human annotation set on a per-category basis; if fewer than half of the 19 categories remain significant in the stereotypical direction for a majority of models, the 'across all models' claim would fail.

Watch

Extended reading notes

Core claim

The central claim is that every LLM examined—from small open-weight models to large safety-tuned systems—produces gender-stereotypical persuasive language: female-targeted messages and arguments are judged as warmer, more communal, more polite, and more pathos-laden, while male-targeted ones are judged as more direct, instrumental, agentic, and logos-laden. The paper reports statistically significant per-category differences (via a signed-rank test) across all models, with the same direction of effect on both interpersonal messages and political arguments, and in 15 additional languages when tested on a single multilingual model. The authors argue these patterns align with well-documented so

Load-bearing premise

The paper's claim that the differences are 'significant across all models' depends on treating the LLM judge's scores as reliable measurements, but judge consistency is only about 75–84% across swapped orders, human annotators disagree substantially (agreement coefficient 0.09–0.2), and no correction is applied for testing 19 categories at once.

Editorial extensions

If this is right

  • LLMs used to draft emails, fundraising appeals, or campaign arguments will systematically code persuasive style by recipient gender, potentially reinforcing traditional gender roles at scale.
  • Because safety-aligned and open models all show the effect, bias mitigation at the alignment layer alone is unlikely to remove gendered persuasive styles.
  • The pairwise framework transfers to other treatment attributes (e.g., noble vs ignoble intent, response language), giving researchers a reusable way to audit which prompt attributes shift generation style.
  • The persistence of the gendered pattern across 16 languages indicates the bias is not an English-specific quirk, likely arising from shared training-data patterns.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the judge reflects what humans perceive, then LLM-based personalisation tools—which already tailor messages to recipients—will amplify stereotypical style differences even when the sender never asked for them; a direct behavioural test would measure whether recipients actually respond differently to the two styles.
  • The larger gender gap in Chinese (versus other languages) despite no correlation with country-level gender inequality suggests training-data composition, rather than societal indices, modulates the bias; auditing a model's pre-training data for gendered language patterns could explain this variation.
  • A testable extension is to apply the same pairwise framework to other binary attributes (e.g., recipient age, socio-economic status, or race) to see whether the communal/agentic split generalises or is specific to gender.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a framework for measuring differences in persuasive language generated by LLMs under pairwise prompt treatments, and applies it to recipient gender (13 LLMs), sender intent framing, and output language (16 languages). Persuasive language is operationalized as 19 categories across five theory-derived dimensions and scored by GPT-4o as an LLM judge, with a symmetric positional swap. The central claim is that all tested LLMs show significant gender-stereotypical differences: female-targeted responses are more affectionate, polite, communal, relational, and pathos-laden, while male-targeted responses are more direct, instrumental, agentic, and logos-laden. The authors include verification measures for one anchor model (LLAMA3.3-70B): human annotation of collapsed Communal+/Agentic+ dimensions, an alternative judge model, gender-term neutralization, and cross-lingual translation checks.

Significance. If the central claim survives re-analysis, this is a practically important result: widely used LLMs automatically produce communal/agentic stylistic splits based solely on the recipient's gender, with potential to reinforce societal stereotypes in everyday persuasive writing. The paper's strengths are its controlled pairwise prompt design, the breadth of the evaluation (13 models, 16 languages, 19 categories), and its explicit verification protocol—especially the alternative-judge, neutralization, and human-annotation checks, which go beyond most LLM-as-judge studies. The paper is also transparent in reporting low human inter-annotator agreement and imperfect positional consistency, which makes the failure to propagate those uncertainties into the headline significance tests the more conspicuous.

major comments (3)
  1. [Sect. 3.2, Eq. (2); App. B; App. D; Figs. 3 and 5] The Wilcoxon signed-rank tests treat the averaged judge scores as exact measurements, but App. B (Table 4) reports positional consistency of only 75–84% for LLAMA3.3, and App. D reports human inter-annotator Krippendorff alpha of 0.09–0.2. Neither source of noise is propagated into the significance tests. With 19 categories × 2 test sets × 13 models (≈494 tests), uncorrected α=0.05 implies a substantial expected number of false positives. Please report FDR-corrected p-values and/or bootstrap intervals that resample the two positional judgments (and, ideally, multiple judge runs). Without this, the Abstract's 'significant ... across all models' claim is not quantitatively supported.
  2. [Sect. 7; App. D; Sect. 4.2] The verification experiments—human annotations, alternative judge, gender-term neutralization, and cross-lingual translation checks—are all performed only on LLAMA3.3-70B responses, and the human annotation collapses the 19 categories into two aggregate questions (Communal+ and Agentic+). The paper's headline conclusion, however, covers 13 models and asserts fine-grained per-category differences, including Cialdini principles and interaction goals. Either extend validation to at least a few additional models and to the full 19-category vector, or restate the conclusion as a narrower claim about the communal/agentic direction, presenting the per-model fine-grained results as exploratory rather than as verified findings.
  3. [App. E, Table 8; Sect. 4.2] Refusal rates are treatment-correlated and high for several models (e.g., GPT-5 female 84.7% vs male 69.3%; CLAUDE-OPUS female 46% vs male 40%). The analysis omits refused prompts and retains only 10 models for the argument condition. Because refusals are not gender-neutral, the subset of prompts could introduce selection bias into the aggregate gender-gap comparisons. Please report whether the pattern in Figs. 4–5 is robust to the omitted prompts (e.g., via sensitivity analysis on the 10-prompt omission list) or explicitly discuss this as a limitation for the argument condition.
minor comments (4)
  1. [Throughout] Typos and grammatical slips: 'catagory' (Eq. 2 context), 'vice-vesa' (Sect. 4.2), 'Simultanously' (Sect. 2), 'wiht key' (App. A.2), 'Krippendorf Alpha ranging from 0.09 to 20.2' (App. D; should be 0.09 to 0.2).
  2. [Fig. 5] In both panels, several model labels are duplicated (Qwen3-235B and Qwen3-30B appear twice each), making it difficult to map rows to models. Please use unique labels.
  3. [Limitations] The sentence 'the strong agreement exhibited these independent communication theoretic frameworks' is ungrammatical and overstates the independence of the categories; the frameworks overlap conceptually (e.g., politeness and communal orientation). A minor rewrite would clarify.
  4. [Sect. 7.1] The phrase 'Sect. 7' appears as a self-reference in the first paragraph; consider replacing with 'this section' for clarity.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity; minor non-load-bearing self-citations only

full rationale

The paper's derivation chain is empirical rather than definitional: paired prompts differing only in the treatment attribute are fed to 13 different generation models; the generated responses are scored by GPT-4O (a separate model) across 19 categories; the mean directional differences Dj (Eq. 2) are tested with a Wilcoxon signed-rank test. The central claim (Sect. 4.2) is an aggregate of these measured scores and could in principle have been null or mixed-direction. The judge is not one of the tested generators, so the inputs and outputs are not the same object. The framework is verified with an alternative judge (Spearman rho=0.852/0.814), gender-term neutralization (rho=0.991/0.987), and human annotations on collapsed Communal+/Agentic+ dimensions (Sect. 7). Self-citations (Pauli et al. 2022, 2025) are used for a definition of persuasive language and related work, but the operationalization is grounded in external sources (Cialdini 2007; Bakan 1966; Wilson & Putnam 2012; Gass & Seiter 2010), so the self-citation is not load-bearing. The skeptical concerns about judge positional consistency, low inter-annotator agreement, and lack of multiple-comparison correction bear on reliability/validity of the measurements, not on whether the derivation reduces to its inputs; they are therefore outside the circularity verdict. No step exhibits Eq.-equals-Eq. reduction, fitted-parameter-renamed-as-prediction, or a uniqueness theorem imported from the authors.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new theoretical entities; its contribution is a measurement framework and empirical findings. The main unverified premise is that an LLM judge provides reliable ground-truth measurements for all models, not just the one anchor model.

assumptions (4)
  • domain assumption The 19 categories (rhetorical appeals, Cialdini principles, agency/communion, interaction goals, tones) coherently operationalize persuasive language.
    Section 3.1 grounds categories in Western theories (Aristotle, Cialdini, Bakan/Abele, Wilson and Putnam); the authors themselves note in Limitations that these may not transfer to non-Western contexts.
  • domain assumption GPT-4O's judgment is a valid and approximately unbiased measurement of these categories in paired texts.
    The entire pipeline (Eq. 1) relies on an LLM judge; human validation is only run for one model (LLAMA3.3, Sect. 7.1), so for the other 12 models this assumption is unverified.
  • domain assumption Paired responses differ only by the gender attribute in the prompt, so observed D_j is attributable to the treatment.
    Prompts are constructed pairwise (Sect. 3.3), but responses are generated independently, so stochastic variation is not controlled; the design assumes no other systematic difference.
  • domain assumption The Wilcoxon signed-rank test is valid on averaged judge scores.
    Scores are averaged over two order-swapped evaluations (Eq. 2), but judge positional inconsistency (App. B, Table 4: 75–84% consistency) is not propagated into the test.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns." pith.science (2026). https://pith.science/paper/PO2AWWPJ

@misc{pith2026260105751,
  author       = {Pith},
  title        = {Pith review of: Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PO2AWWPJ}},
  note         = {Machine review of arXiv:2601.05751}
}
read the original abstract

Large language models (LLMs) are increasingly used for everyday communication tasks, including drafting interpersonal messages intended to influence and persuade. Prior work has shown that LLMs can successfully persuade humans and amplify persuasive language. It is therefore essential to understand how user instructions affect the generation of persuasive language, and to understand whether the generated persuasive language differs, for example, when targeting different groups. In this work, we propose a framework for evaluating how persuasive language generation is affected by recipient gender, sender intent, or output language. We evaluate 13 LLMs and 16 languages using pairwise prompt instructions. We evaluate model responses on 19 categories of persuasive language using an LLM-as-judge setup grounded in social psychology and communication science. Our results reveal significant gender differences in the persuasive language generated across all models. These patterns reflect biases consistent with gender-stereotypical linguistic tendencies documented in social psychology and sociolinguistics.

Figures

Figures reproduced from arXiv: 2601.05751 by the authors.

Figure 1
Figure 1. Example of how LLAMA 3.3 varies persuasive language when the prompt specifies recipient gender. 2023; Salvi et al., 2024). As such, understanding and safeguarding against AI persuasion have be￾come critical cross-disciplinary topics (Burtell and Woodside, 2023; El-Sayed et al., 2024). In this work, we investigate how attributes in the instruction affect the style of persuasive language generated by LLMs. Specificall… view at source ↗
Figure 2
Figure 2. Framework for evaluating differences in LLM-generated persuasive language under pairwise prompt [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Mean differences in persuasive language catagories [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (14 more)
Figure 5
Figure 5. Figure 5: Persuasive language differences per model, [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 8
Figure 8. Figure 8: Gender Gap on languages using GPT5-MINI. pairwise bootstrap tests, we assess whether gender gaps differ significantly between two languages; many do not differ, but for example, the gender gap in Chinese is significantly larger than in all other languages except Englis…
Figure 7
Figure 7. Figure 7: Persuasive language differences per language, [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 9
Figure 9. Figure 9: Aggregated human annotations: Grey: no significant difference, Blue: significantly more male, Red: significantly more female. the input, we manually replace gendered terms in the generated messages and arguments (e.g., man/- woman) with neutral alternatives (e.g., huma…
Figure 10
Figure 10. Figure 10: Distribution over text length (count of char [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: Widget showing a sample of the manual review to replace gender identifier terms with gender-neutral [PITH_FULL_IMAGE:figures/full_fig_p018_11.png]
Figure 12
Figure 12. Figure 12: Judge QWEN2.5-72B: Average over the rated categories Dj on pairwise difference between gender￾treatment responses generated by LLAMA 3.3. The Wilcoxon test is applied to test the significance of the differences. Grey: not significant, Blue: significant in male directi…
Figure 13
Figure 13. Figure 13: Screenshot of the annotaion tool. E Gender treatment experiments across models We check whether the models respond to the re￾quest in the expected format, or whether they refuse to provide the argument or messages. We use a regex expression to find the refusal, and ma…
Figure 15
Figure 15. Figure 15: Arguments test set: scatterplot showing average text length against gender gap across the tested models. To this aim, we select languages from diverse language families for which a language-to-country mapping is, to some extent, reasonably well approximated. To constr…
Figure 16
Figure 16. Figure 16: P-values from bootstrapping analysis over [PITH_FULL_IMAGE:figures/full_fig_p022_16.png]
Figure 17
Figure 17. Figure 17: Scatterplot over Gender Inequality Index and [PITH_FULL_IMAGE:figures/full_fig_p022_17.png]
Figure 18
Figure 18. Figure 18: Average over the rated categories Dj on pairwise difference between the language treatment pair. The Wilcoxon test is applied to test the significance of the differences. 23 [PITH_FULL_IMAGE:figures/full_fig_p023_18.png]
Figure 19
Figure 19. Figure 19: Boostrapping: testing the difference in gender gap in models pairwise (messages). [PITH_FULL_IMAGE:figures/full_fig_p024_19.png]
Figure 20
Figure 20. Figure 20: Boostrapping: testing the difference in gender gap in models pairwise (arguments). [PITH_FULL_IMAGE:figures/full_fig_p025_20.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Not What, But How: A Framework for Auditing LLM Responses across Positioning, Generalization, Anthromorphism, and Maxims

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    Presents FRANZ framework and SQUARE corpus for multi-dimensional audit of LLM response framing on subjective cultural queries, applied to three models to reveal differences and couplings.

  2. Pareto-Guided Teacher Alignment for Fair Personalized Text Generation

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Fairness mitigation in personalized text generation is objective-dependent with methods occupying different regions of the fairness-personalization Pareto frontier rather than any single strategy dominating all objectives.

Reference graph

Works this paper leans on

71 extracted references · 6 canonical work pages · cited by 2 Pith papers

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Andrea E Abele and Bogdan Wojciszke. 2014. Communal and agentic content in social cognition: A dual perspective model. In Advances in experimental social psychology, volume 50, pages 195--255. Elsevier

  4. [4]

    Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.258 MEGA : Multilingual evaluation of generative AI . In Proceedings of the 2023 Conference on Empirical Methods in Natural ...

  5. [5]

    Anthropic. 2025. Claude opus 4.1. https://www.anthropic.com/news/claude-opus-4-1. Large‑language model release announcement

  6. [6]

    Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, and 1 others. 2024. Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), page...

  7. [7]

    David Bakan. 1966. The duality of human existence: An essay on psychology and religion

  8. [8]

    Cassondra Batz-Barbarich, Nicole Strah, and Farhan Masud Ahmed. 2025. Do words matter? the impact of communal and agentic language on women’s application to job opportunities. Journal of Personnel Psychology

Show all 71 references
  1. [9]

    Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fern \'a ndez, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, and 1 others. 2025. LLM s instead of human judges? a large scale empirical study across 20 nlp evaluati...

  2. [10]

    Nimet Beyza Bozdag, Shuhaib Mehri, Gokhan Tur, and Dilek Hakkani-T \"u r. 2025 a . Persuade me if you can: A framework for evaluating persuasion effectiveness and susceptibility among large language models. arXiv preprint arXiv:2503.01829

  3. [11]

    Nimet Beyza Bozdag, Shuhaib Mehri, Xiaocheng Yang, Hyeonjeong Ha, Zirui Cheng, Esin Durmus, Jiaxuan You, Heng Ji, Gokhan Tur, and Dilek Hakkani-T \"u r. 2025 b . Must read: A systematic survey of computational persuasion. arXiv preprint arXiv:2505.07775

  4. [12]

    Simon Martin Breum, Daniel V dele Egdal, Victor Gram Mortensen, Anders Giovanni M ller, and Luca Maria Aiello. 2024. The persuasive power of large language models. In Proceedings of the International AAAI Conference on Web and Social Media, volume 18, pages 152--163

  5. [13]

    S \'e bastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, and 1 others. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712

  6. [14]

    Matthew Burtell and Thomas Woodside. 2023. Artificial Influence: An Analysis Of AI-Driven Persuasion . arXiv e-prints, pages arXiv--2303

  7. [15]

    Evan Chen, Run-Jun Zhan, Yan-Bai Lin, and Hung-Hsuan Chen. 2025. From structured prompts to open narratives: Measuring gender bias in llms through open-ended storytelling. arXiv preprint arXiv:2503.15904

  8. [16]

    Guiming Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024. Humans or llms as the judge? a study on judgement bias. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8301--8327

  9. [17]

    Cheng-Han Chiang and Hung-yi Lee. 2023. https://doi.org/10.18653/v1/2023.acl-long.870 Can large language models be an alternative to human evaluations? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 156...

  10. [18]

    Jaeyoon Choi and Nia Nixon. 2025. Agentic men, communal women?: Exploring gender bias in llm-based leadership identification for collaboration analytics. In Artificial Intelligence in Education, pages 11--18, Cham. Springer Nature Switzerland

  11. [19]

    Cialdini

    Robert B. Cialdini. 2007. Influence: The Psychology of Persuasion. HarperCollins e-books

  12. [20]

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and 1 others. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation ag...

  13. [21]

    Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho-Jui Hsieh. 2024. Or-bench: An over-refusal benchmark for large language models. arXiv preprint arXiv:2405.20947

  14. [22]

    DeepSeek-AI. 2024. https://arxiv.org/abs/2412.19437 Deepseek-v3 technical report . Preprint, arXiv:2412.19437

  15. [23]

    DeepSeek-AI. 2025. https://arxiv.org/abs/2501.12948 Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning . Preprint, arXiv:2501.12948

  16. [24]

    Malika Dikshit, Houda Bouamor, and Nizar Habash. 2024. https://doi.org/10.18653/v1/2024.gebnlp-1.11 Investigating gender bias in STEM job advertisements . In Proceedings of the 5th Workshop on Gender Bias in Natural Language Processing (GeBNLP), pages 179--189, Bangkok, Thaila...

  17. [25]

    Xiangjue Dong, Yibo Wang, Philip Yu, and James Caverlee. 2023. https://openreview.net/forum?id=ZDeEYmKYrR Probing explicit and implicit gender bias through LLM conditional text generation . In Socially Responsible Language Modelling Research

  18. [26]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, and 1 others. 2024. The llama 3 herd of models. arXiv e-prints, pages arXiv--2407

  19. [27]

    Eagly and Steven J

    Alice H. Eagly and Steven J. Karau. 2002. https://doi.org/10.1037/0033-295X.109.3.573 Role congruity theory of prejudice toward female leaders . Psychological Review, 109(3):573--598

  20. [28]

    Eagly and Wendy Wood

    Alice H. Eagly and Wendy Wood. 1991. https://doi.org/10.1177/0146167291173006 Explaining sex differences in social behavior: A meta-analytic perspective . Personality and Social Psychology Bulletin, 17(3):306--315

  21. [29]

    Seliem El-Sayed, Canfer Akbulut, Amanda McCroskery, Geoff Keeling, Zachary Kenton, Zaria Jalan, Nahema Marchal, Arianna Manzini, Toby Shevlane, Shannon Vallor, and 1 others. 2024. A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI . arXiv preprint arX...

  22. [30]

    Miller, Sasha Mitts, Adithya Renduchintala, and 8 others

    Meta Fundamental AI Research Diplomacy Team (FAIR)† FAIR, Meta, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis...

  23. [31]

    Xiyan Fu and Wei Liu. 2025. https://doi.org/10.18653/v1/2025.findings-emnlp.587 How reliable is multilingual LLM -as-a-judge? In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 11040--11053, Suzhou, China. Association for Computational Linguistics

  24. [32]

    Gass and John S

    Robert H. Gass and John S. Seiter. 2010. Persuasion, Social Influence, and Compliance Gaining (4th ed.) . Boston: Allyn & Bacon

  25. [33]

    Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Yuanzhuo Wang, and Jian Guo. 2024. https://api.semanticscholar.org/CorpusID:274234014 A survey on llm-as-a-judge . ArXiv, abs/2411.15594

  26. [34]

    Elizabeth L Haines and Steven J Stroessner. 2019. The role prioritization model: How communal men and agentic women can (sometimes) have it all. Social and Personality Psychology Compass, 13(12):e12504

  27. [35]

    Matthew Honnibal. 2017. spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing. (No Title)

  28. [36]

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276

  29. [37]

    Chuhao Jin, Kening Ren, Lingzhen Kong, Xiting Wang, Ruihua Song, and Huan Chen. 2024. https://doi.org/10.18653/v1/2024.acl-long.92 Persuading across diverse domains: a dataset and persuasion large language model . In Proceedings of the 62nd Annual Meeting of the Association fo...

  30. [38]

    Jones, Frank R

    Bryan D. Jones, Frank R. Baumgartner, Sean M. Theriault, Derek A. Epp, Cheyenne Lee, and Miranda E. Sullivan. 2023. Policy agendas project: Codebook. https://www.comparativeagendas.net/pages/master-codebook

  31. [39]

    Elise Karinshak, Sunny Xun Liu, Joon Sung Park, and Jeffrey T. Hancock. 2023. https://doi.org/10.1145/3579592 Working With AI to Persuade: Examining a Large Language Model's Ability to Generate Pro-Vaccination Messages . Proc. ACM Hum.-Comput. Interact., 7(CSCW1)

  32. [40]

    Hadas Kotek, Rikker Dockum, and David Sun. 2023. https://doi.org/10.1145/3582269.3615599 Gender bias and stereotypes in large language models . In Proceedings of The ACM Collective Intelligence Conference, CI '23, page 12–24, New York, NY, USA. Association for Computing Machinery

  33. [41]

    Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Marie Beckage, Hsuan Su, Hung yi Lee, and Lama Nachman

    Shachi H. Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Marie Beckage, Hsuan Su, Hung yi Lee, and Lama Nachman. 2025. https://openreview.net/forum?id=tIYMiYz6Bf Decoding biases: An analysis of automated methods and metrics for gender bias detec...

  34. [42]

    Robin Lakoff. 1973. Language and woman's place. Language in society, 2(1):45--79

  35. [43]

    Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D

    Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luc...

  36. [44]

    Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E Gonzalez, and Ion Stoica. 2024. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. In Forty-second International Conference on Machine Learning

  37. [45]

    Minqian Liu, Zhiyang Xu, Xinyi Zhang, Heajun An, Sarvech Qadir, Qi Zhang, Pamela J Wisniewski, Jin-Hee Cho, Sang Won Lee, Ruoxi Jia, and 1 others. 2025. Llm can be a dangerous persuader: Empirical study of persuasion safety in large language models. arXiv preprint arXiv:2504.10430

  38. [46]

    Yang Liu. 2024. Quantifying stereotypes in language. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1223--1240

  39. [47]

    Weicheng Ma, Hefan Zhang, Ivory Yang, Shiyu Ji, Joice Chen, Farnoosh Hashemi, Shubham Mohole, Ethan Gearey, Michael Macy, Saeed Hassanpour, and Soroush Vosoughi. 2025. https://doi.org/10.18653/v1/2025.naacl-long.203 Communication makes perfect: Persuasion dataset construction ...

  40. [48]

    Sandra C Matz, Jacob D Teeny, Sumer S Vaid, Heinrich Peters, Gabriella M Harari, and Moran Cerf. 2024. The potential of generative ai for personalized persuasion at scale. Scientific Reports, 14(1):4692

  41. [49]

    Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aim \'e e Kaffee, Tanmay Laud, Anne Lauscher, Roberto...

  42. [50]

    Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. https://doi.org/10.18653/v1/2021.acl-long.416 S tereo S et: Measuring stereotypical bias in pretrained language models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Inte...

  43. [51]

    OpenAI. 2025 a . https://cdn.openai.com/gpt-5-system-card.pdf GPT-5 System Card . Technical report, OpenAI. Technical report; model documentation and safety-evaluation details

  44. [52]

    OpenAI. 2025 b . Introducing gpt-4.1 in the api. https://openai.com/index/gpt-4-1/. [Large language model release announcement]

  45. [53]

    Ruby Ostrow and Adam Lopez. 2025. https://doi.org/10.18653/v1/2025.findings-emnlp.946 LLM s reproduce stereotypes of sexual and gender minorities . In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 17465--17477, Suzhou, China. Association for Comp...

  46. [54]

    Amalie Pauli, Leon Derczynski, and Ira Assent. 2022. https://doi.org/10.18653/v1/2022.nlp4pi-1.11 Modelling persuasion through misuse of rhetorical appeals . In Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI), pages 89--100, Abu Dhabi, United Arab Emirat...

  47. [55]

    Amalie Brogaard Pauli, Isabelle Augenstein, and Ira Assent. 2025. https://doi.org/10.18653/v1/2025.naacl-long.506 Measuring and benchmarking large language models' capabilities to generate persuasive language . In Proceedings of the 2025 Conference of the Nations of the Americ...

  48. [56]

    Yujin Potter, Shiyang Lai, Junsol Kim, James Evans, and Dawn Song. 2024. Hidden persuaders: Llms' political leaning and their influence on voters. arXiv preprint arXiv:2410.24190

  49. [57]

    Till Raphael Saenger, Musashi Hinck, Justin Grimmer, and Brandon M. Stewart. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.913 A uto P ersuade: A framework for evaluating and explaining persuasive arguments . In Proceedings of the 2024 Conference on Empirical Methods in Na...

  50. [58]

    Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West. 2024. On the conversational persuasiveness of large language models: A randomized controlled trial . arXiv preprint arXiv:2403.14380

  51. [59]

    Somesh Singh, Yaman K Singla, Harini SI, and Balaji Krishnamurthy. 2024. Measuring and improving persuasiveness of large language models. arXiv preprint arXiv:2410.02653

  52. [60]

    Guijin Son, Hyunwoo Ko, Hoyoung Lee, Yewon Kim, and Seunghyeok Hong. 2024. LLM -as-a-judge & reward model: What they can and cannot do. arXiv preprint arXiv:2409.11239

  53. [61]

    Shweta Soundararajan and Sarah Jane Delany. 2024. https://aclanthology.org/2024.icnlsp-1.42/ Investigating gender bias in large language models through text generation . In Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024),...

  54. [62]

    Deborah Tannen. 1990. You just don't understand: Women and men. Conversation. New York: Ballantine books

  55. [63]

    Qwen Team. 2025. https://arxiv.org/abs/2505.09388 Qwen3 technical report . Preprint, arXiv:2505.09388

  56. [64]

    Jasper Timm, Chetan Talele, and Jacob Haimes. 2025. Tailored truths: Optimizing llm persuasion with personalization and fabricated statistics. arXiv preprint arXiv:2501.17273

  57. [65]

    Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.243 ``kelly is a warm person, joseph is a role model'': Gender biases in LLM -generated reference letters . In Findings of the Association fo...

  58. [66]

    Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. 2024. https://aclanthology.org/2024.findings-eacl.61/ Do-not-answer: Evaluating safeguards in LLM s . In Findings of the Association for Computational Linguistics: EACL 2024, pages 896--911, St. Julian ' s,...

  59. [67]

    Steven R Wilson and Linda L Putnam. 2012. Interaction goals in negotiation. In Communication yearbook 13, pages 374--406. Routledge

  60. [68]

    An Yang, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoyan Huang, Jiandong Jiang, Jianhong Tu, Jianwei Zhang, Jingren Zhou, Junyang Lin, Kai Dang, Kexin Yang, Le Yu, Mei Li, Minmin Sun, Qin Zhu, Rui Men, Tao He, and 9 others. 2025. Qwen2.5-1m technical report. arXiv prep...

  61. [69]

    Joel Young, Craig H Martell, Pranav Anand, Pedro Ortiz, Henry Tucker Gilbert IV, and 1 others. 2011. A microtext corpus for persuasion detection in dialog. In Analyzing Microtext

  62. [70]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others. 2023. Judging LLM -as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595--46623

  63. [71]

    Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2020. https://doi.org/10.1162/coli_a_00368 The Design and Implementation of X iao I ce, an Empathetic Social Chatbot . Computational Linguistics, 46(1):53--93

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.