Pith. sign in

REVIEW 4 cited by

Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17753 v3 pith:V6YM4MKK submitted 2024-06-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords languagepersuasivellmsdomainsinstructedtextwhenacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We are exposed to much information trying to influence us, such as teaser messages, debates, politically framed news, and propaganda - all of which use persuasive language. With the recent interest in Large Language Models (LLMs), we study the ability of LLMs to produce persuasive text. As opposed to prior work which focuses on particular domains or types of persuasion, we conduct a general study across various domains to measure and benchmark to what degree LLMs produce persuasive language - both when explicitly instructed to rewrite text to be more or less persuasive and when only instructed to paraphrase. We construct the new dataset Persuasive-Pairs of pairs of a short text and its rewrite by an LLM to amplify or diminish persuasive language. We multi-annotate the pairs on a relative scale for persuasive language: a valuable resource in itself, and for training a regression model to score and benchmark persuasive language, including for new LLMs across domains. In our analysis, we find that different 'personas' in LLaMA3's system prompt change persuasive language substantially, even when only instructed to paraphrase.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tailored untruths: How personalisation challenges LLM safeguards

    cs.CL 2025-10 conditional novelty 7.0 of 10

    A 1.6-million-text study of eight LLMs in four languages finds that adding demographic personae to disinformation prompts raises jailbreak rates from 78% to 82%.

  2. It Matters How You Say It: Exploring Rhetorical Patterns for AI-Assisted Information Evaluation

    cs.HC 2026-07 conditional novelty 6.0 of 10

    In a 98-participant fact-checking study, the rhetorical style of AI advice changed accuracy, confidence, and preference, with step-by-step explanations helping most and user preference diverging from performance.

  3. Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Emotional prompts make LLMs produce more cognitively complex language than rational prompts, and models systematically differ in their use of Cialdini influence principles across prompt types.

  4. AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    AssertBench measures how often LLMs keep the same true/false evaluation of a fact across contradictory user framings, and finds most tested models agree with the user's framing more when they do not know the fact.

Pith tools