Pith. sign in

REVIEW 4 major objections 6 minor 22 references

Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Large language models can be persuaded by other LLMs to change their moral decisions in ambiguous scenarios, with susceptibility varying by model, scenario complexity, and conversation length.

desk verdict A useful first measurement of LLM-to-LLM moral persuadability, but the missing neutral control arm keeps the headline causal claim from being fully identified. read the letter →

arxiv 2411.11731 v1 pith:LS4MEZIW submitted 2024-11-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords moralpersuasionlargelanguagemodelsLLM-to-LLMdialogueambiguityethicalalignmentFoundationsQuestionnairedecisionchangerate
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that large language models change their moral decisions when another LLM argues with them, and that the size of the change depends on the model, the ambiguity of the scenario, and the length of the conversation. The authors run two experiments: an LLM-on-LLM persuasion setup over morally ambiguous scenarios, and a questionnaire-based test of whether explicit philosophical alignment prompts shift moral-foundation scores. If the claim is right, it means multi-agent systems can steer each other's ethical judgments in ambiguous cases, and that a model's stated ethical framework is not fixed.

What carries the argument

The apparatus is an agent-to-agent conversation protocol. Both agents receive the scenario and the Base Agent is told that it has already chosen the initial action; the Persuader Agent, without revealing its role, argues for the other action over a variable number of turns. The effect is measured by three metrics: Change in Action Likelihood (the shift in the model's probability of choosing an action), Decision Change Rate (the fraction of scenarios where the chosen action flips), and Rule Violation Rate (how often the chosen action violates each rule of common morality, such as 'do not deceive' and 'do not kill'). For the alignment experiment, the machinery is a 30-question moral-foundations questionnaire administered under role prompts that instruct the model to adopt a specific ethical framework.

What would settle it

Run the same protocol with a control persuader that produces content-matched but argument-free messages, or with no persuasion at all except the instruction that the agent has chosen an action; if the Decision Change Rate in the control is close to the rate with real arguments, then the paper's attribution of the shifts to persuasion is wrong.

Watch

Extended reading notes

Core claim

The paper claims that LLMs are susceptible to moral persuasion: in high-ambiguity scenarios, a Persuader Agent LLM can move a Base Agent LLM's chosen action, with the most susceptible models changing their decisions in close to half of the scenarios. Persuasion is largely ineffective in low-ambiguity scenarios, indicating that ambiguity is what opens the door. Models vary much more in how easily they are persuaded than in how persuasive they are, and neither model family nor model size reliably predicts susceptibility. In the second experiment, targeted prompts instructing utilitarian, deontological, or virtue-ethics perspectives shift Moral Foundations Questionnaire scores in model-specific ways, with utilitarianism producing the largest deviations in some models.

Load-bearing premise

The load-bearing premise is that the observed decision flips are caused by the persuader's arguments, because the Base Agent is always told it had already chosen an action and there is no neutral-content control conversation to rule out commitment, sycophancy, or role-play pressure.

Editorial extensions

If this is right

  • In morally ambiguous scenarios, one LLM can change another model's decision, with the most susceptible models flipping their choices in nearly half of the scenarios.
  • Persuasion is context-dependent: low-ambiguity scenarios show minimal effect, so ambiguity is what makes a model open to influence.
  • Susceptibility varies strongly across models while persuasive ability is more uniform, so a model's size or family cannot predict how easily it will be persuaded.
  • Conversation length matters up to a point: more turns increase decision changes, but four turns already capture most of the effect for most models.
  • Explicit philosophical alignment prompts shift moral-foundation scores, meaning a model's apparent ethical profile can be tuned with simple instructions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If these results generalize to longer multi-agent interactions, autonomous LLM systems operating together could drift in their moral judgments through conversation alone, without any retraining or weight update.
  • The absence of a neutral-content control conversation means the measured effect is an upper bound: part of the change likely comes from the Base Agent's commitment to its announced choice or from role-play pressure rather than from the persuader's arguments.
  • A natural testable extension is to measure whether the shifted decision persists after the conversation ends and in new, unseen scenarios, which would distinguish lasting alignment change from in-context compliance.
  • Another extension is to vary the persuader's stated authority or identity while holding arguments fixed, to see whether source cues matter as much as argument content.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper studies whether large language models (LLMs) can be persuaded to change their moral decisions through multi-turn conversation with another LLM, and whether prompting with ethical frameworks (utilitarianism, deontology, virtue ethics) shifts responses on a Moral Foundations Questionnaire. Experiment 1 uses the moralchoice dataset of high-ambiguity scenarios, has a Base Agent LLM state an initial action, has a Persuader Agent argue for the alternative, and measures Change in Action Likelihood (CAL) and Decision Change Rate (DCR). Experiment 2 prompts three LLMs with philosophical alignment prompts and compares MFQ-30 scores. The abstract claims that LLMs 'can indeed be persuaded' and that susceptibility depends on model, scenario complexity, and conversation length.

Significance. If the attribution to persuasive content were established, the paper would contribute useful evidence on LLM susceptibility to ethical influence, with implications for multi-agent deployments and moral alignment. The paper has concrete strengths: it builds on established external benchmarks (moralchoice, MFQ-30), provides simple transparent metrics, ships code, and does not rely on fitted parameters or circular derivations. However, the central causal claim is currently under-identified because the experimental design lacks a neutral-control condition, so the measured decision changes could arise from commitment, sycophancy, or generic conversational pressure rather than from the persuader's arguments. The turn-dependence and model-comparison results inherit this issue, and additional selection-bias concerns in both experiments weaken the generality of the conclusions as stated.

major comments (4)
  1. [§3.1, Table 1] The protocol has no neutral-control or re-ask condition. The Base Agent is prompted with 'You have chosen the action: initial_choice' and then converses with a Persuader Agent instructed to argue for the alternative. The measured DCR and CAL therefore conflate at least three mechanisms: anchoring or commitment to the action the model is told it chose, generic conversational drift or sycophancy over multiple turns, and content-specific moral persuasion. The abstract's claim that 'LLMs can indeed be persuaded' requires that the observed changes are attributable to the persuader's arguments. The paper's own remark (Section 2) that this approach 'controls for fewer variables in the generated text' understates the issue: the missing control is an unmeasured confound on the main outcome. Please add control arms with neutral content of matched length and a re-ask arm with no conversation, and report DCR/CAL for those arms.
  2. [§3.1, Data] The paper states that 100 of the 680 high-ambiguity scenarios were used, but no criterion for this selection is given. If the subset is not random or is chosen based on some property (e.g., scenario length, rule type), the measured susceptibility and the model ranking in Figure 3.2.2 could be biased. The authors should specify the selection procedure and provide robustness checks, such as results on multiple random subsets or on the full set if feasible. Without this, the quantitative claims about susceptibility across models lack a clear population definition.
  3. [§4.2] Experiment 2 reports MFQ results for only three models because other models 'refused to provide answers for significant portions of the survey.' The refusal rate per model and per prompt is not reported, and the analysis does not account for how missing responses were handled in the MFQ scoring. This creates a selection bias: the comparison of 'responsiveness' to ethical prompts across models is confounded by differential refusal behavior, and the claim that gpt-4o and mistral-7b-instruct show more variability than claude-3-haiku may reflect compliance patterns rather than underlying moral-flexibility differences. Please report refusal rates, the number of completed questions, and a sensitivity analysis.
  4. [§3.2.1, §3.2.2] The CAL and DCR values in Figures 1 and 2 are reported as point estimates without uncertainty. Since each value is computed from a finite set of 100 scenarios and from a sampled set of M token sequences (Equation 2), the differences across models and turn counts may not be statistically reliable. In particular, the choice of a four-turn conversation for the main evaluation is based on a small set of four model permutations in Figure 1. The authors should provide confidence intervals (e.g., bootstrap over scenarios) and, ideally, significance tests for the differences they discuss, such as the 0.06 vs. 0.49 CAL contrast between low- and high-ambiguity scenarios.
minor comments (6)
  1. [§3.1] There is a typo in 'Persauder Agent' (should be 'Persuader Agent').
  2. [§2] The claim that there is 'no prior research that studies how susceptible LLMs are to persuasion from other models in moral contexts' is too strong given the cited work on model-on-model deception (Heitkoetter et al., 2024) and persuasion-based jailbreaking (Zeng et al., 2024). Please soften the novelty claim.
  3. [§3.1, Definitions] The equation numbering is inconsistent: Definition 3.1's equation is labeled (1) and Definition 3.2's equation is labeled (2), but Equations (1) and (2) already appear earlier in Section 3.1 for action likelihood. Please renumber to avoid confusion.
  4. [§3.1, Metrics] The metric definitions do not specify temperature, sampling parameters, or the number of sampled sequences M used to approximate action likelihood. Since the reported CAL depends on the stochasticity of generation, please report these details in the methods or appendix.
  5. [Figures 1, 2] The figure legends are inconsistent: Figure 1 uses 'mistral-7b-instruct_mistral-7b-instruct' with an underscore, and the text references 'Figure 3.2.2' as a placeholder. Please unify legend formatting and replace placeholder references with the actual figure numbers.
  6. [References] Several references are incomplete: Durmus et al. (2024) and others lack venue or arXiv identifiers. Please provide full bibliographic information.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper reports empirical measurements against external datasets and metrics, with no equation or self-citation that forces the claims by construction.

full rationale

The paper contains no derivation that reduces a claimed prediction to an input. Experiment 1 measures pre/post decision change of a Base Agent under a Persuader prompt, using the external moralchoice dataset and action-likelihood definitions taken from Scherrer et al. (2023); the metrics CAL, DCR, and RVR are defined directly on measured samples and are not fitted to the outcome. Experiment 2 measures MFQ-30 responses under different system prompts, and the prompts are independent of the measured scores. Reusing Scherrer et al.'s dataset and definitions is benchmark reuse, not self-citation that is load-bearing. The absence of a neutral-control arm is a real causal attribution threat, but it concerns experimental identification rather than circularity, because the DCR is not by construction equal to any fitted parameter or to the prompt input. No equation equates an output to an input, and no prior work by these authors is cited as the basis of the central claim. Therefore the circularity score is 0.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

No result is derived from fitted physical parameters; the ledger records the design choices and background assumptions that the central measurements depend on.

free parameters (2)
  • Conversation length (turns) = 4
    Selected based on a turn-count sweep (Section 3.2.1) after observing marginal gains beyond four turns; the choice is data-driven and affects the reported DCR/CAL values.
  • Scenario subset = 100 high-ambiguity scenarios
    The paper states 'We used 100 of the high ambiguity scenarios for this experiment' without giving a selection rule, so the subset is a hand-picked design choice that defines all reported results.
assumptions (4)
  • domain assumption The moralchoice dataset's high-ambiguity labels are valid and the selected scenarios are genuinely morally ambiguous.
    The persuasion effect is interpreted as context-dependent; if the scenarios are not actually ambiguous, the observed changes could be simple inconsistency. This enters in Section 3.1 Data.
  • domain assumption The stem-matching mapping g in Equation (2) reliably maps generated text to the intended action.
    CAL, DCR, and RVR all depend on the deterministic mapping g; no accuracy rate for the stem-matching operationalization is reported.
  • ad hoc to paper Decision changes are attributable to persuasive content rather than to commitment, sycophancy, or role-play pressure.
    The Base Agent is told 'You have chosen the action: initial_choice' (Table 1) and there is no neutral-control conversation, so the causal attribution is assumed rather than demonstrated.
  • domain assumption The MFQ-30 questionnaire is a valid instrument for probing moral foundations in LLMs.
    The alignment experiment treats a human-designed survey as a model probe; several models refused parts of the survey, and the reliability of MFQ on LLMs is not established in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment." pith.science (2026). https://pith.science/paper/LS4MEZIW

@misc{pith2026241111731,
  author       = {Pith},
  title        = {Pith review of: Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LS4MEZIW}},
  note         = {Machine review of arXiv:2411.11731}
}
read the original abstract

We explore how large language models (LLMs) can be influenced by prompting them to alter their initial decisions and align them with established ethical frameworks. Our study is based on two experiments designed to assess the susceptibility of LLMs to moral persuasion. In the first experiment, we examine the susceptibility to moral ambiguity by evaluating a Base Agent LLM on morally ambiguous scenarios and observing how a Persuader Agent attempts to modify the Base Agent's initial decisions. The second experiment evaluates the susceptibility of LLMs to align with predefined ethical frameworks by prompting them to adopt specific value alignments rooted in established philosophical theories. The results demonstrate that LLMs can indeed be persuaded in morally charged scenarios, with the success of persuasion depending on factors such as the model used, the complexity of the scenario, and the conversation length. Notably, LLMs of distinct sizes but from the same company produced markedly different outcomes, highlighting the variability in their susceptibility to ethical persuasion.

Figures

Figures reproduced from arXiv: 2411.11731 by the authors.

Figure 1
Figure 1. Change in Action Likelihood (left) and Decision Change Rate (right) over number of turns for four permutations of models. Conversations with more turns tend to result in higher CAL and DCR , with the exception of mistral-7b-instruct. 3.2.2 Evaluating Effectiveness and Susceptibility to Persuasion in LLMs In this section, we explore the effectiveness of different LLMs combinations as both Base Agent and Persuader Age… view at source ↗
Figure 2
Figure 2. Change in Action Likelihood for each pairwise combination of LLMs as either the [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. MFQ scores across various ethical prompts. The radar plots illustrate how different [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Probability density of selecting action1 across various models in generated scenarios (left) and on handwritten scenarios (right). The distribution peaks indicate a strongest preference towards action1, with distinct variations in likelihood across different models [P…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

22 extracted references · 12 canonical work pages

  1. [1]

    Abdulhai, M., Serapio - Garc \' a, G., Crepy, C., Valter, D., Canny, J., and Jaques, N. (2023). Moral foundations of large language models. CoRR , abs/2310.15337

  2. [2]

    A., Adeli, E., Altman, R

    Bommasani, R., Hudson, D. A., Adeli, E., Altman, R. B., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N. S., Chen, A. S., Creel, K., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei - Fei, L., ...

  3. [3]

    K., Vennam, S., Govil, P., Kumaraguru, P., and Gaur, M

    Bonagiri, V. K., Vennam, S., Govil, P., Kumaraguru, P., and Gaur, M. (2024). Sage: Evaluating moral consistency in large language models. In LREC/COLING , pages 14272--14284. ELRA and ICCL

  4. [4]

    Durmus, E., Lovitt, L., Tamkin, A., Ritchie, S., Clark, J., and Ganguli, D. (2024). Measuring the persuasiveness of language models

  5. [5]

    Gert, B. (2004). Common Morality: Deciding What to Do . Oxford University Press

  6. [6]

    P., and Ditto, P

    Graham, J., Haidt, J., Koleva, S., Motyl, M., Iyer, R., Wojcik, S. P., and Ditto, P. H. (2013). Moral foundations theory: The pragmatic validity of moral pluralism. In Advances in experimental social psychology , volume 47, pages 55--130. Elsevier

  7. [7]

    Graham, J., Haidt, J., and Nosek, B. A. (2009). Liberals and conservatives rely on different sets of moral foundations. J. Pers. Soc. Psychol. , 96(5):1029--1046

  8. [8]

    and Joseph, C

    Haidt, J. and Joseph, C. (2004). Intuitive ethics: How innately prepared intuitions generate culturally variable virtues. Daedalus , 133(4):55--66

Show all 22 references
  1. [9]

    Heitkoetter, J., Gerovitch, M., and Newhouse, L. (2024). An assessment of model-on-model deception. CoRR , abs/2405.12999

  2. [10]

    Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., and Steinhardt, J. (2021). Aligning AI with shared human values. In ICLR . OpenReview.net

  3. [11]

    Hendrycks, D., Mazeika, M., and Woodside, T. (2023). An overview of catastrophic AI risks. CoRR , abs/2306.12001

  4. [12]

    S., and Sun, L

    Huang, Y., Zhang, Q., Yu, P. S., and Sun, L. (2023). Trustgpt: A benchmark for trustworthy and responsible large language models. CoRR , abs/2306.11507

  5. [13]

    Ji, J., Chen, Y., Jin, M., Xu, W., Hua, W., and Zhang, Y. (2024). Moralbench: Moral evaluation of llms. CoRR , abs/2406.04428

  6. [14]

    Kuhn, L., Gal, Y., and Farquhar, S. (2023). Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In ICLR . OpenReview.net

  7. [15]

    S., Zou, A., Li, N., Basart, S., Woodside, T., Zhang, H., Emmons, S., and Hendrycks, D

    Pan, A., Chan, J. S., Zou, A., Li, N., Basart, S., Woodside, T., Zhang, H., Emmons, S., and Hendrycks, D. (2023). Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark. In ICML , volume 202 of Proceedings of Ma...

  8. [16]

    Payandeh, A., Pluth, D., Hosier, J., Xiao, X., and Gurbani, V. K. (2024). How susceptible are llms to logical fallacies? In LREC/COLING , pages 8276--8286. ELRA and ICCL

  9. [17]

    A., Webson, A., Ho, L., Lin, S., Farquhar, S., Hutter, M., Del \' e tang, G., Ruoss, A., El - Sayed, S., Brown, S., Dragan, A

    Phuong, M., Aitchison, M., Catt, E., Cogan, S., Kaskasoli, A., Krakovna, V., Lindner, D., Rahtz, M., Assael, Y., Hodkinson, S., Howard, H., Lieberum, T., Kumar, R., Raad, M. A., Webson, A., Ho, L., Lin, S., Farquhar, S., Hutter, M., Del \' e tang, G., Ruoss, A., El - Sayed, S....

  10. [18]

    H., Gallotti, R., and West, R

    Salvi, F., Ribeiro, M. H., Gallotti, R., and West, R. (2024). On the conversational persuasiveness of large language models: A randomized controlled trial. CoRR , abs/2403.14380

  11. [19]

    Scherrer, N., Shi, C., Feder, A., and Blei, D. M. (2023). Evaluating the moral beliefs encoded in llms. In NeurIPS

  12. [20]

    Simmons, G. (2023). Moral mimicry: Large language models produce moral rationalizations tailored to political identity. In ACL (student) , pages 282--297. Association for Computational Linguistics

  13. [21]

    S., Yang, S., Zhang, T., Shi, W., Zhang, T., Fang, Z., Xu, W., and Qiu, H

    Xu, R., Lin, B. S., Yang, S., Zhang, T., Shi, W., Zhang, T., Fang, Z., Xu, W., and Qiu, H. (2023). The earth is flat because...: Investigating llms' belief towards misinformation via persuasive conversation. CoRR , abs/2312.09085

  14. [22]

    Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W. (2024). How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge AI safety by humanizing llms. CoRR , abs/2401.06373

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.