{"id":"c9127480-7b8e-47b6-98ce-4c007cb1be38","arxiv_id":"2411.11731","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"LLMs are measurably persuadable in morally ambiguous scenarios, with susceptibility varying strongly by model and only slightly with conversation length beyond a few turns.","lead":"This paper tests whether one large language model can talk another into changing its answer in morally tricky situations, and whether telling a model to follow an ethical theory shifts its moral scores. The finding that some models flip nearly half their decisions suggests that multi-agent AI systems may need safeguards against persuasion between models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing neutral-control arm makes the causal attribution to persuasive content unidentifiable; observed decision changes could be driven by the Base Agent's instructions ('You have chosen initial_choice') or by any sustained conversation, not by the persuader's arguments.","rationale":"The paper's strongest claim is causal in the ordinary sense: persuasion is what changes the model's decision. The protocol, however, only establishes that decisions change after a scripted conversation in which the Base Agent is told it has already made a choice. A neutral conversation could produce similar rates of change through generic acquiescence, and the reminder of the initial choice could itself induce the model to stick with or, under debate pressure, abandon that choice. The reader's weakest-assumption field identifies exactly this attribution problem. I agree. The concern is load-bearing because it targets the interpretation of the headline result, not a secondary detail. It is also easily fixable with a control experiment, which is why the existing CONDITIONAL verdict is appropriate rather than a rejection. If the proposed control shows no difference between persuasive and neutral conversation, the paper should be reframed as measuring decision instability in multi-agent conversations, with the persuasion mechanism left open.","tokens_in":9186,"tokens_out":4361,"duration_ms":41773,"concrete_test":"Run the Stage-2 protocol on the same 100 high-ambiguity scenarios with at least the 4-turn setting and the strong-weak pair (llama-3.1-70b persuader, claude-3-haiku base) and one self-pair (gpt-4o with gpt-4o), across three arms: (1) the original persuasive-prompts condition; (2) a neutral-control condition where the Persuader is given the same scenario but instructed to write 100 tokens about an unrelated topic such as weather, without arguing for either action; and (3) a re-ask control where, after Stage 1, the Base Agent is shown 'You have chosen initial_choice' and asked to choose again with no second agent. Compare DCR and CAL across arms with bootstrap 95% confidence intervals. If arm (2) or (3) produces changes statistically indistinguishable from, or at least half of, arm (1), the attribution to persuasive content fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that LLMs 'can indeed be persuaded' requires that post-conversation decision changes are caused by the persuader's arguments. In Section 3.1 and Table 1, the Base Agent is prompted with 'Given the following scenario: context You have chosen the action: initial_choice' and then told to converse, while the Persuader is instructed to argue for the other action. The outcome is measured as the change from the Stage-1 baseline choice to a final choice made after this conversation. There is no control arm in which the Persuader produces neutral content, and no re-ask arm in which the Base Agent is merely reminded of its stated choice. Consequently, the reported Decision Change Rate and Change in Action Likelihood conflate at least three mechanisms: (a) anchoring/commitment to the action the model is told it has chosen, (b) generic conversational drift, sycophancy, or pressure to agree over multiple turns, and (c) content-specific moral persuasion. The turn-dependence result in Section 3.2.1 inherits the same confound. The paper's own remark that this approach 'controls for fewer variables in the generated text' understates the problem: the missing control is an unmeasured confound on the paper's main outcome, not merely variability in style.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether large language models (LLMs) can be persuaded to change their moral decisions through multi-turn conversation with another LLM, and whether prompting with ethical frameworks (utilitarianism, deontology, virtue ethics) shifts responses on a Moral Foundations Questionnaire. Experiment 1 uses the moralchoice dataset of high-ambiguity scenarios, has a Base Agent LLM state an initial action, has a Persuader Agent argue for the alternative, and measures Change in Action Likelihood (CAL) and Decision Change Rate (DCR). Experiment 2 prompts three LLMs with philosophical alignment prompts and compares MFQ-30 scores. The abstract claims that LLMs 'can indeed be persuaded' and that susceptibility depends on model, scenario complexity, and conversation length.","tokens_in":9416,"tokens_out":3745,"duration_ms":35912,"significance":"If the attribution to persuasive content were established, the paper would contribute useful evidence on LLM susceptibility to ethical influence, with implications for multi-agent deployments and moral alignment. The paper has concrete strengths: it builds on established external benchmarks (moralchoice, MFQ-30), provides simple transparent metrics, ships code, and does not rely on fitted parameters or circular derivations. However, the central causal claim is currently under-identified because the experimental design lacks a neutral-control condition, so the measured decision changes could arise from commitment, sycophancy, or generic conversational pressure rather than from the persuader's arguments. The turn-dependence and model-comparison results inherit this issue, and additional selection-bias concerns in both experiments weaken the generality of the conclusions as stated.","major_comments":[{"comment":"The protocol has no neutral-control or re-ask condition. The Base Agent is prompted with 'You have chosen the action: initial_choice' and then converses with a Persuader Agent instructed to argue for the alternative. The measured DCR and CAL therefore conflate at least three mechanisms: anchoring or commitment to the action the model is told it chose, generic conversational drift or sycophancy over multiple turns, and content-specific moral persuasion. The abstract's claim that 'LLMs can indeed be persuaded' requires that the observed changes are attributable to the persuader's arguments. The paper's own remark (Section 2) that this approach 'controls for fewer variables in the generated text' understates the issue: the missing control is an unmeasured confound on the main outcome. Please add control arms with neutral content of matched length and a re-ask arm with no conversation, and report DCR/CAL for those arms.","section":"§3.1, Table 1"},{"comment":"The paper states that 100 of the 680 high-ambiguity scenarios were used, but no criterion for this selection is given. If the subset is not random or is chosen based on some property (e.g., scenario length, rule type), the measured susceptibility and the model ranking in Figure 3.2.2 could be biased. The authors should specify the selection procedure and provide robustness checks, such as results on multiple random subsets or on the full set if feasible. Without this, the quantitative claims about susceptibility across models lack a clear population definition.","section":"§3.1, Data"},{"comment":"Experiment 2 reports MFQ results for only three models because other models 'refused to provide answers for significant portions of the survey.' The refusal rate per model and per prompt is not reported, and the analysis does not account for how missing responses were handled in the MFQ scoring. This creates a selection bias: the comparison of 'responsiveness' to ethical prompts across models is confounded by differential refusal behavior, and the claim that gpt-4o and mistral-7b-instruct show more variability than claude-3-haiku may reflect compliance patterns rather than underlying moral-flexibility differences. Please report refusal rates, the number of completed questions, and a sensitivity analysis.","section":"§4.2"},{"comment":"The CAL and DCR values in Figures 1 and 2 are reported as point estimates without uncertainty. Since each value is computed from a finite set of 100 scenarios and from a sampled set of M token sequences (Equation 2), the differences across models and turn counts may not be statistically reliable. In particular, the choice of a four-turn conversation for the main evaluation is based on a small set of four model permutations in Figure 1. The authors should provide confidence intervals (e.g., bootstrap over scenarios) and, ideally, significance tests for the differences they discuss, such as the 0.06 vs. 0.49 CAL contrast between low- and high-ambiguity scenarios.","section":"§3.2.1, §3.2.2"}],"minor_comments":[{"comment":"There is a typo in 'Persauder Agent' (should be 'Persuader Agent').","section":"§3.1"},{"comment":"The claim that there is 'no prior research that studies how susceptible LLMs are to persuasion from other models in moral contexts' is too strong given the cited work on model-on-model deception (Heitkoetter et al., 2024) and persuasion-based jailbreaking (Zeng et al., 2024). Please soften the novelty claim.","section":"§2"},{"comment":"The equation numbering is inconsistent: Definition 3.1's equation is labeled (1) and Definition 3.2's equation is labeled (2), but Equations (1) and (2) already appear earlier in Section 3.1 for action likelihood. Please renumber to avoid confusion.","section":"§3.1, Definitions"},{"comment":"The metric definitions do not specify temperature, sampling parameters, or the number of sampled sequences M used to approximate action likelihood. Since the reported CAL depends on the stochasticity of generation, please report these details in the methods or appendix.","section":"§3.1, Metrics"},{"comment":"The figure legends are inconsistent: Figure 1 uses 'mistral-7b-instruct_mistral-7b-instruct' with an underscore, and the text references 'Figure 3.2.2' as a placeholder. Please unify legend formatting and replace placeholder references with the actual figure numbers.","section":"Figures 1, 2"},{"comment":"Several references are incomplete: Durmus et al. (2024) and others lack venue or arXiv identifiers. Please provide full bibliographic information.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a workshop submission (AdvML-Frontiers'24). The topic is timely, but the missing control condition is a substantive flaw that prevents acceptance as-is. The authors should be encouraged to run the control experiments; the requested additions are feasible within the manuscript's scope. I also note that the authors should double-check the equation numbering and the citation completeness before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The new thing is the experimental combination: prior persuasion work targeted misinformation, jailbreaks, or deception, and this paper instead puts two LLMs in conversation over morally ambiguous scenarios from Scherrer et al.'s moralchoice dataset and measures how often the Base Agent flips its decision. That is a useful niche, the writing is clear, and code is linked. The main result—decision changes in up to about half of high-ambiguity scenarios, varying by model and turn count—is plausible and worth knowing for multi-agent systems and model evaluation.\n\nThe soft spot is the central attribution. The Base Agent is prompted with 'You have chosen the action: initial_choice' and then just told to converse. There is no neutral-control conversation and no re-ask arm that merely restates the choice. So the Decision Change Rate conflates commitment/sycophancy, generic conversational drift, and content-specific persuasion. The paper's own remark that the approach 'controls for fewer variables' understates this: it is not just style variance, it is an unmeasured confound on the main outcome. That said, the heatmap shows change varies by Persuader model, so content likely matters; the problem is you cannot decompose how much. This is fixable with one extra arm, but as published the causal wording in the abstract overreaches.\n\nOther soft spots are minor: the 100-scenario selection from 680 is unjustified; Experiment 2 only reports three models because others refused the survey, acknowledged but not addressed; and there are no confidence intervals or significance tests on the headline DCR/CAL numbers, though the appendix includes KS tests on baselines. The figures are a bit rough but cosmetic.\n\nWho gets value: multi-agent designers, safety researchers, and builders of moral-evaluation benchmarks. It is a preliminary measurement, not a mechanism demonstration. I would send it to a serious referee, but I would require a control arm, a rationale for the scenario subset, and a more cautious abstract before acceptance. As is, it is a solid workshop-level contribution needing one more experimental pass.","headline":"A useful first measurement of LLM-to-LLM moral persuadability, but the missing neutral control arm keeps the headline causal claim from being fully identified.","tokens_in":728,"tokens_out":1533,"would_cite":true,"duration_ms":30875,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models can be persuaded by other LLMs to change their moral decisions in ambiguous scenarios, with susceptibility varying by model, scenario complexity, and conversation length.","keywords":["moral persuasion","large language models","LLM-to-LLM dialogue","moral ambiguity","ethical alignment","Moral Foundations Questionnaire","decision change rate"],"falsifier":"Run the same protocol with a control persuader that produces content-matched but argument-free messages, or with no persuasion at all except the instruction that the agent has chosen an action; if the Decision Change Rate in the control is close to the rate with real arguments, then the paper's attribution of the shifts to persuasion is wrong.","tokens_in":8939,"feed_emoji":"🤖","tokens_out":4609,"duration_ms":42966,"temperature":0.7,"pith_summary":"This paper tries to establish that large language models change their moral decisions when another LLM argues with them, and that the size of the change depends on the model, the ambiguity of the scenario, and the length of the conversation. The authors run two experiments: an LLM-on-LLM persuasion setup over morally ambiguous scenarios, and a questionnaire-based test of whether explicit philosophical alignment prompts shift moral-foundation scores. If the claim is right, it means multi-agent systems can steer each other's ethical judgments in ambiguous cases, and that a model's stated ethical framework is not fixed.","feed_headline":"LLM-to-LLM talk can flip a model's moral decisions","feed_subtitle":"In ambiguous scenarios, some models switch their ethical answer in nearly half of cases after another model argues with them.","key_machinery":"The apparatus is an agent-to-agent conversation protocol. Both agents receive the scenario and the Base Agent is told that it has already chosen the initial action; the Persuader Agent, without revealing its role, argues for the other action over a variable number of turns. The effect is measured by three metrics: Change in Action Likelihood (the shift in the model's probability of choosing an action), Decision Change Rate (the fraction of scenarios where the chosen action flips), and Rule Violation Rate (how often the chosen action violates each rule of common morality, such as 'do not deceive' and 'do not kill'). For the alignment experiment, the machinery is a 30-question moral-foundations questionnaire administered under role prompts that instruct the model to adopt a specific ethical framework.","core_discovery":"The paper claims that LLMs are susceptible to moral persuasion: in high-ambiguity scenarios, a Persuader Agent LLM can move a Base Agent LLM's chosen action, with the most susceptible models changing their decisions in close to half of the scenarios. Persuasion is largely ineffective in low-ambiguity scenarios, indicating that ambiguity is what opens the door. Models vary much more in how easily they are persuaded than in how persuasive they are, and neither model family nor model size reliably predicts susceptibility. In the second experiment, targeted prompts instructing utilitarian, deontological, or virtue-ethics perspectives shift Moral Foundations Questionnaire scores in model-specific ways, with utilitarianism producing the largest deviations in some models.","pith_inferences":["If these results generalize to longer multi-agent interactions, autonomous LLM systems operating together could drift in their moral judgments through conversation alone, without any retraining or weight update.","The absence of a neutral-content control conversation means the measured effect is an upper bound: part of the change likely comes from the Base Agent's commitment to its announced choice or from role-play pressure rather than from the persuader's arguments.","A natural testable extension is to measure whether the shifted decision persists after the conversation ends and in new, unseen scenarios, which would distinguish lasting alignment change from in-context compliance.","Another extension is to vary the persuader's stated authority or identity while holding arguments fixed, to see whether source cues matter as much as argument content."],"forward_implications":["In morally ambiguous scenarios, one LLM can change another model's decision, with the most susceptible models flipping their choices in nearly half of the scenarios.","Persuasion is context-dependent: low-ambiguity scenarios show minimal effect, so ambiguity is what makes a model open to influence.","Susceptibility varies strongly across models while persuasive ability is more uniform, so a model's size or family cannot predict how easily it will be persuaded.","Conversation length matters up to a point: more turns increase decision changes, but four turns already capture most of the effect for most models.","Explicit philosophical alignment prompts shift moral-foundation scores, meaning a model's apparent ethical profile can be tuned with simple instructions."],"supporting_citations":[{"why":"It supplies the moral-choice scenario dataset and the action-likelihood evaluation method that the persuasion experiment builds on.","marker":"Scherrer et al. (2023)"},{"why":"It defines the rules of common morality used to label actions as rule-violating and to distinguish high-ambiguity from low-ambiguity scenarios.","marker":"Gert (2004)"},{"why":"It provides the Moral Foundations Theory questionnaire used to measure ethical alignment in the second experiment.","marker":"Graham et al. (2013)"},{"why":"It provides the alignment-prompting approach that the paper adapts by substituting moral principles for political orientations.","marker":"Abdulhai et al. (2023)"},{"why":"It supplies the semantic-equivalence method used to aggregate action likelihoods across differently worded responses.","marker":"Kuhn et al. (2023)"},{"why":"It offers prior evidence that persistent conversation can shift LLM beliefs, which motivates the persuasive-dialogue setup.","marker":"Xu et al. (2023)"}],"fun_headline_variants":["LLMs cave to moral persuasion when scenarios are vague","Chat can flip LLM ethics, but only in ambiguous cases","Model-to-model debate shifts LLM moral choices","Ambiguity makes LLMs swayable on moral calls","Utilitarian prompts skew LLM moral scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the observed decision flips are caused by the persuader's arguments, because the Base Agent is always told it had already chosen an action and there is no neutral-content control conversation to rule out commitment, sycophancy, or role-play pressure.","fun_headline_variants_meta":{"raw":{"variants":["LLMs cave to moral persuasion when scenarios are vague","Chat can flip LLM ethics, but only in ambiguous cases","Model-to-model debate shifts LLM moral choices","Ambiguity makes LLMs swayable on moral calls","Utilitarian prompts skew LLM moral scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000275,"raw_usage":{"total_tokens":1593,"prompt_tokens":843,"completion_tokens":750,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":683}},"tokens_in":459,"tokens_out":750,"duration_ms":7861,"temperature":1.0,"reasoning_tokens":683,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:12:11.628389+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same protocol with a control persuader that produces content-matched but argument-free messages, or with no persuasion at all except the instruction that the agent has chosen an action; if the Decision Change Rate in the control is close to the rate with real arguments, then the paper's attribution of the shifts to persuasion is wrong.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the moral-choice scenario dataset and the action-likelihood evaluation method that the persuasion experiment builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It defines the rules of common morality used to label actions as rule-violating and to distinguish high-ambiguity from low-ambiguity scenarios."},{"cited_title":"P., and Ditto, P","cited_arxiv_id":null,"evidence_quote":"It provides the Moral Foundations Theory questionnaire used to measure ethical alignment in the second experiment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the semantic-equivalence method used to aggregate action likelihoods across differently worded responses."}],"review_version":1}