{"id":"9568a644-cafc-4c54-ae86-8674f9d0cf6e","arxiv_id":"2506.15928","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Personality prompts change LLM agents' negotiation behavior and communication style, while AI traits such as transparency and adaptability have a smaller but measurable effect on interaction balance.","lead":"Researchers ran thousands of simulated negotiations between AI agents and tested whether giving the agents different personality descriptions changed how they talked and how well they achieved their goals. The work proposes an evaluation framework for checking whether AI assistants can adapt to people with different personality profiles in high-stakes settings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'strong evidence' claim is undercut by LLM-judge contamination: Sotopia-Eval and lexical scores may track prompt labels rather than dialogue behavior, and no human validation is reported.","rationale":"The paper is a substantive simulation study: 8,686 Experiment 1 transcripts, standardized Sotopia scenarios, and trait-level findings directionally consistent with prior personality and negotiation research. Those are real strengths. The load-bearing weakness is not the qualitative direction of the effects but the evidentiary standard: every outcome is produced by LLM-based tools that share priors with the generator and, for Sotopia-Eval, may directly see the manipulation in the character profile. No human validation, no independent behavioral outcome, and no evidence that the evaluator is immune to label leakage are provided. This is the same gap the reader identified, though I would phrase it more specifically as a circularity risk rather than a general 'LLM bias' concern. The causal-discovery analysis is also under-specified (no uncertainty estimates, no multiple-comparison control, and an unclear episode-count sentence in Section 4.1.3), but those issues would affect the quantitative confidence intervals, not the central qualitative claim. The right outcome is the reader's CONDITIONAL verdict, unchanged: the framework is promising and the results plausible, but the 'strong evidence' formulation should be downgraded unless the human-validation check confirms the LLM-judge pattern.","tokens_in":16839,"tokens_out":6477,"duration_ms":80130,"concrete_test":"Conduct a preregistered human-rating study on a stratified random sample of 150–200 Experiment 1 transcripts per trait condition (approximately 1,000–1,200 total). Show raters only the dialogue, with no character profiles or condition labels, and ask them to rate naturalness, empathy, and negotiation success on Likert scales mirroring Sotopia-Eval, plus a forced-choice item guessing the intended Big Five profile. If human ratings do not reproduce the reported trait-condition ordering, or if forced-choice attribution is at chance, the Sotopia-Eval and lexical effects are at least partly LLM-judge artifacts. Compare effect sizes and confidence intervals between human and LLM judges.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim—that personality prompts produce behavior 'consistent with established theoretical predictions' (General Discussion)—depends on Sotopia-Eval and lexical measures applied to LLM-generated dialogues. In this pipeline, the generator and evaluator share the same model family (gpt-4o-mini; Section 4.1.3), and Sotopia episodes include the character's full profile, including the personality prompt (Appendix A). An LLM judge can therefore infer the intended trait directly from the episode input, making Believability and Goal scores partially tautological: a 'high Agreeableness' prompt is rated as believable because the judge already knows the label. The lexical classifiers (DistilBERT-based empathy, emotion, and morality detectors; Table 2) are also prompt-sensitive and are not validated against human annotation. Section 7 acknowledges these limitations, but the General Discussion still calls the result 'strong evidence.' Without human ratings or an independent outcome (e.g., final price relative to target price), the observed trait-effect pattern is equally consistent with genuine personality-consistent behavior and with an evaluator that reproduces trait stereotypes from the prompt. This circularity is the load-bearing gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an evaluation framework for LLM-simulated negotiations. Experiment 1 manipulates Big Five personality prompts in a price-bargaining dialogue between two gpt-4o-mini agents and reports scenario-based (Sotopia-Eval), lexical, and causal analyses of trait effects. Experiment 2 manipulates AI hiring-manager transparency, competence, and adaptability alongside simulated human Agreeableness and Extraversion in a job-negotiation scenario, adding questionnaire measures and transactivity/verbal-equity outcomes. The authors conclude that Agreeableness and Extraversion significantly affect believability, goal achievement, and knowledge acquisition, that personality effects dominate AI-characteristic effects, and that the simulations reproduce established personality-negotiation findings.","tokens_in":17076,"tokens_out":5105,"duration_ms":57966,"significance":"If the causal and validity claims survive scrutiny, the proposed framework would be a scalable, controlled testbed for personality-aware evaluation of agentic AI, addressing a real gap beyond task-completion metrics. The paper is strong in its scope (thousands of simulated episodes across two negotiation settings), its ambition to move from correlation to causal analysis, and its candid acknowledgment of limitations, including prompt-based personality manipulations and missing non-verbal cues. However, two load-bearing issues currently prevent the advertised conclusions from being supported: the reported analyses are SEM weights that are never connected to the described causal-discovery pipeline, and the LLM-based evaluation pipeline is vulnerable to circularity because the same model family generates the dialogues and scores them with access to the personality labels.","major_comments":[{"comment":"The Methods describe causal discovery with CausalNex and causal-forest average treatment effect estimation, but the Results report only \"SEM Weights\" in all figures, with no description of the structural equation model, its specification, estimation method, or fit statistics. The paper never explains how these weights relate to the promised ATEs or to the \"when X increases, we see a decrease in Y\" statements in Section 4.1.4. Because the causal language in the Abstract and General Discussion depends on this analysis, the authors must either report the actual causal-forest ATEs with confidence intervals and intervention definitions or specify the SEM and justify that its coefficients estimate causal, rather than associational, effects.","section":"Section 4.1.4 and Figures 2–10"},{"comment":"Sotopia-Eval scores are produced by an LLM evaluator that appears to be applied to episodes whose input includes the full character profile, including explicit personality labels such as \"Personality Trait: Introversion\" (Appendix A). Since the same model family (gpt-4o-mini) generates the dialogues and scores them, Believability and Goal scores can reflect the prompt label rather than the actual dialogue behavior. This makes the General Discussion's \"strong evidence\" claim (Section 6) vulnerable to circularity. The paper needs to show that the evaluator is blind to the trait labels (for example, by ablating the profile from the evaluation input), or corroborate the scenario-based measures with human judgments, or substantially temper the causal interpretation.","section":"Section 3.2.1 and Appendix A"},{"comment":"The lexical measures (DistilBERT-based empathy, emotion, morality, and subjectivity classifiers) are applied to LLM-generated negotiation dialogues without any validation against human-annotated texts of this type. The observed trait effects on lexical outcome variables could therefore be artifacts of classifier sensitivity to prompt phrasing or genre-specific language rather than genuine personality-consistent communication differences. Please provide validation evidence, at least on a held-out sample of the generated dialogues, or explicitly reframe the lexical results as exploratory rather than confirmatory evidence.","section":"Section 4.1.2 and Table 2"},{"comment":"The General Discussion describes the results as \"strong evidence\" for theory-consistent personality effects, but Section 7 itself acknowledges that prompt-based personality manipulations may not capture the complexity of real human personality, that lexical measures miss non-verbal cues, and that only two scenarios were considered. Combined with the evaluator-circularity and analysis-reporting issues above, the strength of the claim goes beyond what the evidence supports. The conclusions should be recast as proof-of-concept results that require human validation and an evaluator-blindness check.","section":"Section 6 and Section 7"}],"minor_comments":[{"comment":"The sentence \"running a total of 4343 episodes for each treatment combination setting\" is inconsistent with the reported total of 8686 transcripts; please clarify whether 4343 is the total number of episodes or the number per condition and reconcile the transcript count.","section":"Section 4.1.3"},{"comment":"The three AI dimensions are named \"Transparency, Adaptability, and Reliability\" in Section 5.1.1, but Appendix C, the Abstract, and all Results figures refer to \"Competence\" rather than \"Reliability.\" Please use one consistent terminology throughout.","section":"Section 5.1.1"},{"comment":"The text says \"seven original dimensions\" and then lists eight measures: Believability, Financial and Material Benefits, Goal, Knowledge, Overall Score, Relationship, Secret, and Social Rules. Clarify which dimensions are the seven original Sotopia-Eval dimensions and how Overall Score is defined.","section":"Section 4.1.2"},{"comment":"The phrase \"Smaller, positive trait level differences were found for Conscientiousness and Conscientiousness and Openness\" appears to be a typo; it should likely read \"for Conscientiousness and Openness.\"","section":"Section 4.2.5"},{"comment":"The text references \"Figure 8\" for the empathy measures, but the empathy measures are displayed in Figure 7; the figure cross-reference needs correction.","section":"Section 5.2.3"},{"comment":"The caption misspells \"Anticipating\" as \"Ancitipating.\"","section":"Figure 3a caption"},{"comment":"The clause \"which were reversed for for Love and Joy indicators\" contains a duplicated \"for.\"","section":"Section 4.2.4"}],"recommendation":"major_revision","confidential_remarks":"The central finding is plausible and consistent with prior work, but the manuscript in its current form does not supply the evidence needed to distinguish genuine simulated personality effects from evaluator label-reading. A revision with an evaluator-blindness analysis or human validation, plus a full description of the causal estimation, could make this a strong contribution. The scope fit for this journal is acceptable given the emphasis on agentic AI evaluation, but the defense-operations framing is not essential to the technical core."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zoe,\n\nHere's my read. This is a useful engineering contribution wrapped in overstated causal claims. If you treat it as 'we built a high-throughput simulation pipeline and got trait effects that line up with the human literature,' it's a decent paper. If you treat it as 'strong evidence that personality prompts cause LLM agents to behave in theory-consistent ways,' the evidence doesn't support it.\n\nWhat's genuinely new: they extend Huang and Hadfi (2024) by adding CausalNex/causal forest ATE estimation, a much broader lexical battery (empathy, moral foundations, connotation frames, subjectivity), and a human-AI negotiation arm manipulating AI transparency, competence, and adaptability. The scale is large—8,686 episodes in Experiment 1. The direction of effects is reassuring: agreeableness and extraversion are associated with prosocial, engaged behavior; neuroticism with worse outcomes. That internal consistency with prior findings is the paper's best evidence.\n\nSoft spots, in proportion. First, the causal analysis is under-specified. The methods say CausalNex and causal forests, but the results show 'SEM Weights' with no fitted structural equation model, no ATEs, no confidence intervals. As reported, the causal claims are not checkable. Second and more serious, the evaluation is contaminated. Sotopia-Eval is an LLM judge that sees the full episode, including the character's personality prompt (Appendix A). So a 'high Agreeableness' agent gets rated as more believable partly because the judge already knows the label. The lexical measures (DistilBERT-based empathy, morality, emotion classifiers) are similarly unvalidated on LLM-generated negotiation text. The paper acknowledges these issues in Section 7 but then asks the reader to accept 'strong evidence' in the General Discussion. That's overreach.\n\nThe directional findings are probably right—they match a century of personality-negotiation research—so this isn't a fatal flaw. But it means the paper hasn't demonstrated what it claims. The fix is straightforward: add human ratings on a sample of dialogues, report at least one outcome that doesn't depend on the LLM judge (e.g., final price relative to target in Experiment 1), and release code/data.\n\nWho should read this: people building LLM social-simulation testbeds and defense/HCI folks who want a template for pre-deployment agent evaluation. It deserves a careful referee, but it needs revision first.\n\nRecommendation: send to peer review, but with the clear expectation that the causal language gets toned down, the judge's prompt access gets disclosed, and ideally a human-validation substudy is added. I wouldn't cite the quantitative estimates, but I'd consider citing the framework once the pipeline is cleaned up.\n\nBest,\n[Name]","headline":"Useful framework, solid directional findings, but the causal and evidential claims outrun the evaluation design.","tokens_in":17589,"tokens_out":3867,"would_cite":false,"duration_ms":42942,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM agents prompted with Big Five personality traits produce negotiation behavior that matches established personality-psychology predictions, the paper argues.","keywords":["LLM simulation","Big Five personality","negotiation","causal discovery","Sotopia","human-AI teaming","lexical analysis","agentic AI"],"falsifier":"Take a random sample of the generated negotiation transcripts and have independent human raters score them on believability, goal achievement, and empathy without knowing which trait prompt produced them. If the human ratings do not show the same Agreeableness and Extraversion differences reported by the LLM-based Sotopia-Eval and lexical scores, the central claim that prompt-based traits produce theory-consistent simulated behavior would fail. A quicker check is to re-score the same transcripts with a different LLM evaluator and see whether the trait effects survive the change.","tokens_in":16668,"feed_emoji":"🤝","tokens_out":5309,"duration_ms":50717,"temperature":0.7,"pith_summary":"This paper argues that prompting large language model agents with Big Five personality descriptions produces negotiation behavior that tracks established personality psychology: higher Agreeableness and Extraversion reliably shift believability, goal achievement, and knowledge acquisition, while Neuroticism moves outcomes in the opposite direction. The claim is established through two sets of Sotopia simulations—buyer-seller price bargaining and a human-AI job negotiation—analyzed with causal discovery methods that estimate the effect of trait levels on outcomes. The authors also show that lexical markers of empathy, moral foundations, and connotative language shift with trait levels in ways consistent with the human literature. The point of the exercise is practical: if prompt-based personality can steer agent behavior this reliably, then mission-critical AI systems can be stress-tested against diverse operator personalities before deployment.","feed_headline":"LLM agents act out Big Five traits in negotiations","feed_subtitle":"Two simulated experiments show Agreeableness and Extraversion shift believability, empathy, and deal outcomes.","key_machinery":"The machinery is the Sotopia simulation testbed (an LLM-based framework in which two agents converse in pursuit of private social goals) combined with prompt-based trait manipulations drawn from Big Five Inventory items. Outcomes are scored by Sotopia-Eval dimensions such as believability, goal achievement, and knowledge acquisition, and by a suite of lexical analytics for empathy, moral foundations, sentiment, toxicity, connotation frames, and subjectivity. Causal discovery—structural learning via CausalNex followed by average treatment effect estimation with Causal Forests—is what converts the simulated dialogue corpus into the paper's causal claims about trait levels. The same pipeline is applied to Experiment 2 with the addition of AI-agent trait prompts (transparency, competence, adaptability) and post-interaction questionnaires.","core_discovery":"The central claim is that Big Five personality traits, implemented as prompt text derived from BFI questionnaire items, are a valid and controllable lever on LLM-simulated social behavior. Experiment 1 uses 8,686 price-bargaining transcripts between gpt-4o-mini agents and causal discovery (CausalNex DAGs plus Causal Forest average treatment effects) to show that Agreeableness and Extraversion significantly influence Sotopia-Eval scores for believability, goal achievement, and knowledge acquisition, with Neuroticism associated with worse outcomes. Experiment 2 extends the same logic to human-AI negotiation, where simulated human candidates' Agreeableness and Extraversion dominate AI system transparency, competence, and adaptability in shaping questionnaire and lexical measures; AI traits mainly affect conversational balance (transactivity, verbal equity). The authors conclude that LLM-driven social simulation can serve as a valid platform for studying personality-driven negotiation dynamics and for pre-deployment testing of agentic AI.","pith_inferences":["The paper does not test whether the LLM evaluator's scores agree with human judgments; a natural next step is a human-rater study on the same transcripts, which could confirm or overturn the validity of the Sotopia-Eval and lexical measures.","Because both the actors and the evaluators are LLMs, part of the observed trait 'effects' could come from prompt phrase matching rather than genuine behavioral simulation; comparing across different evaluator models would probe this.","The causal-discovery framing implies intervention on trait levels in the prompt, but the prompt is the only intervention; the paper's ATEs should be read as effects of prompt text on output text, not of human personality on negotiation outcomes.","A direct extension would be to run the same two scenarios with a different base LLM (e.g., a smaller open-weight model) and see whether the trait-effect patterns replicate; if they do not, the framework's generality is limited."],"forward_implications":["If prompt-based Big Five manipulations work as claimed, LLM simulations become a scalable substitute for human-subject experiments when probing personality effects on negotiation.","AI agents deployed in high-stakes settings should be evaluated per operator personality profile, since the Agreeableness and Extraversion of the human side dominate the interaction measures.","The dominance of personality over AI characteristics implies that training and design should prioritize personality-aware communication strategies over purely technical AI enhancements.","The framework offers a repeatable pre-deployment test: run the agent across a spectrum of simulated personalities and inspect the causal effects before fielding it.","The lexical measures (empathy, moral foundations, connotation framing) provide an actionable diagnostic for when an agent's dialogue drifts from expected trait-consistent behavior."],"supporting_citations":[{"why":"Supplies the Sotopia simulation testbed and the Sotopia-Eval dimensions used in both experiments.","marker":"[1]"},{"why":"Prior LLM negotiation simulation showing trait-consistent behavior; the direct precedent the paper extends.","marker":"[8]"},{"why":"Provides the CausalNex causal discovery tool used to learn DAGs from simulation outputs.","marker":"[10]"},{"why":"Provides the generalized random forest method behind the Causal Forest average treatment effect estimates.","marker":"[11]"},{"why":"Defines the Big Five five-factor model that the prompt-based trait operationalization draws on.","marker":"[12]"},{"why":"Empirical personality-negotiation findings used as the theoretical benchmark for Experiment 1 results.","marker":"[31]"},{"why":"Supplies the connotation-frame lexical measure used to detect subtle perspective and sentiment in dialogue.","marker":"[42]"},{"why":"Source of the fictitious Craigslist bargaining scenarios used in Experiment 1.","marker":"[61]"}],"fun_headline_variants":["Big Five traits steer LLM negotiation outcomes","Agreeableness and extraversion move LLM deals","Personality beats AI features in LLM negotiation","LLM negotiators follow Big Five, not AI settings","Big Five personality shapes LLM-simulated talks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The scores that measure whether a simulated negotiation was believable, successful, or empathic come from LLM-based tools reading LLM-generated dialogue, with no human raters checking that those scores match real human judgments.","fun_headline_variants_meta":{"raw":{"variants":["Big Five traits steer LLM negotiation outcomes","Agreeableness and extraversion move LLM deals","Personality beats AI features in LLM negotiation","LLM negotiators follow Big Five, not AI settings","Big Five personality shapes LLM-simulated talks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000427,"raw_usage":{"total_tokens":2205,"prompt_tokens":985,"completion_tokens":1220,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":1146}},"tokens_in":601,"tokens_out":1220,"duration_ms":9045,"temperature":1.0,"reasoning_tokens":1146,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:44:24.874432+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the generated negotiation transcripts and have independent human raters score them on believability, goal achievement, and empathy without knowing which trait prompt produced them. If the human ratings do not show the same Agreeableness and Extraversion differences reported by the LLM-based Sotopia-Eval and lexical scores, the central claim that prompt-based traits produce theory-consistent simulated behavior would fail. A quicker check is to re-score the same transcripts with a different LLM evaluator and see whether the trait effects survive the change.","supporting_citations":[{"cited_title":"SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents, March 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the Sotopia simulation testbed and the Sotopia-Eval dimensions used in both experiments."},{"cited_title":"How Personality Traits Influence Negotiation Out- comes? A Simulation based on Large Language Models","cited_arxiv_id":null,"evidence_quote":"Prior LLM negotiation simulation showing trait-consistent behavior; the direct precedent the paper extends."},{"cited_title":"CausalNex, October 2021","cited_arxiv_id":null,"evidence_quote":"Provides the CausalNex causal discovery tool used to learn DAGs from simulation outputs."},{"cited_title":"Generalized random forests","cited_arxiv_id":null,"evidence_quote":"Provides the generalized random forest method behind the Causal Forest average treatment effect estimates."},{"cited_title":"McCrae and Oliver P","cited_arxiv_id":null,"evidence_quote":"Defines the Big Five five-factor model that the prompt-based trait operationalization draws on."},{"cited_title":"Bargainer characteristics in distributive and integrative negotiation","cited_arxiv_id":null,"evidence_quote":"Empirical personality-negotiation findings used as the theoretical benchmark for Experiment 1 results."},{"cited_title":"Decoupling strategy and generation in negotiation dialogues, 2018","cited_arxiv_id":null,"evidence_quote":"Source of the fictitious Craigslist bargaining scenarios used in Experiment 1."}],"review_version":1}