REVIEW 4 major objections 7 minor 1 cited by
Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues
T0 review · 4 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read LLM agents prompted with Big Five personality traits produce negotiation behavior that matches established personality-psychology predictions, the paper argues.
desk verdict Useful framework, solid directional findings, but the causal and evidential claims outrun the evaluation design. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the Sotopia simulation testbed (an LLM-based framework in which two agents converse in pursuit of private social goals) combined with prompt-based trait manipulations drawn from Big Five Inventory items. Outcomes are scored by Sotopia-Eval dimensions such as believability, goal achievement, and knowledge acquisition, and by a suite of lexical analytics for empathy, moral foundations, sentiment, toxicity, connotation frames, and subjectivity. Causal discovery—structural learning via CausalNex followed by average treatment effect estimation with Causal Forests—is what converts the simulated dialogue corpus into the paper's causal claims about trait levels. The same pipeline is applied to Experiment 2 with the addition of AI-agent trait prompts (transparency, competence, adaptability) and post-interaction questionnaires.
What would settle it
Take a random sample of the generated negotiation transcripts and have independent human raters score them on believability, goal achievement, and empathy without knowing which trait prompt produced them. If the human ratings do not show the same Agreeableness and Extraversion differences reported by the LLM-based Sotopia-Eval and lexical scores, the central claim that prompt-based traits produce theory-consistent simulated behavior would fail. A quicker check is to re-score the same transcripts with a different LLM evaluator and see whether the trait effects survive the change.
Extended reading notes
Core claim
The central claim is that Big Five personality traits, implemented as prompt text derived from BFI questionnaire items, are a valid and controllable lever on LLM-simulated social behavior. Experiment 1 uses 8,686 price-bargaining transcripts between gpt-4o-mini agents and causal discovery (CausalNex DAGs plus Causal Forest average treatment effects) to show that Agreeableness and Extraversion significantly influence Sotopia-Eval scores for believability, goal achievement, and knowledge acquisition, with Neuroticism associated with worse outcomes. Experiment 2 extends the same logic to human-AI negotiation, where simulated human candidates' Agreeableness and Extraversion dominate AI system transparency, competence, and adaptability in shaping questionnaire and lexical measures; AI traits mainly affect conversational balance (transactivity, verbal equity). The authors conclude that LLM-driven social simulation can serve as a valid platform for studying personality-driven negotiation dynamics and for pre-deployment testing of agentic AI.
Load-bearing premise
The scores that measure whether a simulated negotiation was believable, successful, or empathic come from LLM-based tools reading LLM-generated dialogue, with no human raters checking that those scores match real human judgments.
Editorial extensions
If this is right
- If prompt-based Big Five manipulations work as claimed, LLM simulations become a scalable substitute for human-subject experiments when probing personality effects on negotiation.
- AI agents deployed in high-stakes settings should be evaluated per operator personality profile, since the Agreeableness and Extraversion of the human side dominate the interaction measures.
- The dominance of personality over AI characteristics implies that training and design should prioritize personality-aware communication strategies over purely technical AI enhancements.
- The framework offers a repeatable pre-deployment test: run the agent across a spectrum of simulated personalities and inspect the causal effects before fielding it.
- The lexical measures (empathy, moral foundations, connotation framing) provide an actionable diagnostic for when an agent's dialogue drifts from expected trait-consistent behavior.
Reading between the lines
- The paper does not test whether the LLM evaluator's scores agree with human judgments; a natural next step is a human-rater study on the same transcripts, which could confirm or overturn the validity of the Sotopia-Eval and lexical measures.
- Because both the actors and the evaluators are LLMs, part of the observed trait 'effects' could come from prompt phrase matching rather than genuine behavioral simulation; comparing across different evaluator models would probe this.
- The causal-discovery framing implies intervention on trait levels in the prompt, but the prompt is the only intervention; the paper's ATEs should be read as effects of prompt text on output text, not of human personality on negotiation outcomes.
- A direct extension would be to run the same two scenarios with a different base LLM (e.g., a smaller open-weight model) and see whether the trait-effect patterns replicate; if they do not, the framework's generality is limited.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an evaluation framework for LLM-simulated negotiations. Experiment 1 manipulates Big Five personality prompts in a price-bargaining dialogue between two gpt-4o-mini agents and reports scenario-based (Sotopia-Eval), lexical, and causal analyses of trait effects. Experiment 2 manipulates AI hiring-manager transparency, competence, and adaptability alongside simulated human Agreeableness and Extraversion in a job-negotiation scenario, adding questionnaire measures and transactivity/verbal-equity outcomes. The authors conclude that Agreeableness and Extraversion significantly affect believability, goal achievement, and knowledge acquisition, that personality effects dominate AI-characteristic effects, and that the simulations reproduce established personality-negotiation findings.
Significance. If the causal and validity claims survive scrutiny, the proposed framework would be a scalable, controlled testbed for personality-aware evaluation of agentic AI, addressing a real gap beyond task-completion metrics. The paper is strong in its scope (thousands of simulated episodes across two negotiation settings), its ambition to move from correlation to causal analysis, and its candid acknowledgment of limitations, including prompt-based personality manipulations and missing non-verbal cues. However, two load-bearing issues currently prevent the advertised conclusions from being supported: the reported analyses are SEM weights that are never connected to the described causal-discovery pipeline, and the LLM-based evaluation pipeline is vulnerable to circularity because the same model family generates the dialogues and scores them with access to the personality labels.
major comments (4)
- [Section 4.1.4 and Figures 2–10] The Methods describe causal discovery with CausalNex and causal-forest average treatment effect estimation, but the Results report only "SEM Weights" in all figures, with no description of the structural equation model, its specification, estimation method, or fit statistics. The paper never explains how these weights relate to the promised ATEs or to the "when X increases, we see a decrease in Y" statements in Section 4.1.4. Because the causal language in the Abstract and General Discussion depends on this analysis, the authors must either report the actual causal-forest ATEs with confidence intervals and intervention definitions or specify the SEM and justify that its coefficients estimate causal, rather than associational, effects.
- [Section 3.2.1 and Appendix A] Sotopia-Eval scores are produced by an LLM evaluator that appears to be applied to episodes whose input includes the full character profile, including explicit personality labels such as "Personality Trait: Introversion" (Appendix A). Since the same model family (gpt-4o-mini) generates the dialogues and scores them, Believability and Goal scores can reflect the prompt label rather than the actual dialogue behavior. This makes the General Discussion's "strong evidence" claim (Section 6) vulnerable to circularity. The paper needs to show that the evaluator is blind to the trait labels (for example, by ablating the profile from the evaluation input), or corroborate the scenario-based measures with human judgments, or substantially temper the causal interpretation.
- [Section 4.1.2 and Table 2] The lexical measures (DistilBERT-based empathy, emotion, morality, and subjectivity classifiers) are applied to LLM-generated negotiation dialogues without any validation against human-annotated texts of this type. The observed trait effects on lexical outcome variables could therefore be artifacts of classifier sensitivity to prompt phrasing or genre-specific language rather than genuine personality-consistent communication differences. Please provide validation evidence, at least on a held-out sample of the generated dialogues, or explicitly reframe the lexical results as exploratory rather than confirmatory evidence.
- [Section 6 and Section 7] The General Discussion describes the results as "strong evidence" for theory-consistent personality effects, but Section 7 itself acknowledges that prompt-based personality manipulations may not capture the complexity of real human personality, that lexical measures miss non-verbal cues, and that only two scenarios were considered. Combined with the evaluator-circularity and analysis-reporting issues above, the strength of the claim goes beyond what the evidence supports. The conclusions should be recast as proof-of-concept results that require human validation and an evaluator-blindness check.
minor comments (7)
- [Section 4.1.3] The sentence "running a total of 4343 episodes for each treatment combination setting" is inconsistent with the reported total of 8686 transcripts; please clarify whether 4343 is the total number of episodes or the number per condition and reconcile the transcript count.
- [Section 5.1.1] The three AI dimensions are named "Transparency, Adaptability, and Reliability" in Section 5.1.1, but Appendix C, the Abstract, and all Results figures refer to "Competence" rather than "Reliability." Please use one consistent terminology throughout.
- [Section 4.1.2] The text says "seven original dimensions" and then lists eight measures: Believability, Financial and Material Benefits, Goal, Knowledge, Overall Score, Relationship, Secret, and Social Rules. Clarify which dimensions are the seven original Sotopia-Eval dimensions and how Overall Score is defined.
- [Section 4.2.5] The phrase "Smaller, positive trait level differences were found for Conscientiousness and Conscientiousness and Openness" appears to be a typo; it should likely read "for Conscientiousness and Openness."
- [Section 5.2.3] The text references "Figure 8" for the empathy measures, but the empathy measures are displayed in Figure 7; the figure cross-reference needs correction.
- [Figure 3a caption] The caption misspells "Anticipating" as "Ancitipating."
- [Section 4.2.4] The clause "which were reversed for for Love and Joy indicators" contains a duplicated "for."
Circularity Check
Believability dimension is self-referential because the character profile contains the personality prompt; Goal and lexical measures retain independent content.
-
self definitional
[Table 1 / Sec. 3.2.1 and Appendix A; interpreted in Sec. 6]
"Believability: How natural, realistic, and consistent the agent’s behavior is with its character profile [0, 10]. ... "personality_and_values": Personality Model: Big 5 Personality Personality Trait: Introversion Task Assignment: Prefers independent tasks and may struggle with collaboration."
The manipulated independent variable is the Big Five trait written into Sotopia's character profile (Appendix A). Sotopia-Eval's Believability dimension is explicitly defined as consistency with that same character profile. Thus, when the trait prompt changes the profile, an agent that follows its own profile will be rated more believable by the definition of the scale; the trait-level-to-Believability association is partially secured by the measurement definition rather than by an independent behavioral outcome. The General Discussion nonetheless cites Believability as part of the 'strong evidence' that trait prompts produce theory-consistent behavior.
full rationale
The paper's strongest claim—that Experiment 1 provides 'strong evidence' that personality prompts produce behavior consistent with personality and negotiation theory—depends on multiple measures. One of those measures, Believability, is self-referential: Sotopia-Eval defines it as consistency with the character profile, and the character profile is exactly where the personality manipulation is inserted. This makes high Believability under a high-trait prompt partly an artifact of the evaluation definition. However, the same conclusion is also supported by Goal, Knowledge, and a suite of external lexical classifiers (empathy, moral foundations, sentiment, connotation frames), which do not reduce to the trait label. The Sotopia-Eval framework is a self-citation ([1] shares an author with this paper), but it is a published, code-released framework and therefore counts as independent support under the code-reproduced criterion. Section 7 honestly acknowledges the absence of human validation and the limitation that prompt-based manipulations may not capture human complexity; those are validity concerns, not circular reductions. Because one headline outcome reduces by construction while the central claim retains independent content, a moderate partial-circularity score is appropriate.
Assumptions & free parameters
assumptions (4)
- domain assumption Sotopia-Eval scores (believability, goal achievement, knowledge) are valid measures of negotiation quality and social outcomes.
- domain assumption Prompt-based Big Five descriptions induce the intended personality traits in LLM agents.
- standard math Causal discovery and causal forest estimates identify true causal effects, requiring correct DAG specification and no unobserved confounding.
- domain assumption LLM-generated dialogue is a faithful proxy for human negotiation communication.
invented entities (1)
-
Human digital twin (HDT) job candidate
Cite this review
Pith. "Pith review of Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues." pith.science (2026). https://pith.science/paper/2ODSS3JV
@misc{pith2026250615928,
author = {Pith},
title = {Pith review of: Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues},
year = {2026},
howpublished = {\url{https://pith.science/paper/2ODSS3JV}},
note = {Machine review of arXiv:2506.15928}
}
read the original abstract
This paper presents an evaluation framework for agentic AI systems in mission-critical negotiation contexts, addressing the need for AI agents that can adapt to diverse human operators and stakeholders. Using Sotopia as a simulation testbed, we present two experiments that systematically evaluated how personality traits and AI agent characteristics influence LLM-simulated social negotiation outcomes--a capability essential for a variety of applications involving cross-team coordination and civil-military interactions. Experiment 1 employs causal discovery methods to measure how personality traits impact price bargaining negotiations, through which we found that Agreeableness and Extraversion significantly affect believability, goal achievement, and knowledge acquisition outcomes. Sociocognitive lexical measures extracted from team communications detected fine-grained differences in agents' empathic communication, moral foundations, and opinion patterns, providing actionable insights for agentic AI systems that must operate reliably in high-stakes operational scenarios. Experiment 2 evaluates human-AI job negotiations by manipulating both simulated human personality and AI system characteristics, specifically transparency, competence, adaptability, demonstrating how AI agent trustworthiness impact mission effectiveness. These findings establish a repeatable evaluation methodology for experimenting with AI agent reliability across diverse operator personalities and human-agent team dynamics, directly supporting operational requirements for reliable AI systems. Our work advances the evaluation of agentic AI workflows by moving beyond standard performance metrics to incorporate social dynamics essential for mission success in complex operations.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
The Language of Bargaining: Linguistic Effects in LLM Negotiations
Language labels shift LLM negotiation outcomes in three games, but the claimed dominance over model choice is inconsistent with the reported model-level gaps.
Reference graph
Works this paper leans on
-
[1]
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents, March 2024
Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents, March 2024
work page 2024
-
[2]
Cohen, Grant Engber- son, Laura Cassani, Trenton W
Svitlana Volkova, Daniel Nguyen, Hsien-Te Kao, Myke C. Cohen, Grant Engber- son, Laura Cassani, Trenton W. Ford, Michael G. Yankoski, Mohammed Almu- tairi, Charles Chiang, Nandini Banerjee, Matthew Belcher, Tim Weninger, and Diego Gomez-Zara. VirTLab: Augmented Intelligence for Modeling and Evaluating Human-AI Teaming through Agent Interactions. in press
-
[3]
Cohen, Hsien-Te Kao, Grant Engberson, Louis Penafiel, Spencer Lynch, and Svitlana Volkova
Daniel Nguyen, Myke C. Cohen, Hsien-Te Kao, Grant Engberson, Louis Penafiel, Spencer Lynch, and Svitlana Volkova. Exploratory Models of Human-AI Teams: Leveraging Human Digital Twins to Investigate Trust Development, November 2024
work page 2024
-
[4]
Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, and Ziran Wang. Human-Autonomy Teaming on Autonomous Vehicles with Large Language Model-Enabled Human Dig- ital Twins. In 2023 IEEE/ACM Symposium on Edge Computing (SEC) , pages 319– 324, December 2023
work page 2023
-
[5]
An LLM-Based Digital Twin for Optimizing Human-in-the Loop Systems, March 2024
Hanqing Yang, Marie Siew, and Carlee Joe-Wong. An LLM-Based Digital Twin for Optimizing Human-in-the Loop Systems, March 2024
work page 2024
-
[6]
Rhyse Bendell, Jessica Williams, Stephen M. Fiore, and Florian Jentsch. Individual and team profiling to support theory of mind in artificial social intelligence.Scientific Reports, 14(1):12635, June 2024
work page 2024
- [7]
-
[8]
How Personality Traits Influence Negotiation Out- comes? A Simulation based on Large Language Models
Yin Jou Huang and Rafik Hadfi. How Personality Traits Influence Negotiation Out- comes? A Simulation based on Large Language Models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024 , pages 10336–10351, Miami, Florida, USA, November
work page 2024
Show all 79 references
-
[9]
Shadish, Thomas D
William R. Shadish, Thomas D. Cook, and Donald T. Campbell. Experimental and Quasi-Experimental Designs for Generalized Causal Inference . Cengage Learning, Belmont, CA, 2nd edition edition, January 2001
2001
-
[10]
CausalNex, October 2021
Paul Beaumont, Ben Horsburgh, Philip Pilgerstorfer, Angel Droth, Richard Oen- taryo, Steven Ler, Hiep Nguyen, Gabriel Azevedo Ferreira, Zain Patel, and Wesley Leong. CausalNex, October 2021
2021
-
[11]
Generalized random forests
Susan Athey, Julie Tibshirani, and Stefan Wager. Generalized random forests. The Annals of Statistics , 47(2):1148–1178, April 2019. 19
2019
-
[12]
McCrae and Oliver P
Robert R. McCrae and Oliver P. John. An Introduction to the Five-Factor Model and Its Applications. Journal of Personality , 60(2):175–215, June 1992
1992
-
[13]
Trait-names: A psycho-lexical study
Gordon W Allport and Henry S Odbert. Trait-names: A psycho-lexical study. Psy- chological monographs, 47(1):i, 1936
1936
-
[14]
Description and measurement of personality
Raymond Bernard Cattell. Description and measurement of personality. 1946
1946
-
[15]
Consistency of the factorial structures of personality ratings from different sources
Donald W Fiske. Consistency of the factorial structures of personality ratings from different sources. The Journal of Abnormal and Social Psychology , 44(3):329, 1949
1949
-
[16]
E. C. Tupes and R. E. Christal. Recurrent personality factors based on trait ratings. USAF ASD Tech. Rep. No. 61-97, US Air Force, Lackland Air Force Base, TX, 1961
1961
-
[17]
Toward an adequate taxonomy of personality attributes: Repli- cated factor structure in peer nomination personality ratings
Warren T Norman. Toward an adequate taxonomy of personality attributes: Repli- cated factor structure in peer nomination personality ratings. The journal of abnor- mal and social psychology , 66(6):574, 1963
1963
-
[18]
The revised neo personality inventory (neo-pi- r)
Paul T Costa and Robert R McCrae. The revised neo personality inventory (neo-pi- r). The SAGE handbook of personality theory and assessment , 2(2):179–198, 2008
2008
-
[19]
LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models, February 2024
Ivar Frisch and Mario Giulianelli. LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models, February 2024
2024
-
[20]
Estimating the Personality of White-Box Language Models, May 2023
Saketh Reddy Karra, Son The Nguyen, and Theja Tulabandhula. Estimating the Personality of White-Box Language Models, May 2023
2023
-
[21]
Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael R. Lyu. Who is ChatGPT? Bench- marking LLMs’ Psychological Portrayal Using PsychoBench, January 2024
2024
-
[22]
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits, April 2024
Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kab- bara. PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits, April 2024
2024
-
[23]
The Power of Personality: A Human Simulation Perspective to Investigate Large Language Model Agents, February 2025
Yifan Duan, Yihong Tang, Xuefeng Bai, Kehai Chen, Juntao Li, and Min Zhang. The Power of Personality: A Human Simulation Perspective to Investigate Large Language Model Agents, February 2025
2025
-
[24]
Is Self-knowledge and Action Consistent or Not: Investigating Large Language Model’s Personality, December 2024
Yiming Ai, Zhiwei He, Ziyin Zhang, Wenhong Zhu, Hongkun Hao, Kai Yu, Lingjun Chen, and Rui Wang. Is Self-knowledge and Action Consistent or Not: Investigating Large Language Model’s Personality, December 2024
2024
-
[25]
Petrov, Gregory Serapio-Garc ´ ıa, and Jason Rentfrow
Nikolay B. Petrov, Gregory Serapio-Garc ´ ıa, and Jason Rentfrow. Limited Ability of LLMs to Simulate Human Psychological Behaviours: A Psychometric Analysis, May 2024
2024
-
[26]
Exploring the Potential of Large Language Models to Simu- late Personality, February 2025
Maria Molchanova, Anna Mikhailova, Anna Korzanova, Lidiia Ostyakova, and Alexandra Dolidze. Exploring the Potential of Large Language Models to Simu- late Personality, February 2025
2025
-
[27]
The Art and Science of Negotiation
Howard Raiffa. The Art and Science of Negotiation . Harvard University Press, 1982. 20
1982
-
[28]
The role of personality in successful negotiating
Roderick W Gilkey and Leonard Greenhalgh. The role of personality in successful negotiating. Negotiation Journal, 2(3):245–256, 1986
1986
-
[29]
Personality and Negotiation Performance: The People Matter, January 2015
Mary Sass and Matthew Liao-Troth. Personality and Negotiation Performance: The People Matter, January 2015
2015
-
[30]
Gerui (Grace) Kang, Lin Xiu, and Alan C. Roline. How do interviewers respond to applicants’ initiation of salary negotiation? An exploratory study on the role of gender and personality. Evidence-based HRM: a Global Forum for Empirical Schol- arship, 3(2):145–158, August 2015
2015
-
[31]
Bargainer characteristics in distributive and integrative negotiation
Bruce Barry and Raymond A Friedman. Bargainer characteristics in distributive and integrative negotiation. Journal of personality and social psychology , 74(2):345, 1998
1998
-
[32]
Amanatullah, Michael W
Emily T. Amanatullah, Michael W. Morris, and Jared R. Curhan. Negotiators who give too much: Unmitigated communion, relational anxieties, and economic costs in distributive and integrative bargaining. Journal of Personality and Social Psychology, 95(3):723–738, 2008
2008
-
[33]
Team coordination dynamics
Jamie C Gorman, Polemnia G Amazeen, and Nancy J Cooke. Team coordination dynamics. Nonlinear dynamics, psychology, and life sciences , 14(3):265–289, July 2010
2010
-
[34]
Amazeen, Nathan J
Mustafa Demir, Polomnia G. Amazeen, Nathan J. McNeese, Aaron Likens, and Nancy J. Cooke. Team Coordination Dynamics in Human-Autonomy Teaming. Pro- ceedings of the Human Factors and Ergonomics Society Annual Meeting , 61(1):236– 236, September 2017
2017
-
[35]
Big Five personality traits in simulated negotiation settings
Pedro Fontes Falc˜ ao, Manuel Saraiva, Eduardo Santos, and Miguel Pina E Cunha. Big Five personality traits in simulated negotiation settings. EuroMed Journal of Business, 13(2):201–213, July 2018
2018
-
[36]
Pennebaker and Laura A
James W. Pennebaker and Laura A. King. Linguistic styles: Language use as an individual difference. Journal of Personality and Social Psychology, 77(6):1296–1312, 1999
1999
-
[37]
Tausczik and James W
Yla R. Tausczik and James W. Pennebaker. The Psychological Meaning of Words: LIWC and Computerized Text Analysis Methods. Journal of Language and Social Psychology, 29(1):24–54, March 2010
2010
-
[38]
Towards emotion-aware agents for improved user satisfaction and partner perception in negotiation dialogues
Kushal Chawla, Rene Clever, Jaysa Ramirez, Gale M Lucas, and Jonathan Gratch. Towards emotion-aware agents for improved user satisfaction and partner perception in negotiation dialogues. IEEE Transactions on Affective Computing , 2023
2023
-
[39]
Graziano, Meara M
William G. Graziano, Meara M. Habashi, Brad E. Sheese, and Ren´ ee M. Tobin. Agreeableness, empathy, and helping: A person × situation perspective. Journal of Personality and Social Psychology , 93(4):583–599, 2007
2007
-
[40]
Does GPT-3 Generate Empa- thetic Dialogues? A Novel In-Context Example Selection Method and Automatic 21 Evaluation Metric for Empathetic Dialogue Generation
Young-Jun Lee, Chae-Gyun Lim, and Ho-Jin Choi. Does GPT-3 Generate Empa- thetic Dialogues? A Novel In-Context Example Selection Method and Automatic 21 Evaluation Metric for Empathetic Dialogue Generation. In Nicoletta Calzolari, Chu- Ren Huang, Hansaem Kim, James Pustejovsky,...
2022
-
[41]
Morality between the lines: Detecting moral sentiment in text
Justin Garten, Reihane Boghrati, Joe Hoover, Kate M Johnson, and Morteza De- hghani. Morality between the lines: Detecting moral sentiment in text. In Proceed- ings of IJCAI 2016 Workshop on Computational Modeling of Attitudes , 2016
2016
-
[43]
Jesse Graham, Jonathan Haidt, and Brian A. Nosek. Liberals and conservatives rely on different sets of moral foundations. Journal of Personality and Social Psychology , 96(5):1029–1046, May 2009
2009
-
[44]
Detoxify, November 2020
Laura Hanu and Unitary team. Detoxify, November 2020
2020
-
[45]
Curhan, Hillary Anger Elfenbein, and Heng Xu
Jared R. Curhan, Hillary Anger Elfenbein, and Heng Xu. What do people value when they negotiate? Mapping the domain of subjective value in negotiation. Journal of Personality and Social Psychology , 91(3):493–512, September 2006
2006
-
[46]
Curhan, Noah Eisenkraft, Aiwa Shirako, and Lucio Baccaro
Hillary Anger Elfenbein, Jared R. Curhan, Noah Eisenkraft, Aiwa Shirako, and Lucio Baccaro. Are Some Negotiators Better Than Others? Individual Differences in Bar- gaining Outcomes. Journal of research in personality , 42(6):1463–1475, December 2008
2008
-
[47]
Conlon, and Remus Ilies
Nikolaos Dimotakis, Donald E. Conlon, and Remus Ilies. The mind and heart (lit- erally) of the negotiator: Personality and contextual determinants of experiential reactions and economic outcomes in negotiation. Journal of Applied Psychology , 97(1):183–193, 2012
2012
-
[48]
The Effect of Virtual Agent Warmth on Human-Agent Negotiation
Pooja Prajod, Mohammed Al Owayyed, and Tim Rietveld. The Effect of Virtual Agent Warmth on Human-Agent Negotiation. 2019
2019
-
[49]
Zhou, Gloria Mark, Jingyi Li, and Huahai Yang
Michelle X. Zhou, Gloria Mark, Jingyi Li, and Huahai Yang. Trusting Virtual Agents: The Effect of Personality. ACM Trans. Interact. Intell. Syst. , 9(2-3):10:1–10:36, March 2019
2019
-
[50]
Sotopia: Interactive evaluation for social intelligence in language agents
Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Zhengyang Qi, Haofei Yu, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. Sotopia: Interactive evaluation for social intelligence in language agents. In- ternational Conference on Learning Rep...
2024
-
[51]
What makes a good conversation? how controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. What makes a good conversation? how controllable attributes affect human judgments. arXiv preprint arXiv:1902.08654, 2019. 22
1902 arXiv
-
[52]
Connotation Frames: A Data- Driven Investigation, August 2016
Hannah Rashkin, Sameer Singh, and Yejin Choi. Connotation Frames: A Data- Driven Investigation, August 2016
2016
-
[53]
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto. Moral foundations theory: The pragmatic validity of moral pluralism. In Advances in experimental social psychology , volume 47, pages 55–130. Elsevier, 2013
2013
-
[54]
Truth of Varying Shades: Analyzing Language in Fake News and Political Fact- Checking
Hannah Rashkin, Eunsol Choi, Jin Yea Jang, Svitlana Volkova, and Yejin Choi. Truth of Varying Shades: Analyzing Language in Fake News and Political Fact- Checking. In Martha Palmer, Rebecca Hwa, and Sebastian Riedel, editors, Proceed- ings of the 2017 Conference on Empirical M...
2017
-
[55]
DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter, March 2020
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter, March 2020
2020
-
[56]
Detoxify
Laura Hanu and Unitary team. Detoxify. Github. https://github.com/unitaryai/detoxify, 2020
2020
-
[57]
DistilBERT for emotion recognition, May 2024
Bhadresh Savani. DistilBERT for emotion recognition, May 2024
2024
-
[58]
Machine intelligence to detect, characterise, and defend against in- fluence operations in the information environment
M Glenski, E Ayton, E Saldanha, J Mendoza, D Arendt, Z Shaw, K Cronk, S Smith, and M Greaves. Machine intelligence to detect, characterise, and defend against in- fluence operations in the information environment. Journal of Information Warfare , 20(2):42–66, 2021
2021
-
[59]
BERT: Pre- training of Deep Bidirectional Transformers for Language Understanding, May 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre- training of Deep Bidirectional Transformers for Language Understanding, May 2019
2019
-
[60]
A Deep Dive into Multilingual Hate Speech Classification
Sai Saketh Aluru, Binny Mathew, Punyajoy Saha, and Animesh Mukherjee. A Deep Dive into Multilingual Hate Speech Classification. In Yuxiao Dong, Georgiana Ifrim, Dunja Mladeni´ c, Craig Saunders, and Sofie Van Hoecke, editors, Machine Learning and Knowledge Discovery in Databas...
2021
-
[61]
Decoupling strategy and generation in negotiation dialogues, 2018
He He, Derek Chen, Anusha Balakrishnan, and Percy Liang. Decoupling strategy and generation in negotiation dialogues, 2018
2018
-
[62]
The Book of Why: The New Science of Cause and Effect
Judea Pearl and Dana Mackenzie. The Book of Why: The New Science of Cause and Effect. Basic Books, New York, 1st edition edition, May 2018
2018
-
[63]
EconML: A python package for ml-based hetero- geneous treatment effects estimation
Keith Battocchi, Eleanor Dillon, Maggie Hei, Greg Lewis, Paul Oka, Miruna Oprescu, and Vasilis Syrgkanis. EconML: A python package for ml-based hetero- geneous treatment effects estimation. https://github.com/py-why/EconML, 2019. Version 0.x
2019
-
[64]
Bradley, John E
Bret H. Bradley, John E. Baur, Christopher G. Banford, and Bennett E. Postleth- waite. Team Players and Collective Performance: How Agreeableness Affects Team Performance Over Time. Small Group Research, 44(6):680–711, December 2013. 23
2013
-
[65]
Driskell, Gerald F
James E. Driskell, Gerald F. Goodwin, Eduardo Salas, and Patrick Gavan O’Shea. What makes a good team player? Personality and team effectiveness. Group Dy- namics: Theory, Research, and Practice , 10(4):249–271, December 2006
2006
-
[66]
Adam M. Grant. Rethinking the Extraverted Sales Ideal: The Ambivert Advantage. Psychological Science, 24(6):1024–1030, June 2013
2013
-
[67]
Lepine, Jason A
Jeffrey A. Lepine, Jason A. Colquitt, and Amir Erez. Adaptability to Changing Task Contexts: Effects of General Cognitive Ability, Conscientiousness, and Openness to Experience. Personnel Psychology, 53(3):563–593, 2000
2000
-
[68]
Suzanne T. Bell. Deep-level composition variables as predictors of team performance: A meta-analysis. The Journal of Applied Psychology , 92(3):595–615, May 2007
2007
-
[69]
K. J. Klein, J. L. Saltz, and D. M. Mayer. HOW DO THEY GET THERE? AN EXAMINATION OF THE ANTECEDENTS OF CENTRALITY IN TEAM NET- WORKS. Academy of Management Journal , 47(6):952–963, December 2004
2004
-
[70]
Miranda A. G. Peeters, Harrie F. J. M. Van Tuijl, Christel G. Rutte, and Isabelle M. M. J. Reymen. Personality and team performance: A meta-analysis. European Journal of Personality , 20(5):377–396, August 2006
2006
-
[71]
The PANAS-X: Manual for the positive and negative affect schedule-expanded form
David Watson and Lee Anna Clark. The PANAS-X: Manual for the positive and negative affect schedule-expanded form. 1994
1994
-
[72]
Robert R. McCrae. NEO-PI-R Data from 36 Cultures. In Robert R. McCrae and J¨ uri Allik, editors, The Five-Factor Model of Personality Across Cultures , pages 105–125. Springer US, Boston, MA, 2002
2002
-
[73]
Habashi, William G
Meara M. Habashi, William G. Graziano, and Ann E. Hoover. Searching for the Prosocial Personality: A Big Five Approach to Linking Personality and Prosocial Behavior. Personality and Social Psychology Bulletin , 42(9):1177–1192, September 2016
2016
-
[74]
Hirsh, Colin G
Jacob B. Hirsh, Colin G. DeYoung, Xiaowen Xu, and Jordan B. Peterson. Com- passionate Liberals and Polite Conservatives: Associations of Agreeableness With Political Ideology and Moral Values. Personality and Social Psychology Bulletin , 36(5):655–664, May 2010
2010
-
[75]
Extraversion and Its Positive Emotional Core
David Watson and Lee Anna Clark. Extraversion and Its Positive Emotional Core. In Handbook of Personality Psychology , pages 767–793. Elsevier, 1997
1997
-
[76]
P. A. Hancock, Theresa T. Kessler, Alexandra D. Kaplan, John C. Brill, and James L. Szalma. Evolving Trust in Robots: Specification Through Sequential and Compar- ative Meta-Analyses. Human Factors, 63(7):1196–1229, November 2021
2021
-
[77]
Schaefer, Jessie Y
Kristin E. Schaefer, Jessie Y. C. Chen, James L. Szalma, and P. A. Hancock. A Meta- Analysis of Factors Influencing the Development of Trust in Automation: Implica- tions for Understanding Autonomy in Future Systems. Human Factors, 58(3):377– 400, May 2016. 24
2016
-
[78]
Hancock, Deborah R
Peter A. Hancock, Deborah R. Billings, Kristin E. Schaefer, Jessie Y. C. Chen, Ewart J. de Visser, and Raja Parasuraman. A Meta-Analysis of Factors Affecting Trust in Human-Robot Interaction. Human Factors: The Journal of the Human Factors and Ergonomics Society , 53(5):517–52...
2011
-
[79]
first_name
Sarah A. Jessup, Tamera R. Schneider, Gene M. Alarcon, Tyler J. Ryan, and August Capiola. The Measurement of the Propensity to Trust Automation. In Jessie Y.C. Chen and Gino Fragomeni, editors, Virtual, Augmented and Mixed Reality. Appli- cations and Case Studies , pages 476–4...
2019
-
[2024]
Association for Computational Linguistics
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.