Pith. sign in

REVIEW 3 major objections 6 minor 39 references

MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper introduces MotiveBench and claims that, measured against human-consensus answers, even the most advanced large language models fall measurably short of human-like motivation and behavior reasoning, with GPT-4o reaching only…

desk verdict A useful new benchmark for motivational reasoning, but the human validation pipeline is not independent enough to support the claim that LLMs fall short of human consensus. read the letter →

arxiv 2506.13065 v1 pith:BY4PC5AX submitted 2025-06-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords MotiveBenchmotivationalreasoningbehaviorMaslow'shierarchyofneedslargelanguagemodelshuman-LLMalignmentquestion-answerbenchmarkchain-of-thought
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large language models are increasingly used as social-simulation agents and companions, but whether they reason about human motivations the way people do has not been systematically measured. This paper introduces MotiveBench, a set of 200 detailed scenarios and 600 multiple-choice questions organized around Maslow's hierarchy of needs and Reiss's 16 basic desires, with three task types per scenario: infer the motivation behind a behavior, infer the behavior behind a motivation, and infer both from a scenario and character profile alone. The authors report that the strongest model tested, GPT-4o, reaches 80.89% accuracy against answer keys that human annotators agreed on, and argue that this gap indicates current LLMs have not achieved human-like motivational reasoning. The benchmark is designed to address limitations of earlier social-commonsense tests, whose contexts are too short, whose answers are often explicitly stated, and which lack a theoretical account of what drives behavior. A sympathetic reader would care because the humanization of LLM agents, for simulations, companions, and decision-support systems, depends on exactly this kind of reasoning.

What carries the argument

The load-bearing machinery is a human-in-the-loop, multi-agent question-generation pipeline organized around a behavioral quadruple: Scenario, Profile, Motivation, and Behavior. For each quadruple, an LLM questioner writes three independent questions, three LLM reviewers critique logical soundness, answer correctness, and difficulty, and the questioner revises for up to five rounds. High-performing LLMs then answer every question so that items can be balanced across easy, medium, and hard difficulty, and the authors manually correct each remaining item. Finally, 15 crowd-sourced annotators from diverse backgrounds score each option and verify that the keyed answer is the human consensus. This mechanism is what turns raw text from Persona-Hub, Amazon, and Blogger into a benchmark whose accuracy numbers are claimed to be indicators of human-like reasoning.

What would settle it

Take the 600 MotiveBench questions to a fresh group of, say, 100 raters drawn from a different cultural or age pool, with the keyed answers hidden. If the modal human choices disagree with the benchmark keys on substantially more than the 7% disagreement already reported, then the "human consensus" reference point is not stable, and the claimed gap between LLMs and humans would need to be re-interpreted.

Watch

Extended reading notes

Core claim

Human-like motivation and behavior reasoning is a distinct capability, and current LLMs have not achieved it. On MotiveBench, which draws 600 questions from 200 detailed scenarios and pre-validates them with 15 annotators whose answers matched human consensus at 93%, the best model, GPT-4o, reaches 80.89% base accuracy, while Qwen2.5-72B reaches 78.61% and smaller models fall well below. Accuracy tracks model size, chain-of-thought prompting generally lowers accuracy rather than raising it, and the clearest weakness is reasoning about "love & belonging" needs. Error analyses of GPT-4o show that it over-rationalizes, falls back on general assumptions instead of case details, assumes people act on idealistic social norms, and underestimates the practical impact of actions.

Load-bearing premise

The load-bearing premise is that the answer key, which starts as LLM-generated options and is then filtered and corrected by a small team of human annotators, is a fair and unbiased proxy for how ordinary people actually reason about motivations and behaviors; if LLM biases survive that review, the reported accuracy numbers measure alignment with an LLM-influenced consensus rather than with human reasoning.

Editorial extensions

If this is right

  • Larger LLMs are better motivators: average base accuracy rises from 57.16% for models under 10B parameters to 71.34% for models over 34B, so scaling is currently the most reliable route toward more human-like motivational reasoning.
  • Chain-of-thought prompting backfires on intrinsic-motivation tasks: most models drop in accuracy, with models of 34B or smaller losing 6.88% compared with 3.14% for larger models, suggesting step-by-step deliberation diverges from human intuition here.
  • The "love & belonging" level of Maslow's hierarchy is the hardest for LLMs, with the best GPT-series models scoring 77.27% and weaker families dropping to the mid-50s, so emotional and relationship reasoning is a targeted weakness.
  • MotiveBench measures a capability that general benchmarks mostly miss: its rankings correlate with LiveBench abilities at an average Pearson coefficient of 0.8175, and the paper argues the Motive dimension adds information about underlying human-like capabilities.
  • LLMs are not yet reliable stand-ins for human annotators in psychological data, since the generation pipeline needed manual correction for logical errors, missed human dynamics, and single-model bias.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the over-rationalization finding generalizes, then LLM-driven social simulations will systematically under-produce emotionally driven, norm-violating, or gut behaviors; designers may need to inject fast, intuitive decision biases rather than only improve logical competence.
  • A testable consequence is that prompting models to answer under time pressure, or to give their first intuitive impression before deliberating, could close part of the measured gap; the paper does not try this.
  • Because the benchmark uses static vignettes, the 80.89% figure is a snapshot of one-shot reasoning; in multi-turn sandbox settings, which the authors name as future work, motivational reasoning could improve with accumulated evidence or degrade under distraction, and the error taxonomy could become a scoring rubric.
  • The 7% disagreement between the predefined answers and the annotators' consensus opens a direct research line: studying which items produce unstable human consensus, since those items cannot cleanly separate human-like from non-human-like reasoning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces MotiveBench, a benchmark of 200 contextual scenarios and 600 multiple-choice questions spanning three reasoning tasks: motivational, behavioral, and combined motive-and-behavior reasoning. The items are generated by a multi-agent LLM framework using LLaMA3.1-70B and Qwen2.5-72B, then filtered by difficulty using high-performing LLMs and manually corrected by the authors. The authors evaluate 29 LLMs from seven model families under base and chain-of-thought prompting, reporting that GPT-4o reaches 80.89% accuracy while "human consensus" agreement is 93%, and they draw several qualitative findings, including particular difficulty with love-and-belonging needs and a tendency toward over-rationalization. The central claim is that current LLMs fall short of human-like motivational reasoning.

Significance. If the benchmark's validation were fully independent, MotiveBench would be a useful new resource: it is theory-grounded in Maslow's hierarchy and Reiss's desires, covers long contexts and varied profiles, evaluates a wide range of open- and closed-source models, and mitigates option-order bias with six permutations. The paper is honest about its limitations, including the reliance on manual correction and the situational QA paradigm. However, the benchmark's central quantitative claim—that LLM accuracy can be read as distance from human motivational reasoning—currently rests on a human-validation procedure that is not independent of the LLM-generated answers, so the headline result is not yet established as a measure of human-like reasoning.

major comments (3)
  1. [Sec. 4.2, Table 2] The human validation does not establish that the 93.00% agreement with predefined answers is a reliable estimate of human consensus on the final benchmark. Appendix C.1 states that annotators were selected "based on their performance in a small set of trial annotations" with a 46.43% rejection rate; if those trials were scored against the LLM-generated predefined answers, the selection procedure screens for annotators who agree with the authors' labels. Moreover, Section 3.4 says "we correct inconsistent or contradictory questions based on annotations," so items that fail to match the predefined answer are removed or revised before the final agreement statistics are reported. These two design choices jointly prevent the reported agreement from being an independent measure of how unselected humans would answer. The claim in Section 4.2 that "MOTIVEBENCH inherently reflect human consensus" therefore needs a direct human test-taker baseline: a fresh sample of unselected participants should answer the final MCQ set under the same conditions as the models, with each participant selecting one option per question, and accuracy should be reported with confidence intervals. Until then, the comparison of GPT-4o's 80.89% to human-level reasoning conflates alignment with LLM-influenced answer distributions with genuine human motivational reasoning.
  2. [Secs. 3.1-3.3, Table 2] There is a circularity burden that is not addressed: the ground-truth answers and candidate options are initially authored and filtered by LLaMA3.1-70B and Qwen2.5-72B (Sections 3.1-3.3), and these same model families appear among the evaluated systems in Table 2. If the human-review step were fully independent, this would be merely a design concern, but as argued above the human validation is not fully independent because it is conditioned on prior agreement with the predefined answers. As a result, high accuracy on MotiveBench may reflect alignment with an LLM-generated answer distribution rather than with human psychology. The authors should either exclude the generation models from the evaluation suite, or quantify the effect by reporting per-model accuracy split by whether the item was generated by that model family, and by demonstrating through a direct human baseline that the post-correction labels are not carrying LLM-specific priors. Mentioning in Section 4.3 Insight 3 that multiple models were used to 'minimize' bias is not sufficient to rule out this systematic effect.
  3. [Sec. 4.2, Table 3] The key findings—such as the 'love & belonging' deficit, the CoT performance drop, and the scaling trend—are reported without any uncertainty quantification. Table 3 shows differences of a few percentage points between need levels (e.g., GPT-4o at 77.27% for Level 3 vs 85.83% for Level 1), but with 600 questions total and no confidence intervals or significance tests, these differences may fall within sampling noise. The three questions generated from the same scenario are also not statistically independent samples, since they share the same context and profile; any analysis that treats all 600 questions as independent observations will overstate precision. The authors should provide scenario-clustered standard errors or bootstrap confidence intervals, and either perform a formal test for the love-and-belonging gap or soften that claim to a qualitative observation. This does not invalidate the overall ranking, but it does undercut the more granular conclusions that are presented as main results.
minor comments (6)
  1. [Sec. 1] The sentence "Motivation is ... at any given scenarios" contains a grammatical error; it should be "in any given scenario."
  2. [Sec. 3.4] The phrase "it taks about 6 minutes per question" contains a typo ('taks' should be 'takes').
  3. [Title/Abstract] The benchmark name is written inconsistently as MotiveBench, MOTIVEBENCH, and sometimes MotiveBench; please standardize the capitalization throughout.
  4. [Sec. 3.3] The text says "retaining 100 scenarios for Person-Hub, 50 for Amazon, and 50 for Blogger," which totals 200, but it calls the Person-Hub subset "scenarios" while Section 1 refers to 200 profiles and scenarios; clarify whether 100 is the number of profiles or scenarios.
  5. [Table 2] The column header "A VG." appears to be a stray or incomplete label, and the bold/† markers are not explained in the caption; please define the notation.
  6. [Sec. 4.3] The text reports "Pearson correlation coefficients ... of rankings," but ranking correlations are typically Spearman correlations; if Pearson is applied to ranks, this should be stated explicitly, and the number of models entering the correlation should be reported.

Circularity Check

1 steps flagged · score 5.0 of 10

The 'human consensus' validation is circular: answer keys are corrected to match annotators and then validated against the same annotator consensus, so the 93% agreement is partly by construction.

  1. fitted input called prediction [Section 3.4, Appendix C.3, Section 4.2]
    "we correct inconsistent or contradictory questions based on annotations, ensuring each final answer aligns accurately with human consensus. ... By aggregating the weights assigned by annotators to each option and selecting the one with the highest cumulative score, we find that consensus answers match our pre-defined correct answers at a rate of 93.00%, thereby validating the accuracy of our predefined answers. ... since our MOTIVEBENCH inherently reflect human consensus, the accuracy of each model serves as an indicator of its capacity for human-like motivation-behavior reasoning."

    The 'pre-defined correct answers' that are being validated are the same answer keys that were already revised after annotation to match human consensus. Computing agreement between aggregated annotator scores and those post-hoc corrected answers is therefore not an independent validation; the 93.00% figure is in part manufactured by the correction step described in Section 3.4. The paper then invokes this figure to assert that MOTIVEBENCH inherently reflects human consensus and to interpret model accuracy as human-like motivational reasoning. Thus the human-consensus property is imposed on the labels by definition rather than demonstrated by an independent measurement, and the model-accuracy interpretation rests on that circular validation.

full rationale

The paper's central model comparison is empirical and not itself derived from its inputs: GPT-4o's 80.89% accuracy is a real measurement against the benchmark. However, the load-bearing claim that MOTIVEBENCH reflects human consensus is supported by a circular validation. The answer key was corrected to align with annotator judgment, and then the same annotator judgment was used to report 93.00% agreement with the answer key, which is partly a restatement of the correction procedure. The use of LLaMA3.1-70B and Qwen2.5-72B both as question generators and as evaluated models compounds the concern, though the human correction step prevents a fully closed loop. No load-bearing self-citation or imported uniqueness theorem is used. The circularity is partial: the human-consensus validation is circular, while the empirical model rankings retain independent content.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The benchmark introduces no mathematical free parameters in the usual sense, but several hand-chosen design constants affect the reported numbers. The main scientific assumptions are psychological taxonomies, the use of annotator consensus as ground truth, and the assumption that LLM-generated content is unbiased after human review.

free parameters (5)
  • Option permutation count = 6
    The paper averages scores over six random option orders in Section 4.1; this hand-chosen hyperparameter affects reported accuracy.
  • Annotator consensus threshold = 3 of 5 annotators
    Questions are considered agreed when at least three of five annotators match, yielding the reported 98.17 percent consistency in Section C.3.
  • Annotator point allocation = 6 points per question
    Each annotator distributes six points across options; this weighting rule defines the consensus answer and the 93.00 percent agreement rate in Section C.2.
  • Difficulty balance targets = approximately equal easy, medium, hard
    Questions are selected so that LLM-assessed difficulty is balanced; this distribution is chosen by hand and changes the composition of the benchmark in Section 3.3.
  • Revision iteration cap = 5 rounds
    The questioner modifies questions up to five times based on reviewer feedback, as described in Appendix F; this cap affects question quality and complexity.
assumptions (5)
  • domain assumption Maslow's hierarchy of needs is a valid and complete taxonomy for human motivations.
    The benchmark's five need levels and the love-and-belonging finding rely on this theory in Section 2.2.
  • domain assumption Reiss's 16 basic desires can be mapped without loss into Maslow's five levels.
    The fine-grained labels are built by treating Reiss categories as subcategories of Maslow levels in Section 2.2.
  • domain assumption Human annotator majority consensus is the correct ground truth for human-like motivational reasoning.
    The paper equates model accuracy with capacity for human-like reasoning based on annotator agreement in Sections 3.4 and 4.2.
  • ad hoc to paper LLM-authored questions and answers, after human review and filtering, are free of systematic LLM bias.
    The construction pipeline starts with LLM-generated quadruples and cannot verify that human correction removes all biases, as described in Sections 3.1 to 3.4.
  • ad hoc to paper The three questions generated from the same scenario are independent samples for statistical purposes.
    Reported accuracies treat each of the 600 questions as an independent observation, but questions sharing a scenario are correlated, as seen in Table 2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?." pith.science (2026). https://pith.science/paper/BY4PC5AX

@misc{pith2026250613065,
  author       = {Pith},
  title        = {Pith review of: MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BY4PC5AX}},
  note         = {Machine review of arXiv:2506.13065}
}
read the original abstract

Large language models (LLMs) have been widely adopted as the core of agent frameworks in various scenarios, such as social simulations and AI companions. However, the extent to which they can replicate human-like motivations remains an underexplored question. Existing benchmarks are constrained by simplistic scenarios and the absence of character identities, resulting in an information asymmetry with real-world situations. To address this gap, we propose MotiveBench, which consists of 200 rich contextual scenarios and 600 reasoning tasks covering multiple levels of motivation. Using MotiveBench, we conduct extensive experiments on seven popular model families, comparing different scales and versions within each family. The results show that even the most advanced LLMs still fall short in achieving human-like motivational reasoning. Our analysis reveals key findings, including the difficulty LLMs face in reasoning about "love & belonging" motivations and their tendency toward excessive rationality and idealism. These insights highlight a promising direction for future research on the humanization of LLMs. The dataset, benchmark, and code are available at https://aka.ms/motivebench.

Figures

Figures reproduced from arXiv: 2506.13065 by the authors.

Figure 1
Figure 1. The difference between the existing traditional [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The statistical overview of MOTIVEBENCH. It contains 200 diverse profiles and real-world scenarios, along with 600 motivational and behavioral reasoning questions, covering multiple finely-grained levels of needs. underexplored: Can current LLMs truly under￾stand and exhibit human-like motivations and behaviors? The complexity of human behavior dy￾namics poses new challenges for LLMs, which are distinct from the cha… view at source ↗
Figure 3
Figure 3. A step-by-step questions generation and correction pipeline using AI-Human collaboration framework. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Generation and extraction of motivations and [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Scaling of various model families across their different versions and sizes in motivational reasoning ability. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: The correlation coefficients between the eval [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: A case study on motivation and behavior reasoning: we analyze GPT-4o’s attention within the question [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Example scenarios for Maslow’s hierarchy of needs in M [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: An example question from Persona-Hub. is logically inconsistent or flawed; 2) Insuf￾ficient background information is provided, making it impossible to derive any meaningful answer; 3) None of the answer options are rea￾sonable or relevant; 4) There are three or more a…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

39 extracted references · 31 canonical work pages

  1. [1]

    Based on the content of the given question, please infer the most likely answer

  2. [2]

    Layla El Asri, Jing He, and Kaheer Suleman

    Out of one, many: Using language mod- els to simulate human samples.Political Analysis, 31(3):337–351. Layla El Asri, Jing He, and Kaheer Suleman. 2016. A sequence-to-sequence model for user simula- tion in spoken dialogue systems.arXiv preprint arXiv:1607.00070. Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu ...

  3. [3]

    Barbara Brehm

    Emergent autonomous scientific research ca- pabilities of large language models.arXiv preprint arXiv:2304.05332. Barbara Brehm. 2014.Psychology of health and fitness. FA Davis. Zhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengt- ing Hu, Yunghwei Lai, Zexuan Xiong, et al. 2024. Tombench: Benchmarking theory of mind...

  4. [4]

    The distractors must be related to certain parts of the information in the question

    Each question should have only one correct answer, along with five distractors. The distractors must be related to certain parts of the information in the question. Please analyze why each option is correct or incorrect

  5. [5]

    This is necessary to ensure each question is challenging

    The question stem must include irrelevant or redundant information that creates distractions and challenges. This is necessary to ensure each question is challenging. The correct answer must not be explicitly stated in the question. Table 8: Detailed prompts of questioner in the multi-agent framework. Prompts of Reviewers You are a strict and discerning p...

  6. [6]

    If they do not, suggest adding the relevant distracting information or modifying the options

    **Clarity and Plausibility of Distractor Options**: Evaluate whether the incorrect options are misleading and whether they correspond to distracting information within the question. If they do not, suggest adding the relevant distracting information or modifying the options

  7. [7]

    **Adequate Distractors and Redundant Information**: Ensure that each question includes enough irrelevant or redundant information to make the question challenging, but without disrupting the logic needed to deduce the correct answer

  8. [8]

    Please provide specific modification suggestions for the question set and give your feedback to the question author in a reasonable tone

    **Objectivity and Neutrality of the Question**: Ensure that the question is presented in a neutral and objective manner, with no implicit suggestion of the correct answer. Please provide specific modification suggestions for the question set and give your feedback to the question author in a reasonable tone. Summarize your evaluation into a single paragra...

Show all 39 references
  1. [9]

    David E Wilkins

    Livebench: A challenging, contamination-free llm benchmark.arXiv preprint arXiv:2406.19314. David E Wilkins. 2014.Practical planning: extending the classical AI planning paradigm. Elsevier. Yang Xiao, Yi Cheng, Jinlan Fu, Jiashuo Wang, Wenjie Li, and Pengfei Liu. 2023. How far...

  2. [10]

    Sotopia: Interactive evaluation for social intelligence in language agents.arXiv preprint arXiv:2310.11667. A Details of the Hierarchy of Needs Maslow’s hierarchy of needs (Maslow, 1943) is a motivational theory proposed by American psychol- ogist Abraham Maslow in 1943, which...

  3. [11]

    (Qwen 2.5, Qwen 2, and Qwen), the Phi se- ries (Abdin et al., 2024) (Phi 3.5 and Phi 3), the GLM series (GLM et al., 2024) (ChatGLM 3 and GLM 4), as well as other models like Baichuan 2 (Yang et al., 2023) and Yi 1.5 (Young et al., 2024). These models span a wide spectrum of a...

  4. [14]

    Question: {Question_Content} Options: {Options} CoT Prompt for Evaluation The following is a {Question_Type}

    The result can only return **one character without any other explanation**. Question: {Question_Content} Options: {Options} CoT Prompt for Evaluation The following is a {Question_Type}. You should: {Type_Interpretation}. Carefully read the given question, fully immerse yoursel...

  5. [15]

    Based on the content of the given question, please think step by step and infer the most likely answer

  6. [16]

    A, B, C, D, E, F

    You must select one answer from the given options: "A, B, C, D, E, F" as the most likely choice. Even if the question does not provide sufficient information to determine the correct answer, you should randomly choose one option as your output

  7. [17]

    Please first think through the question step by step, analyze the reasoning process for the possible answers, and finally output the most likely answer’s letter. **The last line of your reply should only contain one character of your final choice.** Question: {Question_Content...

  8. [18]

    The question should not contain any direct description related to the predicted motivation

    Motivational Reasoning Question: Given a complex scenario, a specific profile, and a given behavior, infer the most likely motivation behind the character’s behavior. The question should not contain any direct description related to the predicted motivation

  9. [19]

    The question should not contain any direct description related to the predicted behavior

    Behavioral Reasoning Question: Given a complex scenario, a specific profile, and a given motivation, infer the most likely behavior the character will perform based on that motivation. The question should not contain any direct description related to the predicted behavior

  10. [20]

    The question should only include the complex scenario and the character’s profile

    Motive&Behavior Reasoning Question: This is a more advanced test. The question should only include the complex scenario and the character’s profile. Using only the complex scenario and specific profile, infer the most likely motivation the character will have and the correspon...

  11. [21]

    Please rewrite this scenario by cor- recting any logical inconsistencies, and add relevant details to make the scenario, profile, motivation, and behavior more vivid and complex

    You will be provided with a simple scenario description. Please rewrite this scenario by cor- recting any logical inconsistencies, and add relevant details to make the scenario, profile, motivation, and behavior more vivid and complex

  12. [22]

    However, ensure that the motivation and behavior are only related to real human needs, not to any POIs or products in the text

    Choose the most appropriate motivation and behavior to create the questions. However, ensure that the motivation and behavior are only related to real human needs, not to any POIs or products in the text

  13. [23]

    Therefore, please ensure that each question has enough rich and complex scenario and profile information to support correct reasoning

    The three questions are independent of each other and should be answered separately, meaning that each question should only rely on its own stem and not contain any information from the others. Therefore, please ensure that each question has enough rich and complex scenario an...

  14. [26]

    The motivation reasoning question should include additional behavioral information about the character

    **Reasonableness of the Question Information and Type**: Specifically, all three questions should contain a concrete scenario and character profile. The motivation reasoning question should include additional behavioral information about the character. The behavior reasoning q...

  15. [27]

    **Logical Consistency and Reasonableness of the Four-Tuple**: Assess whether, in the given scenario, a character with a specific profile would logically perform the stated behavior based on the provided motivation

  16. [28]

    If not, suggest modifications to the scenario or character profile to make the information clearer or more comprehensive

    **Sufficiency of Information to Derive the Correct Answer**: Examine whether the information provided in each question is enough to infer the correct answer. If not, suggest modifications to the scenario or character profile to make the information clearer or more comprehensive

  17. [29]

    **Challenge and Difficulty of the Question**: Evaluate whether the question presents an appropriate level of difficulty and challenge for the respondent

  18. [30]

    **Correct Answer Must Not Be Explicitly Stated**: Ensure that the correct answer does not appear explicitly in the question information and can only be deduced through reasoning steps

  19. [34]

    The question should not include any description related to the predicted motivation

    Motivational Reasoning Question: Based on a complex scenario, a specific character profile, and a given behavior, deduce the most likely motivation behind the character’s action. The question should not include any description related to the predicted motivation

  20. [35]

    The question should not include any description related to the predicted behavior

    Behavioral Reasoning Question: Based on a complex scenario, a specific character profile, and a given motivation, deduce the most likely behavior the character will perform based on that motivation. The question should not include any description related to the predicted behavior

  21. [36]

    Motive&Behavior Reasoning Question: The question should only include a complex scenario and character profile. This is a more difficult question type, where the respondent must deduce the most likely behavior and corresponding motivation of the character based solely on the sc...

  22. [37]

    Carefully consider each suggestion based on the given questions and selectively make rea- sonable changes to the questions

  23. [38]

    Do not delete the distracting information related to the incorrect answers, as this is necessary to ensure the questions remain challenging

  24. [39]

    Respondents should only reason based on the question provided, without seeing any other information

    The three questions are independent of each other and are to be answered separately. Respondents should only reason based on the question provided, without seeing any other information. Therefore, ensure that each question has sufficiently rich and complex scenario and charact...

  25. [40]

    self-promotion

    After making revisions, analyze each option to determine why it is correct or incorrect. If there are any issues, modify the question again to ensure the uniqueness of the correct answer. Table 10: Detailed prompts of modifier in the multi-agent framework. G.2 Detailed Case St...

  26. [2006]

    InProceed- ings of 2006 AAAI spring symposium on computa- tional approaches for analyzing weblogs, volume 1

    Effects of age and gender on blogging in pro- ceedings of 2006 aaai spring symposium on computa- tional approaches for analyzing weblogs. InProceed- ings of 2006 AAAI spring symposium on computa- tional approaches for analyzing weblogs, volume 1. Natalie Shapira, Mosh Levy, Se...

  27. [2010]

    InProceedings of the SIGDIAL 2010 Conference, pages 116–123

    Parameter estimation for agenda-based user simulation. InProceedings of the SIGDIAL 2010 Conference, pages 116–123. Florian Kreyssig, Iñigo Casanueva, Pawel Budzianowski, and Milica Gasic. 2018. Neu- ral user simulation for corpus-based policy optimisation for spoken dialogue ...

  28. [2016]

    InWorkshops at the Thirtieth AAAI Conference on Artificial Intelligence

    An overview of affective motivational collab- oration theory. InWorkshops at the Thirtieth AAAI Conference on Artificial Intelligence. Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Du...

  29. [2019]

    Revisiting the evaluation of theory of mind through question answering. InProceedings of the 2019 Conference on Empirical Methods in Natu- ral Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5872–5877. Sangmin...

  30. [2023]

    InInternational Conference on Machine Learning, pages 337–371

    Using large language models to simulate mul- tiple humans and replicate human subject studies. InInternational Conference on Machine Learning, pages 337–371. PMLR. Lisa P Argyle, Ethan C Busby, Nancy Fulda, Joshua R Gubler, Christopher Rytting, and David Wingate

  31. [2024]

    Maarten Sap, Ronan LeBras, Daniel Fried, and Yejin Choi

    Emobench: Evaluating the emotional intel- ligence of large language models.arXiv preprint arXiv:2402.12071. Maarten Sap, Ronan LeBras, Daniel Fried, and Yejin Choi. 2022. Neural theory-of-mind? on the limits of social intelligence in large lms.arXiv preprint arXiv:2210.13312. ...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.