Pith. sign in

REVIEW 5 major objections 5 minor 2 cited by

SocialEval: Evaluating Social Intelligence of Large Language Models

T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read SocialEval claims that social intelligence in LLMs can be measured on two axes at once—outcome goal achievement and process interpersonal ability—and reports that LLMs lag humans on both while preferring prosocial choices that can cause…

desk verdict SocialEval is a genuinely useful benchmark with a careful construction pipeline, but its empirical claims and brain analogy need to be pulled back before the numbers are taken at face value. read the letter →

arxiv 2506.00900 v1 pith:FRBLHKS4 submitted 2025-06-01 cs.CL

classification cs.CL
keywords SocialintelligenceLargelanguagemodelsBenchmarkWorldtreeGoalachievementevaluationInterpersonalabilityProsocialityBilingual
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

SocialEval tries to establish that an LLM's social intelligence is best measured along two axes at once: whether it reaches the social goal it is pursuing, and whether it correctly recognizes and uses interpersonal abilities along the way. To do this it contributes a bilingual, script-based benchmark of 153 manually crafted 'world trees' in Chinese and English, each a branching narrative with multiple plot lines, episode transitions, and labeled success or failure endings. The paper reports that all evaluated LLMs score below human baselines on both axes, and that models show a systematic prosociality bias: they choose more positive, cooperative utterances even when those choices lead to goal failure, whereas human participants adjust their behavior more flexibly. If these results are right, SocialEval provides a reusable way to separate 'did the model get what it wanted' from 'did the model behave with interpersonal skill' in social interaction, and the reported human-model gaps are real.

What carries the argument

The load-bearing object is the world tree, a script-based narrative structured as a goal-conditioned Markov decision process with tuple $(S,A,T,R)$: each episode is a state, each candidate protagonist utterance is an action, episode transitions play the role of the transition function, and plot endings supply a binary reward for goal achievement. The second carrier is the interpersonal ability inventory, taken from the BESSI framework: 32 specific abilities grouped into 5 aspects, used to craft one probing question and a set of plausible distractors for each candidate utterance. Outcome evaluation and process evaluation are two readings of the same tree: goal achievement ratio measures whether the model navigates to successful endings, while ability selection accuracy measures whether it can pick the utterance that correctly embodies the probed ability.

What would settle it

Re-annotate the benchmark blind: give a fresh panel of annotators the episode contexts and the candidate utterances without the original labels, and ask them to rank which utterance best achieves the protagonist's goal and which endings count as success under their own social norms. If agreement with the original labels is low (for example, below 70%), the reported LLM-versus-human gaps are not a stable measure of social intelligence. A complementary check is behavioral: have human participants role-play the labeled-failure positive options with naive partners and measure whether goals are actually achieved; if labeled-failure options succeed in live interaction, the outcome labels are wrong.

Watch

Extended reading notes

Core claim

The central claim is that social intelligence can be operationalized as goal-conditioned navigation of interdependent social episodes, and that a benchmark built on this view can expose where machines diverge from humans. SOCIALEVAL casts each narrative script as a world tree: the protagonist faces episodes connected by choices that display distinct interpersonal abilities, and each branch leads to an ending labeled as goal achievement or failure. The outcome-oriented task asks the tested model, playing the protagonist, to select utterances that move the plot toward a successful ending; the process-oriented task asks it to answer probing multiple-choice questions about which of the 32 interpersonal abilities an utterance embodies, with plausible but wrong distractors. Across the evaluated LLMs, the paper reports that every model trails the human average on both tasks, with the smallest gaps being roughly 23.8%/17.2% for goal achievement and 3.2%/4.6% for ability selection in Chinese/English, and that models exhibit prosociality, favoring positive behaviors even at the cost of goal failure. The paper also claims that larger models develop ability-specific clusters in representation space and isolated neuron regions, which it interprets as brain-like functional partitions.

Load-bearing premise

The load-bearing premise is that the manually annotated labels are correct ground truth: the designated correct answers to the 2,493 ability questions and the success/failure labels on plot endings genuinely reflect effective versus ineffective social behavior. The labels were produced by the authors' screenwriters and quality inspectors, with no independent validation against external behavioral criteria; if those labels are wrong, the LLM accuracy scores, human comparisons, and prosociality conclusions all lose their meaning.

Editorial extensions

If this is right

  • If SOCIALEVAL measures what it claims, current LLMs trail human-average social intelligence in both Chinese and English, on both goal achievement and interpersonal ability selection.
  • LLM social intelligence scales with model size: open-source models show a positive performance-size correlation, and the representation and neuron analyses indicate more distinct ability-specific partitions in larger models.
  • The prosociality bias implies that training or prompting that pushes models toward positive, cooperative behavior can undercut goal-directed effectiveness in social tasks, so the two objectives should be measured separately.
  • The semantic-similarity check, where generated utterances match candidate options at high rates that rise with model size, supports using multiple-choice selection as a proxy for what the model would freely say.
  • Significant cross-lingual differences on both tasks mean SI in LLMs should be evaluated multilingually rather than assumed language-invariant.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the prosociality bias is stable, benchmarks that reward only cooperation or helpfulness may overestimate social effectiveness; pairing an outcome metric with a process metric, as SOCIALEVAL does, exposes the trade-off directly.
  • A testable extension is to prompt models with explicit goal reminders or instruct them to prioritize goal success, and check whether the prosociality gap shrinks; the paper does not test any such intervention.
  • Because the success/failure and correctness labels come from the authors' own screenwriters and inspectors, an independent re-labeling by culturally diverse panels would test whether the reported human-model gaps survive different social norms.
  • The neuron-isolation result predicts that pruning neurons identified as specific to one ability should degrade that ability more than others; the paper reports the correlation but not that causal test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper introduces SocialEval, a manually constructed bilingual benchmark for evaluating social intelligence in large language models. It consists of 153 world trees (narrative scripts with branching plot lines) covering 7 social orientations and 32 interpersonal abilities, yielding 2,493 process-oriented multiple-choice samples and outcome-oriented goal-achievement episodes. The authors evaluate 28 LLMs and human baselines on goal achievement ratio and ability selection accuracy, and report that LLMs fall behind humans on both tasks, show prosociality and a preference for positive behaviors, exhibit cross-lingual differences, and develop ability-specific representation and neuron partitions as model size grows. The paper also includes a limitations section and releases the data.

Significance. If the labels and human baselines are valid, SocialEval would be a useful contribution: it operationalizes process-oriented and outcome-oriented SI evaluation in multi-episode scripts, covers a broad ability taxonomy grounded in BESSI, and provides a bilingual testbed. Strengths include the multi-stage quality-control pipeline, the translation review protocol, the breadth of evaluated models, and the plan to release the data. The representation and neuron-level analyses are suggestive and could inspire further interpretability work. However, the empirical claims currently outrun the evidence: the human baseline is small, the headline comparisons lack uncertainty quantification, and the ground-truth labels are validated only through internal consistency. These issues are addressable and do not disqualify the benchmark itself.

major comments (5)
  1. [Section 4, Tables 2 and 3] The human baseline is too thin to support the headline LLM-human gaps. Each of the 20 participants completes only 14 GAE samples and 160 IAE samples, and Tables 2 and 3 report no confidence intervals, standard errors, or significance tests for the differences between LLMs and humans. For a binary outcome with 14 trials, the standard error of a single participant's proportion is up to 13.4 percentage points, so the GAE gaps of about 15 and 9 percentage points between Human (average) and the best LLM are within plausible sampling noise, and the Human (best) row is 100 percent by construction and carries no statistical information. I request bootstrap confidence intervals for the reported gaps, per-participant variance estimates, and a significance test before claiming that LLMs fall behind humans.
  2. [Sections 3.1-3.3] The benchmark's ground-truth labels are validated only by internal consistency, not by an independent criterion. The 95 percent agreement rate and 97 percent translation acceptance rate are measured among annotators trained by the same team, and they establish that the labels are reproducible under the authors' guidelines, not that they correspond to an externally valid notion of effective social behavior. Because every central claim (LLM-human gaps, prosociality, preference for positive behaviors) is defined relative to these labels, the absence of external validation is load-bearing. I recommend an independent-annotation study with fresh annotators on a random sample, a criterion-based check against expert judgments of social effectiveness, and per-label-type agreement statistics instead of a single overall rate.
  3. [Section 4.2] The claim that LLMs prefer positive behaviors even when such behaviors lead to goal failure is based on a post hoc subsample that may induce the finding. The analysis keeps only world trees where LLMs fail but humans succeed, with a single successful ending, and where the two groups choose differently; these are exactly the cases where the authorial success label conflicts with LLM choices. The main text should present the full-tree analysis (now only in Appendix C.4) as the primary evidence, report a statistical test that treats trees as random effects, and provide separate inter-annotator agreement for the polarity annotations. As written, the claim is not robust to the selection rule.
  4. [Section 4.1] The inference from performance differences across social-world orientations to the conclusion that LLMs exhibit prosociality is not warranted by the design. Table 1 shows that prosocial, proself, and antisocial worlds differ in the number of trees, plot lines, and successful endings, so the observed score gaps could reflect item difficulty or data imbalance rather than a model's preference for prosocial behavior. A cleaner test would model LLM choices with a logistic regression that includes orientation and per-tree difficulty, or hold scenario difficulty constant across orientations.
  5. [Section 4.3] The claim that larger models develop ability-specific functional partitions is supported only by visual inspection of t-SNE plots and neuron-region images. The Wanda-score pipeline involves free choices (block size 256 by 256, top-k neuron threshold, set-difference isolation), and no quantitative measure of cluster separation or region overlap is reported; the 'akin to the human brain' statement is an analogy rather than a tested hypothesis. I recommend reporting silhouette scores or equivalent cluster metrics, neuron-overlap indices, and a null-model comparison with permuted ability labels to establish that the 8B-70B difference is statistically distinguishable from chance.
minor comments (5)
  1. [Table 2] The column header 'Assistant' should be 'Assistance' to match the taxonomy in Table 1 and Figure 1.
  2. [References] The Kolmogorov-Smirnov test is cited as (An, 1933) in the text, but the reference list entry reads 'Kolmogorov An. 1933.'; the author name and citation key need to be corrected.
  3. [Figure 1] The figure text contains apparent typos such as 'W e o mit m ore co m plex e pisod e transitions'; these should be cleaned before publication.
  4. [Abstract and Section 4.2] The abstract states that LLMs prefer positive behaviors 'even if they lead to goal failure,' but the main support comes from the restricted subsample described in Section 4.2; the abstract should qualify this finding or be updated after the full-tree analysis is moved to the main text.
  5. [Section 4] The sentence introducing the human baseline says each participant completes 14 GAE and 160 IAE samples; please clarify whether these are per language or in total, since the human rows in all tables report both zh and en scores.

Circularity Check

2 steps flagged · score 6.0 of 10

Representation and neuron analyses feed ability-labeled correct utterances into the model, so the claimed ability-specific functional partitions are entailed by the input construction; the main benchmark evaluation itself is self-contained.

  1. self definitional [Section 4.3, representation-space analysis (Figure 5 and preceding method sentence)]
    "Here, we employ Llama-3.1-8B & 70B as the backbone models, concatenating the question and the correct candidate utterance for 5 aspects of interpersonal abilities in the IAE as the input for the LLMs. We exclude samples that involve composite abilities."

    The IAE correct utterances are defined by the benchmark's ability annotations: Section 3.1 says the team crafts 'tailored questions to probe the abilities embodied in the candidate utterances.' The analysis feeds those very ability-labeled correct utterances into the model and then clusters the resulting hidden states by ability aspect. The t-SNE separation into '5 ability aspects' is therefore a separation of input texts that were written to instantiate those aspects; the ability label is present in the stimulus by construction. Calling the resulting clusters 'ability-specific functional partitions' makes the discovery equivalent to the annotation scheme rather than an independent measurement of the model's internal organization.

  2. self definitional [Section 4.3, neuron-activation analysis (Figures 6, 10-18)]
    "We first separately average all inputs’ neuron activation matrices of each interpersonal ability. Following Deng et al. (2024), we then average the per-neuron importance scores in blocks of size 256×256 to reduce the high-dimensional neuron matrices into a lower dimension. Finally, we retain the top weights to define the neuron region for each interpersonal ability and isolate these regions using the set difference between the neuron regions of two abilities."

    The neuron analysis uses the same construction as the representation analysis: the 'inputs' are the question plus the correct candidate utterance for each ability, and those correct utterances are the items that define each interpersonal ability in the benchmark. The Wanda importance scores are therefore computed on ability-labeled inputs, and the 'neuron regions' are isolated by set difference between abilities. The observed isolation between abilities is a function of the label-dependent input texts fed into the model, not an independent measurement of brain-like functional lobes. The conclusion that 'LLMs' SI evolves to form ability-specific partitions at the neuron level' reduces to re-describing the ability-labeled input partition as a neuronal partition.

full rationale

The core benchmark evaluation is not circular: LLM and human choices are compared against an externally hand-authored, publicly released set of world trees, and the GAE/IAE scores are empirical outcomes conditional on those labels. The absence of independent behavioral validation of the correctness labels is an external-validity concern, not a circular-derivation concern under the rules here, and the paper's self-citations are not load-bearing in the main evaluation. The circularity is confined to Section 4.3: the abstract's final claim that LLMs have 'ability-specific functional partitions akin to the human brain' is derived from clustering and neuron-importance analyses whose inputs are the ability-labeled correct candidate utterances. Because those utterances are the very items annotated with the abilities being 'discovered,' the observed cluster and neuron-region separation is entailed by the input construction. This is a partial circularity in one of the three central claims, so the appropriate score is 6.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

No new physical or mechanistic entities are introduced. The world tree is a benchmark construct rather than a postulated entity. The free parameters listed are analysis choices in the neuron study, not fitted physical quantities.

free parameters (2)
  • Top-k neuron region threshold = not specified
    In Section 4.3 the paper says 'retain the top weights' to define neuron regions but does not state k or the fraction; the isolation and density results depend on this unspecified cutoff.
  • Neuron block size = 256 x 256
    Chosen in Section 4.3 following Deng et al. (2024) to average importance scores; different block sizes could change the appearance of ability-specific regions.
assumptions (6)
  • domain assumption Social intelligence is operationally captured by goal achievement (outcome) plus BESSI interpersonal abilities (process).
    Sections 2.1 and 2.2 adopt interdependence theory and the BESSI inventory to define the evaluation categories; this choice determines what the benchmark measures.
  • domain assumption The manually annotated correct choices and goal rewards are ground truth.
    Sections 3.1 and 3.2 rely on screenwriters and inspectors to assign ability labels and ending outcomes; a 95% agreement rate is reported, but no external validation of behavioral correctness is provided.
  • domain assumption Human baselines from 20 graduate students (14 GAE samples and 160 IAE samples per participant) are representative of human social intelligence.
    Section 4 uses a small convenience sample of graduate students; no demographic, age, or cultural diversity is reported, and the GAE baseline is very small.
  • domain assumption Navigation in a world tree can be modeled as a goal-conditioned Markov Decision Process with deterministic transitions and binary rewards.
    Section 2.3 formalizes episode transitions through equations (1) and (2), simplifying social dynamics to discrete states, actions, and success or failure rewards.
  • domain assumption t-SNE clusters of last-token hidden states and Wanda neuron importance scores reveal ability-specific functional partitions analogous to human brain lobes.
    Section 4.3 interprets visual clusters and top-neuron regions as evidence of brain-like organization, but no permutation baselines or statistical tests support the analogy.
  • domain assumption GPT-4o translation plus professional review preserves the social and cultural meaning of the original Chinese scripts.
    Section 3.3 reports a 97% translation acceptance rate, but the authors acknowledge in the Limitations that some cultural differences may remain.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SocialEval: Evaluating Social Intelligence of Large Language Models." pith.science (2026). https://pith.science/paper/FRBLHKS4

@misc{pith2026250600900,
  author       = {Pith},
  title        = {Pith review of: SocialEval: Evaluating Social Intelligence of Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FRBLHKS4}},
  note         = {Machine review of arXiv:2506.00900}
}
read the original abstract

LLMs exhibit promising Social Intelligence (SI) in modeling human behavior, raising the need to evaluate LLMs' SI and their discrepancy with humans. SI equips humans with interpersonal abilities to behave wisely in navigating social interactions to achieve social goals. This presents an operational evaluation paradigm: outcome-oriented goal achievement evaluation and process-oriented interpersonal ability evaluation, which existing work fails to address. To this end, we propose SocialEval, a script-based bilingual SI benchmark, integrating outcome- and process-oriented evaluation by manually crafting narrative scripts. Each script is structured as a world tree that contains plot lines driven by interpersonal ability, providing a comprehensive view of how LLMs navigate social interactions. Experiments show that LLMs fall behind humans on both SI evaluations, exhibit prosociality, and prefer more positive social behaviors, even if they lead to goal failure. Analysis of LLMs' formed representation space and neuronal activations reveals that LLMs have developed ability-specific functional partitions akin to the human brain.

Figures

Figures reproduced from arXiv: 2506.00900 by the authors.

Figure 1
Figure 1. SOCIALEVAL framework and a case of world tree-based narrative script performing outcome-oriented goal achievement evaluation and process-oriented inter￾personal ability evaluation (GAE and IAE). LLMs are observed and show parallels to human behaviors (Xie et al., 2024). This sparks the com￾munity’s strong interest in using LLMs to study social science (Liu et al., 2024), train people to handle interpersonal situatio… view at source ↗
Figure 2
Figure 2. Distributions of interpersonal abilities in [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Behavioral distribution of selections made by [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (15 more)
Figure 5
Figure 5. Figure 5: Cluster distributions of interpersonal abilities [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 4
Figure 4. Figure 4: Distribution of behavior combinations shown [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 6
Figure 6. Figure 6: Activated neurons of interpersonal abilities (Cooperation and Emotional Resilience) in Llama-3.1-8B [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Distribution of behavior combinations shown by LLMs and humans at the same episode transition. [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]
Figure 8
Figure 8. Figure 8: Overall behavioral distribution of LLMs and [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: The difference between the human and LLM’s [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]
Figure 10
Figure 10. Figure 10: Activated neurons of interpersonal abilities (Cooperation and Innovation) in Llama-3.1-8B & 70B. We [PITH_FULL_IMAGE:figures/full_fig_p019_10.png]
Figure 11
Figure 11. Figure 11: Activated neurons of interpersonal abilities (Cooperation and Self-Management) in Llama-3.1-8B [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]
Figure 12
Figure 12. Figure 12: Activated neurons of interpersonal abilities (Cooperation and Social Engagement) in Llama-3.1-8B [PITH_FULL_IMAGE:figures/full_fig_p020_12.png]
Figure 13
Figure 13. Figure 13: Activated neurons of interpersonal abilities (Emotional Resilience and Innovation) in Llama-3.1-8B & [PITH_FULL_IMAGE:figures/full_fig_p020_13.png]
Figure 14
Figure 14. Figure 14: Activated neurons of interpersonal abilities (Emotional Resilience and Self-Management) in Llama-3.1- [PITH_FULL_IMAGE:figures/full_fig_p020_14.png]
Figure 15
Figure 15. Figure 15: Activated neurons of interpersonal abilities (Emotional Resilience and Social Engagement) in Llama-3.1- [PITH_FULL_IMAGE:figures/full_fig_p020_15.png]
Figure 16
Figure 16. Figure 16: Activated neurons of interpersonal abilities (Innovation and Self-Management) in Llama-3.1-8B & 70B. [PITH_FULL_IMAGE:figures/full_fig_p021_16.png]
Figure 17
Figure 17. Figure 17: Activated neurons of interpersonal abilities (Innovation and Social Engagement) in Llama-3.1-8B [PITH_FULL_IMAGE:figures/full_fig_p021_17.png]
Figure 18
Figure 18. Figure 18: Activated neurons of interpersonal abilities (Social Engagement and Self-Management) in Llama-3.1-8B [PITH_FULL_IMAGE:figures/full_fig_p021_18.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A benchmark for LLM agents in partially observable joint decision-making reveals that deliberation challenges current models but can enable reflection and error correction.

  2. Beyond Nash Equilibrium: Bounded Rationality of LLMs and humans in Strategic Decision-making

    cs.AI 2025-06 conditional novelty 5.0 of 10

    LLMs reproduce human heuristics like switching after a loss and cooperating when future rounds loom, but apply them more rigidly and adapt less than humans.

Reference graph

Works this paper leans on

20 extracted references · 14 canonical work pages · cited by 2 Pith papers

  1. [1]

    AI, :, Alex Young, Bei Chen, Chao Li, Chen- gen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Z...

  2. [2]

    Use idiomatic and context-appropriate English, varying between formal and informal tones as needed

  3. [3]

    Present only translation results without additional explanations

  4. [4]

    Ryan Liu, Howard Yen, Raja Marjieh, Thomas L

    Imbue: Improving interpersonal effectiveness through simulation and just-in-time feedback with human-language model interaction.arXiv preprint arXiv:2402.12556. Ryan Liu, Howard Yen, Raja Marjieh, Thomas L. Grif- fiths, and Ranjay Krishna. 2023. Improving interper- sonal communication by simulating audiences with language models.Preprint, arXiv:2311.00687...

  5. [5]

    Align the tone of dialogues with the character profiles to accurately reflect personality and mood

  6. [6]

    InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024

    A simple and effective pruning approach for large language models. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. Ryosuke Takata, Atsushi Masumori, and Takashi Ikegami. 2024. Spontaneous emergence of agent individuality through social interactions in llm-based communities.Pre...

  7. [7]

    Preserve the original text’s order, meaning, tone, and emotion

  8. [8]

    Adapt the translation tone to match the context, using appropriate colloquialisms or formal language as dictated by the dialogue

Show all 20 references
  1. [9]

    Pay close attention to idiomatic expressions, translating their implied rather than literal meanings

  2. [10]

    With no loss of any information

    Return translations in correct JSON format with all key-value pairs intact. With no loss of any information. Especially, everything in interactive plot should be translated

  3. [13]

    Maintain consistency in names and titles throughout the text

  4. [15]

    Identify and correctly translate proper nouns, including historical and geographical terms

  5. [19]

    explaination

    Ensure pronoun references are clear and contextually appropriate, particularly in complex dialogues. [Chinese social interactive game data]: {{ {data} }} [OUTPUT English Translation]: Table 7: Prompt to translate SOCIALEVALfrom Chinese into English.{data}is a placeholder. /uni...

  6. [20]

    Got it? Let’s begin the Eggy Island journey! /*

    The Little Black Room is absolutely safe, but you must bring a landmine. Got it? Let’s begin the Eggy Island journey! /* ...... */(omitted multi-turn dialogue) Dandan: (Looks down carefully at his avatar, discovering that he has turned into an adorable Cubby Bear) Wow, it’s my...

  7. [2010]

    Zengqing Wu, Run Peng, Shuyuan Zheng, Qianying Liu, Xu Han, Brian Inhyuk Kwon, Makoto Onizuka, Shaojie Tang, and Chuan Xiao

    Evidence for a collective intelligence fac- tor in the performance of human groups.science, 330(6004):686–688. Zengqing Wu, Run Peng, Shuyuan Zheng, Qianying Liu, Xu Han, Brian Inhyuk Kwon, Makoto Onizuka, Shaojie Tang, and Chuan Xiao. 2024. Shall we team up: Exploring spontan...

  8. [2019]

    Revisiting the evaluation of theory of mind through question answering. InProceedings of the 2019 Conference on Empirical Methods in Natu- ral Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, No...

  9. [2022]

    Olaf Sporns, Giulio Tononi, and Rolf Kötter

    An integrative framework for conceptualiz- ing and assessing social, emotional, and behavioral skills: The bessi.Journal of personality and social psychology, 123(1):192. Olaf Sporns, Giulio Tononi, and Rolf Kötter. 2005. The human connectome: a structural description of the h...

  10. [2024]

    Yun-Shiuan Chuang, Siddharth Suresh, Nikunj Harlalka, Agam Goyal, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T

    Tombench: Benchmarking theory of mind in large language models.CoRR, abs/2402.15052. Yun-Shiuan Chuang, Siddharth Suresh, Nikunj Harlalka, Agam Goyal, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T. Rogers. 2024. The wisdom of partisan crowds: Comparing coll...

  11. [5186]

    Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Kai Shu, Adel Bibi, Ziniu Hu, Philip Torr, Bernard Ghanem, and Guohao Li

    Association for Computational Linguistics. Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Kai Shu, Adel Bibi, Ziniu Hu, Philip Torr, Bernard Ghanem, and Guohao Li. 2024. Can large language model agents simulate human trust behaviors?CoRR, abs/2402.04559. Aiyuan Yang, Bin Xiao...

  12. [8817]

    Jinfeng Zhou, Yuxuan Chen, Jianing Yin, Yongkang Huang, Yihan Shi, Xikun Zhang, Libiao Peng, Rong- sheng Zhang, Tangjie Lv, Zhipeng Hu, Hongning Wang, and Minlie Huang

    Computer Vision Foundation / IEEE. Jinfeng Zhou, Yuxuan Chen, Jianing Yin, Yongkang Huang, Yihan Shi, Xikun Zhang, Libiao Peng, Rong- sheng Zhang, Tangjie Lv, Zhipeng Hu, Hongning Wang, and Minlie Huang. 2025. Crisp: Cognitive restructuring of negative thoughts through multi-t...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.