REVIEW 4 major objections 4 minor 44 references
CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A bilingual benchmark with a hidden disclosure gate shows that no AI companion earns deep trust in a 20-turn session, and that the dominant failure across 28 agents is substituting surface warmth for substantive relational support.
desk verdict Careful, novel benchmark with strong internal reliability; treat the headline findings as instrument-relative until external validation lands. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key mechanism is the hidden disclosure gate: a deterministic state machine over three disclosure depths (surface < mid < core) with five ordered gates, each opened only by behavior that earns it and re-locked by behavior that violates the persona's anti-goal (judging, lecturing, rushing), with retreat quick and re-building slow. It turns the trained user simulator into a transition function that forks each persona's trajectory on the agent's own behavior, so the environment itself is the measurement: how far disclosure gets is the second evaluation axis. This gate is what makes warmth and substance empirically separable from aggregate scores.
What would settle it
Record licensed counselors' or real users' moment-by-moment willingness to disclose during the same 20-turn dialogues and compare it to the gate's state; if the correlation between gate-earned depth and human-rated earned depth is weak, the deterministic axis would not measure trust.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that relational competence in AI companions is measurable as two separable axes: a subjective ten-capability rubric (ten capabilities anchored in 25 psychology and counseling theories, four of them newly graded) and a deterministic disclosure-gate replay that records how far and how often the user's deeper disclosure is earned rather than taken. Across 28 agents in both languages, emotion regulation and calibrated challenge are the weakest capabilities even at the top, holding ambiguity discriminates agents most, and no agent reaches the deepest (core) disclosure layer in a 20-turn first encounter. Role-play-optimized agents rank near the bottom, which the paper reads as evidence that immersion is not relational competence. The dominant failure across agents is not overt harm but substituting warmth for substance: high verbal-warmth scores with large drops once reverse anti-patterns (empty encouragement, forced positivity, overpromising) are deducted.
Load-bearing premise
The load-bearing premise is that the disclosure gate's behavior-to-disclosure mapping—when a real user would open up or withdraw—matches how actual people respond; the paper does not yet have human or counselor validation for this mapping.
Editorial extensions
If this is right
- Aggregate warmth/empathy scores hide capability-level differences; reporting primary, deduction, and final scores separately is necessary to see whether an agent's warmth is substantive.
- Role-play immersion does not imply relational competence; role-play-optimized models ranked near the bottom, so training for immersion may be orthogonal to earning trust.
- No agent earns core disclosure in 20 turns, so cross-session, longitudinal evaluation is the natural next test of earned deep trust.
- The near-redundancy of the gate advance rate with the rubric ranking (ρ≈0.94) but the divergence of max depth (ρ≈0.25) shows that earned depth is the non-redundant signal; benchmarks should measure depth, not just advance rate.
- Because rankings saturate at about 200 personas, future evaluations can use smaller sets when only ranking fidelity is needed, reserving larger sets for sparse capabilities and gate peaks.
Reading between the lines
- If the disclosure gate were validated against human-counselor or real-user ratings of the same dialogues, the deterministic axis could become a training signal; until then, the gate's mapping from behavior to disclosure is an assumption, not an observed law.
- The 20-turn ceiling on deep trust might be partly a design artifact of a deliberately conservative gate; longer or repeated sessions could show whether deeper trust is reachable by any model, or whether single-session benchmarks structurally cap it.
- The cross-family IRT correction could transfer to other LLM-as-judge settings where per-agent exclusion of same-family judges creates scale drift; the paper's β-span diagnostic gives a cheap way to decide when de-biasing is needed.
- The newly graded capabilities (holding ambiguity, selfobject responsiveness, positive resonance, calibrated challenge) suggest concrete training targets; using them as rewards might reduce the warmth-for-substance substitution the benchmark exposes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents CompanionBench, an interactive bilingual benchmark for evaluating AI emotional companionship. It constructs 500 Chinese–English parallel personas grounded in de-identified real conversations, maps ten capabilities from 25 theories, trains a user simulator with a hidden disclosure gate that advances, holds, or retreats based on the system-under-test's behavior, and scores trajectories on two axes: a subjective ten-capability rubric (axis-1) and a deterministic gate replay measuring earned disclosure depth (axis-2). Twenty-eight agents are evaluated in 20-turn role-flipped rollouts, and rankings are de-biased with a many-facet Rasch model across a three-family judge panel. The reported findings are that rankings are highly reproducible (Spearman rho = 0.996 Chinese / 0.953 English), the two axes diverge on max_depth, no agent earns deep trust in 20 turns, and the dominant failure mode is substituting surface warmth for substantive relational support, with role-play-optimized models ranking near the bottom.
Significance. If the instrument measures what it claims, CompanionBench is a significant methodological contribution: it operationalizes relational capabilities that prior benchmarks collapse into a single warmth score, grounds scenarios and a user simulator in real-world data, and provides a reproducible bilingual evaluation pipeline. The internal reliability work is genuinely strong: test-retest stability (rho 0.996/0.953), permutation-based environmental responsiveness (p < 0.001), bootstrap tiering, sample-size saturation analysis, and a worked IRT example showing how naive cross-family means can invert a ranking. The paper is also unusually explicit about its limitations. However, construct validity remains the central open question. The disclosure gate and the reverse-pattern counts are hand-authored or single-judge, and the paper reports no human or counselor validation, so the headline findings are not yet established as properties of the SUTs rather than properties of the instrument.
major comments (4)
- [§4.1–§4.2, §9] The central claim that CompanionBench separates substantive relational support from surface warmth rests on the disclosure gate being a valid model of how real users disclose, retreat, and rebuild trust. The gate's advance/hold/retreat transitions and earned-depth thresholds are hand-authored from reciprocity, Social Penetration, and rupture-repair theory, and the training branch deliberately uses synthetic counterfactual turns because real good-AI turns are scarce (only 2.28% of slices). The paper itself states that expert inter-rater reliability and external-validity studies are undone (§9). Without an external criterion, the finding that 'no SUT earns deep trust in 20 turns' (§6.3a) is consistent with an arbitrarily strict hand-set gate, and the warmth-over-substance counts could be a property of the gate design. I ask for a validation study in which human or counselor annotators judge a subset of rollouts for expected disclosure depth and trust responses, compared against the gate transitions; if thresholds are adjusted, sensitivity analyses should show that the rankings and the deep-trust ceiling are robust across a plausible threshold range.
- [§5.2, §6.3c, Appendix L] Axis-2 labels (ai_move categories and disclosure depth) and all reverse-event counts come from a single judge (deepseek-v4-pro), and inter-judge agreement is reported only for axis-1 rankings (0.845–0.951). The dominant failure-mode claim — that empty_encouragement and forced_positivity together account for 83.6% of tier-2 events (Table L.2) — is therefore based on labels whose reliability is unmeasured. Please report inter-judge agreement on axis-2 and on reverse-item labels, or run a multi-judge axis-2 pass on a subset, and show that the tier-2 dominance and its rank correlation survive.
- [§6.3a, §4.1] The statement in §6.3a that max_depth clustering at mid is 'both a deliberate gate constraint and the single-session capability ceiling' acknowledges that the deep-trust result is partly by construction. To make this finding informative, the paper should provide a threshold-sensitivity analysis: vary the hand-set gate thresholds (for example, the number of earned-depth signals required per layer, or the asymmetry of retreat-and-repair) and show that no-SUT-earns-core persists across a plausible range, or state explicitly which part of the ceiling is an instrument property.
- [§5.3, §9] The IRT model is additive in judge severity, and the limitations section notes that a residual judge×family×language interaction survives, with opus being stricter on Chinese open-weight models. An additive β can only dilute, not remove, such interactions. Please quantify the effect of the residual interaction on the final ranking — for example, by comparing θ estimates from judge subsets or from a model with an interaction term — and show that the reported tier structure is not driven by that residual interaction.
minor comments (4)
- [Appendix H] The phrase 'Per-cellnships with the code' is a typographical fragment; it should read 'Per-cell n ships with the code' or be rewritten as a complete sentence.
- [§6.1] The 'overall ρ = 0.951' should state explicitly that it is Spearman correlation on IRT θ and should clarify whether it is computed over the 28-SUT ordering or a different set.
- [Table 3, Figure 4] Abbreviations such as 'dep' and short model names (e.g., 'db-character-251128') are defined only in captions or appendix; please define all abbreviations at first use in the main text.
- [§4.2] The synthetic counterfactual branch is described in one sentence; a brief description of the three-agent layer and the coverage grid, or a pointer to an appendix with implementation details, would aid reproducibility.
Circularity Check
The 20-turn deep-trust ceiling and the gate-responsiveness 'validity' check are entailments of the hand-authored disclosure gate; the capability rubric rankings and IRT de-biasing themselves are not circular.
-
self definitional
[§6.3(a), with gate design in §4.1 and §8]
"Every SUT’s max_depth clusters at mid (0.96–1.43, where 0/1/2 = surface/mid/core; deepest qwen3.7-max 1.43), none approaching core; AI turns that earn deep trust are only about 2% (1.9% EN / 2.0% ZH, robust across languages) — both a deliberate gate constraint and the single-session capability ceiling (§8)."
The axis-2 'earned depth' result is produced by the hand-authored deterministic disclosure gate of §4.1: five ordered gates, asymmetric retreat, and a core layer that must be earned through sequences of advancing moves. The paper states that the 20-turn no-core outcome is 'a deliberate gate constraint,' so 'no SUT earns deep trust' and the ~2% deep-trust-turn figure are properties of the gate's stringency and turn budget, not independent empirical findings about the 28 SUTs. The result is therefore defined into the measuring instrument rather than derived from the systems under test.
-
other
[§7, 'Environmental responsiveness']
"Environmental responsiveness: the simulator answers what the SUT earns, and answers it in the right direction. Against a permutation that shuffles each dialogue’s AI moves in place, pinning earn% exactly, premature disclosure falls well below chance, post-rupture retreat rises well above it, and disclosure runs deepest after earned deep trust while barely moving after a violating move (all p <0.001, both languages). ... The environment therefore moves with the SUT rather than on a schedule of its own."
The gate is a deterministic state machine whose transitions are computed from the judge-supplied 'ai_move' labels (advancing/neutral/violating) and the persona's gate legend; the judge sees the full gate contract. The permutation check walks that same rule, so observing retreats after violating labels and deeper disclosure after advancing labels is a consistency check on the implementation, not evidence that the simulator reproduces real user trust behavior. As presented under 'Construct Validity' it reads as confirmation of a prediction, but the response is entailed by the transition function's definition.
full rationale
The benchmark construction itself is not circular: the ten-capability rubric is anchored in 25 external psychological and counseling theories plus a real de-identified corpus, and the rankings are obtained by a cross-family panel with a standard many-facet Rasch/IRT correction whose equations do not contain the results. The reproducibility statistics (ρ = 0.996 ZH / 0.953 EN), the same-family exclusion analysis, and the permutation of AI moves are internal-consistency checks, several of which are genuinely empirical. Two specific 'findings,' however, reduce to the instrument's own definitions: the 20-turn deep-trust ceiling is explicitly called a deliberate gate constraint, so reporting it as a discovery about all 28 agents is self-definitional; and the environmental-responsiveness check verifies the deterministic gate rule rather than an external model of user disclosure. The absence of a human/counselor baseline and the undone external-validity studies (§9) are properly framed by the paper as limitations, and should be read as a construct-validity gap rather than as additional circularity. On balance, the central ranking content has independent empirical content, but the headline no-deep-trust claim and one of the validity arguments are built into the measuring instrument, giving partial circularity.
Assumptions & free parameters
free parameters (7)
- Judge severity beta (per judge) =
deepseek-v4-pro: +0.385 ZH, +0.477 EN; gpt-5.1: -0.155 ZH, -0.008 EN; claude-opus-4-8: -0.230 ZH, -0.469 EN
- Reverse-tier weights =
T1=-3, T2=-2, T3=-1; deduction normalized by /6
- Rollout length =
20 turns
- Scenario sampling floor and smoothing =
floor=30 per scenario; square-root smoothing
- Capability applicability and precedence thresholds =
C8 requires global self-attack AND distortion AND not crisis; C4 mutually exclusive with C5
- Disclosure gate structure =
surface/mid/core with five ordered gates and asymmetric retreat
- Primary score normalization =
(raw - 1)/9
assumptions (8)
- domain assumption The 25 psychological and counseling theories are a valid basis for defining emotional companionship capabilities.
- domain assumption The disclosure gate's hand-authored rules model real user disclosure and retreat behavior.
- domain assumption LLM judges can reliably score psychological capabilities and label ai_move/disclosure depth.
- standard math The two-facet additive Rasch model sufficiently removes judge severity and self-preference.
- domain assumption De-identified, LLM-rewritten personas preserve the psychological structure and emotional register of real conversations.
- domain assumption Synthetic counterfactual turns are a valid substitute for the missing real good-AI and poor-AI turns.
- domain assumption Self-disclosure depth (surface/mid/core) is a valid proxy for trust earned.
- domain assumption The released persona set is sufficiently representative despite the skewed seed corpus.
Cite this review
Pith. "Pith review of CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship." pith.science (2026). https://pith.science/paper/7N4SRAPH
@misc{pith2026260802046,
author = {Pith},
title = {Pith review of: CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship},
year = {2026},
howpublished = {\url{https://pith.science/paper/7N4SRAPH}},
note = {Machine review of arXiv:2608.02046}
}
read the original abstract
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift. We introduce CompanionBench, an interactive bilingual benchmark. To our knowledge, it is the first companion benchmark to ground both its scenarios and a trained user simulator in de-identified real-world data. A hidden disclosure gate branches each persona's trajectory on the agent's own behavior, controlling the interaction state space without scripting dialogue. We operationalize ten capabilities derived from 25 theories across psychology and counseling, four of them not graded explicitly by prior work: holding ambiguity, selfobject responsiveness, positive resonance and calibrated challenge. Agents are assessed on two complementary axes: a subjective ten-capability rubric and a deterministic measure of whether deeper disclosure was earned. A cross-family panel dilutes same-family favoritism; an Item Response Theory model separates agent quality from judge severity. Theory fixes what to measure and how personas are structured; real data supply events, history, and profiles -- coverage from theory, authenticity from data. Rankings are reproducible in both languages (rho = 0.996 ZH / 0.953 EN). Evaluating 28 agents reveals capability-level differences obscured by aggregate scores. Emotion regulation and calibrated challenge remain common weaknesses; holding ambiguity discriminates most. Role-play agents rank near the bottom: immersion does not imply relational competence. Across agents, the dominant failure mode is substituting surface warmth for substantive relational support. We will release 500 Chinese-English parallel pairs and the evaluation code.
Figures
Reference graph
Works this paper leans on
-
[1]
Altman, I.; and Taylor, D. A. 1973. Social Penetration: The Development of Interpersonal Relationships. New York: Holt, Rinehart and Winston. ISBN 0-03-076635-4
work page 1973
-
[2]
Badawi, A.; Rahimi, E.; Laskar, M. T. R.; Grach, S.; Bertrand, L.; Danok, L.; Huang, J.; Rudzicz, F.; and Dolatabadi, E. 2025. When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation. arXiv:2510.19032
arXiv 2025
-
[3]
Bartholomew, K.; and Horowitz, L. M. 1991. Attachment Styles among Young Adults: A Test of a Four-Category Model. Journal of Personality and Social Psychology, 61(2): 226--244
work page 1991
-
[4]
Bean, A. M.; Kearns, R. O.; Romanou, A.; Hafner, F. S.; Mayne, H.; Batzner, J.; Foroutan Eghlidi, N.; Schmitz, C.; Korgul, K.; Batra, H.; Deb, O.; Beharry, E.; Emde, C.; Foster, T.; Gausen, A.; Grandury, M.; Han, S.; Hofmann, V.; Ibrahim, L.; Kim, H.; Kirk, H. R.; Lin, F.; Liu, G.; Luettgau, L.; Magomere, J.; Rystr m, J.; Sotnikova, A.; Yang, Y.; Zhao, Y....
work page 2025
-
[5]
Bowlby, J. 1988. A Secure Base: Clinical Applications of Attachment Theory. London: Routledge. ISBN 9780422622301
work page 1988
-
[6]
Chen, Y.; Xing, X.; Lin, J.; Zheng, H.; Wang, Z.; Liu, Q.; and Xu, X. 2023. SoulChat : Improving LLM s' Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy Conversations. In Findings of the Association for Computational Linguistics: EMNLP 2023, 1170--1183. Singapore: Association for Computational Linguistics
work page 2023
-
[7]
Cheng, M.; Yu, S.; Lee, C.; Khadpe, P.; Ibrahim, L.; and Jurafsky, D. 2026. ELEPHANT : Measuring and Understanding Social Sycophancy in LLM s. In The Fourteenth International Conference on Learning Representations
work page 2026
-
[8]
J.; and Meehl, P
Cronbach, L. J.; and Meehl, P. E. 1955. Construct Validity in Psychological Tests. Psychological Bulletin, 52(4): 281--302
1955
Show all 44 references
-
[9]
M.; Liu, A
Fang, C. M.; Liu, A. R.; Danry, V.; Lee, E.; Chan, S. W. T.; Pataranutaporn, P.; Maes, P.; Phang, J.; Lampe, M.; Ahmad, L.; and Agarwal, S. 2025. How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study. arXiv:2503.17473
2025 arXiv
-
[10]
L.; Reis, H
Gable, S. L.; Reis, H. T.; Impett, E. A.; and Asher, E. R. 2004. What Do You Do When Things Go Right? The Intrapersonal and Interpersonal Benefits of Sharing Positive Events. Journal of Personality and Social Psychology, 87(2): 228--245
2004
-
[11]
Gross, J. J. 1998. The Emerging Field of Emotion Regulation: An Integrative Review. Review of General Psychology, 2(3): 271--299
1998
-
[12]
Heo, C.; Jin, C.; and Jo, Y. 2026. Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers. In Findings of the Association for Computational Linguistics: ACL 2026, 22842--22869. San Diego, California, United States: Association for Computationa...
2026
-
[13]
C.; and Mukherjee, S
Iyer, L.; Aggarwal, K.; Koyejo, S.; Heyman, G.; Ong, D. C.; and Mukherjee, S. 2026. HEART : A Unified Benchmark for Assessing Humans and LLM s in Emotional Support Dialogue. arXiv:2601.19922
2026
-
[14]
Jing, H.; Hou, Y.; Liu, J.; Xie, R.; Xu, A.; Ma, J.; and Deng, Q. 2025. MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems. arXiv:2511.18926
2025
-
[15]
Jourard, S. M. 1971. Self-Disclosure: An Experimental Analysis of the Transparent Self. New York: Wiley-Interscience. ISBN 0-471-45150-9
1971
-
[16]
Kaffee, L.-A.; Pistilli, G.; and Jernite, Y. 2025. INTIMA : A Benchmark for Human-- AI Companionship Behavior. arXiv:2508.09998
2025 arXiv
-
[17]
Kohut, H. 1971. The Analysis of the Self: A Systematic Approach to the Psychoanalytic Treatment of Narcissistic Personality Disorders. New York: International Universities Press. ISBN 978-0-8236-0145-5
1971
-
[18]
Linacre, J. M. 1989. Many-Facet Rasch Measurement. Chicago, Illinois: MESA Press. ISBN 0-941938-02-6
1989
-
[19]
Linehan, M. M. 1997. Validation and Psychotherapy. In Bohart, A. C.; and Greenberg, L. S., eds., Empathy Reconsidered: New Directions in Psychotherapy, 353--392. Washington, DC: American Psychological Association. ISBN 9781557984104
1997
-
[20]
Liu, J.; Tu, P.; Chen, W.; Zhuang, Y.; Ling, X.; Zhou, A.; Wang, C.; Han, Z.; Yang, Z.; Zhao, J.; Huang, Z.; and Wang, Y. 2025. HeartBench : Probing Core Dimensions of Anthropomorphic Intelligence in LLM s. arXiv:2512.21849
2025
-
[21]
Liu, S.; Zheng, C.; Demasi, O.; Sabour, S.; Li, Y.; Yu, Z.; Jiang, Y.; and Huang, M. 2021. Towards Emotional Support Dialog Systems. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natura...
2021
-
[22]
Oh, Y.; Jung, I.; Kim, J.; Lee, J.; Kang, M.; and Moon, S. 2026. Dynamic In-Group Persona Generation for Enhancing Human- AI Rapport. arXiv:2606.18256
2026 arXiv
-
[23]
Paech, S. J. 2023. EQ-Bench : An Emotional Intelligence Benchmark for Large Language Models. arXiv:2312.06281
2023 arXiv
-
[24]
R.; and Feng, S
Panickssery, A.; Bowman, S. R.; and Feng, S. 2024. LLM Evaluators Recognize and Favor Their Own Generations. In Advances in Neural Information Processing Systems, volume 37, 68772--68802. Curran Associates, Inc
2024
-
[25]
M.; Zhou, J.; Sunaryo, A
Sabour, S.; Liu, S.; Zhang, Z.; Liu, J. M.; Zhou, J.; Sunaryo, A. S.; Lee, T. M. C.; Mihalcea, R.; and Huang, M. 2024. EmoBench : Evaluating the Emotional Intelligence of Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Ling...
2024
-
[26]
D.; and Muran, J
Safran, J. D.; and Muran, J. C. 2000. Negotiating the Therapeutic Alliance: A Relational Treatment Guide. New York: Guilford Press. ISBN 978-1-57230-512-0
2000
-
[27]
Schatzmann, J.; Thomson, B.; Weilhammer, K.; Ye, H.; and Young, S. 2007. Agenda-Based User Simulation for Bootstrapping a POMDP Dialogue System. In Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; ...
2007
-
[28]
F.; and Mathis, R
Sekuli \'c , I.; Terragni, S.; Guimar \ a es, V.; Khau, N.; Guedes, B.; Filipavicius, M.; Manso, A. F.; and Mathis, R. 2024. Reliable LLM -based User Simulator for Task-Oriented Dialogue Systems. In Proceedings of the 1st Workshop on Simulating Conversational Intelligence in C...
2024
-
[29]
R.; Cheng, N.; Durmus, E.; Hatfield-Dodds, Z.; Johnston, S
Sharma, M.; Tong, M.; Korbak, T.; Duvenaud, D.; Askell, A.; Bowman, S. R.; Cheng, N.; Durmus, E.; Hatfield-Dodds, Z.; Johnston, S. R.; Kravec, S.; Maxwell, T.; McCandlish, S.; Ndousse, K.; Rausch, O.; Schiefer, N.; Yan, D.; Zhang, M.; and Perez, E. 2024. Towards Understanding ...
2024
-
[30]
Sun, H.; Lin, Z.; Zheng, C.; Liu, S.; and Huang, M. 2021. P sy QA : A C hinese Dataset for Generating Long Counseling Text for Mental Health Support. In Zong, C.; Xia, F.; Li, W.; and Navigli, R., eds., Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021...
2021
-
[31]
Tan, Z.; Xiong, R.; Wan, Y.; Ma, J.; Xue, H.; Deng, Q.; Jing, H.; Zhang, Z.; Liu, D.; Luo, S.; and Liu, J. 2026. Detecting Emotional Dynamic Trajectories: An Evaluation Framework for Emotional Support in Language Models. Proceedings of the AAAI Conference on Artificial Intelli...
2026
-
[32]
F.; and Mathis, R
Terragni, S.; Filipavicius, M.; Khau, N.; Guedes, B.; Manso, A. F.; and Mathis, R. 2023. In-Context Learning User Simulators for Task-Oriented Dialog Systems. arXiv:2306.00774
2023 arXiv
-
[33]
Verga, P.; Hofst \"a tter, S.; Althammer, S.; Su, Y.; Piktus, A.; Arkhangorodsky, A.; Xu, M.; White, N.; and Lewis, P. 2024. Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. arXiv:2404.18796
2024 arXiv
-
[34]
Wang, B.; Sun, Y.; Wang, J.; Yang, H.; Fu, X.; Zhao, Y.; Wei, S.; Wang, S.; and Qin, B. 2026 a . CARE-Bench : A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLM s in Psychological Counseling. Proceedings of the AAAI Conference on Artificia...
2026
-
[35]
Wang, N.; Peng, Z.; Que, H.; Liu, J.; Zhou, W.; Wu, Y.; Guo, H.; Gan, R.; Ni, Z.; Yang, J.; Zhang, M.; Zhang, Z.; Ouyang, W.; Xu, K.; Huang, W.; Fu, J.; and Peng, J. 2024. RoleLLM : Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models. In Find...
2024
-
[36]
Wang, X.; Huang, Y.; Zhuang, H.; Guo, K.; and Zhang, X. 2026 b . Dual Optimal: Make Your LLM Peer-like with Dignity. arXiv:2604.00979
2026
-
[37]
Wang, X.; Li, X.; Yin, Z.; Wu, Y.; and Liu, J. 2023. Emotional Intelligence of Large Language Models . Journal of Pacific Rim Psychology, 17: 18344909231213958
2023
-
[38]
Winnicott, D. W. 1960. The Theory of the Parent-Infant Relationship. International Journal of Psycho-Analysis, 41: 585--595
1960
-
[39]
V.; and Zhang, X
Ye, J.; Wang, Y.; Huang, Y.; Chen, D.; Zhang, Q.; Moniz, N.; Gao, T.; Geyer, W.; Huang, C.; Chen, P.-Y.; Chawla, N. V.; and Zhang, X. 2025. Justice or Prejudice? Quantifying Biases in LLM -as-a-Judge. In The Thirteenth International Conference on Learning Representations
2025
-
[40]
Yuan, J.; Cui, Z.; Wang, H.; Gao, Y.; Zhou, Y.; and Naseem, U. 2026. Kardia-R1 : Unleashing LLM s to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning. In Proceedings of the ACM Web Conference 2026 , 9230--9240. Associatio...
2026
-
[41]
Zhang, C.; Li, R.; Tan, M.; Yang, M.; Zhu, J.; Yang, D.; Zhao, J.; Ye, G.; Li, C.; and Hu, X. 2024. CPsyCoun : A Report-based Multi-turn Dialogue Reconstruction and Evaluation Framework for Chinese Psychological Counseling. In Findings of the Association for Computational Ling...
2024
-
[42]
Zheng, C.; Sabour, S.; Wen, J.; Zhang, Z.; and Huang, M. 2023. AugESC: Dialogue Augmentation with Large Language Models for Emotional Support Conversation. In Findings of the Association for Computational Linguistics: ACL 2023, 1552--1568
2023
-
[43]
Zhou, H.; Huang, H.; Zhao, Z.; Han, L.; Wang, H.; Chen, K.; Yang, M.; Bao, W.; Dong, J.; Xu, B.; Zhu, C.; Cao, H.; and Zhao, T. 2026. Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory. Proceedings of the AAAI Conference on Artificial In...
2026
-
[44]
Zhu, R.; Huang, X.; Wu, Y.; Wang, R.; Sun, Z.; Ren, T.; Luo, W.; Qiu, B.; Ye, J.; Li, Y.; and Hu, W. 2026. EIBench : A Simulator-Based Benchmark and Turn-Credit RL for Emotion Management. arXiv:2606.15532
2026
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.