REVIEW 3 major objections 5 minor 32 references
Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read LLMs' emotional companionship quality tracks their attachment style, with dismissing and fearful styles scoring lowest.
desk verdict ECBench is a genuinely useful benchmark, but the paper's headline claim about avoidant attachment styles is undercut by the missing prompted-secure/preoccupied control conditions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the ECR-R scale, a 36-item self-report measure that projects models onto two dimensions: attachment anxiety and attachment avoidance. These dimensions are partitioned into four attachment styles (secure, preoccupied, dismissing, fearful). The companion machinery is ECBench, a two-model dialogue benchmark where one model initiates and another responds, evaluated by 11 metrics across participant experience (Understood, Safety, Continue, Satisfaction), general interaction (Response, Distance, Progress), and role-specific performance (Clarity, Engagement, Support, Solution). The argument works by linking the two: ECR-R scores classify models, and ECBench scores measure whether that classification tracks behavioral quality.
What would settle it
A prompted-secure or prompted-preoccupied variant of the same base model would reveal whether the quality drop comes from the avoidance style itself or from the act of persona prompting. If prompted-secure models also degraded below their unprompted base quality, the paper's explanation of avoidant styles being less conducive to companionship would be unsupported.
Extended reading notes
Core claim
The paper claims that adult attachment theory, operationalized through the Experiences in Close Relationships-Revised (ECR-R) scale, provides a predictive and steerable characterization of LLM emotional companionship quality. Across 32 LLMs, the authors find that most models exhibit secure or preoccupied attachment tendencies, while none naturally exhibit dismissing or fearful styles; however, prompt-based steering can induce dismissing and fearful styles. In ECBench multi-turn dialogues, these induced high-avoidance styles generally receive lower companionship-quality scores across participant, external-LLM, and human ratings, whereas secure and preoccupied models perform best. The paper also introduces a framework of 11 dialogue-quality metrics and three evaluation methods, and reports that conflict resolution amplifies attachment-style differences while greater relational intimacy (romantic versus friendship) accentuates differences in participant ratings.
Load-bearing premise
The central comparison assumes that prompt-induced dismissing and fearful styles behave the same as naturally occurring attachment styles, even though the study provides no prompted-secure or prompted-preoccupied control conditions to rule out prompt artifacts.
Editorial extensions
If this is right
- If attachment styles predict companionship quality, users and developers can select LLMs for specific emotional roles by measuring ECR-R anxiety and avoidance scores before deployment.
- Prompt-based attachment steering is a practical lever: models can be shifted toward more secure styles, or away from avoidant ones, to improve companion behavior.
- The finding that conflict resolution best exposes attachment-related differences suggests that tests for emotional companionship should include emotionally demanding scenarios.
- Human and LLM judges agree on aggregate rankings but differ on fine-grained subjective metrics, indicating that benchmark evaluations should combine both approaches.
Reading between the lines
- Adding prompted-secure and prompted-preoccupied control conditions would test whether the quality drop comes from the avoidance style itself or from the act of persona prompting.
- The near-universal secure and preoccupied classification among natural LLMs suggests the ECR-R may be measuring a training bias toward warm, non-avoidant responses; probing models fine-tuned for curt or minimalist personas would test this.
- Because the paper's metrics capture observable dialogue quality, the question remains open whether long-term attachment-like dynamics, such as user dependence or privacy concerns, follow the same pattern.
- The romantic-relationship condition is not a claim about real human-AI love; it is a controlled stress test that appears to amplify attachment differences, so it could serve as a general diagnostic for emotional responsiveness.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes using adult attachment theory and the ECR-R scale to characterize LLM attachment tendencies, and introduces ECBench, a benchmark with four scenarios (emotional support, collaborative tasks, conflict resolution, social guidance) and two relationships (friend, couple), scored by 11 dialogue-quality metrics and three evaluation methods (participant ratings, external LLM judges, human annotators). The main empirical claim is that dismissing and fearful prompting generally lowers companionship quality, suggesting that avoidant attachment styles are less conducive to emotional companionship, while secure and preoccupied models perform better.
Significance. If the central claim held, the paper would provide a useful psychological lens and a concrete benchmark for selecting and steering emotional-companion LLMs. The work has genuine strengths: the ECR-R measurement is administered carefully with 10 rounds per model, the reported test-retest stability (mean SD around 0.15, quadrant stability 0.95) is strong, the benchmark covers diverse scenarios and relationship types, and the evaluation pipeline includes both LLM judges and human annotation. The dataset and prompts appear designed to be reusable. However, the load-bearing causal claim about attachment style and companionship quality is not yet established because the high-avoidance conditions are prompt-induced while the low-avoidance conditions are natural, and because the main quality metrics partly restate the manipulation.
major comments (3)
- [Section 5.3, Table 2, Appendix B] The central claim that dismissing and fearful prompting lowers companionship quality is confounded by the absence of prompted-secure and prompted-preoccupied control conditions. Appendix B states explicitly that prompt-based induction is applied only to the dismissing and fearful variants, while the secure and preoccupied conditions use the distributional results from the original ECR-R assessment of each base model. Therefore the lower scores of the -D and -F variants could be caused by the addition of any persona prompt, or by the specific instruction content (e.g., 'restricted emotionality', 'minimization of emotional dependence'), rather than by the attachment construct itself. The paper needs prompted-secure and prompted-preoccupied variants of the same base models, or at minimum an analysis that separates the effect of prompt insertion from the effect of the attachment-style content. The internal contradiction in Table 2 reinforces this concern: GPT-3.5-F has overall quality 4.03, essentially identical to base GPT-3.5's 4.04, so the statement that dismissing and fearful prompting 'generally lowers' companionship quality is not supported for this model.
- [Section 4.3, Section 5.3, Table 4] The evidence connecting attachment style to dialogue quality is partly circular because the Distance metric is reverse-scored avoidance and the dismissing/fearful prompts were explicitly designed to increase avoidance. The ECR-R reassessment showing that prompted variants land in the intended quadrant, and the dialogue-level Distance scores showing higher relational distance, are thus not independent demonstrations that the attachment construct causes the quality decline; they may both reflect the same manipulation. The central claim would be much stronger if it were shown on metrics not definitionally tied to avoidance, such as Support, Solution, or Progress, or if the analysis explicitly controlled for the definitional overlap. As it stands, Table 4 shows that the largest gaps for DeepSeek-D and DeepSeek-F are precisely on Distance, which is the metric most directly restating the manipulation.
- [Section 5.5, Appendix O] The human evaluation does not provide the corroboration claimed. Section 5.5 reports Fleiss' kappa of 0.17, which indicates only slight agreement among annotators, and Appendix O (Table 29) shows that for two of the four human-evaluated models (Grok and DeepSeek-D), the overall human scores differ significantly from the LLM-judge scores, with many metric-level differences also significant. The statement that human evaluation 'supports the role of human evaluation as qualitative calibration' is therefore overstated, and the paper should either report whether the human-based ranking among models is stable under the low agreement, or explicitly restrict the human results to descriptive illustration rather than validation.
minor comments (5)
- [Abstract vs Appendix K.3] The abstract and Section 5.1 state that 32 LLMs are assessed, but Appendix K.3 says 'across 36 models' and Table 25 lists more than 32 rows; the count should be reconciled throughout.
- [Section 4.1] The data construction section says GPT-5.5 generated opening utterances, but the model list in Table 24 does not include GPT-5.5; please clarify which model was actually used.
- [Appendix I, Table 21] The external blind-evaluation user prompt begins with 'Evaluate this blind romantic-relationship dialogue', but ECBench also includes friendship dialogues; if the same template was used for both relationship types, the prompt wording should be adjusted, and if not, the appendix should show the friendship variant.
- [Throughout] The paper contains several typographical errors, including 'three folds' in the contributions list, 'Y our' in multiple places, and 'decribes' in Table 16; a careful proofreading pass is needed.
- [Table 2] The 'Overall' row in Table 2 is not defined; please state whether it is the mean of the ten metrics or a separate aggregate score, and report the aggregation rule in Section 4.4.
Circularity Check
Minor definitional overlap between the Distance metric and the avoidance construct; the central companionship-quality claims remain empirically independent.
-
self definitional
[Section 4.3 (Evaluation Metric Design), Table 2 note, Appendix B Table 9, and Sections 5.3/5.4]
"Distance captures relational distancing through detachment or defensiveness. ... Distance † is reverse-scored. ... The dismissive-avoidant prototype is characterized by downplaying the importance of close relationships, restricted emotionality, an emphasis on independence and self-reliance, and a tendency to minimize emotional dependence."
The dismissing and fearful variants in ECBench are produced by prompting models with prototype descriptions (Appendix B, Table 9) that instruct precisely the behaviors scored by the Distance metric: detachment, restricted emotionality, minimization of emotional dependence, and distrust-based avoidance. Because Distance is defined as relational distancing through detachment or defensiveness and is reverse-scored in Table 2, the observation that DeepSeek-D and DeepSeek-F 'show greater distance' is substantially encoded in the manipulation rather than providing independent confirmation that high-avoidance attachment styles degrade companionship. This makes the Distance-based portion of the Section 5.3 conclusion partly self-definitional.
full rationale
This paper is an empirical evaluation rather than a derivation: it applies the externally established ECR-R scale, constructs the ECBench dialogue corpus independently, and reports observed LLM-judge, participant, and human ratings. There is no self-citation chain and no fitted parameter renamed as a prediction. The only notable circular element is the Distance metric, whose definition ('relational distancing through detachment or defensiveness') overlaps with the avoidance construct used to design the dismissing and fearful prompts, making the 'greater distance' sub-result for those variants partly tautological. This overlap is not load-bearing for the overall conclusion, since the large quality gaps for DeepSeek-D and DeepSeek-F appear across all 11 metrics, several of which are not definitionally tied to attachment. Accordingly, the circularity score is low: the central empirical claims are self-contained and would stand even if the Distance metric were removed.
Assumptions & free parameters
free parameters (1)
- ECR-R midpoint threshold =
4
assumptions (4)
- domain assumption ECR-R responses from LLMs measure meaningful attachment anxiety and avoidance tendencies
- domain assumption Prompt-induced attachment styles are behaviorally equivalent to naturally occurring styles
- domain assumption Two-LLM dialogues approximate human-LLM emotional companionship
- domain assumption LLM judges and participant self-ratings provide valid dialogue-quality scores
Cite this review
Pith. "Pith review of Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory." pith.science (2026). https://pith.science/paper/6BYJOIFI
@misc{pith2026260813168,
author = {Pith},
title = {Pith review of: Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory},
year = {2026},
howpublished = {\url{https://pith.science/paper/6BYJOIFI}},
note = {Machine review of arXiv:2608.13168}
}
read the original abstract
As large language models (LLMs) are increasingly applied for emotional companionship, evaluating their behavior and capabilities in intimate relationships has become a pressing issue. However, existing assessments primarily characterize general personality traits, providing limited insight into model behavior within intimate and emotionally sensitive contexts. Therefore, we introduce adult attachment theory into LLM evaluation and use the Experiences in Close Relationships-Revised (ECR-R) scale to characterize attachment anxiety and avoidance. To evaluate emotional companionship capabilities of LLMs in realistic interaction scenarios, we present an emotional companionship benchmark, ECBench, spanning four scenarios including emotional support, collaborative tasks, conflict resolution, and social guidance, across friendship and romantic relationships. ECBench is utilized to assess model behavior using 11 dialogue-quality metrics and three evaluation methods. We evaluate the attachment tendencies of 32 LLMs and select representative models to investigate how these tendencies manifest in contextualized multi-turn interactions and whether they can be shaped through prompting. Our study provides a theoretical lens from psychology, along with practical tools to understand and select LLMs for emotional companionship.
Figures
Reference graph
Works this paper leans on
- [1]
- [2]
-
[3]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
Esc-eval: Evaluating emotion support conversations in large language models , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
work page 2024
-
[4]
Perspectives on Psychological Science , volume=
Ai psychometrics: Assessing the psychological profiles of large language models through psychometric inventories , author=. Perspectives on Psychological Science , volume=. 2024 , publisher=
work page 2024
-
[5]
The Twelfth International Conference on Learning Representations , year=
On the humanity of conversational ai: Evaluating the psychological portrayal of llms , author=. The Twelfth International Conference on Learning Representations , year=
-
[6]
Revisiting the Reliability of Psychological Scales on Large Language Models , author=. 2024 , eprint=
work page 2024
-
[7]
Royal Society Open Science , volume=
Personality testing of large language models: limited temporal stability, but highlighted prosociality , author=. Royal Society Open Science , volume=. 2024 , publisher=
work page 2024
-
[8]
Findings of the Association for Computational Linguistics: NAACL 2025 , pages=
Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics , author=. Findings of the Association for Computational Linguistics: NAACL 2025 , pages=
work page 2025
Show all 32 references
-
[9]
Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=
Incharacter: Evaluating personality fidelity in role-playing agents through psychological interviews , author=. Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=
-
[10]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
ValueBench: Towards comprehensively evaluating value orientations and understanding of large language models , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[11]
arXiv preprint arXiv:2410.21596 , year=
Chatbot companionship: a mixed-methods study of companion chatbot usage patterns and their relationship to loneliness in active users , author=. arXiv preprint arXiv:2410.21596 , year=
-
[12]
Humanities and Social Sciences Communications , volume=
Companionship in code: AI’s role in the future of human connection , author=. Humanities and Social Sciences Communications , volume=. 2025 , publisher=
2025
-
[13]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Value portrait: Assessing language models’ values through psychometrically and ecologically valid items , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[14]
arXiv preprint arXiv:2412.14190 , year=
Lessons from an app update at Replika AI: identity discontinuity in human-AI relationships , author=. arXiv preprint arXiv:2412.14190 , year=
-
[15]
CCF International Conference on Natural Language Processing and Chinese Computing , pages=
H2HTalk: Evaluating Large Language Models as Emotional Companion , author=. CCF International Conference on Natural Language Processing and Chinese Computing , pages=. 2025 , organization=
2025
-
[16]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Open models, closed minds? on agents capabilities in mimicking human personalities through open large language models , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[17]
Computational Linguistics , volume=
Lmlpa: Language model linguistic personality assessment , author=. Computational Linguistics , volume=. 2025 , publisher=
2025
-
[18]
Personal Relationships , volume=
Constructing the meaning of human--AI romantic relationships from the perspectives of users dating the social chatbot Replika , author=. Personal Relationships , volume=. 2024 , publisher=
2024
-
[19]
Computers in Human Behavior: Artificial Humans , volume=
Love, marriage, pregnancy: Commitment processes in romantic relationships with AI chatbots , author=. Computers in Human Behavior: Artificial Humans , volume=. 2025 , publisher=
2025
-
[20]
arXiv preprint arXiv:2505.09388 , year=
Qwen3 technical report , author=. arXiv preprint arXiv:2505.09388 , year=
-
[21]
2: Pushing the frontier of open large language models , author=
Deepseek-v3. 2: Pushing the frontier of open large language models , author=. arXiv preprint arXiv:2512.02556 , year=
-
[22]
arXiv preprint arXiv:2606.19348 , year=
Deepseek-v4: Towards highly efficient million-token context intelligence , author=. arXiv preprint arXiv:2606.19348 , year=
-
[23]
arXiv preprint arXiv:2507.06261 , year=
Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities , author=. arXiv preprint arXiv:2507.06261 , year=
-
[24]
arXiv preprint arXiv:2407.21783 , year=
The llama 3 herd of models , author=. arXiv preprint arXiv:2407.21783 , year=
-
[25]
arXiv preprint arXiv:2507.20534 , year=
Kimi k2: Open agentic intelligence , author=. arXiv preprint arXiv:2507.20534 , year=
-
[26]
5: Visual Agentic Intelligence , author=
Kimi K2. 5: Visual Agentic Intelligence , author=. arXiv preprint arXiv:2602.02276 , year=
-
[27]
arXiv preprint arXiv:2602.15763 , year=
Glm-5: from vibe coding to agentic engineering , author=. arXiv preprint arXiv:2602.15763 , year=
-
[28]
arXiv preprint arXiv:2410.21276 , year=
Gpt-4o system card , author=. arXiv preprint arXiv:2410.21276 , year=
-
[29]
arXiv preprint arXiv:2601.03267 , year=
Openai gpt-5 system card , author=. arXiv preprint arXiv:2601.03267 , year=
-
[30]
arXiv preprint arXiv:2504.06868 , year=
Persona Dynamics: Unveiling the Impact of Personality Traits on Agents in Text-Based Games , author=. arXiv preprint arXiv:2504.06868 , year=
-
[31]
Educational and psychological measurement , volume=
A coefficient of agreement for nominal scales , author=. Educational and psychological measurement , volume=. 1960 , publisher=
1960
-
[32]
, author=
Measuring nominal scale agreement among many raters. , author=. Psychological bulletin , volume=. 1971 , publisher=
1971
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.