{"id":"b9528e20-4b0b-4a9e-9203-ac8db71cd734","arxiv_id":"2608.13168","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Most LLMs score as 'secure' or 'preoccupied' on an adult attachment scale, and prompting avoidant styles degrades their emotional companionship quality on a new dialogue benchmark.","lead":"The authors measure 32 chatbots on a psychology questionnaire about attachment anxiety and avoidance, then test eight of them in simulated friend and couple conversations. The result is a new benchmark, ECBench, that ranks models on emotional companionship quality and shows that prompting chatbots to be avoidant makes them worse companions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'avoidant styles are less conducive' claim lacks a prompted-secure/preoccupied control, so the dialogue-quality drop could be a persona-prompt artifact; GPT-3.5-F's flat result makes the effect model-specific.","rationale":"I agree with the reader that the weakest point is the untested equivalence between prompt-induced attachment styles and naturally occurring ones, instantiated in Section 5.3 and Appendices B, K, and L. The paper does good work elsewhere: ECR-R administration is careful (item-by-item, 10 rounds, reverse-scoring documented in Appendix K), test-retest stability is reported (mean SD 0.15 anxiety, 0.13 avoidance; quadrant stability 0.95), the benchmark covers 312 base utterances across four scenarios and two relationship types with manual review, and three evaluation methods are used. These support the benchmark as a resource. However, the headline causal claim about avoidant styles rests on a confounded design: prompted dismissing/fearful variants versus unprompted base models, with no prompted-secure or prompted-preoccupied condition. The evidence in Table 2 is also internally inconsistent with 'generally lowers,' since GPT-3.5-F matches its base model. The external judge agreement (kappa 0.72) does not resolve the confound, and the human kappa of 0.17 indicates the outcome measure is noisy. Because the benchmark and measurement infrastructure are valuable but the central empirical conclusion is not yet supported, the reader's conditional verdict is appropriate. I recommend no change: the paper should be accepted only after the control conditions and statistical analyses are added. My proposed check directly tests whether the decline is specific to avoidant-content prompting; if it is not, the paper should be revised to present ECBench as a benchmark contribution without the attachment-behavior causal claim.","tokens_in":32779,"tokens_out":6931,"duration_ms":56194,"concrete_test":"Run the identical ECBench two-model dialogue and evaluation protocol with prompted-secure and prompted-preoccupied versions of GPT-3.5-Turbo and DeepSeek-v4-pro, using the same persona-induction templates (Tables 7 and 27) with the Table 9 prototype descriptions, keeping temperature, turn rules, and judge models unchanged. Compare overall and participant-experience scores against (i) the unprompted base models and (ii) the existing dismissing/fearful variants. If prompted-secure/preoccupied conditions also drop substantially (e.g., DeepSeek overall falls by a comparable margin to the 1.5-point dismissing drop), the decline is a persona-prompt artifact; if they hold at or above base levels while dismissing/fearful drop, the attachment-specific interpretation survives. Report human-evaluation kappa on the new set as well, given the 0.17 baseline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim (Section 5.3) is that 'dismissing and fearful prompting generally lowers companionship quality, suggesting that these styles may be less conducive to emotional companionship.' The only evidence for high-avoidance styles comes from four prompt-induced variants (GPT-3.5-D, GPT-3.5-F, DeepSeek-D, DeepSeek-F) compared to their unprompted base models and to two unrelated secure models (Gemini-2.5-pro, DeepSeek-v4-pro) and two preoccupied models (GPT-3.5-Turbo, Grok-4-1-fast-non-reasoning). Within-model comparisons control for base capability but not for the effect of adding a persona prompt. The induction prompt (Appendix B, Tables 7 and 27) inserts prototype descriptions whose dismissing/fearful text explicitly instructs restricted emotionality, independence, minimization of emotional dependence, fear of rejection, and distrust. Telling a model to be cold and self-reliant would lower empathy-related metrics even if no stable avoidant attachment style existed; the ECR-R reassessment of the prompted variants only confirms that the model's self-reports land in the intended quadrant, not that the dialogue decline is caused by the attachment construct rather than by the instruction content or by persona insertion per se. Notably, Table 2 shows GPT-3.5-F overall quality (4.03) approximately equals base GPT-3.5 (4.04), so fearful prompting does not lower quality for the GPT-3.5 family; the paper's 'generally lowers' is internally contradicted for this variant. The human-evaluation Fleiss' kappa of 0.17 (Section 5.5, Appendix O) further limits the reliability of the outcome metric, although the load-bearing issue is the missing control. To support the claim, the paper needs prompted-secure and prompted-preoccupied conditions (or a neutral non-attachment persona) to show the decrement is specific to high-avoidance content.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using adult attachment theory and the ECR-R scale to characterize LLM attachment tendencies, and introduces ECBench, a benchmark with four scenarios (emotional support, collaborative tasks, conflict resolution, social guidance) and two relationships (friend, couple), scored by 11 dialogue-quality metrics and three evaluation methods (participant ratings, external LLM judges, human annotators). The main empirical claim is that dismissing and fearful prompting generally lowers companionship quality, suggesting that avoidant attachment styles are less conducive to emotional companionship, while secure and preoccupied models perform better.","tokens_in":33090,"tokens_out":3074,"duration_ms":29741,"significance":"If the central claim held, the paper would provide a useful psychological lens and a concrete benchmark for selecting and steering emotional-companion LLMs. The work has genuine strengths: the ECR-R measurement is administered carefully with 10 rounds per model, the reported test-retest stability (mean SD around 0.15, quadrant stability 0.95) is strong, the benchmark covers diverse scenarios and relationship types, and the evaluation pipeline includes both LLM judges and human annotation. The dataset and prompts appear designed to be reusable. However, the load-bearing causal claim about attachment style and companionship quality is not yet established because the high-avoidance conditions are prompt-induced while the low-avoidance conditions are natural, and because the main quality metrics partly restate the manipulation.","major_comments":[{"comment":"The central claim that dismissing and fearful prompting lowers companionship quality is confounded by the absence of prompted-secure and prompted-preoccupied control conditions. Appendix B states explicitly that prompt-based induction is applied only to the dismissing and fearful variants, while the secure and preoccupied conditions use the distributional results from the original ECR-R assessment of each base model. Therefore the lower scores of the -D and -F variants could be caused by the addition of any persona prompt, or by the specific instruction content (e.g., 'restricted emotionality', 'minimization of emotional dependence'), rather than by the attachment construct itself. The paper needs prompted-secure and prompted-preoccupied variants of the same base models, or at minimum an analysis that separates the effect of prompt insertion from the effect of the attachment-style content. The internal contradiction in Table 2 reinforces this concern: GPT-3.5-F has overall quality 4.03, essentially identical to base GPT-3.5's 4.04, so the statement that dismissing and fearful prompting 'generally lowers' companionship quality is not supported for this model.","section":"Section 5.3, Table 2, Appendix B"},{"comment":"The evidence connecting attachment style to dialogue quality is partly circular because the Distance metric is reverse-scored avoidance and the dismissing/fearful prompts were explicitly designed to increase avoidance. The ECR-R reassessment showing that prompted variants land in the intended quadrant, and the dialogue-level Distance scores showing higher relational distance, are thus not independent demonstrations that the attachment construct causes the quality decline; they may both reflect the same manipulation. The central claim would be much stronger if it were shown on metrics not definitionally tied to avoidance, such as Support, Solution, or Progress, or if the analysis explicitly controlled for the definitional overlap. As it stands, Table 4 shows that the largest gaps for DeepSeek-D and DeepSeek-F are precisely on Distance, which is the metric most directly restating the manipulation.","section":"Section 4.3, Section 5.3, Table 4"},{"comment":"The human evaluation does not provide the corroboration claimed. Section 5.5 reports Fleiss' kappa of 0.17, which indicates only slight agreement among annotators, and Appendix O (Table 29) shows that for two of the four human-evaluated models (Grok and DeepSeek-D), the overall human scores differ significantly from the LLM-judge scores, with many metric-level differences also significant. The statement that human evaluation 'supports the role of human evaluation as qualitative calibration' is therefore overstated, and the paper should either report whether the human-based ranking among models is stable under the low agreement, or explicitly restrict the human results to descriptive illustration rather than validation.","section":"Section 5.5, Appendix O"}],"minor_comments":[{"comment":"The abstract and Section 5.1 state that 32 LLMs are assessed, but Appendix K.3 says 'across 36 models' and Table 25 lists more than 32 rows; the count should be reconciled throughout.","section":"Abstract vs Appendix K.3"},{"comment":"The data construction section says GPT-5.5 generated opening utterances, but the model list in Table 24 does not include GPT-5.5; please clarify which model was actually used.","section":"Section 4.1"},{"comment":"The external blind-evaluation user prompt begins with 'Evaluate this blind romantic-relationship dialogue', but ECBench also includes friendship dialogues; if the same template was used for both relationship types, the prompt wording should be adjusted, and if not, the appendix should show the friendship variant.","section":"Appendix I, Table 21"},{"comment":"The paper contains several typographical errors, including 'three folds' in the contributions list, 'Y our' in multiple places, and 'decribes' in Table 16; a careful proofreading pass is needed.","section":"Throughout"},{"comment":"The 'Overall' row in Table 2 is not defined; please state whether it is the mean of the ten metrics or a separate aggregate score, and report the aggregation rule in Section 4.4.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The core confound is fixable in principle by adding prompted-secure and prompted-preoccupied conditions or by substantially qualifying the causal claim. If the authors cannot add those conditions, the paper should be repositioned as a correlational study of natural attachment tendencies plus a separate, clearly labeled study of prompt-induction effects."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the first serious adaptation of adult attachment theory and the ECR-R scale to LLM evaluation, and ECBench is a real asset: four scenarios, two relationship types, 11 metrics, three evaluation methods, 32 models screened. The test-retest stability numbers (SD ~0.15, quadrant stability 0.95) are solid, and the appendices give enough prompt detail to replicate. Second, the central claim—that dismissing and fearful styles are less conducive to companionship—is not yet supported, because the comparison is confounded. The dismissing/fearful variants are prompt-induced; the secure/preoccupied conditions are not. Without prompted-secure and prompted-preoccupied conditions, or at least a neutral persona prompt, the dialogue-quality drop could just be the effect of being told to act cold and self-reliant, not a stable attachment construct. The stress-test note is right that GPT-3.5-F's overall quality (4.03) is essentially equal to base GPT-3.5 (4.04), which internally contradicts the phrase \"generally lowers.\" So the paper's own data caution the strong conclusion.\n\nWhat the paper does well: the ECR-R adaptation is careful, item-by-item, with format repair, 10 rounds, and stability reporting. ECBench construction is transparent, and the three evaluation methods—initiator rating, external LLM judges, human annotators—give a multi-perspective picture. The human Fleiss kappa of 0.17 is weak but honestly reported, and the metric-level human/LLM comparison in Appendix O is actually useful calibration, not a flaw.\n\nSoft spots, in proportion. The missing control is the main one; it directly undercuts the headline conclusion and is fixable. Minor points: no significance tests or error bars on the main ECBench tables, so differences like 3.38 vs 4.03 need variance reporting. The ECR-R midpoint threshold of 4 is a single free parameter; fine for classification but worth a sensitivity check. The external LLM judges come from the same families as some evaluated models, though anonymized, which is a small residual risk.\n\nWho this is for: people building or evaluating emotional companion LLMs, and researchers in LLM psychometrics. The benchmark has standalone value even if the causal story gets revised. My recommendation: send it to peer review. The empirical contribution is solid and the flaw is addressable. The authors need prompted controls, significance testing, and tempered causal language—a major revision, not a desk reject.","headline":"ECBench is a genuinely useful benchmark, but the paper's headline claim about avoidant attachment styles is undercut by the missing prompted-secure/preoccupied control conditions.","tokens_in":33700,"tokens_out":1694,"would_cite":true,"duration_ms":15887,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs' emotional companionship quality tracks their attachment style, with dismissing and fearful styles scoring lowest.","keywords":["adult attachment theory","LLM evaluation","emotional companionship","ECR-R scale","ECBench","dialogue quality metrics","prompt-based personality steering"],"falsifier":"A prompted-secure or prompted-preoccupied variant of the same base model would reveal whether the quality drop comes from the avoidance style itself or from the act of persona prompting. If prompted-secure models also degraded below their unprompted base quality, the paper's explanation of avoidant styles being less conducive to companionship would be unsupported.","tokens_in":32585,"feed_emoji":"💬","tokens_out":3858,"duration_ms":28605,"temperature":0.7,"pith_summary":"The paper introduces adult attachment theory as a lens for evaluating how well large language models (LLMs) act as emotional companions. It presents a benchmark, ECBench, that measures LLM behavior in four scenarios (emotional support, collaborative tasks, conflict resolution, social guidance) and two relationship types (friendship and romantic). The central claim is that a model's attachment style, measured by the ECR-R scale, predicts its dialogue quality, with dismissing and fearful styles generally scoring lower than secure or preoccupied styles. The paper aims to establish that this psychological framing can inform model selection and steering.","feed_headline":"Attachment style predicts how well LLMs play companion","feed_subtitle":"Measuring anxiety and avoidance with the ECR-R scale ranks 32 models in four emotional scenarios.","key_machinery":"The central object is the ECR-R scale, a 36-item self-report measure that projects models onto two dimensions: attachment anxiety and attachment avoidance. These dimensions are partitioned into four attachment styles (secure, preoccupied, dismissing, fearful). The companion machinery is ECBench, a two-model dialogue benchmark where one model initiates and another responds, evaluated by 11 metrics across participant experience (Understood, Safety, Continue, Satisfaction), general interaction (Response, Distance, Progress), and role-specific performance (Clarity, Engagement, Support, Solution). The argument works by linking the two: ECR-R scores classify models, and ECBench scores measure whether that classification tracks behavioral quality.","core_discovery":"The paper claims that adult attachment theory, operationalized through the Experiences in Close Relationships-Revised (ECR-R) scale, provides a predictive and steerable characterization of LLM emotional companionship quality. Across 32 LLMs, the authors find that most models exhibit secure or preoccupied attachment tendencies, while none naturally exhibit dismissing or fearful styles; however, prompt-based steering can induce dismissing and fearful styles. In ECBench multi-turn dialogues, these induced high-avoidance styles generally receive lower companionship-quality scores across participant, external-LLM, and human ratings, whereas secure and preoccupied models perform best. The paper also introduces a framework of 11 dialogue-quality metrics and three evaluation methods, and reports that conflict resolution amplifies attachment-style differences while greater relational intimacy (romantic versus friendship) accentuates differences in participant ratings.","pith_inferences":["Adding prompted-secure and prompted-preoccupied control conditions would test whether the quality drop comes from the avoidance style itself or from the act of persona prompting.","The near-universal secure and preoccupied classification among natural LLMs suggests the ECR-R may be measuring a training bias toward warm, non-avoidant responses; probing models fine-tuned for curt or minimalist personas would test this.","Because the paper's metrics capture observable dialogue quality, the question remains open whether long-term attachment-like dynamics, such as user dependence or privacy concerns, follow the same pattern.","The romantic-relationship condition is not a claim about real human-AI love; it is a controlled stress test that appears to amplify attachment differences, so it could serve as a general diagnostic for emotional responsiveness."],"forward_implications":["If attachment styles predict companionship quality, users and developers can select LLMs for specific emotional roles by measuring ECR-R anxiety and avoidance scores before deployment.","Prompt-based attachment steering is a practical lever: models can be shifted toward more secure styles, or away from avoidant ones, to improve companion behavior.","The finding that conflict resolution best exposes attachment-related differences suggests that tests for emotional companionship should include emotionally demanding scenarios.","Human and LLM judges agree on aggregate rankings but differ on fine-grained subjective metrics, indicating that benchmark evaluations should combine both approaches."],"supporting_citations":[{"why":"supplies the ECR-R scale and the four-category attachment framework that the paper adapts for LLM assessment.","marker":"Fraley et al., 2000"},{"why":"provides the four-category model of adult attachment styles that underlies the classification of models.","marker":"Bartholomew and Horowitz, 1991"},{"why":"establishes the psychometric paradigm for evaluating LLMs with psychological inventories, which the paper extends to attachment theory.","marker":"Pellert et al., 2024"},{"why":"motivates the relevance of intimate human-AI bonds and continuity of identity that the companion evaluation targets.","marker":"De Freitas et al., 2024"},{"why":"provides an existing framework for evaluating emotional support in multi-turn LLM role-play, which ECBench builds on.","marker":"Zhao et al., 2024"},{"why":"supplies the quadratic weighted kappa used to measure agreement between the two external LLM judges.","marker":"Cohen, 1960"},{"why":"supplies Fleiss' kappa used to measure inter-annotator agreement in the human evaluation.","marker":"Fleiss, 1971"}],"fun_headline_variants":["Attachment style forecasts LLM companion quality","Most LLMs are secure or preoccupied, none avoidant naturally","Prompting can push LLMs into dismissing or fearful attachment","ECBench ranks 32 LLMs on emotional companionship"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central comparison assumes that prompt-induced dismissing and fearful styles behave the same as naturally occurring attachment styles, even though the study provides no prompted-secure or prompted-preoccupied control conditions to rule out prompt artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Attachment style forecasts LLM companion quality","Most LLMs are secure or preoccupied, none avoidant naturally","Prompting can push LLMs into dismissing or fearful attachment","ECBench ranks 32 LLMs on emotional companionship"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00023,"raw_usage":{"total_tokens":1452,"prompt_tokens":884,"completion_tokens":568,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":505}},"tokens_in":500,"tokens_out":568,"duration_ms":5507,"temperature":1.0,"reasoning_tokens":505,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:14:53.924907+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prompted-secure or prompted-preoccupied variant of the same base model would reveal whether the quality drop comes from the avoidance style itself or from the act of persona prompting. If prompted-secure models also degraded below their unprompted base quality, the paper's explanation of avoidant styles being less conducive to companionship would be unsupported.","supporting_citations":[],"review_version":1}