Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory
Junkai Zhou, Shiting Guan, Zhaoyi Zhang
cs.CL
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: Based on the paper "Which LLM Is Your Ideal Companion? Evaluating Emotional Companionship Capabilities of LLMs Based on Adult Attachment Theory," here is a detailed summary: Introduction and
Terminology
Summary
Based on the paper Which LLM Is Your Ideal Companion? Evaluating Emotional Companionship Capabilities of LLMs Based on Adult Attachment Theory,
here is a detailed summary:
Introduction and Motivation
The paper addresses the growing use of large language models (LLMs) in emotional companionship, noting that existing evaluations primarily focus on general personality traits like MBTI and the Big Five, which provide limited insight into model behavior within intimate and emotionally sensitive contexts.
The authors argue that the interaction styles of LLMs in sustained emotional companionship along with their implications for interaction quality remain underexplored.
To address this gap, they introduce adult attachment theory as a psychological lens for evaluating LLM behavior in close relationships.
Theoretical Framework
The paper applies adult attachment theory, which posits that internal working models of self and others shape behavior in intimate relationships, giving rise to four attachment styles: secure, preoccupied, dismissing, and fearful attachment.
These styles are characterized along two dimensions: attachment anxiety (concerns about rejection, neglect, or relationship instability
) and attachment avoidance (discomfort with intimacy, dependence, and emotional disclosure
). The four styles are defined as: secure (low anxiety, low avoidance), preoccupied (high anxiety, low avoidance), dismissing (low anxiety, high avoidance), and fearful (high anxiety, high avoidance).
Methodology
The study has two main components:
-
Attachment Style Assessment: The authors use the Experiences in Close Relationships-Revised (ECR-R) scale, a 36-item self-report measure, to assess attachment anxiety and avoidance in 32 LLMs from 11 developers. Each model completes the scale across 10 runs, and scores are mapped onto the two-dimensional attachment space to classify each model into one of the four attachment styles.
-
ECBench Benchmark: To evaluate behavior in realistic interactions, the authors introduce ECBench, a dialogue benchmark covering four scenarios: emotional support, collaborative tasks, conflict resolution, and social guidance. These scenarios are instantiated in both friendship and romantic relationship settings. The benchmark uses a two-model interaction protocol where one model initiates and the other responds. Dialogue quality is assessed using 11 metrics across three categories: participant experience (Understood, Safety, Continue, Satisfaction), general interaction (Response, Distance, Progress), and role-specific performance (Clarity, Engagement, Support, Solution). Evaluation is conducted via three methods: participant ratings, external LLM judges, and human annotators.
Key Findings
-
Attachment Tendencies of LLMs: The ECR-R assessment revealed that
26 LLMs are secure and 6 are preoccupied, while none are dismissing or fearful.
This indicates thatmost LLMs exhibit low avoidance, however, there are significant differences in their anxiety levels.
The authors then used attachment-style prompts to steer GPT-3.5-turbo and DeepSeek-v4-pro toward dismissing and fearful styles, successfully eliciting response patterns consistent with the target styles. -
Overall Performance on ECBench: Eight representative models were selected for dialogue evaluation: secure (Gemini-2.5-pro, DeepSeek-v4-pro), preoccupied (GPT-3.5-turbo, Grok-4-1-fast-non-reasoning), and dismissing/fearful prompted variants of GPT-3.5-turbo and DeepSeek-v4-pro. Results showed that "Secure and preoccupied models perform best overall, led by Gemini and GPT-3.5. Dismissing and fearful prompting generally lowers companionship quality, suggesting that these styles may be less conducive to emotional companionship." However, the effect varied by base model, with GPT-3.5 variants remaining competitive while DeepSeek variants declined markedly.
-
Participant Experience Metrics: Secure and preoccupied models performed well across scenarios and relationships. Model differences were greatest in conflict resolution, where
higher emotional demands amplify attachment-related variation.
Participant ratings were higher in friendship than in romance, suggestinggreater intimacy accentuates differences in emotional responsiveness and relationship maintenance.
Dismissing and fearful variants, particularly DeepSeek-D and DeepSeek-F, received significantly lower scores. -
General Interaction Metrics: External LLM evaluations showed that "secure and preoccupied models perform better in response quality, relational distance, and problem progress. Dismissing and fearful variants, especially DeepSeek-D and DeepSeek-F, show greater distance and less progress, indicating weaker emotional responsiveness.
Most models performed best in collaborative tasks, while emotional support and conflict resolution amplified attachment-related differences. Most models also performed better as responders than initiators,
possibly reflecting the assistant-oriented design of LLMs." -
Role-specific Metrics: Secure and preoccupied models expressed needs more clearly and sustained engagement as initiators, while providing more consistent emotional support as responders. Dismissing and fearful models performed worse in both roles. Interestingly, external LLM judges rated romance higher than friendship, contrary to participant ratings, a discrepancy the authors attribute to different perspectives:
external judges emphasize observable response quality, while participants are more influenced by relational expectations and their own interaction experience.
-
Human Evaluation: Human evaluation of four representative models confirmed the general trends. The secure model Gemini achieved the highest scores, while the dismissing DeepSeek-D scored lowest. The paper notes that
conflict resolution yields larger cross-model differences, indicating that emotionally demanding scenarios are likely to make attachment-style differences more pronounced.
Conclusion and Contributions
The paper concludes that its framework "provides a new psychological perspective to understand the emotional interaction patterns of LLMs, while offering practical guidance for selecting emotional companionship LLMs that more closely align with the needs and interaction preferences of users." The three main contributions are: (1) introducing adult attachment theory and the ECR-R scale into LLM evaluation; (2) presenting ECBench, a comprehensive benchmark for evaluating emotional companionship capabilities; and (3) analyzing dialogue performance to provide a theoretical basis for understanding and selecting LLMs in emotional companionship scenarios.
Improvements for AI systems
Improvements to AI Systems:
-
Attachment-Style-Aware Response Modulation: Implement an adaptive dialogue module that dynamically adjusts an LLM’s emotional responsiveness based on the user’s inferred attachment style (from ECR-R-like signals in conversation). The system can shift between secure, preoccupied, dismissing, or fearful response patterns—e.g., increasing emotional validation and reassurance for high-anxiety users, while reducing intrusive closeness for high-avoidance users—to optimize long-term companionship quality.
-
Conflict-Resolution Specialization: Train or fine-tune LLMs specifically on conflict-resolution dialogues (as defined in ECBench) with reinforcement learning from human feedback (RLHF) targeting the 11 metrics (e.g., Understood, Progress, Solution). The improved system will detect emotional escalation and proactively de-escalate by using secure attachment strategies (e.g., acknowledging feelings, offering collaborative solutions) rather than defaulting to generic supportive responses.
-
Role-Aware Interaction Optimization: Develop a dual-mode architecture where the LLM explicitly switches between “initiator” and “responder” roles. As an initiator, it will practice clearer need expression and sustained engagement (based on role-specific metrics like Clarity and Engagement); as a responder, it will prioritize consistent emotional support (Support metric). This addresses the observed weakness in initiator performance and improves proactive emotional maintenance.
-
Intimacy-Adaptive Emotional Calibration: Incorporate a context-sensitive emotional intensity controller that adjusts response depth based on relationship type (friendship vs. romance). The system will recognize that romance requires higher emotional attunement and will reduce avoidance behaviors (e.g., distancing language) in intimate settings, as the paper shows intimacy amplifies attachment-related differences.
-
Anxiety-Level Tuning for Companionship: Use the ECR-R assessment results to calibrate baseline anxiety levels in LLMs. Since all models are low-avoidance, the system can be tuned to maintain low anxiety (secure) for general use, but offer a “preoccupied mode” for users who prefer more emotionally expressive and attentive companionship, while avoiding dismissing/fearful configurations that degrade performance.
Capabilities of the Improved AI System:
-
Personalized Emotional Companion: The AI can infer a user’s attachment style from conversational cues and adapt its responses to reduce relational friction, improving user satisfaction and perceived understanding (as measured by participant experience metrics).
-
Superior Conflict Resolution: It can handle emotionally charged disagreements with higher scores on Progress and Solution, reducing relational distance and increasing the likelihood of constructive outcomes.
-
Proactive and Responsive Balance: It can seamlessly switch between initiating conversations (e.g., checking in on user’s emotional state) and responding empathetically, maintaining engagement without being intrusive.
-
Context-Aware Intimacy Handling: It can modulate emotional depth appropriately for friendships (lower intensity) vs. romantic relationships (higher attunement), avoiding both coldness and over-closeness.
-
Reliable Long-Term Interaction: By maintaining a secure attachment baseline, it avoids the performance degradation seen in dismissing/fearful variants, ensuring consistent quality across scenarios and over sustained interactions.
Sources
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Lessons From an App Update at Replika AI: Identity Discontinuity in Human-AI Relationships
- The Llama 3 Herd of Models
- OpenAI GPT-5 System Card
- Kimi K2: Open Agentic Intelligence
- GPT-4o System Card
- Revisiting the Reliability of Psychological Scales on Large Language Models
- Persona Dynamics: Unveiling the Impact of Personality Traits on Agents in Text-Based Games
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Qwen3 Technical Report
- Chatbot Companionship: A Mixed-Methods Study of Companion Chatbot Usage Patterns and Their Relationship to Loneliness in Active Users
- GLM-5: from Vibe Coding to Agentic Engineering
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering