CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship

Published
Source
arXiv
Paper number
798
Field
LLMs / NLP
arXiv ID
2608.02046

Key points

  • The capability system measures 10 capabilities extracted from 25 theory papers on a 1-point to 10-point scale, with 42 sub-items in total, and it adds a reverse-layer list of 18 negative patterns with three severity levels.
  • It adds four capabilities that prior work did not score explicitly: tolerating ambiguity, self-object responsiveness, positive resonance, and appropriately calibrated challenge.
  • The second axis is self-disclosure gate replay, which measures how often disclosure is obtained and how deep it goes, and it also counts disclosures that were elicited without being originally given.
  • It is highly reproducible, with Spearman correlations of 0.996 in Chinese and 0.953 in English for the ranking of the 10 capability scores under independent reruns, and 8 of the 10 capabilities are at or above 0.93; the lowest are positive resonance at 0.829 and calibrated challenge at 0.863.
  • The gate behavior really follows the agent's behavior, because the rate of obtained disclosures and the rate of early self-disclosure are strongly negatively correlated, with Pearson correlations of -0.938 in English and -0.979 in Chinese.
  • Across the 28-model leaderboard, gpt-5.5 ranks first in English with an IRT score of 8.27, while the lowest, doubao-char-251128, scores 4.28; three models specialized for roleplay sit near the bottom, and the biggest deficits are in tolerating ambiguity, calibrated challenge, and acknowledgment.
  • The numbers show that warmth and substance are different dimensions, because claude-haiku-4-5 drops from 0.780 before penalty rules to 0.548 after penalties are applied, and the lowest model drops from 0.395 to 0.017.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)