ResearchPod Summary
AI emotional companions are increasingly used in high-stakes personal settings, yet existing evaluation methods often fail to distinguish between genuine relational support and superficial, sycophantic warmth. The authors introduce CompanionBench, a bilingual, theory-anchored benchmark designed to measure relational competence through a two-layer approach: a top-down theoretical framework derived from 25 psychology and counseling theories, and a bottom-up grounding in de-identified real-world conversation data.
To evaluate agents, the researchers developed a trained user simulator featuring a hidden 'disclosure gate.' This mechanism forces the interaction to branch based on the agent's behavior, allowing the system to test whether an agent can earn deeper user disclosure over time rather than simply responding with generic empathy. The benchmark assesses agents on ten specific capabilities—including four previously unmeasured ones like 'holding ambiguity' and 'calibrated challenge'—using both a subjective rubric and a deterministic measure of earned disclosure depth.
Evaluating 28 agents in both Chinese and English, the study demonstrates that aggregate empathy scores often obscure significant capability-level weaknesses. While many models excel at surface-level warmth, they frequently fail to provide substantive support. Notably, role-play agents, which prioritize immersion, ranked near the bottom of the evaluation, suggesting that high immersion does not equate to relational competence. The most common failure mode identified across models was the tendency to prioritize warmth over the measured, sometimes challenging, responses required to build genuine trust.
As AI companions become more integrated into users' daily lives, the ability to distinguish between a performance of empathy and actual relational support is critical for user safety and well-being. By moving beyond single-score metrics and incorporating theory-driven, interactive simulations, CompanionBench provides a more rigorous framework for developers to identify where their models fail to provide meaningful, long-term emotional support.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.