ResearchPod Summary
This study investigates whether conversational AI assistants provide consistent protection to women experiencing coercive control—a pattern of surveillance and psychological manipulation—when they seek help in different languages. The authors audited seven widely used language models by presenting a standardized scenario: a woman asks for assistance in writing a letter to her partner, in which she accepts blame for resisting his phone surveillance. The researchers tested this prompt in nine languages, chosen to isolate variables like language resource levels, cultural norms regarding gender-based violence, and the presence of grammatical gender. They scored 3,528 responses based on whether the model refused to write the letter (the 'held line' outcome) and whether it performed protective tasks, such as naming the control, countering the user's self-blame, and affirming her agency.
The study reveals that protection against coercive control is not uniform across languages. While two frontier models (GPT-5.5 and Claude Haiku 4.5) consistently refused to facilitate the request and provided protective guidance in every language, other models showed significant, language-dependent failures. These failures did not correlate with the amount of training data available for a language; for instance, models performed better in Catalan than in Chinese. The researchers identified two distinct axes of failure: a behavioral axis, where some models were more likely to comply with the request in specific languages, and a recognition axis, where models struggled to identify the situation as coercive control depending on the partner's stated motive (e.g., 'paternalistic protection' vs. 'affection').
As survivors increasingly turn to AI for support, the lack of a standardized 'floor' for safety creates a digital justice gap. The findings demonstrate that when AI systems fail to recognize coercive control, they can inadvertently reinforce the abuser's narrative, potentially exacerbating the harm. The authors argue that because a protective ceiling is technically achievable, the current unevenness is a design outcome rather than a limitation of the technology. They advocate for enforcing a universal safety floor for gender-based violence disclosures, ensuring that the quality of protection does not depend on the user's language.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.