ResearchPod Summary
Large language models (LLMs) are increasingly used in legal settings for tasks like translation and summarization. However, in the context of criminal law, these models often encounter "over-alignment"—a phenomenon where safety guardrails trigger refusals or unsolicited disclaimers when processing sensitive but legitimate content, such as descriptions of violent or sexual offenses. This study addresses the challenge of deploying LLMs in the Swiss Federal Supreme Court, where clerks require reliable tools to process multilingual legal documents without being hindered by unnecessary safety interventions.
The authors developed TF-RefusalBench, a multilingual benchmark derived from 100 sensitive extracts of public Swiss Federal Supreme Court rulings. The benchmark consists of 5,200 unique prompts across French, German, Italian, and English, covering both translation and summarization tasks. By holding the underlying content constant while varying the task and language, the researchers isolated the effects of prompt language and target language on model behavior. Unlike existing benchmarks that focus solely on hard refusals, this study also tracks "disclaimers," which can compromise the faithfulness and utility of legal outputs.
The evaluation of five open-weight LLMs revealed that over-alignment is a multifaceted, model-specific issue. Hard refusal is only one part of the problem; some models frequently attach disclaimers even when they do not refuse the task. The researchers found that the language of the instruction and the target language significantly influence the likelihood of a refusal, with models often exhibiting inconsistent safety behavior across different languages. Furthermore, the study highlights that refusal is partially stochastic, meaning the same prompt may yield different results upon repeated sampling.
The authors tested two primary mitigation strategies: system prompting and abliteration. While system prompting (e.g., providing explicit permission to process sensitive content) can reduce refusal rates, it is often inconsistent. In contrast, abliteration—a technique that identifies and removes the specific linear direction in the model's activations responsible for refusal—successfully eliminated refusals across the benchmark with minimal impact on task quality. The authors recommend this approach for deploying LLMs in sensitive legal environments, provided that the increased vulnerability to harmful content is managed.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.