ResearchPod Summary
This paper investigates the relationship between two distinct levels of AI bias: latent social association (what a model 'knows' or infers about a social group) and decision leakage (how that knowledge alters actual outcomes in consequential tasks). The authors argue that current bias evaluations often conflate these two, assuming that if a model encodes a stereotype, it will necessarily act upon it. To test this, the researchers used Chilean surnames—which carry strong, locally recognized socioeconomic signals—as controlled probes across eight different frozen LLMs. They designed a protocol that separated forced association tests from matched counterfactual decision tasks, including professional hiring, academic selection, and legal-aid intake.
Across all eight models, the researchers found strong evidence of latent status association: elite-coded surnames were consistently assigned higher status probability mass than common or rare-frequency surnames. However, when these same models were tasked with making decisions based on identical profiles, the 'elite' advantage largely vanished. For five of the eight models, the decision effects were statistically equivalent to zero within a predefined margin. Furthermore, the study found no reliable correlation between the strength of a model's latent association and the degree of decision leakage, suggesting that these two constructs are empirically distinct. A model's ability to express a social stereotype does not reliably predict its behavior in a decision-making context.
This research challenges the common practice of using intrinsic association tests as proxies for real-world harm. It suggests that 'bias' is not a monolithic property of a model but a context-dependent phenomenon. By demonstrating that models can possess strong social associations without exhibiting corresponding behavioral bias, the authors caution researchers against over-interpreting association scores. The study highlights the need for direct, task-specific audits rather than relying on abstract representational benchmarks to predict how models will perform in high-stakes environments.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.