ResearchPod Summary
This paper introduces LEX-EC, a black-box auditing framework designed to evaluate how Large Language Models (LLMs) infer personality traits from text. Because closed-source models do not allow for internal inspection of weights or activations, the authors treat LLMs as behavioral systems. The framework uses a combination of prevalence diagnostics, item-level agreement checks, and a controlled lexical ablation process. By masking topical and demographic content—such as course identifiers or biographical details—the researchers isolate whether personality predictions persist based solely on function words, affective terms, and cognitive-style vocabulary.
The study reveals that LLM personality predictions are highly sensitive to the genre and length of the input text. In long-form essays, models show a broad but weak signal. In contrast, short-form texts like Facebook statuses provide little to no stable evidence for personality, suggesting a lower bound of content required for meaningful inference. When topical and demographic information is removed, the predictive accuracy for many traits collapses, indicating that models often rely on superficial content shortcuts rather than deep psycholinguistic patterns. Furthermore, while linguistic prompting can shift the content of model-generated explanations, it does not eliminate the model's reliance on topical cues.
As LLMs are increasingly deployed for personality assessment in high-stakes environments like hiring and education, understanding the basis of these inferences is critical. This work demonstrates that "plausible" personality labels are not necessarily valid; they may instead reflect demographic priors or topical associations rather than the author's actual personality. The LEX-EC framework provides a necessary tool for researchers and developers to audit these systems, helping to identify when personality-labeling outputs are artifacts of the input data rather than genuine trait assessments.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.