ResearchPod Summary
Recent research identified "emotion vectors" in Claude Sonnet 4.5—internal linear directions that encode emotional concepts and influence model behavior. This study investigates whether these emotion vectors are a universal feature of Large Language Models (LLMs) or specific to certain architectures. The authors specifically examine how these representations emerge across different layers and how the choice of input text (story corpus) influences the extraction of these vectors.
The researchers analyzed two open-weight models: APERTUS-8B and GEMMA-4-E4B. They generated synthetic story datasets for 171 emotions and used a contrastive method to isolate emotion-specific activations from general linguistic features. By applying Principal Component Analysis (PCA) to these activations, they identified directions corresponding to valence and arousal. They then used Centered Kernel Alignment (CKA) to compare the representational geometry across layers and models.
These findings suggest that while LLMs share a common geometric structure for emotions, the computational "path" to these representations is not uniform. This has significant implications for mechanistic interpretability: researchers cannot assume that a specific layer in one model performs the same function as the corresponding layer in another. Furthermore, the sensitivity to corpus choice highlights that interpretability results may be as much a reflection of the probing data as the model itself.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.