ResearchPod Summary
This study investigates how Large Language Models (LLMs) internally represent essay quality, moving beyond the 'black-box' perception of these models. The authors analyze eight instruction-tuned LLMs (including Llama, Qwen, and Phi families) across three datasets (ASAP++, CSEE, and ENEM) to determine whether scoring ability stems from superficial statistical cues or structured, high-level representations. The researchers employ linear probing, cross-prompt generalization, dimensionality reduction (PCA), and neuron-level intervention to map how essay quality information is distributed and utilized within the models.
The study provides consistent evidence that essay quality is encoded in a linearly decodable form. Key findings include:
Understanding the internal mechanisms of LLM-based Automated Essay Scoring (AES) is critical for moving toward more reliable and interpretable educational assessment tools. By demonstrating that LLMs encode structured, transferable representations of writing quality, this work provides a foundation for auditing LLM-based grading systems. It suggests that these models are not merely relying on superficial patterns but are developing latent features that align with human-defined quality metrics, which is a necessary step for their deployment in high-stakes educational environments.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.