ResearchPod Summary
Deep decoder-only Transformers traditionally use the Post-Norm architecture, where normalization follows the residual connection. However, these models are notoriously difficult to train, often suffering from rank collapse—a phenomenon where token representations become nearly identical. This paper provides a theoretical framework to explain why this happens and why standard training dynamics fail to recover from it.
The authors identify a two-stage process that leads to failure. First, at initialization, the causal attention mechanism acts as a prefix-averaging operator. This operation forces token representations to become increasingly similar across the depth of the network. While the SwiGLU feed-forward layers provide a minor damping effect, they are insufficient to counteract the similarity growth driven by attention.
Second, once the model enters a high-similarity regime, the residual norms grow. In the Post-Norm design, this growth causes the RMSNorm backward factor to become contractive. Consequently, gradients flowing back to earlier layers decay geometrically. Because the network is already in a state of high similarity, the optimization process lacks the signal strength required to "repair" the representations, effectively trapping the model in a suboptimal state.
The study also characterizes the behavior of a fully collapsed network. In this state, the model's best possible prediction is the frequency distribution—simply predicting tokens based on their global occurrence counts in the training data. The authors demonstrate that once a network reaches this state, the parameter gradients in the collapsed layers vanish, meaning the model reaches a stationary point that is fundamentally incapable of learning complex patterns.
This work bridges the gap between empirical observations of training instability and theoretical understanding. By identifying that Post-Norm collapse is a combination of forward similarity amplification and backward repair incapacity, the authors explain why Pre-Norm variants are more robust: they avoid the specific gradient contraction caused by the placement of normalization. These insights provide a clearer rationale for architectural choices in large-scale language model design.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.