ResearchPod Summary
Deep neural networks, including Transformers, face a fundamental challenge: as signals propagate through successive layers, the effective rank of both gradients and representations tends to collapse. This means that the network's internal state becomes confined to a low-dimensional subspace, limiting its expressive power. While skip connections and normalization are traditionally viewed as tools for controlling signal magnitude to prevent vanishing or exploding gradients, this paper reinterprets them as critical mechanisms for preserving the rank of the gradient across depth.
The paper demonstrates that skip connections function as a bypass for the rank-reducing residual branch. Because the identity path of a skip connection is inherently full-rank, it restores the rank lost during the matrix multiplications and nonlinear activations within the feedforward block. This creates a characteristic sawtooth pattern in the rank across depth: rank decreases through the residual branch and is restored at each skip. The authors show that the branch-to-skip ratio—determined by the branch scale and weight initialization—governs a fundamental tradeoff. If the branch dominates, the network suffers from rank collapse; if the skip dominates, the network behaves more like an ensemble of shallow layers, potentially sacrificing the benefits of deep composition.
The internal structure of the feedforward block is also optimized to prevent rank loss. The two-matrix structure (up-projection and down-projection) is not merely for parameter efficiency; the width expansion between these matrices ensures that the branch Jacobian remains full-rank. By applying the rank-reducing activation function in an expanded hidden space, the network retains enough directions to span the original input space. Furthermore, the second matrix acts to decorrelate the residual branch, preventing the formation of a coherent mean spike that would otherwise cause the representation to collapse onto a single direction.
The placement of normalization layers—Pre-Norm versus Post-Norm—is shown to control the branch-to-skip ratio across depth. Pre-Norm allows the rank to plateau at a stable level, whereas Post-Norm leads to more aggressive rank collapse. This perspective unifies various findings in the literature regarding normalization placement and depth scaling, suggesting that the success of modern architectures is largely due to how they navigate the intrinsic tension between rank preservation, ensemble-like behavior, and parameter count.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.