ResearchPod Summary
Modern language models are built from stacks of identical layers, distributing parameter capacity uniformly across depth. However, research suggests that layers contribute non-uniformly to model outputs, with later layers primarily refining the residual stream rather than performing complex transformations. This paper investigates whether language models can be improved by aligning parameter capacity with this observed layer-wise asymmetry.
The authors introduce Tapered Language Models (TLMs), an architectural principle where a parameter-bearing component—specifically the intermediate dimension of the MLP—is monotonically tapered across depth. By keeping the total parameter budget and FLOPs constant, the authors compare uniform-width baselines against three smooth decay schedules: linear, cosine, and sigmoid. They evaluate this approach across four distinct architectural families (Transformer, Gated Attention, Hope-attention, and Titans) and three model scales (440M, 760M, and 1.3B parameters).
The study finds that front-loading capacity in early layers significantly improves performance. Across all tested architectures and scales, the cosine-tapered models consistently outperformed uniform baselines on both perplexity and downstream commonsense reasoning benchmarks. The authors demonstrate that MLP outputs become increasingly aligned with the residual stream at greater depths, providing a mechanistic justification for why reducing capacity in later layers does not harm—and often improves—overall model performance.
This work establishes depth-aware capacity allocation as a "free" design lever. Because the tapering is implemented by simply adjusting the width of existing MLP layers, it requires no additional training compute or inference latency. It suggests that the standard practice of uniform layer sizing is suboptimal and that architectural design should explicitly account for the varying roles of layers across the depth of a neural network.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.