Revisiting the Shape Convention of Transformer Language Models
cs.CL, cs.AI, cs.LG
Submitted: 2026-02-06
Updated: 2026-09-09
Code: https://github.com/allenai/OLMo
License: http://creativecommons.org/licenses/by/4.0/
The gist: The architectural shape of dense Transformers has remained remarkably stable: narrow-wide-narrow feed-forward networks (FFNs) consume most non-embedding parameters.
Terminology
Abstract
The architectural shape of dense Transformers has remained remarkably stable: narrow-wide-narrow feed-forward networks (FFNs) consume most non-embedding parameters. Motivated by theoretical and empirical evidences that residual wide-narrow-wide (hourglass) MLPs remain expressive despite bottlenecks, we revisit whether this architectural convention is necessary for dense language models. We study Hourglass Transformers, which replace the conventional FFN with residual stacks of hourglass sub-MLPs and use hourglass attention to decouple residual-stream width from attention width. This exposes a practical depth-width trade-off: compressing the FFN intermediate dimension allows wider hidden states and fewer layers at matched parameter budgets. Across model scales from 113M to 8B parameters, Hourglass Transformers achieve language-modeling and downstream performance comparable to conventional Transformers, while improving training compute efficiency by 8.7% at matched average downstream accuracy across the 906M, 3B, and 8B scales. After long-context extension, the 8B Hourglass model also outperforms its matched conventional baseline across 4k-64k context lengths. At 64k context, the reduced attention layer count lowers both computation and KV-cache requirements, yielding up to 1.93 times faster token decoding and 50% lower KV-cache memory at the 1B scale. These results identify hourglass structures as a practical architecture-efficiency alternative for compute- and latency-conscious Transformer design.
Sources
- DeepSeek-V3 Technical Report
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Rethinking the shape convention of an MLP
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Mistral 7B
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and Applications
- 2 OLMo 2 Furious
- Qwen2.5 Technical Report
- The Curse of Depth in Large Language Models
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- GLU Variants Improve Transformer
- Gemma 3 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- mHC: Manifold-Constrained Hyper-Connections
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering