You Do Not Fully Utilize Transformer's Representation Capacity
cs.LG, cs.CL
Submitted: 2025-02-13
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026 (Main Conference)
Code: https://github.com/corl-team/lime
License: http://creativecommons.org/licenses/by/4.0/
The gist: In contrast to RNNs, which compress their history into a single hidden state, Transformers can attend to all past tokens directly.
Terminology
Abstract
In contrast to RNNs, which compress their history into a single hidden state, Transformers can attend to all past tokens directly. However, standard Transformers rely solely on the hidden state from the previous layer to represent the entire context. We show that this design creates pressure toward representation collapse and can degrade performance. To address this issue, we introduce Layer-Integrated Memory (LIMe), a lightweight extension that leverages existing key-value buffers and learns per-head, per-layer routing weights to integrate representations from previous layers. Across language modeling, synthetic reasoning, and deep architectures, LIMe improves perplexity per FLOP in the studied regimes and yields strong gains on synthetic tasks while preserving higher value-vector entropy and token separability. Finally, learned routing weights reveal systematic reuse of local and long-distance features, showing how LIMe enriches attention-time memory without increasing hidden-state size. Code is available at https://github.com/corl-team/lime.
Sources
- Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
- Training Deeper Neural Machine Translation Models with Transparent Attention
- Transformers need glasses! Information over-squashing in language tasks
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- The Llama 3 Herd of Models
- Training Large Language Models to Reason in a Continuous Latent Space
- Deep Residual Learning for Image Recognition
- Densely Connected Convolutional Networks
- Mistral 7B
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- Qwen2.5 Technical Report
- Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
- Highway Networks
- Hyper-Connections
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks