What Does Layer-Importance Reveal About Transformers and State-Space Models?
cs.LG, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Code: https://github.com/chandar-lab/layer-importance-ssm-vs-transformers
License: http://creativecommons.org/licenses/by/4.0/
The gist: Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical knowledge built for transformers transfers to SSMs.
Terminology
Abstract
Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical knowledge built for transformers transfers to SSMs. We address this through the lens of layer importance which underpins compression, selective fine-tuning, and interpretability across both families. We decompose layer importance into two distinct notions. Necessity captures how much the pretrained model depends on a layer's existing contribution, measured by the loss increase from bypassing it. Plasticity captures where the model absorbs new information during fine-tuning, measured by the magnitude of task-specific weight updates. Our analysis reveals that the two families behave fundamentally differently: in every evaluated residual transformer up to 14 B parameters, Necessity and Plasticity anti-align across depth, whereas in the evaluated Mamba-style SSMs they point to overlapping regions. The sign of this alignment also predicts downstream adaptation behavior. In the evaluated transformers, concentrating updates in the most plastic layers increases catastrophic forgetting, while this tier-dependent effect disappears in the evaluated Mamba-style SSMs.
Sources
- Sanity Checks for Saliency Maps
- Towards better understanding of gradient-based attribution methods for Deep Neural Networks
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Parameter-Efficient Fine-Tuning of State Space Models
- The Zamba2 Suite: Technical Report
- The Llama 3 Herd of Models
- The Unreasonable Ineffectiveness of the Deeper Layers
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- LoRA: Low-Rank Adaptation of Large Language Models
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- MambaLRP: Explaining Selective State Space Sequence Models
- Understanding Layer Significance in LLM Alignment
- Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
- Axiomatic Attribution for Deep Networks
- Locating and Editing Factual Associations in GPT
- Understanding and Guiding Layer Placement in Parameter-Efficient Fine-Tuning of Large Language Models
- Attention Is All You Need
- Qwen3 Technical Report
- TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
- Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks