What Does Layer-Importance Reveal About Transformers and State-Space Models?

arXiv:2609.16537 · cs.LG, cs.AI · Submitted 2026-09-15 · Read on arXiv

cs.LG, cs.AI

Submitted: 2026-09-15

Updated: 2026-09-15

Code: https://github.com/chandar-lab/layer-importance-ssm-vs-transformers

License: http://creativecommons.org/licenses/by/4.0/

The gist: Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical knowledge built for transformers transfers to SSMs.

Terminology

Abstract

Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical knowledge built for transformers transfers to SSMs. We address this through the lens of layer importance which underpins compression, selective fine-tuning, and interpretability across both families. We decompose layer importance into two distinct notions. Necessity captures how much the pretrained model depends on a layer's existing contribution, measured by the loss increase from bypassing it. Plasticity captures where the model absorbs new information during fine-tuning, measured by the magnitude of task-specific weight updates. Our analysis reveals that the two families behave fundamentally differently: in every evaluated residual transformer up to 14 B parameters, Necessity and Plasticity anti-align across depth, whereas in the evaluated Mamba-style SSMs they point to overlapping regions. The sign of this alignment also predicts downstream adaptation behavior. In the evaluated transformers, concentrating updates in the most plastic layers increases catastrophic forgetting, while this tier-dependent effect disappears in the evaluated Mamba-style SSMs.

Sources

Related papers