Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence

arXiv:2609.18612 · cs.LG, cs.CL · Submitted 2026-09-16 · Read on arXiv

cs.LG, cs.CL

Submitted: 2026-09-16

Updated: 2026-09-16

Comments: Accepted to EMNLP 2026. Supersedes arXiv:2505.17936

Code: https://github.com/sjgerstner/RW_functionalities

Project page: https://gluscope.github.io

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

The gist: We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs).

Terminology

Abstract

We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs). We propose a simple analysis method: For each neuron, we compute the cosine similarities between its input (reading) and output (writing) weight vectors. In this scheme, a strong negative cosine similarity indicates the neuron weakens the direction it detects in the residual stream, so we call this a weakening neuron. This allows us to gain a number of novel insights. First, we show that nine different LLMs have similar patterns: weakening neurons appear mostly in late layers whereas their counterparts, (conditional) strengthening neurons, are frequent in early-middle layers. Second, we find that weakening neurons display surprising behavior: even though there are few, they activate often and have a large influence on model behavior. Third, weakening neurons have a strong effect on model output when gate values are negative -- which is surprising since negative gate values are not expected to encode functionality.

Sources

Related papers