Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence
cs.LG, cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: Accepted to EMNLP 2026. Supersedes arXiv:2505.17936
Code: https://github.com/sjgerstner/RW_functionalities
Project page: https://gluscope.github.io
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs).
Terminology
Abstract
We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs). We propose a simple analysis method: For each neuron, we compute the cosine similarities between its input (reading) and output (writing) weight vectors. In this scheme, a strong negative cosine similarity indicates the neuron weakens the direction it detects in the residual stream, so we call this a weakening neuron. This allows us to gain a number of novel insights. First, we show that nine different LLMs have similar patterns: weakening neurons appear mostly in late layers whereas their counterparts, (conditional) strengthening neurons, are frequent in early-middle layers. Second, we find that weakening neurons display surprising behavior: even though there are few, they activate often and have a large influence on model behavior. Third, weakening neurons have a strong effect on model output when gate values are negative -- which is surprising since negative gate values are not expected to encode functionality.
Sources
- Yi: Open Foundation Models by 01.AI
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Gaussian Error Linear Units (GELUs)
- Mistral 7B
- Language Models (Mostly) Know What They Know
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- The Remarkable Robustness of LLMs: Stages of Inference?
- Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
- Olmo 3
- GPT-4 Technical Report
- A Robust Evaluation of Probe Robustness: Lessons for Reliable OOD Uncertainty Quantification
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- GLU Variants Improve Transformer
- OPT: Open Pre-trained Transformer Language Models
- Qwen3 Technical Report
- Qwen2 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks