Transformer MLP Gate Thresholds Are Couplings to a Carried Reference Direction
cs.CL, cs.LG
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/EvidentSolutions/llm-interp
Terminology
Sources
- The Discrete Charm of the MLP: Binary Routing of Continuous Signals in Transformer Feed-Forward Layers
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
- Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization
- PaLM: Scaling Language Modeling with Pathways
- Hidden Dynamics of Massive Activations in Transformer Training
- Understanding Gated Neurons in Transformers from Their Input-Output Functionality
- Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator
- The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers
- Not a nuisance but a useful heuristic: Outlier dimensions favor frequent tokens in language models
- Steered LLM Activations are Non-Surjective
- A Refined Analysis of Massive Activations in LLMs
- A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Transformer Training
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
- Qwen2.5 Technical Report
- A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models
- Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes
- On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
- Massive Activations in Large Language Models
- The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering