Output Dilution: Redundant but Fragile Representations in MoE Models
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-25
Updated: 2026-08-25
Comments: 18 pages, 4 figures. Code and outputs at https://github.com/deepsteer/deepsteer
Code: https://github.com/deepsteer/deepsteer
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mixture-of-Experts (MoE) models appear to encode moral content as robustly as dense models, yet prove far more fragile in their encoding.
Terminology
Abstract
Mixture-of-Experts (MoE) models appear to encode moral content as robustly as dense models, yet prove far more fragile in their encoding. In OLMoE-1B-7B, linear probes recover moral valence from nearly every expert-layer combination, with mean peak-layer accuracy above 90%. But these representations collapse under levels of activation noise that a dense model of matched size easily tolerates, with a 4.2-fold difference in robustness. We trace this to output dilution. Because the MoE block averages across active experts before contributing to the residual stream, the feedforward signal reaching downstream layers is nearly two orders of magnitude smaller than in a dense MLP. Moral information, our interest, survives aggregation intact but at a scale trivially overwhelmed by perturbation. Routing itself remains stable under noise while the vulnerability originates entirely in the diluted aggregate. Checkpoint trajectories confirm this is architectural, not learned. Experts never specialize and accuracy saturates within the first few thousand steps. In sparse architectures, redundant encoding does not imply robust encoding.
Sources
- On the Representation Collapse of Sparse Mixture of Experts
- What you can cram into a single vector: Probing sentence embeddings for linguistic properties
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- OLMo: Accelerating the Science of Language Models
- Mixtral of Experts
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Locating and Editing Factual Associations in GPT
- Insights on representational similarity in neural networks with canonical correlation
- OLMoE: Open Mixture-of-Experts Language Models
- 2 OLMo 2 Furious
- When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks