High-probability guarantees for linear accessibility in feature superposition
stat.ML, cs.AI, cs.IR, cs.LG, math.PR
Submitted: 2026-09-09
Updated: 2026-09-25
Comments: preprint
License: http://creativecommons.org/licenses/by/4.0/
The gist: Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features.
Terminology
Abstract
Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, we derive high-probability bounds for fixed supports under subgaussian noise, proving the sufficient dimension scales linearly (d=O epsilon(k m)) rather than prior worst-case quadratic limits. We then validate these bounds across system parameters through Gaussian-tail approximations. These results quantify the geometric constraints of the linear representation hypothesis, providing a framework for evaluating sparse autoencoders, compositional generalization, and neural interpretability.
Sources
- Lower Bounds for Sparse Recovery
- Decoding by Linear Programming
- How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
- When Does LeJEPA Learn a World Model?
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- The Scaling Properties of Implicit Deductive Reasoning in Transformers
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey