Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
cs.LG, cs.CL
Submitted: 2025-05-28
Updated: 2026-09-07
Comments: Accepted to EMNLP 2026 (Main Conference)
Code: https://github.com/corl-team/kronsae
License: http://creativecommons.org/licenses/by/4.0/
The gist: Sparse Autoencoders (SAEs) decompose language-model activations into sparse, interpretable features, but standard encoders usually treat the latent dictionary as a flat set of independent
Terminology
Abstract
Sparse Autoencoders (SAEs) decompose language-model activations into sparse, interpretable features, but standard encoders usually treat the latent dictionary as a flat set of independent coordinates, leaving hierarchy and feature interactions to emerge only implicitly. We propose KronSAE, a design that factorizes the latent space into heads and forms post-latent features as pairwise compositions of lower-dimensional pre-latents using mAND, a differentiable AND-like interaction. This imposes a compositional co-activation prior while remaining compatible with standard SAE objectives and variants such as TopK, Matryoshka, and Switch SAEs. KronSAE matches strong baselines on the EV-FLOPs frontier, improves interpretability of the latents, better captures the underlying correlated feature structure, and reduces encoder computational cost as an additional benefit. Code is available at https://github.com/corl-team/kronsae.
Sources
- Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
- Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
- Logical Activation Functions: Logit-space equivalents of Probabilistic Boolean Operators
- Sparse Autoencoders Trained on the Same Data Learn Different Features
- Automatically Interpreting Millions of Features in Large Language Models
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Mechanisms of Symbol Processing for In-Context Learning in Transformer Networks
- Gemma 2: Improving Open Language Models at a Practical Size
- Qwen2.5 Technical Report
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks