Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
cs.LG, cs.CL
Submitted: 2026-07-22
Updated: 2026-08-29
Project page: https://skylion007.github.io/OpenWebTextCorpus
Terminology
Sources
- Evolution of SAE Features Across Layers in LLMs
- The Llama 3 Herd of Models
- Improving Steering Vectors by Targeting Sparse Autoencoder Features
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
- Not All Language Model Features Are One-Dimensionally Linear
- Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
- Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
- Gemma 3 Technical Report
- The Geometry of Concepts: Sparse Autoencoder Feature Structure
- Sparse Autoencoders Trained on the Same Data Learn Different Features
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
- k-Sparse Autoencoders
- SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks