Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression?
cs.LG
Submitted: 2026-07-19
Updated: 2026-09-13
Code: https://github.com/aniket-desh/decoder-preserving-sae
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- BatchTopK Sparse Autoencoders
- Revisiting End-To-End Sparse Autoencoder Training: A Short Finetune Is All You Need
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
- Task-Driven Dictionary Learning
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks