When Can We Trust the Sparse Lens? A Certification Framework for SAE Faithfulness
cs.LG, cs.CL
Submitted: 2026-06-16
Updated: 2026-09-26
Code: https://github.com/newcodevelop/SAE-Faithfulness
Terminology
Sources
- Stronger generalization bounds for deep nets via a compression approach
- PIQA: Reasoning about Physical Commonsense in Natural Language
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- The Llama 3 Herd of Models
- Non-Vacuous Generalization Bounds for Large Language Models
- Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
- Uniform convergence may be unable to explain generalization in deep learning
- WinoGrande: An Adversarial Winograd Schema Challenge at Scale
- Gemma 2: Improving Open Language Models at a Practical Size
- LogitLens4LLMs: Extending Logit Lens Analysis to Modern Large Language Models
- HellaSwag: Can a Machine Really Finish Your Sentence?
- Understanding deep learning requires rethinking generalization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks