SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs
summary
The gist
Cascaded Sparse Autoencoders (CSAEs) are introduced as a novel framework for learning hierarchical visual concepts in Multimodal Large Language Models (MLLMs), addressing the interpretability
In short
Cascaded Sparse Autoencoders (CSAEs) learn hierarchical visual concepts in Multimodal Large Language Models by training a second-level SAE directly on the decoder weights of the first level. This approach creates 'concepts of concepts,' improving coherence and enabling effective group-level interventions in MLLM outputs compared to existing methods.
Key concepts
- Level-1 SAE
- This is the low-level sparse autoencoder that decomposes MLLM activations into fine-grained, sparse features. It learns basic visual concepts from the model's internal representations.
- Level-2 SAE
- The high-level autoencoder trained on the decoder weights of Level-1. It learns 'concepts of concepts' by treating each low-level concept direction as a data point, allowing for higher abstraction.
- Hierarchical Mono-Semanticity (HMS) score
- A metric used to quantify success. A high HMS score means that the concepts assigned to the same high-level parent are semantically close in embedding space, indicating coherent grouping.
- Shared-prefix coupling
- A limitation found in Matryoshka SAEs where early directions are reused across multiple levels. This reuse can amplify local semantic errors by linking different abstraction levels inappropriately.
Terminology used across episodes
This episode discusses
- SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs · Paper Radio
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Gemma: Open Models Based on Gemini Research and Technology
- Hallucination of Multimodal Large Language Models: A Survey
- Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
- Interpreting the linear structure of vision-language model embedding spaces
- GPT-4 Technical Report
- Gemini: A Family of Highly Capable Multimodal Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Qwen2.5-VL Technical Report
- Qwen3-VL Technical Report
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- Pixtral 12B
- Gemma 3 Technical Report
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- BatchTopK Sparse Autoencoders
- HyperNetworks
The paper
SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs · Read on arXiv
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs".
Tom: Cascaded Sparse Autoencoders (CSAEs) are introduced as a novel framework for learning hierarchical visual concepts in Multimodal Large Language Models (MLLMs),
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, to recap on "SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs," the main thesis is that existing SAE architectures are often limited because they tend to recover only flat feature dictionaries. This paper introduces cascaded sparse autoencoders as a way to learn hierarchical visual concepts within MLLMs.
Jane: They claim that by training a second-level SAE on the decoder weights from the first level, we can achieve higher-level abstraction over these concepts, effectively learning "concepts of concepts." This is presented as a solution to getting more explicit and interpretable concept hierarchies in models like Qwen3-VL, Gemma-three and LLaVA <ref:2606.16193#pg0,Qwen3-VL, Gemma-3, and>.
Lu: What's really compelling about their approach is how they avoid the known issues found in other methods. They specifically state they are avoiding drawbacks from the shared-prefix coupling of nesting, Matryoshka-style hierarchies, and the bottlenecks of naively stacked SAEs <ref:2606.16193#pg1>.
Meng: Avoiding those specific limitations sounds practical because it suggests a more robust way to build these hierarchical structures without introducing new structural failures like capacity constraints or semantic contamination.
Lalam: If they can learn these concepts of concepts, it means the model's internal visual knowledge won't be as messy; instead, we could have clearer pathways for understanding what a model is actually seeing and how it groups those things together.
Tom: Exactly! It’s about moving beyond just flat feature dictionaries to actual organizational structures inside the activations. This paper sets up a new way to think about applying SAEs for mechanistic interpretability in MLLMs.
Jane: It really matters because it provides a principled method for imposing a hierarchy on the visual features that are otherwise dense and hard to visualize.
Conclusion: Tom: Looking at "SAE++: Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs," the authors are showing how to explicitly learn a concept hierarchy using this cascaded approach, which is quite different from what's been done before.
Jane: The implication here is that we can gain a much more structured view into the visual knowledge of large language and vision models by creating these clear parent-child relationships between concepts.
Lu: The paper suggests this framework could be generalized to an L-level cascade, allowing for progressively higher-order semantic abstraction, which opens up a lot of creative avenues for how we structure interpretability studies.
Meng: For practical application, if we can steer these concept groups effectively with group-level interventions as the paper suggests, it means we could precisely guide the model's output in complex multimodal tasks without having to retrain the entire thing from scratch.
Lalam: I think this has huge potential for our culture here; being able to map out exactly how a model organizes visual ideas makes debugging and fine-tuning much more targeted and efficient.
Tom: So, the big picture here is moving from just seeing what's there to understanding how those features are organized in layers, which gives us explicit control over the model's internal organization.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization