Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders".
Jane: Sparse autoencoders (SAEs) are presented as an instrument for open-ended feature discovery from foundation model representations, enabling systematic exploration of learned patterns beyond pre-specified targets.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we're looking at "Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders," and the authors are Samuel Stevens, Jacob Beattie, Tanya Berger-Wolf, and Yu Su from The Ohio State University. That title really sets the stage for what they are trying to do here.
Jane: It’s a very clear goal; they aren't aiming to build a new task-specific predictor; they are looking at how we can use sparse autoencoders to uncover patterns that were never explicitly taught during the model's initial training.
Lu: The authors frame it as investigating whether sparse autoencoders can enable this kind of discovery from foundation model representations, which is a very direct and challenging question for the field. It moves away from using these models only for confirmation tasks, like just checking if an image is a cat or not.
Meng: So, instead of asking the model to confirm a label we already have, they are asking it to reveal what kinds of features it actually has stored internally that might be useful for science. That’s a shift in purpose.
Lalam: It really shows that these foundation models contain latent information beyond the immediate tasks, and this paper offers a methodology to systematically pull that out. It’s about unlocking hidden structure within the representations themselves.
The paper's summary: Tom: In essence, the core idea of "Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders" is that foundation models usually just give us dense, opaque embeddings when we feed them an image. This paper proposes using sparse autoencoders to decompose those dense representations into a library of interpretable semantic concepts.
Jane: That means instead of getting one big blob of data, we get a structured vocabulary where each entry corresponds to a specific pattern the model has learned, and that’s what allows us to investigate those patterns systematically.
Lu: The methodology involves training an overcomplete dictionary that reconstructs the original model activations through sparse linear combinations, which is encouraging sparsity in individual features to capture coherent semantic units. They stress that this whole process happens without needing any concept supervision during the training phase itself.
Meng: I see what you mean; they are teaching the model to organize its internal knowledge into these distinct, sparse building blocks based purely on reconstruction accuracy and sparsity penalties. That’s a smart way to force interpretability onto a massive, complex structure.
Lalam: It means we get per-example activation maps for each concept, which is powerful because it lets us see exactly where in an image the model is activating on that specific learned feature. It’s like having a scientific microscope for the model's internal thinking.
The paper's improvements: Tom: One of the main improvements they highlight is moving away from methods that require researchers to specify concepts or tasks in advance, which limits them to confirmatory analysis. They propose using linear probes and concept vectors as alternatives that can discover new traits instead of just checking for known ones.
Jane: That's a big improvement because it opens up the possibility of finding unknown patterns, like some subtle textural choices or specific layer selections within a model’s architecture that we wouldn't have thought to look for initially.
Lu: The paper shows that their approach supports open-ended feature discovery by surfacing the model’s learned vocabulary in a structured format without any supervision, which is explicitly contrasted against methods like task-specific SAEs or simple labeled concept vectors.
Meng: So, the improvement is moving from a closed system where we define the question to an open system where we let the model's internal structure dictate what questions we can ask of it. That’s a fundamental methodological improvement for scientific exploration.
Lalam: The authors mention that by using these features, they can discover new traits related to texture, layer selection, and even dataset characteristics, which suggests the potential for discovering things entirely outside the original training objectives.
Conclusion: Tom: So to wrap up on "Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders," the main point is that sparse decomposition of foundation model representations can systematically surface semantic structure as a prerequisite step toward genuine scientific discovery.
Jane: It’s about using these SAEs not just to confirm what the model knows, but to generate hypotheses about unexpected structures and establish reproducible measurements across different studies.
Lu: The protocol itself is domain-agnostic, which means it’s applicable to any setting where you have foundation model representations, whether that's vision or something else entirely.
Meng: I wonder how this translates into a practical pipeline; the authors validate their approach on two vision domains, but getting it running seamlessly across completely different types of data sources needs some real engineering work.
Lalam: It’s exciting because this framework gives us a universal instrument to investigate what patterns models captured from data, allowing researchers to generate new ideas and establish measurements that are reproducible across studies.
Tom: Fantastic stuff, team; the paper really solidifies the idea that SAEs are a valuable tool for systematic exploration, and I'm really looking forward to seeing how this gets integrated into our workflows.
Samuel Stevens, Jacob Beattie, Tanya Berger-Wolf, Yu Su
The Ohio State University
cs.CV
Submitted: 2025-11-21
Updated: 2026-09-29
Importance score: 89/100
The gist: Sparse autoencoders (SAEs) are presented as an instrument for open-ended feature discovery from foundation model representations, enabling systematic exploration of learned patterns beyond
Key concepts
- Foundation Model Representations
- These are the high-dimensional internal patterns learned by massive neural networks trained on huge datasets across various fields like vision or genomics. They represent compressed knowledge but are often used only for tasks they were originally trained for.
- Sparse Autoencoders (SAEs)
- SAEs learn an overcomplete dictionary of features that reconstruct the model's activations using only a few active components at a time. This sparsity forces each learned feature to capture a distinct, coherent semantic unit from the model's knowledge.
- Concept Alignment Probes
- These are binary logistic regression classifiers trained on the SAE features to test how well they align with known or unknown scientific concepts. They help researchers systematically investigate which learned features correspond to specific structures or patterns in the data.
Terminology
Summary
Sparse autoencoders (SAEs) are presented as an instrument for open-ended feature discovery from foundation model representations, enabling systematic exploration of learned patterns beyond pre-specified targets. This approach is important because current methods excel at confirmation but fail to support the discovery of unknown patterns in high-dimensional scientific data.
The gist
Sparse decomposition provides a practical instrument for exploring what scientific foundation models have learned, an important prerequisite for moving from confirmation to genuine discovery.
Foundation Model and the Problem
Large neural networks trained on massive datasets across genomics, ecology, and vision cover vast amounts of data but are predominantly used in science as task-centric tools—frozen feature extractors or fine-tuned predictors. These models learn compressed representations optimized for specific tasks, which limits them to confirmatory applications where target concepts are pre-specified. Scientific discovery requires finding unexplained phenomena,
such as morphological features or unknown correlations, which current interpretability methods (like saliency maps or Concept Activation Vectors) cannot perform because they require specifying concepts in advance.
The SAE Mechanism
To address this gap, the authors study sparse autoencoders (SAEs) to decompose foundation model representations into interpretable features without concept supervision. SAEs learn an overcomplete dictionary that reconstructs model activations through sparse linear combinations, with sparsity encouraging individual features to capture coherent semantic units. Unlike supervised methods, SAEs require no concept labels during training.
Each learned feature provides a per-example activation score together with its decoding direction and exemplar evidence,
allowing for the systematic investigation of what patterns the model represents.
Experimental Validation and Evidence
The method was evaluated in controlled rediscovery studies using two settings: (1) ADE20K scene segmentation, containing 150 object classes, and (2) ecological images testing recovery of anatomical body parts from the FishVista dataset. The experiments compared SAEs against decomposition baselines like k-means clustering and PCA. Results indicated that SAEs reliably extract semantic concepts that were never explicitly labeled in DINOv3’s self-supervised training objective.
Specifically, Matryoshka SAEs showed a substantially higher concept alignment
than standard SAEs on ADE20K (reaching 7.9% higher class coverage).
Domain-Specific Discovery
The approach was applied to ecological imagery, where it successfully surfaced fine-grained anatomical structure without access to segmentation or part labels,
providing a scientific case study with ground-truth validation. In the FishVista case study, Matryoshka SAEs reliably find domain-specific concepts,
showing that higher concept prevalence improves label-free rediscovery.
Furthermore, testing foundation model size effects showed that while larger models are harder to reconstruct,
they lead to better downstream concept alignment metrics (mAP), suggesting that strong foundation models in combination with SAEs can rediscover semantic concepts.
Conclusion and Application
The study concludes that sparse decomposition of foundation model representations can systematically surface semantic structure, acting as a prerequisite step toward scientific discovery.
While the work validates SAEs as an instrument for open-ended feature discovery, it is not intended to deliver novel biological or physical discoveries itself. The findings suggest that SAE-based decomposition enables researchers to investigate what patterns models captured from data, generate hypotheses about unexpected structures, and establish reproducible measurements across studies.
The protocol is domain-agnostic and applicable to any setting with foundation model representations.
How it works
The core mechanism involves training a ReLU SAE on patch-level activations extracted from a pre-trained Vision Transformer (ViT), such as DINOv3. The SAE maps an input activation vector to a sparse representation using the objective function: L(θ) = x − ˆx2 / 2 + λS(f(x))
where the training objective minimizes reconstruction error while encouraging sparsity, controlled by the penalty coefficient λ.
Matryoshka SAEs for Feature Splitting
To avoid feature splitting, the authors employ Matryoshka SAEs. This technique learns a single SAE by training multiple nested dictionaries simultaneously, sampling random prefixes of latents and reconstructing inputs using each prefix. The objective function is modified to account for this nesting: L(θ) = X m∈M x − ˆx0:m2 / 2 + λS(f(x)) (6)
This forces early latents to capture more general features, and the resulting latents are then used for concept alignment probes.
Concept Alignment Probes
The learned SAE features are tested for alignment with semantic concepts using binary logistic regression classifiers. For each concept, a 1-D logistic regression is trained on each SAE latent's activations to record the best loss across all latents.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed this paper, Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders.
The core contribution is demonstrating that Sparse Autoencoders (SAEs) can systematically recover unlabeled semantic structure from large foundation model representations without prior concept supervision, effectively enabling open-ended feature discovery.
Here are the specific improvements to AI systems and what they can achieve:
-
The integration of SAE decomposition into the standard foundation model pipeline (e.g., DINOv3).
-
The training of Matryoshka SAEs to learn nested, hierarchical dictionaries from a single model's activations.
-
The use of performance probes (Probe R, mAP, Purity@k, Coverage@τ) to systematically measure the alignment between learned latent features and known semantic concepts (e.g., segmentation classes or anatomical parts).
Specific Improvements:
-
A new module for
Scientific Feature Extraction
that takes the high-dimensional embeddings from a vision foundation model (like DINOv3) and decomposes them into a structured, sparse dictionary of interpretable features using SAEs. -
The implementation of a Matryoshka SAE architecture to capture multi-level semantic hierarchies within the model's latent space, allowing for the discovery of both coarse and fine-grained patterns simultaneously.
-
A concept alignment engine that uses binary logistic regression probes (Eq. 7) to automatically identify which specific SAE latents best represent a target concept (e.g.,
head,
dorsal fin
) without requiring pre-labeled datasets for the probe itself.
What the Improved AI System Can Do:
-
The system can perform unsupervised, open-ended feature discovery by identifying previously unknown semantic patterns within massive, unlabeled scientific datasets (e.g., genomic sequences or ecological imagery).
-
It can systematically surface fine-grained anatomical structures in images (like fish species) without needing human segmentation labels or part-labels during the initial training phase.
-
It can generate measurable, quantifiable features that are directly comparable against existing scientific annotations, serving as a prerequisite for hypothesis generation and scientific discovery rather than just task confirmation.
-
It can be applied to domain-agnostic tasks—such as protein structure analysis or climate modeling—by simply replacing the vision backbone with a relevant foundation model's activations.
Sources
- Understanding intermediate layers using linear classifier probes
- Perception Encoder: The best visual embeddings are not at the output of the network
- Learning Multi-Level Features with Matryoshka Sparse Autoencoders
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Towards scientific discovery with dictionary learning: Extracting biological concepts from microscopy foundation models
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
- World Models
- Training Compute-Optimal Large Language Models
- Open-Endedness is Essential for Artificial Superhuman Intelligence
- Scaling Laws for Neural Language Models
- Adam: A Method for Stochastic Optimization
- Sparse Autoencoders Do Not Find Canonical Units of Analysis
- Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images
- Position: Use Sparse Autoencoders to Discover Unknowns
- Localizing Objects with Self-Supervised Transformers and no Labels
- DINOv3
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- Emergent Abilities of Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models