An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders".
Jane: The paper was written by Scott C. Lowe, Joakim Bruslund Haurum, Sageev Oore, Thomas B. Moeslund and Graham W. Taylor from Vector Institute, Canada and Aalborg University, Denmark and Pioneer Centre for AI, Denmark and Dalhousie University, Canada and University of Guelph, Canada.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Initial Implications: Jane: The title itself promises a lot, focusing on how these self-supervised encoders perform when they are tested on datasets they were never trained for, which is a concept known as zero-shot transfer.
Tom: And the authors confirm that this generalization isn't just theoretical; they have a comprehensive suite of benchmarking experiments to prove it.
Lu: This suggests that the underlying principles used to learn features—like identifying edges or textures—are universal across domains, even if the training set didn's looked like different kinds of natural scenes.
Meng: The most impactful implication is that we’ don't necessarily need to retrain a huge model for every new task; we can just leverage the rich knowledge it already possesses from its original pretraining phase.
Lalam: We are moving toward understanding the data not just by its rigid category, but by how its physical properties naturally interact with our perception, allowing us to see hidden relationships.
Tom: It's an amazing way to frame this whole idea of "unseen datasets" and setting the stage for a deeper look into what these specific models can actually achieve in the next section.
Summary of Key Findings: Jane: The summary provides a clear picture, showing that while SSL encoders generally perform worse than supervised models on datasets similar to their training data, there's a surprising shift when we look at truly novel information.
Tom: Specifically, those "Far-OOD" groups are where the SSL models start outperforming their traditional counterparts by a noticeable margin in terms of performance.
Lu: This indicates that when we look at completely unfamiliar data, these SSL encoders are actually more capable of discovering underlying structure because they're not constrained by the specific biases of the training set.
Meng: The authors also found that using manifold-based reduction, like UMAP, was essential for making those complex embeddings usable for clustering; it makes the raw data tractable.
Lalam: It’s about establishing that the foundational knowledge learned by these models is stable, ensuring that we are not just seeing a random alignment of points but truly seeing emergent structure.
Tom: And beyond just performance, they have discovered that measuring an AI's ability to produce well-clustered embeddings gives us a powerful new way to evaluate its performance, independent of traditional kNN accuracy.
Jane: The researchers also observed that Agglomerative Clustering performed best among the methods they tested, though it was noted that the effect size was relatively small compared to some other techniques.
Lu: This allows us to investigate how clusterings can be further analyzed to see which stimulus attributes an encoder is prioritizing—be it color or texture—giving us deeper insight into the model's internal decision-making process.
Meng: This method helps in identifying biases, providing a clear way for us to evaluate if an AI is working as intended, regardless of whether we are dealing with complex art styles or natural landscapes.
Lalam: We are seeing the data not just as rigid categories, but as interconnected structures that inherently reflect human perception of the world's qualities.
Methodology and Improvements: Jane: The authors were very careful about how they tested this generalization, ensuring the results weren't just a lucky coincidence by performing an exhaustive parameter search.
Tom: They did this by using subsets of ImageNet-1k and other smaller datasets to find stable settings for all clusterers across different levels of data granularity.
Lu: This is essential because it ensures that even if we are dealing with small datasets or those with complex labels, the inherent structure can still be identified through the method itself.
Meng: The practical approach here is using a staggered sweep over relevant parameters, which makes the entire pipeline highly reliable for deployment on real-world data streams where inputs are often messy and inconsistent.
Lalam: It’s about establishing that the foundational knowledge learned by these models is stable, ensuring that we are not just seeing a random alignment of points but truly seeing emergent structure.
Tom: The paper also provided detailed analysis of how different types of SSL learning—for instance, comparing Contrastive Learning versus Masked Image Modeling—contribute to the final result. They showed that the way we train the model dictates what it sees.
Jane: They found that MAE-trained models performed particularly poorly in certain scenarios because they need fine-tuning to achieve success on whole images, which is a limitation we must keep in mind for zero-shot tasks.
Lu: This tells us that the training paradigms are critical; if a model is trained on local features, it focuses on texture, but if trained differently, its focus shifts to global forms.
Meng: The study helps us understand which SSL paradigm is best suited for a task by showing how each method handles different levels of complexity and noise in real-world data.
Lalam: It’s about moving towards seeing the data not as rigid categories, but as interconnected structures that are inherently meaningful to our perception.
Conclusion and Final Takeaways: Tom: So, we've seen that "An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders" really shows us that these modern AI systems possess a surprising level of generalization when facing totally unfamiliar data.
Jane: It’s fascinating to think about how they were able to perform this zero-shot transfer, even when the data was completely outside the original training set.
Lu: I see this as a massive step toward what we might call emergent generalization, where the model is identifying intrinsic relationships that go beyond confirming its initial training data patterns.
Meng: From a practical standpoint, it proves we can use these generalized, pre-trained representations as a solid foundation for clustering without needing to retrain the models from scratch for every single new task.
Lalam: It’s really about recognizing that AI has already learned patterns in the world—patterns of form and color—and that we can finally harness those intrinsic groupings in ways that truly reflect human perception.
Tom: That is a powerful way to put it, Lalam; you're talking about leveraging latent knowledge for discovery in this zero-shot learning space, which is a major leap forward.
Jane: And Meng is correct, making those generalized representations usable without retraining is a huge practical win for almost every industry using AI today. The utility of this research is undeniable.
Lu: The finding that SSL encoders often outperform supervised models when moving into the Far-Out-of-Domain space really shows us where these systems are headed—beyond just mastering what they know.
Meng: That trend of performing better outside the training distribution is a signal that we should be looking at these models as powerful discovery tools, not just prediction engines. They are finding patterns we haven't seen yet.
Lalam: We’re moving toward seeing the data not just by its category, but by how its physical properties naturally interact with our own perception of the world.
Tom: This study has really opened up a lot of conversation about what we expect from AI versus what we are actually seeing in this zero-shot learning space, Jane.
Jane: It’s a lot of information to process, but it’s exciting to see how these models find meaning in unexpected data. We hope that "An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders" provides a clear path forward for the world that's waiting for the next major breakthrough.
Lu: I think the future is looking very bright with these findings, especially for global pattern recognition across different cultures and environments.
Meng: And I hope this gives us a clear guide to real-world data analysis, utilizing "An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders" as a practical roadmap.
Lalam: We are moving toward a future where AI doesn't just fit things into boxes, but helps us see the interconnected web of data in ways that feels more organic and insightful.
Vector Institute, Canada · Aalborg University, Denmark · Pioneer Centre for AI, Denmark · Dalhousie University, Canada · University of Guelph, Canada
cs.LG, cs.AI, cs.CV
Submitted: 2024-06-04
Updated: 2026-09-04
Comments: Published in Transactions on Machine Learning Research (08/2026)
Journal ref: Transactions on Machine Learning Research (2026)
Code: https://github.com/scottclowe/zs-ssl-clustering
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: " The study addresses the question, "Can pretrained models generalize to new datasets without any retraining?" It investigates whether embeddings from self-supervised learning (SSL) encoders can form
Key concepts
- Zero-Shot Transfer
- This concept refers to the ability an AI model has to perform tasks or generalize on datasets it was never trained on. The study demonstrates this by showing how self-supervised encoders can successfully apply their learned knowledge to completely unfamiliar data.
- Self-Supervised Encoders (SSL)
- These are types of AI models that learn features from vast amounts of data without explicit human labeling. The research highlights that these models possess foundational knowledge, allowing them to identify underlying structure in new datasets.
- Far-OOD Groups
- This term refers to completely unfamiliar or 'out-of-distribution' data. The study found that SSL models often outperform traditional methods when analyzing these novel groups of information.
Terminology
Summary
"
The study addresses the question, Can pretrained models generalize to new datasets without any retraining?
It investigates whether embeddings from self-supervised learning (SSL) encoders can form meaningful clusters when applied to image datasets that were not seen during training.
While prior work suggested SSL features were suitable for clustering, the paper notes a lack of investigation into this zero-shot transfer-learning task,
especially given the high cost of compute involved in fine-tuning large pretrained models.
** Feature Encoders:** The researchers compare various self-supervised paradigms, including:
-
Contrastive Learning: MoCo-v3 (Chen et al., 2021)
-
Self-Distillation: DINO (Caron et al., 2021)
-
Canonical Correlation Analysis: VICReg (Bardes et al., 2022)
-
Masked Image Modelling: MAE (He et al., 2022).
The models are trained on ImageNet-1k (IN-1k) and use two backbone architectures: ResNet-50 and ViT-B.
** Clustering Methods:** A suite of classical clustering methods is employed, including K-Means, Spectral Clustering, Agglomerative Clustering (AC), Affinity Propagation (AP), and HDBSCAN.
** Datasets:** The experiments on a diverse set of datasets are organized into five groups:
-
In-domain (ID): Images and labels within the IN-1k domain.
-
Domain-shifted (DS): Class labels aligned with IN-1k, but images changed (e.g., background removed).
-
Near-out-of-domain (Near-OOD): Images similar to IN-1k, but with new classes and distributional shift.
-
Fine-grained near-out-of-domain (FG): Natural images resembling a subdomain of IN-1k, but labeled at a much finer level of granularity.
-
Far out of domain (Far-OOD): Images outside the domain, such as textures, text, faces, or microscopy slides.
** Evaluation Metrics:** Performance is measured using two metrics: adjusted mutual information (AMI) and silhouette score (S).
** Dimensionality Reduction:** Due to the curse of dimensionality,
various reduction processes were tested: PCA (linear), UMAP (manifold-based), and PaCMAP. The researchers found that embeddings with pretrained networks were typically best clustered with some form of manifold-based reduction.
** Clustering Performance:**
-
Among the clusterers, "AC w/ C performs best (p < 0.05; Wilcoxon signed-rank test versus each other clusterer)."
-
"HDBSCAN performed worst (p < 10−33), due to its use of a noise class instead of trying to place every sample in a cluster."
** SSL vs. Supervised Encoders:**
-
The performance of the SSL encoders is lower than that of the supervised network on in-domain, domain-shifted, and near-OOD datasets...
-
"Taken together, these results demonstrate that supervised encoders perform better at clustering unseen datasets similar to the training data, but as the data moves further from the training dataset, performance of supervised networks decreases and SSL encoders increases such that they become better."
** Effect of Dataset Granularity:**
- The study observed a correlation between performance and granularity. Specifically,
we find that for all methods the AMI score peaks at a medium-grained level at either the class or order taxonomic level, while drastically decreasing when the labels are too coarse or fine-grained.
** Background Challenge (Foreground vs. Background):**
- The study found that "Supervised and fine-tuned networks consistently had more information about the background of the images (FGC and BG), congruent with the widely held belief that supervised networks learn to exploit information in the background of images. Surprisingly then, we find SSL-encoders have nearly twice as large a BG-gap than their supervised counterparts."
** Fine-Tuned SSL Encoders:**
-
Fine-tuning significantly improves performance:
As shown in Figure 3, we found fine-tuning unsurprisingly increases performance on in-domain and domain-shifted datasets...
-
In the Far-OOD category, however,
the performance of SSL-encoders became worse than supervised encoders.
** Orthogonal Evaluation:**
- The study established that clustering provides an evaluation method
orthogonal to kNN,
finding only amoderate positive correlation between kNN-probing accuracy and AMI (0.33 ≤ ρ ≤ 0.54).
** Manifold Learning:**
- The findings suggest that the output of trained encoders lie on a low-dimensional, non-linear manifold:
embeddings with pretrained networks were typically best clustered with some form of manifold-based reduction.
The study concludes that while supervised encoders are superior for data similar to their training distribution, SSL encoders offer better performance when dealing with out-of-domain data. The research emphasizes the importance of the alignment between the model’s training task and its downstream applications, noting that fine-tuning improves performance but this gain comes at a cost, greatly reducing the performance on Far-OOD data.
Improvements for AI systems
Based on the detailed empirical study presented in this paper, we can derive several specific and actionable improvements for designing robust AI systems that require feature representation and grouping (clustering) without relying on ground-truth labels (zero-shot learning).
The core of the improvement lies in leveraging the superior intrinsic clustering properties of Self-Supervised Learning (SSL) embeddings when applied to unseen data, specifically by optimizing the dimensionality reduction and selection of clusterers.
Improvement: Instead of relying on raw high-dimensional embeddings or linear dimensionality reduction (like PCA), the system must mandate a non-linear, manifold-based dimension reduction step—specifically Uniform Manifold Approximation and Projection (UMAP)—before clustering.
System Capability: The improved system will be able to successfully cluster data points that lie on complex, non-linear manifolds within the embedding space. This allows for the discovery of legitimate
groupings (e.g, grouping all 'bird' images regardless of specific species) even if the underlying dataset is entirely novel and lacks ground-truth labels.
Improvement: The system should prioritize Agglomerative Clustering with a fixed number of clusters (AC w/ C) over other non-parametric methods like HDBSCAN or Affinity Propagation, as the empirical results showed AC w/ C consistently yields the highest average Adjusted Mutual Information (AMI) and Silhouette score across various datasets.
Improvement: The system must integrate the Silhouette Score (S) as a primary, intrinsic validation metric for any clustering algorithm that does not rely on ground-truth labels. This score should be calculated specifically within the UMAP-reduced feature space.
Improvement: The system must be architecturally designed to transfer pretraining knowledge from a single source (e.g., ImageNet-1k) and then apply the resulting feature embeddings directly to target datasets without any per-dataset parameter tuning or fine-tuning, enabling a true zero-shot transfer.
Abstract
Can pretrained models generalize to new datasets without any retraining? We deploy pretrained image models on datasets they were not trained for, and investigate whether their embeddings form meaningful clusters. Our suite of benchmarking experiments uses encoders pretrained solely on ImageNet-1k with either supervised or self-supervised training techniques, deployed on image datasets that were not seen during training, and clustered with conventional clustering algorithms. This evaluation provides new insights into the embeddings of self-supervised models, which prioritize different features to supervised models. We find evidence that supervised encoders offer more utility than SSL encoders within the training domain, and vice-versa far outside of it. However, fine-tuning SSL encoders for ImageNet-1k classification results in the opposite behaviour, with better performance than supervised-only models on in-domain and decreased performance on far out of domain data - worse at far-OOD than either SSL-only or supervised-only models. Clustering provides a way to evaluate the utility of self-supervised learnt representations orthogonal to existing feature quality estimation methods. Additionally, we find the silhouette score when measured in a UMAP-reduced space is highly correlated with clustering performance, and can therefore be used as a proxy for clustering performance on data with no ground truth labels. Our code implementation is available at https://github.com/scottclowe/zs-ssl-clustering/.
Sources
- A Cookbook of Self-Supervised Learning
- Measuring Dataset Granularity
- CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at Scale
- Self-supervised Pretraining of Visual Features in the Wild
- Fine-Grained Visual Classification of Aircraft
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- DINOv2: Learning Robust Visual Features without Supervision
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Noise or Signal: The Role of Image Backgrounds in Object Recognition
- LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop
- iBOT: Image BERT Pre-Training with Online Tokenizer
- A Comprehensive Survey on Deep Clustering: Taxonomy, Challenges, and Future Directions
- Deep Clustering with Features from Self-Supervised Pretraining
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks