The Rashomon Effect for Visualizing High-Dimensional Data
cs.LG
Submitted: 2026-04-01
Updated: 2026-08-26
Comments: The paper is accepted in AISTATS 2026
Code: https://github.com/berenslab/contrastive-ne
Project page: https://pair-code.github.io/understanding-umap
License: http://creativecommons.org/licenses/by/4.0/
The gist: Dimension reduction (DR) is inherently non-unique: multiple embeddings can preserve the structure of high-dimensional data equally well while differing in layout or geometry.
Terminology
Abstract
Dimension reduction (DR) is inherently non-unique: multiple embeddings can preserve the structure of high-dimensional data equally well while differing in layout or geometry. In this paper, we formally define the Rashomon set for DR -- the collection of `good' embedding -- and show how embracing this multiplicity leads to more powerful and trustworthy representations. Specifically, we pursue three goals. First, we introduce PCA-informed alignment to steer embeddings toward principal components, making axes interpretable without distorting local neighborhoods. Second, we design concept-alignment regularization that aligns an embedding dimension with external knowledge, such as class labels or user-defined concepts. Third, we propose a method to extract common knowledge across the Rashomon set by identifying trustworthy and persistent nearest-neighbor relationships, which we use to construct refined embeddings with improved local structure while preserving global relationships. By moving beyond a single embedding and leveraging the Rashomon set, we provide a flexible framework for building interpretable, robust, and goal-aligned visualizations.
Sources
- TriMap: Large-scale Dimensionality Reduction Using Triplets
- Automatic Selection of t-SNE Perplexity
- Clustering with UMAP: Why and How Connectivity Matters
- Representation Learning with Contrastive Predictive Coding
- Median Consensus Embedding for Dimensionality Reduction
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks