CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data
Fenosoa Randrianjatovo, Maya Saleh, Simon Girard, Amadou Barry
Institut national de la recherche scientifique · Centre Armand-Frappier Santé Biotechnologie · Université du Québec à Chicoutimi
q-bio.GN, cs.LG, stat.CO, stat.ME
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: 35 pages; 18 figures
Code: https://github.com/FenosoaRandrianjatovo/CosMAP-dr
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data Summary This paper introduces CosMAP (Contrastive Manifold Approximation and
Terminology
Summary
CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data
Summary
This paper introduces CosMAP (Contrastive Manifold Approximation and Projection), a novel graph-based unsupervised dimensionality-reduction method designed to produce faithful and interpretable low-dimensional embeddings of high-dimensional, sparse, and noisy data, particularly omics datasets such as single-cell RNA sequencing data and genealogical data.
Problem Statement: The authors note that existing dimensionality-reduction methods may distort local neighbourhoods, global organization, or the cohesion of meaningful populations. A central limitation of many graph-based DR methods lies in the construction of the high-dimensional neighbourhood graph, which is often fixed and treated as reliable. However, in very high-dimensional spaces, pairwise distances may lose contrast (curse of dimensionality), making nearest-neighbor relations unstable and potentially introducing spurious edges. This is particularly problematic for gene-expression and single-cell data.
Proposed Approach and Hypotheses: CosMAP is based on four key hypotheses:
-
High-dimensional data approximately lie on a low-dimensional manifold, and cosine similarity is advantageous for constructing neighbourhood graphs in high-dimensional settings because it reduces the influence of vector magnitude differences and better captures expression pattern similarities than Euclidean distance.
-
Contrastive principles can strengthen graph-based dimensionality reduction when integrated into both affinity graph construction and embedding optimization.
-
A method based on cosine k-NN graphs offers a favorable balance between embedding quality and computational efficiency compared to more complex metrics like the Dimension Insensitive Euclidean Metric (DIEM).
-
High-dimensional neighbourhood graphs may contain unstable or spurious edges, so CosMAP introduces a refinement strategy that first learns an intermediate representation and then reconstructs a more reliable graph before optimizing the final low-dimensional embedding.
Methodology:
-
High-dimensional similarity: CosMAP constructs a k-NN graph using cosine similarity by default. It defines a local, temperature-scaled neighbourhood distribution using a temperature parameter τ that controls the sharpness of the similarity distribution. The symmetric high-dimensional edge strength is computed as the average of directional probabilities.
-
Low-dimensional similarity: CosMAP uses a heavy-tailed kernel (UMAP-style) to model pairwise similarities in the embedding space, which helps alleviate the crowding problem.
-
Optimization: The embedding is learned by minimizing a binary cross-entropy objective with attractive and repulsive components. Attractive updates are applied to positive graph edges, while repulsive updates use negative sampling for scalability.
-
Two-phase refinement strategy: CosMAP first learns an intermediate higher-dimensional representation (e.g., r=30), then reconstructs the neighbourhood graph from this representation and initializes the final low-dimensional embedding (d=2) using a coordinate projection of the intermediate representation. The second phase uses fewer epochs since most structural information is already captured.
Results:
-
MNIST and USPS handwritten-digit benchmarks: CosMAP produced the most visually interpretable cluster organization among evaluated methods, with compact and well-separated digit classes. LocalMAP was the closest competitor. CosMAP provided clearer separation of visually similar digits (e.g., 3, 5, 8 and 4, 7, 9).
-
Retina dataset (mouse retinal cells): CosMAP achieved the highest average Normalized Mutual Information (NMI) score, indicating strongest agreement between clusters recovered from the embedding and annotated cell-type labels. It clearly separated the rod bipolar cell (RBC) cluster while maintaining readable organization of other cell types. The partial overlap between BC3B and BC4 was noted as biologically plausible given their transcriptional similarity.
-
Cortex dataset (mouse cortical cells): CosMAP provided competitive visualization with well-separated cell-type clusters, particularly distinguishing astrocyte/ependymal and endothelial-mural populations. However, PaCMAP achieved the highest NMI score, followed by NCVis and CosMAP. The paper notes that the two-phase refinement can sometimes over-refine small datasets, potentially fragmenting coherent classes.
-
Genealogical dataset (BALSAC-CARTaGENE kinship matrix): CosMAP produced the clearest and most coherent visual organization of regional structures, distinguishing Greater Montréal, Saguenay–Lac-Saint-Jean/Charlevoix, Bas-Saint-Laurent/Côte-Nord/Gaspésie/Îles-de-la-Madeleine, and Québec City regions. CosMAP attained the best average NMI score, outperforming all other methods.
Conclusion: CosMAP is a competitive and interpretable framework for unsupervised visualization of high-dimensional data. It is particularly effective at producing compact, well-separated clusters while maintaining readable global organization. The paper notes that the optimal choice of metric is dataset-dependent: cosine similarity suits sparse, high-dimensional data like scRNA-seq and images, while Euclidean distance works better for kinship data. Future work will focus on extending CosMAP to additional metrics and developing criteria for selecting between one-phase and two-phase refinement. The implementation is publicly available at https://github.com/FenosoaRandrianjatovo/CosMAP-dr under a BSD 2-Clause license.
Improvements for AI systems
Improvements to AI Systems:
- Adaptive Metric Selection for Dimensionality Reduction
-
Implement a dynamic metric-selection module that automatically chooses between cosine similarity (for sparse, high-dimensional omics/image data) and Euclidean distance (for dense, low-dimensional genealogical data) based on data sparsity and intrinsic dimensionality estimates.
-
The improved system can self-tune its neighborhood graph construction, reducing user bias and improving robustness across heterogeneous datasets.
- Contrastive Graph Refinement for Unsupervised Learning
-
Integrate CosMAP’s two-phase refinement (intermediate representation → graph reconstruction → final embedding) into existing contrastive learning frameworks (e.g., SimCLR, BYOL) to denoise spurious edges in high-dimensional feature spaces.
-
The improved system can produce more stable and semantically coherent clusters in unsupervised tasks like cell-type identification, anomaly detection, or community detection in social networks.
- Temperature-Scaled Similarity for Robust k-NN Construction
-
Replace fixed k-NN graphs with CosMAP’s temperature-parameterized local neighborhood distribution (τ) in graph neural networks (GNNs) and manifold learning pipelines.
-
The improved system can adaptively sharpen or flatten similarity distributions, improving performance on noisy or imbalanced datasets where nearest-neighbor relations are unreliable.
- Heavy-Tailed Embedding Kernel for Crowding Mitigation
-
Adopt CosMAP’s UMAP-style heavy-tailed kernel in autoencoders or variational autoencoders (VAEs) for low-dimensional latent spaces.
-
The improved system can generate more interpretable latent representations that preserve both local and global structure, reducing the “crowding problem” in generative models and enabling better visualization of high-dimensional data.
- Hybrid Objective with Attractive/Repulsive Loss for Graph-Based Learning
-
Incorporate CosMAP’s binary cross-entropy objective (attractive positive edges + repulsive negative sampling) into graph embedding or node representation learning models (e.g., node2vec, GraphSAGE).
-
The improved system can learn more discriminative embeddings that separate meaningful populations (e.g., disease subtypes, geographic origins) while maintaining computational scalability via negative sampling.
- Automatic One-Phase vs. Two-Phase Refinement Selection
-
Develop a meta-learning criterion (e.g., based on dataset size, intrinsic dimensionality, or cluster coherence) to decide whether to apply CosMAP’s two-phase refinement or a simpler one-phase embedding.
-
The improved system can avoid over-fragmentation on small datasets (as observed in the cortex dataset) while maximizing structure preservation on larger, noisier datasets.
- Interpretable Cluster Cohesion Scoring for Biological Validation
-
Use CosMAP’s NMI-based evaluation framework to build a post-hoc validation tool that quantifies the biological plausibility of clusters (e.g., partial overlap of transcriptionally similar cell types like BC3B/BC4).
-
The improved system can flag over-segmentation or biologically implausible groupings, aiding researchers in hypothesis generation and quality control.
- Scalable Negative Sampling for High-Dimensional Omics
-
Integrate CosMAP’s negative sampling strategy into existing single-cell analysis pipelines (e.g., Scanpy, Seurat) to accelerate embedding of millions of cells without sacrificing cluster fidelity.
-
The improved system can process large-scale scRNA-seq datasets in real-time, enabling interactive exploration of cell atlases.
- Cross-Domain Transfer of Graph Refinement
-
Apply CosMAP’s graph-reconstruction step to other graph-based AI tasks (e.g., knowledge graph completion, recommendation systems) where initial similarity graphs are noisy.
-
The improved system can learn more reliable relational structures, improving link prediction and recommendation accuracy in sparse, high-dimensional feature spaces.
- Uncertainty-Aware Embedding for Rare Population Detection
-
Extend CosMAP’s temperature parameter to output uncertainty estimates for each embedding point, flagging low-confidence regions (e.g., transitional cell states or rare subtypes).
-
The improved system can guide targeted experimental validation in biology or identify outlier communities in social data, enhancing exploratory data analysis.
Sources
- Attention Is All You Need
- Dimension Reduction with Locally Adjusted Graphs
- The Shape of Attraction in UMAP: Exploring the Embedding Forces in Dimensionality Reduction
- Distributed Representations of Words and Phrases and their Compositionality
- Surpassing Cosine Similarity for Multidimensional Comparisons: Dimension Insensitive Euclidean Metric
- A survey of dimensionality reduction techniques
- Visualizing Large-scale and High-dimensional Data
- TriMap: Large-scale Dimensionality Reduction Using Triplets
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- From $t$-SNE to UMAP with contrastive learning
- Understanding How Dimension Reduction Tools Work: An Empirical Approach to Deciphering t-SNE, UMAP, TriMAP, and PaCMAP for Data Visualization
- Auto-Encoding Variational Bayes
- NCVis: Noise Contrastive Approach for Scalable Visualization
- Node Embeddings via Neighbor Embeddings
- Scikit-learn: Machine Learning in Python
- Unsupervised visualization of image datasets using contrastive learning
- Hierarchical Nearest Neighbor Graph Embedding for Efficient Dimensionality Reduction
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Array Programming with NumPy
- The Faiss library