MOE-Enhanced Explanable Deep Manifold Transformation for Complex Data Embedding and Visualization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "MOE-Enhanced Explanable Deep Manifold Transformation for Complex Data Embedding and Visualization".
Jane: The paper was written by Zelin Zang, Yuhao Wang, Jinlin Wu, Hong Liu, Yue Shen et al. from Centre for Artificial Intelligence and Robotics, HKISI-CAS and Westlake University and Hangzhou City University and Academy of Edge Intelligence, Hangzhou City University and State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA and School of Artificial Intelligence, University of Chinese Academy of Sciences and Ant Group.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back, everyone. I'm Tom, and alongside me is the wonderful Jane. We're diving into a fresh paper today, and the title alone is a mouthful: "MOE-Enhanced Explainable Deep Manifold Transformation for Complex Data Embedding and Visualization."
Jane: Tom, that title is a real puzzle box. Let's crack it open for our listeners. Basically, this is about dimensionality reduction—taking super high-dimensional data, like images with thousands of pixels or gene expression profiles, and squishing it down to something we can actually see in 2D or three dee.
Tom: Right, and the "MOE" part stands for Mixture of Experts. That's the clever twist here. Instead of one giant neural network trying to learn everything, you have a bunch of smaller, specialized networks, and a router decides which one handles each piece of data.
Jane: It's like having a team of specialists instead of one generalist. If you're looking at a picture of a cat, you might want the "whisker expert" to chime in, while the "fur texture expert" handles another part. This paper, from authors like Zelin Zang and Yuhao Wang, is saying that this team approach makes the whole process both more accurate and easier to understand.
Tom: And that's the "Explainable" part of the title. Usually, these deep learning models are black boxes. You feed them data, and they spit out an embedding, but you have no idea why. This paper wants to open that box and show you exactly which features each expert is using.
Jane: Exactly. So we're not just talking about a better way to visualize data; we're talking about a way to visualize data *and* understand what the model is seeing. That's a huge deal for fields like biology, where you want to know *why* certain cells cluster together, not just that they do.
Tom: So, Jane, if I'm a researcher with a messy dataset, this could be the tool that finally lets me see the forest *and* the trees, while also telling me which trees matter most. I'm already excited to see how they pulled this off.
Jane: Me too, Tom. Let's get into the meat of how they actually built this thing in the next segment.
Paper discussion segment 2: Tom: So, Jane, we've got the title decoded. Now, how does this "MOE-Enhanced Explainable Deep Manifold Transformation" actually work under the hood? The paper, which we'll just call DMT-ME for short, has a few key parts.
Jane: Right. The first piece is the "Multiple Gumbel Matchers." That's their way of deciding which features go to which expert. It's not a random split. The model learns to assign, say, a specific set of pixels or genes to a specific expert based on what it's seeing.
Tom: And then there's the "Hyperbolic Mapper." This is where things get really interesting. Instead of embedding data in a flat, Euclidean space, they use a curved, hyperbolic space. Think of it like a saddle shape or a Pringle chip.
Jane: A Pringle chip, Tom? That's a new one for me.
Tom: Hey, it works! The point is, hyperbolic space is great for representing hierarchical data, like a family tree or a cell differentiation pathway. A flat space runs out of room, but a curved space can hold that tree structure much more naturally.
Jane: That's a great way to put it. And this is where the "Deep Manifold Transformation" comes in. They're not just mapping points; they're trying to preserve the underlying structure, the manifold, of the data. They have a special loss function, the Sub-Manifold Matching loss, that makes sure the relationships between points in the high-dimensional space are kept intact in the low-dimensional one.
Tom: And they don't stop there. They also have an "Expert Exclusive Loss" to make sure the experts don't all end up doing the same thing. It's like telling your team of specialists, "Hey, you two, stop copying each other. Focus on your own area."
Jane: The whole system is a balancing act. You want each expert to be good at its job, but you also want them to be different. And the results seem to show it works. They tested it on everything from MNIST digits to complex biological datasets like the Human Cell Landscape.
Tom: And the numbers are impressive. On CIFAR-one hundred a notoriously hard image dataset, they're getting classification accuracy in the high 70s, while a classic method like t-SNE is stuck in the single digits. That's not a small jump.
Jane: It's a massive jump. But the real question for me, Tom, is whether this complexity is worth it. Is it just a performance boost, or does the explainability part actually deliver? Let's dig into that in the next segment.
Paper discussion segment 3: Tom: Welcome back. We've talked about the architecture of DMT-ME, but the real headline here is the explainability. The paper claims you can look at the experts and see what they're focusing on. Let's bring in Lu and Meng to get their take on this.
Lu: Thanks, Tom. This is the part that gets me excited. The paper shows that on the K-MNIST dataset, which has handwritten Kanji characters, different experts literally specialize in different stroke patterns. You can see which expert is responsible for which part of the character. That's not just a black box giving you an answer; it's a model showing its work.
Meng: And from an engineering standpoint, that's incredibly valuable for debugging. If the model is making a mistake, you can look at which expert is firing and see if it's focusing on the wrong features. That's a huge step up from trying to reverse-engineer a monolithic network. But I have to ask, what's the computational cost of running ten experts?
Jane: That's a fair question, Meng. The paper actually addresses that. They show that while DMT-ME has a higher upfront training cost, it's actually faster than many non-parametric methods on large datasets. On the HCL dataset, it took about five and a half minutes, while t-SNE took over thirteen. So the parallel nature of the experts pays off.
Lu: And the explainability isn't just for images. They did a case study on the Human Cell Landscape data, and they were able to link specific experts to specific tissue types and even to specific genes. That's a potential tool for unsupervised biomarker discovery. You could find new genes that are important for a cell type without having any prior labels.
Meng: So it's not just about visualization. It's about generating testable hypotheses. You see a cluster, you look at which expert is responsible, and you see which genes it's weighting heavily. That gives you a lead to go validate in the lab.
Tom: So, Lu and Meng, you're both saying this could be a game-changer for how we interact with complex data. It's not just a better picture; it's a more transparent and actionable picture.
Jane: And that's the key difference from the older methods. They give you a picture, but DMT-ME gives you a picture with a legend and a map. Now, let's wrap this up and see what the big-picture impact could be.
Conclusion: Tom: We've reached the end of our discussion on the "MOE-Enhanced Explainable Deep Manifold Transformation for Complex Data Embedding and Visualization" paper. Jane, what's the final verdict?
Jane: I think the biggest takeaway is that this paper tackles the classic trade-off between performance and interpretability. For a long time, you had to choose: either a fast, simple method that gives you a decent picture, or a powerful deep learning model that's a black box. DMT-ME shows you can have both.
Tom: And it does it with a clever combination of ideas. The Mixture of Experts for specialization, the hyperbolic space for hierarchy, and the loss functions to keep everything in check. It's a well-engineered solution.
Jane: The implications are huge. For a biologist looking at single-cell data, this could mean finding new cell types and the genes that define them. For someone working with images, it could mean understanding exactly what features a model uses to make a decision. It's about building trust in the models we use.
Tom: And it's not just for experts. The fact that the model can explain itself makes it more accessible to people who aren't machine learning specialists. That's a win for everyone.
Jane: Absolutely. We've said goodbye to this paper, but I have a feeling the ideas behind it are going to stick around. The push for explainable, high-performance models is only going to grow.
Tom: Well said, Jane. That's all the time we have for this one. Thanks to Lu and Meng for joining us. And to our listeners, stay curious, and we'll see you for the next paper.
Jane: Take care, everyone.
Zelin Zang, Yuhao Wang, Jinlin Wu, Hong Liu, Yue Shen, Zhen Lei, Stan Z. Li
Centre for Artificial Intelligence and Robotics, HKISI-CAS · Westlake University · Hangzhou City University · Academy of Edge Intelligence, Hangzhou City University · State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA · School of Artificial Intelligence, University of Chinese Academy of Sciences · Ant Group
cs.LG, cs.AI
Submitted: 2026-08-16
Updated: 2026-08-18
Comments: 14 pages, 8 figures
Code: https://github.com/zangzelin/code_dmtme
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 47/100
Key concepts
- Mixture of Experts (MOE)
- Instead of one large neural network, this system uses several smaller, specialized networks or 'experts.' A router decides which specific expert should handle each piece of data. This allows the model to function like a team of specialists—for example, having one expert focus on texture while another focuses on shape.
- Hyperbolic Space
- This is a curved mathematical space used for embedding data. Unlike flat Euclidean space, hyperbolic geometry is excellent for representing hierarchical structures, such as family trees or cell differentiation pathways. It allows the model to naturally hold complex relationships that would run out of room in a standard 2D or 3D visualization.
- Explainable AI
- This concept moves beyond 'black box' deep learning models. The this system is designed so that users can see exactly which features or experts are being used to make a decision. This transparency is crucial for researchers, allowing them to understand the underlying logic of complex data clustering.
- Deep Manifold Transformation (DMT)
- This process maps high-dimensional data into a lower dimension while preserving the underlying structure, or 'manifold,' of that data. It uses specific loss functions, like Sub-Manifold Matching loss, to ensure that the relationships between points in the original space are kept intact in the final visualization.
Terminology
Summary
Summary
This paper introduces the MOE-based Explainable Deep Manifold Transformation (DMT-ME), a novel dimensionality reduction (DR) method designed to address the trade-off between high DR accuracy and strong explainability, particularly for complex high-dimensional data such as images, tabular data, and text. The authors state: To address these limitations in global structure modeling and interpretability, we propose the mixture of experts (MOE)-based explainable deep manifold transformation (DMT-ME).
The proposed approach integrates hyperbolic embeddings, which effectively capture complex hierarchical structures,
with Mixture of Experts (MOE) models, which dynamically allocate tasks based on input features.
The main contributions of the work are threefold: First, a submanifold-based dynamic matching loss function that enhances DR accuracy by capturing global structures more effectively.
Second, a MOE strategy that strengthens explainability and stability by linking input features, embeddings, and key components explicitly.
Third, extensive evaluations on global and local performance, time efficiency, and additional metrics to demonstrate the advantages of DMT-ME.
The DMT-ME model architecture consists of four key components: an Allocator, a Backbone, a Hyperbolic Mapper, and an Interpreter. The Allocator employs multiple Gumbel operator-based matchers to assign different data segments to expert networks in the MOE backbone, promoting diverse task allocation regulated by an orthogonal loss.
The Backbone extracts features via expert networks while preserving the underlying manifold structure. The Hyperbolic Mapper projects data into hyperbolic space to capture non-Euclidean relationships. The Interpreter enhances explainability and visualization by refining outputs with a visualization manifold loss.
A key innovation is the multiple Gumbel operator,
an extension of Gumbel-Softmax that allows each expert to dynamically select a sparse, task-relevant subset of features. The authors note: This routing strategy promotes disentangled expert behaviors and inherently supports explanation sparsity.
The MOE network consists of 10 expert models, each structured as a 4-layer MLP with 512 neurons per hidden layer. The multiple Gumbel matchers is implemented as a two-layer MLP with 100 hidden neurons, and the value of O (number of features assigned to each expert) is set to ⌈0.9 × D⌉.
The model uses a Sub-Manifold Matching (SMM) loss function, which the authors theoretically justify as providing provably more stable gradient dynamics than the KL-divergence used in t-SNE (see Lemma 1 and Lemma 2), resulting in smoother optimization and faster convergence.
The SMM loss is defined as: LSMM(h, h+):= 1/2 Σi Σj log S hh ij + Σi Σj log S h+h+ ij − γ · Σi log diag(S hh+ ii), where γ > 0 is an exaggeration factor. The similarity matrices are computed using a t-distribution kernel. The overall loss function is L:= LSMM(h, h+) + LSMM(ẽ, ẽ+) + λLExc(e), where LExc is an expert exclusive loss that promotes diversity among expert representations via cosine similarity penalties.
Experimental results demonstrate DMT-ME's superior performance. On global structure preservation (SVM classification accuracy), DMT-ME achieves the highest accuracy across all ten benchmark datasets. For instance, on CIFAR-10 and CIFAR-100, DMT-ME attains 77.5%/74.9% (train/test) and 39.1%/38.9%, respectively, substantially outperforming UMAP and t-SNE, which perform below 6% on CIFAR-100.
On MNIST and EMNIST, DMT-ME achieves 97.8%/97.0% and 69.9%/69.0%, respectively. On biological datasets, DMT-ME achieves 83.8%/76.8% on HCL and 85.4%/76.4% on MCA. For local structure preservation, DMT-ME achieves the highest average Trustworthiness score of 81.7% and the highest average KNN accuracies of 80.4% (training) and 76.6% (testing).
Time efficiency comparisons show that DMT-ME achieves the fastest runtimes on E-MNIST (6 min 2 s) and HCL (5 min 34 s), highlighting its scalability. Ablation studies confirm that removing any core module (MOE, Gumbel matcher, hyperbolic mapper, or SMM loss) degrades performance, and the full DMT-ME configuration consistently achieves the highest Trustworthiness and kNN accuracy. Hyperparameter studies show optimal performance with K=5 neighbors, 20 MOE experts, and hyperbolic parameter ν=0.2.
Case studies on K-MNIST and HCL datasets demonstrate DMT-ME's explainability. On K-MNIST, the model produces well-organized hyperbolic representations, forming distinct clusters that correspond to different character classes
and shows robustness to label noise
by embedding mislabeled samples closer to their true categories. On HCL, DMT-ME enhances explainability by explicitly linking tissue clusters with relevant gene markers
through a tri-partite mapping from tissue types to specialized experts and their associated genes, showing potential of unsupervised target discovery
for identifying biologically significant genes.
The authors conclude: "This paper proposes DMT-ME, a novel MOE-based hyperbolic explainable deep manifold transformation method for DR, which addresses the limitations of traditional methods by integrating hyperbolic embeddings with an MoE framework. DMT-ME improves both effectiveness and interpretability by preserving complex data structures and enabling adaptive task allocation across diverse datasets."
Improvements for AI systems
Based on the scientific paper, here are the specific improvements I can make to AI systems, along with the resulting capabilities:
Improvement: Integrate the Sub-manifold Matching (SMM) loss function into existing representation learning frameworks. Unlike contrastive losses (InfoNCE) or KL-divergence (t-SNE), SMM provides provably bounded and stable gradients (Theorem 1), preventing gradient explosion during training. It aligns full similarity matrices rather than anchor-positive pairs, preserving both local neighborhoods and global inter-cluster distances.
Resulting Capability: The AI system can generate low-dimensional embeddings (2D/3D) that retain hierarchical and non-Euclidean structures, making it suitable for visualizing complex datasets (e.g., single-cell genomics, high-resolution images) with superior cluster separation and biological interpretability. It also scales to large datasets (up to 300K samples) without memory overflow, a common failure of parametric-free methods.
Improvement: Implement the multiple Gumbel matchers to allocate disjoint, task-relevant feature subsets to each expert network. This creates a provably faithful additive explanation (Lemma 4) where the model output is exactly the sum of expert contributions, with bounded explanation complexity (C(f) ≤ s). The orthogonal loss (LExc) ensures expert diversity, preventing redundancy.
Improvement: Replace the final Euclidean projection layer with a Hyperbolic MLP (HMLP) using exponential/logarithmic maps. This allows the model to represent tree-like and hierarchical structures with lower distortion than Euclidean spaces, as demonstrated by improved performance on datasets like CIFAR-100 and HCL.
Improvement: Adopt the neighbor-based interpolation augmentation for tabular data (Eq. 3) and standard image augmentations (crop, flip, color jitter) combined with the SMM loss. This improves model generalization and stability, particularly for small datasets where deep learning methods typically underperform.
Improvement: Use the Gumbel-Softmax-based routing to dynamically assign input samples to specialized experts based on their features. This avoids the bottleneck of a single monolithic model and improves computational efficiency (e.g., 2x faster than DMT-EV on HCL dataset).
The improved AI system can:
-
Visualize complex, high-dimensional data with clear separation of clusters and preservation of both local and global structures, even for datasets with 100K+ samples.
-
Explain its embeddings by linking input features to expert modules, providing feature importance scores and group-level insights (e.g., which genes drive cell-type separation).
-
Scale to large datasets without memory overflow, using efficient GPU-parallelized training.
-
Generalize to unseen data via a parametric model, unlike t-SNE/UMAP which require re-computation.
-
Adapt to diverse data types (image, tabular, biological) with minimal hyperparameter tuning, using the recommended settings (ν=0.1, K=5-20, 10-20 experts).
Abstract
Dimensionality reduction (DR) plays a crucial role in various fields, including data engineering and visualization, by simplifying complex datasets while retaining essential information. However, achieving both high DR accuracy and strong explainability remains a fundamental challenge, especially for users dealing with high-dimensional data. Traditional DR methods often face a trade-off between precision and transparency, where optimizing for performance can lead to reduced explainability, and vice versa. This limitation is especially prominent in real-world applications such as image, tabular, and text data analysis, where both accuracy and explainability are critical. To address these challenges, this work introduces the MOE-based Explainable Deep Manifold Transformation (DMT-ME). The proposed approach combines hyperbolic embeddings, which effectively capture complex hierarchical structures, with Mixture of Experts (MOE) models, which dynamically allocate tasks based on input features. DMT-ME enhances DR accuracy by leveraging hyperbolic embeddings to represent the hierarchical nature of data, while also improving explainability by explicitly linking input data, embedding outcomes, and key features through the MOE structure. Extensive experiments demonstrate that DMT-ME consistently achieves superior performance in both DR accuracy and model explainability, making it a robust solution for complex data analysis. The code is available at https://github.com/zangzelin/code dmtme
Sources
- GNUMAP: A Parameter-Free Approach to Unsupervised Dimensionality Reduction via Graph Neural Networks
- Hyperbolic domains in real Euclidean spaces
- A Survey on Mixture of Experts in Large Language Models
- Parametric UMAP embeddings for representation and semi-supervised learning
- Categorical Reparameterization with Gumbel-Softmax
- Higher-dimensional Euclidean and non-Euclidean structures in planar circuit quantum electrodynamics
- Representation Learning with Contrastive Predictive Coding
- A Review of BioTree Construction in the Context of Information Fusion: Priors, Methods, Applications and Trends
- GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks