Machine learning Majorana topology using unsupervised and supervised learning

arXiv:2512.13825 · cond-mat.dis-nn, cond-mat.mes-hall, cs.LG · Submitted 2025-12-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.

Kai: Today's paper: "Machine learning Majorana topology using unsupervised and supervised learning".

Mira: Unsupervised and supervised machine learning techniques are employed to identify Majorana topology in unlabeled data from realistic short disordered nanowires,

Kai: First, who's behind it and why it matters.

Title and authors: Kai: So, we're diving into this paper now titled "Machine learning Majorana topology using unsupervised and supervised learning." Mira, what are your initial thoughts on that title? Does it sound like something the field is actually ready for?

Mira: I think the title immediately signals a methodological shift; it moves away from needing pre-labeled data, which is a big deal when we're dealing with experimental physics where topology isn't always obvious. It suggests they are trying to use the machine learning structure itself to uncover what the underlying physical state actually is.

Lev: From my side, I’m interested in how much of this relies on just pattern recognition versus needing a strong theoretical underpinning for those patterns to be physically meaningful. If the AI finds a cluster, does that automatically mean we have confirmed Majorana physics?

Kai: Exactly, Lev. It sounds like they are trying to automate the identification process itself, bypassing some of the traditional hurdles in experimental characterization where you have to manually decide what labels go where. It’s about letting the data speak for itself initially.

Mira: And that's precisely what pages one and two explain: they use unsupervised learning with an autoencoder on Majorana energy splitting measurements to see if it can naturally sort the data into topological or trivial categories without any prior knowledge of which is which.

Lev: The fact that they found two clusters when the wire length is large and disorder is small tells me a lot about where the system is most stable in terms of topology, but what about those other regimes where things get messy?

Kai: That's what page one points to; when the wire length isn't necessarily large and disorder isn't necessarily small, that’s where they found three distinct patterns emerging from the unsupervised learning. That intermediate regime is a key area for us to focus on.

Mira: Right, and that intermediate regime is where things get tricky because the topology might just be ill-defined or dominated by Andreev Bound State contamination, which is a major physical concern in these systems.

Lev: So they're essentially mapping out the landscape of what’s physically possible based on how well their AI can discern those hidden patterns. That feels like it could help us prioritize which experimental parameters are most promising to probe further.

Kai: It gives us a map, Lev, not just a single point. We get to see where the system sits in terms of its topological stability based on the measurable inputs they are feeding into the model.

The paper's summary: Mira: So, let’s look at what they actually did in "Machine learning Majorana topology using unsupervised and supervised learning." Essentially, they combined an autoencoder for dimensionality reduction followed by k-means clustering to see if unlabeled data from short disordered nanowires could reveal hidden topological patterns.

Kai: That sounds like a very hands-on approach for the experimentalist; they took raw measurements of the energy splitting and topological visibility and fed them into this network to try and sort them out. It’s about using these ML techniques to find structure where traditional analysis might miss it because you don't have those topology labels yet.

Lev: I see the setup, but what is the actual input data they are working with? Are we talking about just one parameter, or a set of measurable quantities that define the system state? That determines how much information the autoencoder has to work with.

Mira: They use two main inputs: Majorana energy splitting, E s(V z), and topological visibility, TV(V z). The autoencoder is designed to compress this data into a fifteen-dimensional latent space using three one-dimensional convolutional layers before applying k-means clustering <ref:2512.13825#pg2,into a 15-dimensional latent>.

Kai: Fifteen dimensions sounds manageable for a compressed representation of what might be complex physical behavior in these nanowires. I wonder if those convolutional layers are capturing the necessary spatial or energy dependencies correctly without losing crucial information about the disorder effects they mentioned earlier.

Mira: The paper indicates that this unsupervised approach is effective, finding two clusters when wire length is large and disorder is small, which points to explicit topological or trivial phases, and three clusters when wire length and disorder aren't perfectly constrained, revealing an intermediate regime.

Lev: That finding about the three patterns being found in the unlabeled data seems powerful because it directly addresses the problem of identifying topology without prior labeling. If they can reliably find that intermediate phase, it opens up a new kind of experimental signature we might not have thought to look for.

Kai: It really does; this suggests that instead of just looking for a sharp transition, we should be looking for those complex crossover behaviors where the physics is transitioning between known states.

Mira: Exactly. The paper summarizes the unsupervised part by showing that the UML can discern patterns in realistic samples, even though it requires significant computing resources and doesn't guarantee finding more than three distinct patterns without prompting.

Lev: I think those computational costs are a real barrier to entry for anyone trying to implement this on standard experimental setups right away, though the theoretical payoff could be significant if we get the results we expect.

The paper's improvements: Kai: Now, let's talk about what they suggest as improvements in "Machine learning Majorana topology using unsupervised and supervised learning." They introduce a dual approach: first, the unsupervised clustering to find patterns, and then using a supervised learning component to actually label those clusters.

Mira: That’s the key improvement; they use the results of the UML to inform a supervised model designed to predict things we can't easily measure directly, like topological visibility TV(V z) and disorder strength sigma. The supervised network uses global context vectors as a conditional input to its decoder via a FiLM layer.

Lev: That’s interesting because it turns the problem around; instead of just classifying based on the raw data clusters, they are using one model to predict the missing physical parameters, which then helps label the topology. It makes sense that they needed that topological visibility input for the clustering to be meaningful at all.

Kai: The paper highlights how crucial including TV(V z) is; without it, as page one points out, the clustering becomes trivial and just based on wire length alone, so adding it provides the necessary physical context for the AI to actually learn topology.

Mira: And in terms of accuracy, they reported really good results: an R2 of zero point eight six one for predicting TV(V z) and zero point nine two two for predicting sigma(V z) when testing across the full disorder range, and it got even better when restricting that disorder to the range sigma in zero one meV.

Lev: A prediction of R squared values like those suggests that this model is actually quite robust at extracting physically relevant information from the inputs and predicting parameters that are theoretically accessible in experiments <ref:2512.13825#pg0>. If we can use it to predict sigma, it gives us a better handle on the material quality, which is always an important factor.

Kai: So, what does this mean practically? It means we can take experimental measurements of E s(V z) and L and use this AI to estimate the topological visibility and disorder strength, giving us a much richer dataset for future analysis.

Conclusion: Mira: To wrap up "Machine learning Majorana topology using unsupervised and supervised learning," they show that by combining UML to find underlying patterns with SML to provide topological labels, they can comprehensively identify distinct phases in the physical parameter space. They found a phase boundary separating the weak-disorder topological phase from the strong-disorder trivial phase from their two-cluster classification.

Kai: And for their three-cluster classification, they identified an additional cluster corresponding to that intermediate crossover disorder regime where topology is ill-defined, which is very telling about where we need to focus our attention experimentally. This tool helps us classify Majorana experimental data.

Lev: From a practical standpoint, this method provides a framework for ruling out the parameter regimes where topology might not manifest in realistic nanowires, which is a necessary step before we can even attempt to build more complex devices like those needed for error correction.

Mira: It’s essentially giving us a sophisticated way to classify the data and point us toward the most physically interesting regions of parameter space based on what the machine learning has learned from the unlabeled samples. This paper, "Machine learning Majorana topology using unsupervised and supervised learning," offers a method for automated pattern identification in these systems.

Kai: It’s been fascinating watching how they use this combined approach to move from raw data to a classification of physical phases, and I think this technique has real potential for classifying the experimental results we're getting from these nanowires.

Lev: I just think if we can get the hardware to actually run on these predictions reliably, then this tool moves us much closer to having a viable path forward for realizing topological quantum computing.

Mira: Indeed, the findings suggest that this method can be useful in identifying topology in Majorana nanowires by providing a way to classify experimental data based on the patterns discovered by their learning framework.

Condensed Matter Theory Center and Joint Quantum Institute, Department of Physics, University of Maryland · Department of Physics and Astronomy, Center for Materials Theory, Rutgers University

cond-mat.dis-nn, cond-mat.mes-hall, cs.LG

Submitted: 2025-12-15

Updated: 2026-10-06

Comments: 14 pages, 12 figures

Journal ref: Phys. Rev. B 114, 165418 (2026)

DOI: 10.1103/ywsw-29xc

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 79/100

The gist: Unsupervised and supervised machine learning techniques are employed to identify Majorana topology in unlabeled data from realistic short disordered nanowires, providing a potential breakthrough for

Key concepts

Majorana Zero Modes (MZMs)
These are localized excitations in superconducting systems that are crucial for topological quantum computing. Identifying them is the main goal of the research because disorder and short wire lengths often suppress their existence, making detection challenging.
Unsupervised Learning (UML)
This initial step uses an autoencoder to compress input data into a lower-dimensional latent space and then applies k-means clustering. This process discovers hidden patterns in the raw data without needing pre-labeled topological information, revealing underlying physical structures.
Supervised Learning (SML)
After UML identifies patterns, SML is used to predict unknown quantities like topological visibility and disorder strength based on measurable inputs (splitting energy and wire length). This model provides the necessary 'topological labels' to classify the discovered phases.
Topological Visibility (TV(Vz))
This is a quantity that helps determine if a system exhibits topological properties. The research found that including this factor in the unsupervised clustering was crucial; without it, the initial clustering results were trivial and not meaningful.

Terminology

Summary

Unsupervised and supervised machine learning techniques are employed to identify Majorana topology in unlabeled data from realistic short disordered nanowires, providing a potential breakthrough for confirming topological order in Majorana nanowires.

The Problem and Motivation

Topological quantum computing relies on localized Majorana zero modes (MZMs), but identifying them in realistic systems is difficult due to disorder suppressing the superconducting gap and short wire lengths potentially suppressing anyonic nature. While supervised machine learning requires labeled training data, experimental data lacks explicit topology labels, creating a challenge for traditional methods. The motivation is to develop an unsupervised or self-supervised ML algorithm that learns to label data by discerning hidden patterns without prior topological labels, and then use supervised learning to provide the necessary topological labels.

The Unsupervised Learning Framework

The unsupervised learning component uses an autoencoder neural network for dimensionality reduction followed by k-means clustering in the latent space. The input data consists of two ingredients: the Majorana energy splitting, denoted as Es(Vz), and the topological visibility, TV(Vz). The autoencoder architecture employs 1D convolutional layers to compress the input data of Es(Vz) and TV(Vz) into a 15-dimensional latent space. This latent space is then subjected to k-means clustering. The analysis reveals that "the UML-discerned hidden pattern only when the wire length is not necessarily large and disorder is not necessarily small, and the UML then finds three distinct patterns in the unlabeled data: topological and trivial as well as an intermediate regime in between where the topology is ill-defined."

The Supervised Learning Component

The success of the unsupervised learning crucially depends on including TV(Vz) in the input data; otherwise, the clustering is trivially based on the system size L alone. Therefore, a supervised learning framework is designed to predict inaccessible quantities. This model takes experimentally measurable quantities as inputs: the Majorana energy splitting Es(Vz) and the wire length L, and predicts theoretically accessible quantities: the topological visibility TV(Vz) and the disorder strength σ. The supervised network utilizes a novel architecture incorporating global context vectors (GCV) into an encoder-decoder setup, where the GCV is used as a conditional input to the decoder via a FiLM layer. This model achieved generally good prediction accuracy with an R2 of 0.861 for predicting TV(Vz) and 0.922 for predicting σ(Vz) in the full disorder range tested, with significantly improved accuracy when restricting the disorder range to σ ∈ [0, 1] meV.

The Combined Analysis and Findings

The combination of UML and SML allows for a comprehensive identification of topological phases. For the two-cluster classification derived from UML, a phase boundary separating the weak-disorder topological phase (Cluster 1 in green) and the strong-disorder trivial phase (Cluster 0 in blue) was found. Furthermore, for the three-cluster classification, an additional cluster (Cluster 2 in orange) that corresponds to an intermediate crossover disorder regime was identified. The work concludes that this method can be useful in the classification of Majorana experimental data, particularly in the context of ruling out the parameter regimes where topology may not manifest in realistic nanowires.

Model and Data Details

The physical system is modeled using a semiconductor-superconductor single-band Hamiltonian (Equation 1), incorporating terms for HSM, HZ, HSC, and Hdis. The dataset consists of Es(Vz) and TV(Vz) as functions of the Zeeman field Vz. The autoencoder architecture uses three 1D convolutional layers with [32, 64, 128] filters, compressing data into a bottleneck latent space of size 15. The supervised network input vector length is 512. Key metrics used to evaluate clustering quality include the Silhouette score, where a higher score indicates better-defined clusters. The study demonstrates that the inclusion of TV(Vz) in the autoencoder input is crucial for a meaningful clustering result, as omitting it results in trivial clustering. The supervised model's performance is compared against ground truth values along the diagonal line, showing good correlation.

Conclusion

The paper successfully introduces a comprehensive method to do unsupervised ML for MZM topology identification by first using UML to identify underlying patterns and then utilizing SML as an input to label the topology. This combined approach effectively identifies distinct phases in the physical parameter space, providing a tool for classifying Majorana experimental data. The findings suggest that this technique can be useful in identifying topology in Majorana nanowires.


The gist

Unsupervised and supervised machine learning techniques are employed to identify Majorana topology in unlabeled data from realistic short disordered nanowires, providing a potential breakthrough for confirming topological order in Majorana nanowires.

How it works

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed this paper's methodology—combining Unsupervised Learning (UML) for pattern discovery in unlabeled data with Supervised Learning (SML) for topological labeling via a predictive model—and the underlying physics of Majorana zero modes (MZMs).

The core improvement lies in creating a robust, automated pipeline that can rapidly classify experimental nanowire data as topological or trivial and map the crossover behavior in parameter space.

Here are the specific improvements to an AI system derived from this research:


  1. The improved system will be a dual-stage architecture combining an Autoencoder (for dimensionality reduction/pattern extraction) followed by a Supervised Neural Network (for topological classification).

  2. This system can perform the following specific tasks:

Ease of Use and Automation:

The system can ingest raw experimental data from semiconductor-superconductor nanowire experiments, which are inherently unlabeled. Instead of requiring manual theoretical simulation for labeling (as in traditional supervised ML), the system performs automated pattern recognition on the unlabeled data.

Topological Phase Classification:

The primary function is to classify a given physical realization of a nanowire (defined by its measurable properties like Majorana splitting energy, Es(Vz), and wire length, L) into one of three distinct topological regimes:

a. Explicit Topological Phase (identified by the Cluster 1 in the two-cluster model).

b. Explicit Trivial Phase (identified by the Cluster 0 in the two-cluster model).

c. Intermediate/Crossover Regime (identified by the Cluster 2 in the three-cluster model), indicating a region where topology is ill-defined or dominated by Andreev Bound States (ABS) contamination.

Mapping Topological Phase Diagrams:

The system can generate unsupervised phase diagrams in the physical parameter space (disorder strength, σ, vs. wire length, L). This allows researchers to visualize:

a. The precise boundary between the weak-disorder topological phase and the strong-disorder trivial phase.

b. How this crossover shifts as a function of wire length (e.g., demonstrating that longer wires are more robust against disorder).

Predictive Topological Visibility Estimation:

A crucial capability is using the Supervised Learning component to predict theoretically inaccessible quantities directly from experimental measurements:

a. The system can predict the topological visibility, TV(Vz), which is essential for determining if MZMs are actual zero-energy anyonic states or just low-energy trivial fermionic bound states.

b. It can also estimate the disorder strength, σ, using only measurable data (Es(Vz) and L), providing a proxy for the underlying material conditions.

Robustness Assessment:

The system can assess the reliability of its own classifications by evaluating clustering quality (using metrics like Silhouette Score) and prediction accuracy on withheld test sets, ensuring that the identified patterns are physically meaningful rather than artifacts of the learning process.

In summary, this improved AI system transforms experimental data analysis from a manual guess-and-check process into an automated scientific discovery tool capable of identifying topological states and defining their phase boundaries with high fidelity.

Sources

Related papers