A Code-Agnostic Graph Neural Network Decoder from the Detection Error Model

arXiv:2610.01683 · quant-ph · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.

Kai: I'm Kai, and with me are Mira and Lev, guest researcher.

Mira: Today's paper: "A Code-Agnostic Graph Neural Network Decoder from the Detection Error Model".

Kai: The gist The POLYMECHANON,

Mira: First, who's behind it and why it matters.

Title and authors: Kai: So we're looking at this paper called "A Code-Agnostic Graph Neural Network Decoder from the Detection Error Model." It sounds like they've built something that doesn't need to know the specific type of quantum code you are working with.

Mira: That’s what it claims. The focus here is on using a graph neural network, or GNN, as a decoder whose only input is the detection error model of a quantum code under some noise conditions.

Kai: So, the title suggests this isn't tied to one particular stabilizer code structure; it’s designed to be general.

Lev: It’s interesting because traditionally you need a decoder tailored for your specific code, but this approach tries to decouple that part from the architecture itself >

Mira: Exactly. The authors are proposing a way where the network learns from how errors happen—the detection error model—rather than just learning how to decode one specific code structure directly >

The paper's summary: Kai: So, what’s the core idea behind this POLYMECHANON decoder, if we can call it that?

Mira: They represent the detection error model as a tripartite graph. It builds vertices out of detectors, error mechanisms, and logical observables >

Kai: A tripartite graph sounds complicated. How does that actually translate into decoding something?

Lev: The idea is to explicitly show the belief-propagation structure of the problem on this graph >

Mira: They use this structure so that decoding involves figuring out the probabilities for those error mechanisms and logical observables based only on the syndrome bits from the detectors and some prior probabilities they have about those mechanisms >

The paper's improvements: Kai: The paper lists several ways they’ve improved upon existing methods. What are the practical advantages here?

Mira: One major improvement is showing that their approach can decode unseen codes, even if their size is similar to what they trained on >

Lev: That suggests a degree of generalizability across different code families, which is always hard to achieve in these kinds of models >

Kai: They also talk about making it suitable for real-time decoding. How does that happen practically?

Mira: They use a fixed computational graph structure so that the cost doesn't increase as the noise strength goes up, which makes it good for neutral-atom processors >

Conclusion: Kai: So, putting it all together, what’s the big picture here regarding this paper on "A Code-Agnostic Graph Neural Network Decoder from the Detection Error Model"?

Mira: It shows that by focusing on the error model structure rather than just a fixed code layout, you can build something that is adaptable across different stabilizer codes and noise models >

Lev: For real hardware, the cost being independent of the noise strength is huge because you don't have to redesign your entire decoder when the physical conditions change >

Kai: It also addresses a limitation we often run into—the need to know exactly what code you are using upfront. This POLYMECHANON decoder seems designed to handle that flexibility >

Mira: Yes, and they even suggest using confidence-based post-selection to lower logical error rates significantly while still keeping most of your experimental shots >

Lev: It’s a solid framework for tackling the complexity of different noise environments without having to rewrite the entire decoding algorithm every time a new code is introduced >

Federico Alberto Astolfi, *, Guido Pupillo

University of Strasbourg and CNRS · QPerfect SAS

quant-ph

Submitted: 2026-10-01

Updated: 2026-10-01

Comments: 19 pages, 8 figures, 6 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 84/100

The gist: The gist The POLYMECHANON, a graph neural network (GNN) decoder for quantum error correction whose only input is the detection error model (DEM) of a quantum code under a given noise model,

Key concepts

Detection Error Model (DEM)
The DEM describes how errors are detected in a quantum code when subjected to noise. The paper represents this model as a tripartite graph where nodes represent detectors, error mechanisms, and logical observables. This structure is the input for the decoder.
Tripartite Graph G
This is the core representation of the DEM, structured with three types of nodes: D-nodes (detectors), EM-nodes (error mechanisms), and L-nodes (logical observables). Edges connect these nodes to explicitly show how detectors relate to errors and errors relate to logical outcomes.
Belief Propagation Structure
This refers to a probabilistic decoding method, similar to belief propagation used in other contexts. The tripartite graph structure makes this structure explicit, meaning the decoder determines the probabilities of error mechanisms and logical states based on syndrome bits and prior probabilities.
Code-Agnostic Approach
The POLYMECHANON is designed to work for any stabilizer code without needing to change its internal structure or decoder design. It achieves this by basing its input entirely on the detection error model, making it universally applicable across different quantum codes.

Terminology

Summary

The gist The POLYMECHANON, a graph neural network (GNN) decoder for quantum error correction whose only input is the detection error model (DEM) of a quantum code under a given noise model, represents a code-agnostic approach that decodes any stabiliser code without redesigning the decoder on page 1.

How it works

The core innovation is representing the detection error model (DEM) of a quantum code under a given noise model as a tripartite graph built out of detectors, error mechanisms and logical observables on page 1. Every input feature is computed by using the quantum code as data rather than design choice on page 1. This structure allows the same architecture to decode in principle any stabiliser code, under any noise model that can be expressed as a detection error model, once trained on the DEM on page 1.

The tripartite graph G is explicitly represented with vertices V = D ∪ EM ∪ L and edges E = ED,EM ∪ EEM,L on page 7. The nodes are defined as:

**- D-nodes (detectors): one per detector, carrying the syndrome bit 0/1; EM-nodes (error mechanisms): one per fault mechanism; L-nodes (logical observables): one per logical qubit on page 7. The edges ED,EM and EEM,L encode the matrices H and L previously defined on page 7. This representation makes the belief-propagation structure [72] of the problem explicit on page 7. Decoding consists in determining the probabilities of EM-nodes and L-nodes in G, having as inputs solely the syndrome bits for the D-nodes and the prior probabilities in the EM-nodes and in the EM-nodes and the L-nodes on page 7. The probabilistic information depending only on the noise model and on the specific faults on page 7. The DEM treats error mechanisms as independent between each others, so that the prior probability over error configurations e = (e1,..., eNEM) ∈ F NEM 2 factorises as P(e) = N YEM j=1 p ej j (1 − pj) 1−ej, (3) where pj is the prior probability of the single j-th mechanism on page 7. Each mechanism is equivalently described by the logit of its prior probability, λj = logpj/(1 − pj), (4) the quantity carried by the corresponding EM-node, which we define later in this paragraph on page 7. Given a syndrome σ, a maximum-likelihood decoder returns the most probable of two scenarios, that the error e triggered a logical error or that it did not on page 7. The decoded outcome should be ˆl = arg maxl P(l σ), i.e. the logical operator ˆl that for a given syndrome σ is most likely to have happened on page 7. However, the sum runs over an exponentially large number of error configurations, making the exact maximization intractable in general on page 7. This unfeasible regime gives the motivation to consider approximate decoders such as belief propagation [71, 72] and machine-learned decoders as the one presented in this work on page 7. The DEM has a natural tripartite structure that we make explicit by representing it as a tripartite graph G = (V, E) — V and E being the vertices and the edges of G respectively on page 7. Specifically, V = D ∪ EM ∪ L and E = ED,EM ∪ EEM,L) (Fig. 3) as described in what follows on page 7. As for the error mechanisms, they are initially endowed with the logits of their prior probabilities on page 7. Finally, edges ED,EM and EEM,L encode respectively the matrices H and L previously defined on page 7. This representation makes the belief-propagation structure [72] of the problem explicit on page 7. Decoding consists in determining the probabilities of EM-nodes and L-nodes in G, having as inputs solely the syndrome bits for the D-nodes and the prior probabilities in the EM-nodes and in the EM-nodes and the L-nodes on page 7. The DEM has a natural tripartite structure that we make explicit by representing it as a tripartite graph G = (V, E) — V and E being the vertices and the edges of G respectively on page 7. Specifically, V = D ∪ EM ∪ L and E = ED,EM ∪ EEM,L) (Fig. 3) as described in what follows on page 7. As for the error mechanisms, they are initially endowed with the logits of their prior probabilities on page 7. Finally, edges ED,EM and EEM,L encode respectively the matrices H and L previously defined on page 7. This representation makes the belief-propagation structure [72] of the problem explicit on page 7. Decoding consists in determining the probabilities of EM-nodes and L-nodes in G, having as inputs solely the syndrome bits for the D-nodes and the prior probabilities in the EM-nodes and in the EM-nodes and the L-nodes on page 7. The DEM has a natural tripartite structure that we make explicit by representing it as a tripartite graph G = (V, E) — V and E being the vertices and the edges of G respectively on page 7. Specifically, V = D ∪ EM ∪ L and E = ED,EM ∪ EEM,L) (Fig. 3) as described in what follows on page 7. As for the error mechanisms, they are initially endowed with the logits of their prior probabilities on page 7. Finally, edges ED,EM and EEM,L encode respectively the matrices H and L previously defined on page 7. This representation makes the belief-propagation structure [72] of the problem explicit on page 7. Decoding consists in determining the probabilities of EM-nodes and L-nodes in G, having as inputs solely the syndrome bits for the D-nodes and the prior probabilities in the EM-nodes and in the EM-nodes and the L-nodes on page 7. The DEM has a natural tripartite structure that we make explicit by representing it as a tripartite graph G = (V, E) — V and E being the vertices and the edges of G respectively on page 7. Specifically, V = D ∪ EM ∪ L and E = ED,EM ∪ EEM,L) (Fig. 3) as described in what follows on page 7. As for the error mechanisms, they are initially endowed with the logits of their prior probabilities on page 7. Finally, edges ED,EM and EEM,L encode respectively the matrices H and L previously defined on page 7. This representation makes the belief-propagation structure [72] of the problem explicit on page 7. Decoding consists in determining the probabilities of EM-nodes and L-nodes in G, having as inputs solely the syndrome bits for the D-nodes and the prior probabilities in the EM-nodes and in the EM-nodes and the L-nodes on page 7.

Improvements for AI systems

  1. Improve decoder generalizability by training a single model across multiple code families, as demonstrated by a single model trained across codes of different families outperforms uncorrelated MWPM on the graph-like codes seen during training. This enables the system to decode unseen codes whose size is comparable to those seen in training, although accuracy drops as the target graph grows beyond the largest one seen during training.

  2. Enhance real-time decoding capability by leveraging a fixed computational graph structure where the cost of our decoder does not grow with the noise strength, making it suitable for real-time decoding on neutral-atom processors. This allows for a deterministic latency, which is crucial for keeping pace with the syndrome stream.

  3. Implement confidence-based post-selection to lower logical error rates by more than an order of magnitude on the surface code while keeping more than 90% of the shots, using a confidence metric defined as c = min k=1,..., NL max(ˆpk, 1 − pˆk). This strategy allows for rejecting high-weight error events, which are precisely those where the decoder can fail.

  4. Optimize decoding cost by utilizing the fact that the GNN's per-shot latency is flat in physical error rate, as its cost is set by the size of graph and of the neural architecture, not by the noise. This makes it cheaper per shot on every code of the family, by a factor between 2.5 and 5.1 at the lowest noise.

  5. Develop a framework for adapting to diverse error models beyond simple detection error models, as the decoder only requires a DEM, which is all that an optimal decoder should see in order to operate its correction. This allows the system to address atom loss, where the DEM changes from shot to shot, and other noise types without requiring redesigning the architecture.

Abstract

We present POLYMECHANON, a graph neural network (GNN) decoder for quantum error correction whose only input is the detection error model (DEM) of a quantum code under a given noise model. We represent the DEM as a tripartite graph of detectors, error mechanisms and logical observables, where every input feature is computed by using the quantum code as data rather than design choice. In this way, the same architecture decodes in principle any stabiliser code, under any noise model that can be expressed as a detection error model, once trained on the DEM. We test this approach along four directions. Firstly, on the rotated surface code the decoder outperforms correlated MWPM under both phenomenological and circuit-level noise, with up to 25% fewer logical failures. Secondly, on a family of high-rate qLDPC codes it matches BP+OSD on the smaller codes and surpasses it on the larger ones, with up to 16% fewer logical failures on [![130,4,6]!]. Thirdly, its decoding time does not depend on the physical error rate, whereas that of BP+OSD grows with it: on [![130,4,6]!] its effective cost per shot is of the order of a few ms on a single GPU against a few tens of milliseconds for BP+OSD on a single CPU core, and its fixed computational graph makes it a candidate for real-time decoding on neutral-atom processors. Finally, a single model trained across codes of different families outperforms uncorrelated MWPM on the graph-like codes seen during training and stays within 10% of BP+OSD on the others, while it tends to fail to generalize to unseen codes, particularly larger ones. Its probabilistic output enables confidence-based post-selection, lowering the logical error rate by more than an order of magnitude on the surface code while keeping more than 90% of the shots. Since only the DEM changes, new codes and noise models can be decoded without redesigning the decoder.

Sources

Related papers