Code-space recovery for sample-based quantum diagonalization beyond native symmetry constraints
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.
Kai: I'm Kai, and with me are Mira and Lev, guest researcher.
Mira: Today's paper: "Code-space recovery for sample-based quantum diagonalization beyond native symmetry constraints".
Kai: Sample-based quantum diagonalization (SQD) diagonalizes Hamiltonians by sampling configurations and using native constraints, such as particle-number symmetry, to recover noisy samples.
Mira: First, who's behind it and why it matters.
Title and authors: Kai: So, we're talking about the paper titled "Code-space recovery for sample-based quantum diagonalization beyond native symmetry constraints," and the authors are Byeongyong Park, Sanha Kang, Doyeol Ahn, and Keunhong Jeong. The title itself suggests they're tackling a problem where the usual shortcuts don't apply.
Mira: I agree; when you see "beyond native symmetry constraints," it immediately makes me think about Hamiltonians that don't have those neat particle-number symmetries we usually rely on in quantum chemistry. It points toward a much broader applicability for these kinds of sampling methods.
Lev: From an error correction standpoint, that’s huge because it means the recovery mechanism isn't tied to a specific physical property of the Hamiltonian; it could theoretically work for any general Hermitian operator, which is what we really need.
Kai: Exactly! I was looking at how they set up this new framework, and they use a dual-rail representation to engineer structure right into the sampling process itself rather than hoping the target problem has some inherent symmetry.
Mira: That's a clever move; instead of assuming the physics gives us a constraint, they are designing an encoding where violations of that design become visible in our measured data. It shifts the burden from knowing the problem to engineering its structure in the circuit.
Lev: If we think about running this on real hardware, it means we have to build these specific encoding circuits into our sampler, which adds complexity, but if it works, it gives us a way to recover useful information even when those native constraints are missing.
Kai: Right. So they're not just applying existing techniques; they’re fundamentally changing how we extract the necessary structure from noisy quantum samples by injecting an engineered constraint.
Mira: It really makes me think about how much of our current methodology is built around these inherent symmetries, and this paper offers a pathway to use encoding as a general tool for recovery when those symmetries are absent.
Lev: That opens up possibilities for applying SQD ideas to classes of problems that we currently can't even approach with standard symmetry-based recovery procedures.
Kai: It sounds like they’re pointing us toward making the sampling and recovery process more robust against the lack of natural problem constraints, which is a key challenge in scaling up these kinds of experiments.
The paper's summary: Kai: So, to summarize what they’re doing with "Code-space recovery for sample-based quantum diagonalization beyond native symmetry constraints," they are taking the existing SQD workflow and adapting it for a much wider class of eigenvalue problems where the usual particle-number symmetry constraint isn't guaranteed.
Mira: They achieve this by introducing a dual-rail representation, mapping each logical qubit to a physical pair, and then using this encoding to create constraints at the sample level that are directly observable in the measured bitstrings.
Lev: The summary emphasizes that they use this encoding not just for storing states, but specifically to expose violations of the intended code space within the SQD samples, which is a novel way to get recovery information.
Kai: And then they detail a self-consistent recovery procedure where they cluster these samples into NC clusters using a weighted clustering method like the Bernoulli mixture model, and then iteratively repair invalid pairs.
Mira: That iterative repair process involves sampling which rail to flip based on a modified ReLU scoring rule, phi rho, delta(u), and reassigning initial bitstrings to the nearest reference vector from the previous iteration using Manhattan distance.
Lev: It sounds like a very intricate dance between the clustering, the mapping to reference vectors, and that iterative repair mechanism, which is precisely where I worry about implementation complexity on real quantum hardware.
Kai: The authors claim that this workflow yields lower projected Ritz energies compared to standard unencoded sample-support baselines, even when looking at smaller projected-subspace dimensions.
Mira: That result strongly suggests that even with the added resource overhead from the dual-rail encoding, the quality of the recovered subspace can actually be better than what you get from direct unencoded sampling.
Lev: If they can maintain that quality improvement despite doubling the qubit count in the register and adding circuit overhead, then it really validates this approach as a potentially useful method for extracting spectral information under difficult conditions.
Kai: It seems like they’ve shown that engineered constraints can provide recovery information even when the target problem itself doesn't have an obvious native constraint to exploit.
The paper's improvements: Mira: Looking at the specific improvements in "Code-space recovery for sample-based quantum diagonalization beyond native symmetry constraints," the main advancement is their shift from relying on inherent problem constraints to engineering recoverable structure through encoding.
Kai: They introduce a dual-rail representation where logical qubits are mapped to physical pairs, which assigns a constraint that 'Hamming weight one' means one and ten are valid encoded outcomes, while zero and eleven signal code-space violations.
Lev: That specific encoding is key because it makes the violation immediately visible in the samples, providing a direct signal for recovery when the Hamiltonian lacks its own symmetry to exploit.
Mira: Furthermore, they replace each logical operation in the sampling circuit with an "encoded counterpart A' satisfying A'V = VA," which ensures that this constraint is preserved in the noiseless limit while noise-induced violations supply the necessary information for recovery.
Kai: The self-consistent recovery procedure itself is a major improvement, involving initial clustering into NC clusters using a Bernoulli mixture model, followed by stochastic repair maps and iterative reassignment to nearest reference vectors.
Lev: From an error correction view, that iterative refinement step where you reassign bitstrings based on Manhattan distance to the previous iteration's reference vector sounds like a sophisticated way to handle the noise and bring the recovered state closer to a valid logical basis.
Mira: The paper’s results show that even fully code-valid bitstrings only account for about zero point zero two three six percent–four point seven seven zero zero percent, which demonstrates that this recovery mechanism is effectively utilizing imperfect encoded samples to build useful logical bases, not just relying on perfect ones.
Kai: So the core improvement is using encoding as a way to expose structure violations in noisy hardware samples as a direct source of recovery information, even when the target problem doesn't have its own constraints.
Conclusion: Mira: So, wrapping up this discussion on "Code-space recovery for sample-based quantum diagonalization beyond native symmetry constraints," the authors conclude that their engineered pair constraint provides a way to recover structure without needing native symmetries.
Kai: They showed that logical subspaces recovered from these encoded samples yield lower projected Ritz energies than the corresponding unencoded sample-support baselines, even at smaller projected-subspace dimensions.
Lev: This means we can get better quality projections for the ground state energy estimates in general Hermitian operators, which is a significant step forward if you're working with Hamiltonians that don't fit standard symmetry classes.
Mira: The implication is that engineered codespace constraints can indeed offer recovery information that boosts projected-subspace quality, despite the resource overhead introduced by doubling the qubit count and adding circuit complexity.
Kai: It really shows this approach can complement native constraints rather than compete with them, suggesting a more versatile toolbox for spectral analysis in quantum simulations.
Lev: If we look at what this means for real hardware execution, it suggests that while you’ll need those extra qubits and overhead, the gain in projection quality is significant enough to make the effort worthwhile.
Byeongyong Park, Sanha Kang, Doyeol (David) Ahn, Keunhong Jeong
Center for Quantum Information Processing, University of Seoul · Singularity Quantum Inc. · Department of Chemistry, Sogang University
quant-ph
Submitted: 2026-07-11
Updated: 2026-09-29
Comments: 18 pages, 4 figures; supplementary information provided as an ancillary file
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: Sample-based quantum diagonalization (SQD) diagonalizes Hamiltonians by sampling configurations and using native constraints, such as particle-number symmetry, to recover noisy samples.
Key concepts
- Sample-based quantum diagonalization (SQD)
- SQD diagonalizes Hamiltonians by sampling configurations and using native constraints, like particle-number symmetry, to recover noisy samples. This method is being adapted here to work when those native constraints are missing.
- Dual-rail representation
- This encoding maps each logical qubit to a physical pair of qubits. This specific encoding is used to create observable constraints at the sample level, where violations signal code-space issues even if the Hamiltonian lacks inherent symmetry.
- Self-consistent recovery procedure
- This involves clustering samples into NC clusters using a Bernoulli mixture model, followed by stochastic repair maps and iterative reassignment to reference vectors based on Manhattan distance. This process iteratively repairs invalid pairs to build a valid logical basis.
Terminology
Summary
Sample-based quantum diagonalization (SQD) diagonalizes Hamiltonians by sampling configurations and using native constraints, such as particle-number symmetry, to recover noisy samples. This paper introduces code-space recovery,
a novel strategy that engineers recoverable structure through encoding rather than relying on inherent problem constraints. This approach extends the applicability of SQD to a broader class of eigenvalue problems where native constraints are absent, suggesting that engineered structure can improve projected subspace quality even with increased circuit overhead.
Motivation and Problem Statement
Large-scale eigenvalue problems are computationally demanding, and while SQD is a promising method for extracting spectral information from quantum samples, its performance often relies on native constraints like particle-number symmetry. The paper notes that For a broad class of eigenvalue problems, however, no analogous constraint is guaranteed,
which limits the applicability of existing SQD recovery procedures. The core limitation addressed is the dependence on native constraints for sample recovery; this dependence restricts the direct extension of the SQD workflow to general eigenvalue problems lacking such a constraint.
Code-Space Encoding and Constraint Engineering
The authors introduce a dual-rail representation to engineer a new constraint at the sample level.
-
Each logical qubit is mapped to a physical pair:
0⟩ → 01⟩ and 1⟩ → 10〉.
-
This encoding assigns each logical qubit to a physical pair constrained to have
Hamming weight one, so that 01 and 10 represent valid encoded outcomes whereas 00 and 11 are identifiable code-space violations.
-
The encoding is used not for protecting a stored state but
to expose violations of the intended code space in SQD samples.
-
Each logical operation in the sampling circuit is replaced by an
encoded counterpart A′ satisfying A′V = V A,
ensuring that the constraint is preserved in the noiseless limit and noise-induced violations provide recovery information.
Self-Consistent Recovery Procedure
The paper details a self-consistent recovery algorithm applied to these encoded samples. The procedure involves:
-
Initial clustering of weighted samples into
NC clusters in the physical 2n-bit encoded space using a weighted clustering method suitable for binary data, such as a Bernoulli mixture model (BMM).
-
Each cluster is represented by a
pair-normalized reference vector n(t)κ,
where the initial reference vectors are derived from theweighted rail occupancies within cluster κ.
-
A
stochastic recovery map R(xe; n)
repairs invalid pairs (00 or 11) by sampling which rail to flip based on a modified-ReLU scoring rule, defined as ϕρ,δ(u). -
Iterative recovery involves reassigning initial bitstrings to the nearest reference vector from the preceding iteration in Manhattan distance:
g(t)s = arg min κ=1,…,NC xe(0)s − n(t−1)κ.
-
The process continues until a stopping criterion is met, such as when the
global-best Ritz value has not strictly improved for Nconv consecutive iterations,
resulting in the final projected subspace defined by B⋆.
Benchmarking and Performance
The strategy was benchmarked on Ising-type spin Hamiltonians (1D and 2D TFIMs/MFIMs) up to 36 spin sites, which lack the U(1) symmetry used in conventional SQD recovery.
Across all benchmarks, logical subspaces recovered from encoded samples yield lower projected Ritz energies than the corresponding unencoded sample-support baselines.
These results indicate that engineered codespace constraints can provide recovery information that improves projected-subspace quality, despite the resource overhead introduced by encoding.
The comparison showed that the encoded workflow consistently yielded lower Ritz energy estimates than direct unencoded sampling baselines across smaller projected basis sizes. Furthermore, fully code-valid bitstrings accounted for only 0.0236%–4.7700%,
demonstrating that the recovery mechanism effectively utilizes imperfect encoded samples to construct useful logical bases.
Resource Trade-off and Conclusion
The paper acknowledges a clear resource trade-off: the dual-rail encoding doubles the number of qubits in the sampling register, and implementing the sampling circuits in encoded form also incurs additional circuit overhead.
Despite this overhead, the results suggest that engineered pair constraint can provide recovery information that offsets this added circuit overhead at the level of projected-subspace quality.
The framework demonstrates that engineered constraints can extend SQD beyond native constraints by providing a mechanism to recover structure even when the target problem lacks an analogous native constraint. While extension requires problem-specific design for encoded operator realizations, the core idea is to use encoding to expose violations of the intended code space in noisy hardware samples
as a source of recovery information. The paper also notes that engineered constraints can complement native constraints rather than compete with them.
Improvements for AI systems
Here are the specific improvements to AI systems that can be derived from this research, along with what these improved systems can achieve:
-
The core improvement is a new class of quantum algorithm called
code-space recovery for sample-based quantum diagonalization
(SQD). This moves beyond relying on problem-specific native symmetries (like particle number conservation) to engineer constraints via encoding. -
The improved AI system can perform highly accurate, low-energy eigenvalue estimation for large, general Hermitian operators in complex physical systems (e.g., condensed matter Hamiltonians).
Specifically, the improved AI system can:
-
Perform spectral analysis on quantum many-body Hamiltonians that lack obvious native symmetries (like the transverse-field Ising models studied), which are common in materials science and chemistry simulations.
-
Extract the lowest Ritz pair (ground state energy estimate) with high precision, even when the available quantum samples are noisy or do not perfectly align with a native symmetry sector.
-
Achieve lower projected Ritz energies than standard sample-support diagonalization methods, suggesting improved subspace quality in variational quantum algorithms.
-
Function as a robust sampler and recovery mechanism: It can take noisy physical measurements (bitstrings) and use an engineered dual-rail encoding constraint to detect errors, repair them self-consistently, and reconstruct the logical basis states for accurate projection.
-
Operate effectively in hybrid quantum-classical workflows: The system is designed to interface with classical optimization loops by iteratively refining the logical subspace based on recovered samples, allowing for a more efficient use of limited quantum hardware resources compared to methods that discard noisy data.
-
Scale its applicability: By engineering constraints through encoding, the framework can be applied to any eigenvalue problem where the target state is sparse or compressible in a measurement basis accessible to the quantum sampler, broadening its scope beyond traditional quantum chemistry problems.
Sources
- Quantum measurements and the Abelian Stabilizer Problem
- Cluster-Adaptive Sample-Based Quantum Diagonalization for Strongly Correlated Systems
- Quantum-Centric Algorithm for Sample-Based Krylov Diagonalization
- Quantum chemistry with provable convergence via randomized sample-based Krylov quantum diagonalization
- Physics-Informed Generative Machine Learning for Accelerated Quantum-centric Supercomputing
- Crossing the 12,000-atom barrier with heterogeneous quantum-classical supercomputing: quantum chemistry of protein-ligand complexes
- Quantum computing with Qiskit
Related papers
- Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more
- Encrypted clones can leak: Classification of informative subsets in Quantum Encrypted Cloning
- Polynomial-time classical and quantum simulation of quantum impurity models
- Theory of quantum-enhanced interferometry with general Markovian light sources
- A convergent hierarchy of spectral gap certificates for qubit Hamiltonians
- Universal Bound and Phase Transition in Many-Body Fermionic Non-Gaussianity