Quantum convolutional neural networks for jet images classification
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.
Kai: Today's paper: "Quantum convolutional neural networks for jet images classification".
Mira: Quantum convolutional neural networks (QCNNs) are investigated as a quantum machine learning approach to classify high-energy physics jet images, specifically for top-quark tagging,
Kai: First, who's behind it and why it matters.
Title and authors: Kai: We've been looking at the title and authors of "Quantum convolutional neural networks for jet images classification," focusing on how the team is approaching this problem. The paper addresses the specific challenge of classifying complex jet images that arise from top quark decays using quantum machine learning techniques.
Mira: I think the authors have set up a very clear framework by comparing their QCNN against a classical CNN using a simulator, which establishes a solid benchmark for what we are trying to compare against in this field. This comparative approach is key because it lets us isolate the performance gains of the quantum aspect itself.
Lev: From my perspective as someone focused on error correction, I'm paying attention to how much of this comparison relies on that "classical noiseless simulator"; if we move to real hardware, that simulation might not perfectly capture the noise profile we encounter.
Kai: That’s a fair point, Lev. But for now, the paper establishes a clear methodology for testing different quantum components against a known classical baseline before worrying about hardware noise.
Mira: The authors then explore several encoding schemes and convolutional unit types, like SO(four) and SU(four), which shows they are systematically testing the fundamental building blocks of this QCNN architecture.
Lev: When you test different unitary circuits, how do you account for the inherent difficulty in building those complex quantum gates reliably on physical hardware?
Kai: They test these circuits to see which ones yield better results, but that's usually done in simulation first before attempting to implement them on actual cooling systems.
Mira: The paper also investigates the loss functions used, including cross-entropy and MSE, which shows they are considering how different objectives influence the learned classification outcome.
Lev: And those loss functions matter because they dictate what kind of optimization landscape the AI will navigate; a poor choice can lead to a very difficult training process, even if we have perfect hardware.
Kai: So, in essence, they're mapping out the performance surface by systematically changing these variables to find the best combination for classifying those jet images.
Mira: It’s a comprehensive look at how varying these elements impacts the final accuracy on page zero of that work; it gives us a good picture of what kind of setup is most promising.
Lev: That systematic variation is useful, but it reminds me that physical implementation constraints, like coherence times and gate fidelity, are the ultimate deciding factors for real-world viability.
Kai: Exactly, Lev. But for now, we're focusing on the theoretical optimization before we get bogged down in the practical limitations of cooling systems.
The paper's summary: Kai: Now that we've looked at the structure, let’s go back to what they actually found in the "Quantum convolutional neural networks for jet images classification" study. What is the main finding regarding this QCNN approach?
Mira: The main finding is that QCNN setups with proper configurations generally perform better than their classical CNN counterparts, especially when those convolution blocks are kept relatively small in terms of parameters.
Lev: That's encouraging, but I need to know how much of that performance edge is actually due to the quantum mechanics versus just the clever classical data encoding they are using.
Kai: The authors found that QCNN accuracy results were generally better across most choices for data encoding, batch size, and loss function when compared to their classical CNN counterparts.
Mira: They provided a concrete example on page zero of that work: for an SO(four) QCNN with thirty parameters and MSE loss along with HEE1 encoding, the accuracy was reported as ninety-nine point one three ± zero point one seven versus the CNN's ninety-five point four seven ± one point four two, which is a substantial gap in performance metrics for this task on top-quark tagging images.
Lev: A jump from about ninety-five percent to nearly a hundred percent accuracy is significant, but I wonder if that performance holds up when we introduce the inherent randomness of quantum states into the training loop.
Kai: They also looked at batch sizes and saw that for both SO(four) and SU(four), accuracy tends to decrease as the batch size increases, with a batch size of sixteen showing faster convergence compared to larger sizes.
Mira: And they didn't ignore the parameter count issue, showing how DEA can help reduce those trainable parameters, finding that for SU(four), removing "seventeen of them parameters were found to be redundant," leaving thirty-one parameters in the optimized circuit on page zero of that work.
Lev: So, if we can manage to keep the parameter count manageable while seeing these accuracy gains, then it becomes a more viable path toward building useful quantum models.
Kai: That really suggests that optimizing the structure is just as important as finding the right quantum gates; it’s about making sure we aren't training an unnecessarily massive model.
The paper's improvements: Mira: Focusing on the improvements, they didn't just stop at reporting results; they suggested a method called Dimensional Expressivity Analysis, or DEA, as a way to optimize the QCNN structure itself.
Lev: I know we talked about DEA before; can you elaborate on what that analysis actually entails in terms of its computational cost? Is it feasible to run this optimization on current quantum simulators?
Kai: The paper defines DEA as checking if the derivative of the Parameter Quantum Circuit map with respect to each parameter is linearly dependent on others, which is a way to ensure maximal expressivity with the fewest trainable parameters.
Mira) That process ensures that we're not wasting computational resources on redundant parts of the circuit, which directly addresses one of the core issues in quantum model design. [Lev: It sounds like a necessary step for making these models practical, but what about the physical constraints? Does this optimization change how we view things when mapping it to physical qubits?
Kai: It helps because if we can reduce parameters significantly, the resulting circuit is smaller, which makes mapping it onto hardware much more feasible.
Mira) That reduced size has direct implications for circuit depth and gate count, which are both crucial metrics when considering running on real quantum systems. [Lev: So, does this optimization lead to a model that is just theoretically better but practically useless if the required circuit depth becomes unmanageable?
Kai: The paper suggests that DEA can reduce the number of trainable parameters while maintaining good overall accuracy, which leads to smaller, more practical quantum models.
Conclusion: Kai: So we’ve covered a lot regarding "Quantum convolutional neural networks for jet images classification" and what these findings mean for the field. To summarize, this paper shows that QCNNs can outperform classical CNNs in certain conditions when properly structured and optimized with techniques like DEA.
Mira: The key takeaway seems to be that the combination of careful circuit selection, encoding choices, and parameter reduction via DEA offers a strong pathway for building more efficient quantum models for complex image classification tasks.
Lev: From an error correction standpoint, I think the focus on reducing parameters is what really matters because it simplifies the theoretical work needed to handle noise.
Kai: I agree; if we can build smaller, more parameter-efficient models that still maintain accuracy, that’s a practical step forward for quantum hardware experimentation.
Mira: Overall, this paper contributes by showing how quantum structure can be leveraged effectively in machine learning for high-energy physics data analysis.
Lev: It gives us a roadmap for moving from theoretical concepts to models that are more manageable when considering the noise floor of physical systems.
Kai: We've explored the specifics of "Quantum convolutional neural networks for jet images classification," and I think this sets a good direction for future research into making these quantum models more robust.
Deutsches Elektronen-Synchrotron DESY · Institut f¨ur Physik, Humboldt-Universit¨at zu Berlin
quant-ph, hep-ph
Submitted: 2024-08-16
Updated: 2026-10-06
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: Quantum convolutional neural networks (QCNNs) are investigated as a quantum machine learning approach to classify high-energy physics jet images, specifically for top-quark tagging, aiming to surpass
Key concepts
- QCNN
- A quantum machine learning model that uses quantum operations to perform convolution and pooling on data encoded in a quantum state. It aims to classify complex images, like jet images, using quantum computation methods.
- Jet Images Classification
- The task of classifying complex images generated by high-energy particle collisions (QCD jets) into specific types, such as top quark jets or background noise. This is relevant in physics research to identify rare particles.
- Dimensional Expressivity Analysis (DEA)
- A technique used to optimize the QCNN structure by checking if any parameters in the quantum circuit are redundant. If a parameter's effect can be explained by others, it can be removed, leading to a smaller, more efficient model with similar performance.
Terminology
Summary
Quantum convolutional neural networks (QCNNs) are investigated as a quantum machine learning approach to classify high-energy physics jet images, specifically for top-quark tagging, aiming to surpass classical convolutional neural networks (CNNs) in accuracy while addressing limitations like excessive trainable parameters.
How it works
The study focuses on classifying quantum chromodynamics (QCD) jet images for top quark jet tagging. This task is relevant in beyond the Standard Model physics, where highly energetic particles from top quark decays can form complex jet images that are challenging for classical CNNs to classify accurately due to the high complexity of the image features. The methodology involves comparing a Quantum Convolutional Neural Network (QCNN) against a classical CNN using a classical noiseless simulator
for fair comparison.
The QCNN architecture is constructed by leveraging quantum operations to accomplish convolution and pooling techniques, with data first encoded into a quantum state: The encoding Venc can be done in several ways.
The overall process involves encoding the classical input data as an entangled state, processing this state via parametrized unitaries composed of convolutional and pooling blocks,
and then measuring the output using a PauliZ measurement to determine the classification.
Model Components and Architectures
The researchers compared various setups for the QCNN by varying several key components:
-
The convolutional circuit (testing unitary two-qubit circuits like SO(4) and SU(4)).
-
The type of encoding, testing schemes such as
tensor product embedding (TPE),
hardware efficient embedding (HEE),
andClassically Hard Embedding (CHE).
-
The loss function, which included the cross-entropy loss, Mean-Square-Error (MSE), and Hinge loss.
-
Batch sizes for training.
For the convolutional circuits, they tested unitary two-qubit circuits:
- SO(4) circuit:
- SU(4) circuit:
The study also employed a technique called Dimensional Expressivity Analysis (DEA) to optimize the QCNN structure by identifying redundant parameters, which is defined as checking if the derivative of the Parameter Quantum Circuit (PQC) map with respect to each parameter is linearly dependent on others. This process ensures maximal expressivity with the fewest trainable parameters.
Dataset Preparation and Preprocessing
The analysis utilized publicly available JetNet library datasets, specifically the TopTagging dataset, which contains hadronic top jets as signal and QCD jets as background. The initial jet four-momentum vector is transformed into a two-dimensional histogram image. This transformation involves:
-
Re-scaling the jet four-momentum such that its mass is mB and applying a Lorentz boost (choosing γB = 10).
-
Using the Gram-Schmidt method to construct an orthonormal basis from the jet constituents, resulting in two-dimensional coordinates (Xi, Yi) for each constituent.
-
The final image coordinates are related to the four-momentum components through specific formulae, where the histogram is filled with the weight ωi = p0i EB.
To reduce dimensionality and increase efficiency for both models, Principal Component Analysis (PCA) was applied to reduce the 28x28 images down to a 2x2 representation, resulting in only four pixels as inputs for the QCNN and CNN models.
Results and Findings
The results indicate that QCNN with proper setups tend to perform better than their CNN counterparts, especially when the convolution block has a lower number of parameters.
The study found that QCNN accuracy results are generally better than classical CNN accuracy across most choices of data encoding, batch-size, and loss function. Specifically:
- For SO(4) QCNN with 30 parameters and MSE loss along with HEE1 encoding, the accuracy was 99.13 ± 0.17 compared to the CNN's 95.47 ± 1.42.
- The DEA circuit for SU(4) required removing 17 of them [parameters] were found to be redundant, leaving a total of 31 parameters in the optimized circuit,
which showed comparable accuracy values
to the original 48-parameter circuit. This suggests that DEA can reduce the number of trainable parameters while maintaining strong performance.
The dependence on batch size showed that for both SO(4) and SU(4), accuracy tends to decrease as the batch size increases, with a batch size of 16 often showing faster convergence compared to larger sizes. The overall conclusion is that DEA can reduce the number of trainable parameters while maintaining good overall accuracy.
Conclusion and Outlook
The paper concludes that QCNNs perform better than corresponding CNN models in the absence of noise for most setups. The key takeaway is that DEA demonstrates the ability to significantly reduce the number of trainable parameters while maintaining good overall accuracy,
which could have a "great impact when one is aiming for noisy hardware implementation.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems, based on the provided research, and what those improved systems could achieve:
-
Enhance performance in high-energy physics (HEP) jet image classification tasks (specifically top-quark tagging).
-
Implement Quantum Convolutional Neural Networks (QCNNs) as a quantum-inspired machine learning architecture for complex image classification problems where classical CNNs struggle with highly energetic and dense features.
-
Develop QCNN architectures optimized using Dimensional Expressivity Analysis (DEA) to reduce the number of trainable parameters while preserving maximal expressive capability, leading to more efficient and potentially lower-complexity quantum circuits suitable for Noisy Intermediate-Scale Quantum (NISQ) devices.
-
Improve feature extraction in image classification by leveraging quantum entanglement within QCNNs, allowing the model to learn correlations between features across the entire image rather than relying solely on local information.
-
Create more robust and generalizable machine learning models by systematically exploring and optimizing various quantum circuit components (convolutional circuits like SO(4) and SU(4), different encoding schemes like TPE, HEE1, HEE2, CHE, loss functions like Cross-Entropy, MSE, and Hinge loss).
-
Improve model training stability by investigating the impact of batch size variations on QCNN performance and identifying optimal batch sizes that balance convergence speed with final accuracy for specific quantum circuits.
-
Design more parameter-efficient models by applying DEA to prune redundant parameters in QCNN circuits, resulting in smaller, more practical quantum models that maintain comparable or improved accuracy over classical counterparts with similar parameter counts.
-
Develop
Equivariant Quantum Neural Networks
(EQNNs) that automatically satisfy data symmetries (like rotational symmetry inherent in jet images), potentially allowing the use of higher-dimensional input data without requiring dimensionality reduction techniques like PCA, thus avoiding information loss and enabling the use of more complex, high-resolution input images.
Abstract
Recently, interest in quantum computing has significantly increased, driven by its potential advantages over classical techniques. Quantum machine learning (QML) exemplifies one of the important quantum computing applications that are expected to surpass classical machine learning in a wide range of instances. This paper addresses the performance of QML in the context of high-energy physics (HEP). As an example, we focus on the top-quark tagging, for which classical convolutional neural networks (CNNs) have been effective but fall short in accuracy when dealing with highly energetic jet images. In this paper, we use a quantum convolutional neural network (QCNN) for this task and compare its performance with CNN using a classical noiseless simulator. We compare various setups for the QCNN, varying the convolutional circuit, type of encoding, loss function, and batch sizes. For every quantum setup, we design a similar setup to the corresponding classical model for a fair comparison. Our results indicate that, using a classical simulator, QCNN with proper setups tend to perform better than their CNN counterparts, especially when the convolution block has a lower number of parameters. For the higher parameter regime, the QCNN circuit was adjusted according to the dimensional expressivity analysis (DEA) to lower the parameter count while preserving its optimal structure. The DEA circuit demonstrated improved results over the comparable classical CNN model.
Sources
- Classification with Quantum Neural Networks on Near Term Processors
- Parameterized quantum circuits as machine learning models
- Quantum Machine Learning in High Energy Physics
- Quantum-inspired Machine Learning on high-energy physics data
- The Machine Learning Landscape of Top Taggers
- A Survey on State-of-the-art Deep Learning Applications and Challenges
- Quantum Convolutional Neural Networks are Effectively Classically Simulable
- Quantum Convolutional Neural Networks for High Energy Physics Data Analysis
- PennyLane: Automatic differentiation of hybrid quantum-classical computations
- An Introduction to Convolutional Neural Networks
- Adam: A Method for Stochastic Optimization
- Subtleties in the trainability of quantum machine learning models
- Dimensional Expressivity Analysis, best-approximation errors, and automated design of parametric quantum circuits
- A robust anomaly finder based on autoencoders
Related papers
- Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more
- Encrypted clones can leak: Classification of informative subsets in Quantum Encrypted Cloning
- Polynomial-time classical and quantum simulation of quantum impurity models
- Theory of quantum-enhanced interferometry with general Markovian light sources
- A convergent hierarchy of spectral gap certificates for qubit Hamiltonians
- Universal Bound and Phase Transition in Many-Body Fermionic Non-Gaussianity