Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling
summary
The gist
The VQE algorithm was implemented and benchmarked for computing the potential energy surface of hydrogen molecules across 100 bond lengths using an HPC cluster featuring 4× NVIDIA H100 GPUs,
In short
The VQE algorithm was tested for hydrogen molecule energy surfaces across 100 bond lengths using an HPC cluster with four NVIDIA H100 GPUs. By combining JIT compilation, GPU acceleration, and MPI parallelization, the total speedup reached 117×, reducing computation time from nearly 10 minutes to just over 5 seconds. This demonstrates that VQE is practical for high-performance quantum chemistry on modern hardware.
Key concepts
- VQE (Variational Quantum Eigensolver)
- VQE is a hybrid quantum-classical algorithm used to find the lowest energy state of molecules. It uses a parameterized quantum circuit (ansatz) on a quantum computer and a classical optimizer to iteratively adjust the parameters until the system finds an approximate ground state energy. It is useful because it can run on current noisy hardware.
- JIT Compilation (Just-In-Time Compilation)
- JIT compilation, using tools like JAX and Catalyst, optimizes quantum circuits during runtime. This process translates the high-level circuit description into highly efficient machine code specifically for the target hardware (like GPUs). This significantly reduces execution time by optimizing how the quantum operations are performed on the computer.
- MPI Parallelization
- MPI (Message Passing Interface) is a standard method for allowing different computers in a cluster to work together. In this study, MPI was used to parallelize the outer loop of VQE, meaning calculations for different bond lengths were run simultaneously across multiple processors. This strategy allowed the workload to be distributed efficiently across the HPC cluster.
- Multi-GPU Scaling
- This refers to using several high-performance GPUs (like NVIDIA H100s) together to speed up computations. The paper showed that scaling from one GPU to four GPUs provided significant performance gains, confirming that VQE parameter sweeps are highly efficient when deployed across multiple accelerators.
Terminology used across episodes
This episode discusses
- Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling · Paper Radio
- PennyLane: Automatic differentiation of hybrid quantum-classical computations
The paper
Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling · Read on arXiv
Department of Physical Sciences, Embry-Riddle Aeronautical University
The Variational Quantum Eigensolver (VQE) is a hybrid quantum-classical algorithm for ground-state energies of molecular systems. We compute the potential energy surface of the hydrogen molecule (H2) at 100 bond lengths with PennyLane on the ERAU Vega cluster (AMD EPYC 9654 CPUs, 4 NVIDIA H100 GPUs) and compare a serial baseline, an Optax and Catalyst JIT version, a lightning.gpu version, and an MPI version. Early stopping, a larger learning rate, and warm starts reduce the iteration count from 30,000 to 7,033 and the runtime from 593.95 s to 143.80 s, while the time per iteration stays between 0.0198 s and 0.0204 s. At 4 qubits, lightning.gpu lowers the time per iteration to 0.0145 s. MPI runs with 2 to 32 ranks finish in 8.45 s to 5.04 s; the serial and MPI jobs ran under different core allocations, so we do not report their ratio as a parallel speedup. In a 4 to 26 qubit study with different CPU and GPU software stacks, the CPU/GPU time ratio ranges from 3.5 to 81.5. One H100 runs up to 29 qubits.
Transcript
Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.
Kai: Today's paper: "Parallelizing the Variational Quantum Eigensolver".
Mira: The VQE algorithm was implemented and benchmarked for computing the potential energy surface of hydrogen molecules across 100 bond lengths using an HPC cluster featuring 4× NVIDIA H100 GPUs,
Kai: First, who's behind it and why it matters.
Title and authors: Kai: So, this paper, "Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling," it looks at taking VQE and making it run much faster on real hardware.
Mira: Right, Kai. It’s about how they took the VQE algorithm, which is that hybrid quantum-classical thing for finding molecular energies, and they looked at how to parallelize it so you can actually get results in a reasonable time on a cluster.
Lev: I'm curious if this stuff is practical right now. Because when we talk about real hardware, we always have to worry about the overhead of things like running the circuit versus actual measurement time, especially with noisy devices.
Kai: Exactly, Lev. The paper sets up a whole study on how they parallelized it across four phases: first JIT compilation and optimizer work together for a four point one three times speedup, then GPU acceleration on the device itself, then MPI for distributed computing across cores, and finally multi-GPU scaling which got them to three point nine eight times speedup with ninety-nine point four percent parallel efficiency one <ref:2601.09951#pg1>.
Mira: That number of efficiency is pretty impressive for a workload like this. It shows that the way they structured the problem—where each bond length calculation is independent—is really well suited for this kind of scaling one <ref:2601.09951#pg1>.
Lev: And what they found was that when you combine all those phases, from JIT compilation to MPI and multi-GPU scaling, you get a total speedup of one hundred seventeen times for the hydrogen molecule potential energy surface, cutting the time from fifty-nine point ninety five seconds down to just over five seconds one <ref:2601.09951#pg1>.
Kai: That jump is huge. It really shows that the combined effect of all those parallelization strategies—JIT, GPU acceleration, MPI—is what gets you from a run that takes nearly ten minutes down to something interactive.
Mira: And they confirmed the core claim of their initial work: that this single-parameter double excitation ansatz is actually sufficient to get good accuracy for H2 near its equilibrium bond length one <ref:2601.09951#pg1>. That’s the physical chemistry part we care about.
Lev: From a hardware perspective, what I find interesting is that they did a CPU versus GPU scaling study from four qubits all the way up to twenty-six qubits, and they found that the GPU advantage keeps growing, jumping from ten times speedup at four qubits to over eighty times speedup at twenty-six qubits one <ref:2601.09951#pg1>.
Kai: That's a big indicator for anyone thinking about using larger systems. The paper shows that for this specific VQE setup on their cluster, the GPU is definitely the way to go when you start scaling up the number of qubits you can handle.
Mira: But we have to keep an eye on the limitations they pointed out. They mentioned that a single H100 GPU with eight gigabytes of memory really limits them to twenty-nine qubits, and going beyond that requires distributed methods because thirty qubits would need about sixty-four gigabytes one <ref:2601.09951#pg1>.
Lev: That’s a crucial piece of information for anyone trying to plan experiments on current Noisy Intermediate-Scale Quantum devices. It gives you a hard constraint on how big your simulation can realistically be without needing more complex distribution setups.
Kai: So, what they’re saying is that the VQE framework itself is solid for these small molecules, but the performance gains are almost entirely dependent on how cleverly you parallelize the classical optimization steps alongside the quantum circuit execution.
Mira: The paper highlights three major improvements they suggest: first, using JIT compilation to get adaptive early convergence, which helps reduce those two hundred VQE iterations we need for each bond length one. Second, making sure you establish a proper baseline before measuring speedup to make sure you’re isolating the right factors.
Lev: And third, the decomposition they do of the total speedup—breaking it down into Optimizer plus JIT, GPU device acceleration, MPI parallelization, and multi-GPU scaling—is really helpful for figuring out exactly where your bottlenecks are in a future run one <ref:2601.09951#pg1>.
Kai: So to wrap up this paper on "Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling," it’s showing us that you can take a VQE setup and apply smart parallelization across multiple layers, resulting in an eleven-seventy times speedup for the hydrogen molecule surface one <ref:2601.09951#pg1,Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling>.
Mira: It confirms that for simple molecules like H2, the ansatz works well and the variational principle holds up under these kinds of computational scaling tests one <ref:2601.09951#pg1>.
Lev: From a researcher standpoint, it tells us that if we want to run these kinds of calculations on real hardware, we need to be very deliberate about choosing the right parallelization strategy for our specific problem size one <ref:2601.09951#pg1>.
Kai: It’s clear that the real lesson here is adopting JIT compilation and using MPI for distributed computing when you’re pushing the limits of what a single machine can do one <ref:2601.09951#pg1>.
Mira: So we'll leave you with this study on "Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling," showing how smart parallelization can turn a long calculation into something that actually runs in seconds.
The paper's summary: Kai: So we’re talking about how they took VQE, that whole hybrid quantum thing for finding molecular energies, and they looked at making it run way faster by layering different speedups on top of each other.
Mira: Right, so the main takeaway is this combined approach—JIT compilation plus GPU acceleration plus MPI parallelization—gets you a total speedup of one hundred seventeen times for calculating the potential energy surface of hydrogen.
Lev: That number is huge because it shows that for problems like this, where you have a lot of independent calculations, the classical optimization steps can be made incredibly efficient when you use the right tools.
Kai: It means they managed to get a run that used to take nearly ten minutes down to just five seconds using all these different tricks working together.
Mira: And what’s really interesting is how they broke down that total speedup, showing it came from four main things, which is JIT compilation, GPU device acceleration, MPI for cores, and multi-GPU scaling.
Lev: That decomposition helps us understand exactly where the bottlenecks are if we want to run these kinds of experiments on real hardware later.
Kai: It also confirms that the single-parameter double excitation ansatz they used actually gets good enough accuracy for hydrogen near its stable bond length, which is the physical chemistry validation here.
Mira: But they also flagged a few things, like the memory limits of an H100 GPU and how you need more complex distributed methods when you try to push the number of qubits too high.
Lev: So it's practical for today’s Noisy Intermediate-Scale Quantum devices, but it gives us clear limits on how big our simulations can get before we need bigger hardware setups.
Kai: The authors are suggesting that using JIT compilation helps with early convergence, which means the optimizer doesn't have to run as many rounds of work before it finds a good answer.
Mira: That adaptive control is a smart way to handle these parameter updates, and they emphasize setting up a solid baseline first so you can actually measure the speedup accurately.
Lev: If we look at what this means for the broader field, it shows that VQE isn't just an academic curiosity; it’s definitely a viable method if you are smart about parallelizing the classical parts of the quantum-classical loop.
Kai: It really suggests that we should focus on building systems that can handle these multi-GPU setups because they showed a pretty high efficiency rate, nearly ninety-nine point four percent.
Mira: So the implication is that for those who want to do actual quantum chemistry work on HPC clusters, this paper gives them a concrete roadmap for how to maximize their compute time.
Lev: And it sets the stage for future work by clearly defining what we need—like better memory management and more efficient ways to handle those larger qubit counts you mentioned.
The paper's improvements: Kai: So, we’re talking about how they’ve suggested specific ways to actually make this VQE setup faster in the future, moving beyond just what they did on their cluster.
Mira: The big suggestion is leaning into JIT compilation even harder, specifically tying it to adaptive early convergence so the optimizer doesn't waste time running two hundred iterations when it could find an answer sooner.
Lev: That means for real hardware, we need to design the circuit structure in a way that lets the AI adjust its search strategy dynamically based on what’s happening during the simulation.
Kai: And they really pushed for establishing a solid baseline before you measure any speedup factors because it makes sure you aren't just measuring some random noise.
Mira: They also stress that when you look at performance, you have to break it down into those four parts—JIT, GPU acceleration, MPI cores, and multi-GPU scaling—to really pinpoint what’s slowing things down.
Lev: That kind of clear breakdown is useful for error correction research because it tells us if the issue is in the quantum circuit itself or in how we're distributing the classical optimization tasks across different processors.
Kai: It also pointed out that using MPI over something like Ray was a better choice for their specific setup because it had less overhead when dealing with this kind of static workload.
Mira: So, the practical implication is that if you build a quantum chemistry workflow, you should probably always think about how to layer these optimization techniques on top of each other for maximum benefit.
Lev: It's about making sure the underlying hardware—like those H100s we talked about—is being used in the most efficient way possible, which is a major consideration when planning experiments.
Kai: And they also gave some specific advice on what kind of hardware you need to consider if you want to go beyond a single GPU setup for these simulations.
Mira: They highlighted that exceeding eight gigabytes of memory on one card means you absolutely have to switch to more complex distributed methods because the state vectors just get too big.
Lev: So it’s not just about having more qubits; it’s about having the right distribution strategy ready before you try to scale up.
Kai: This study really pushes us toward using JIT compilation as a standard practice for any quantum algorithm that involves many iterative classical steps, even if the hardware is small.
Conclusion: Kai: So we’re wrapping up on "Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling," which essentially shows how you can stack different parallelization techniques to get a massive speedup for calculating molecular energies.
Mira: Right, so the main point is that this combined approach—JIT compilation plus GPU acceleration plus MPI cores—gets you a total speedup of one hundred seventeen times for calculating the hydrogen molecule potential energy surface.
Lev: That number is huge because it shows that for problems where you have a lot of independent calculations, the classical optimization steps can be made incredibly efficient when you use the right tools.
Kai: It means they managed to get a run that used to take nearly ten minutes down to just five seconds using all these different tricks working together.
Mira: And what’s really interesting is how they broke down that total speedup, showing it came from four main things—JIT, GPU acceleration, MPI for cores, and multi-GPU scaling.
Lev: That decomposition helps us understand exactly where the bottlenecks are if we want to run these kinds of experiments on real hardware later.
Kai: It also confirms that the single-parameter double excitation ansatz they used actually gets good enough accuracy for hydrogen near its stable bond length, which is the physical chemistry validation here.
Mira: But they also flagged a few things, like the memory limits of an H100 GPU and how you need more complex distributed methods when you try to push the number of qubits too high.
Lev: So it’s practical for today’s Noisy Intermediate-Scale Quantum devices, but it gives us clear limits on how big our simulations can get before we need bigger hardware setups.
Kai: The authors are suggesting that using JIT compilation helps with early convergence, which means the optimizer doesn't have to run as many rounds of work before it finds a good answer.
Mira: That adaptive control is a smart way to handle those parameter updates, and they emphasize setting up a solid baseline first so you can actually measure the speedup accurately.
Lev: If we look at what this means for the broader field, it shows that VQE isn't just an academic curiosity; it’s definitely a viable method if you are smart about parallelizing the classical parts of the quantum-classical loop.
Kai: It really suggests that we should focus on building systems that can handle these multi-GPU setups because they showed a pretty high efficiency rate, nearly ninety-nine point four percent.
Mira: So the implication is that for those who want to do actual quantum chemistry work on HPC clusters, this paper gives them a concrete roadmap for how to maximize their compute time.
Lev: And it sets the stage for future work by clearly defining what we need—like better memory management and more efficient ways to handle those larger qubit counts you mentioned.
More episodes
- 2610.01068-Learned Parallel Bit-Flipping Sequential Belief Propagation Decoding of Quantum LDPC Codes
- 2610.01074-The stationarity test: a framework for learning quantum many-body systems from their thermal states
- 2610.01094-Quantum synchronization in atom-cavity coupled systems
- 2610.01402-Transport theory for a generic two-arm co-propagating Majorana interferometer with Majorana fermion and edge vortex tunneling
- 2610.01167-Vector chiral order and dynamical quantum phase transitions in an Ising chain with dimerized anisotropic Gamma interaction
- 2610.01163-Robustness hierarchy of bipartite quantum correlations under noisy dynamics
- 2610.01183-Additive solid immersion lenses for enhanced collection efficiency of shallow NV centers by pulsed laser deposition and structurization of high-k amorphous oxides
- 2610.01112-Dissipation-Sensitivity Trade-Off in Dissipative Bosonic Systems
- 2610.01099-Constant-Per-Layer-Depth MPS-Pretrained Ansatz for Noisy Distributed Quantum Processors
- 2610.01141-Classical Hardness of Learning Functions of Hamiltonians