Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling

arXiv:2601.09951 · quant-ph · Submitted 2026-01-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.

Kai: Today's paper: "Parallelizing the Variational Quantum Eigensolver".

Mira: The VQE algorithm was implemented and benchmarked for computing the potential energy surface of hydrogen molecules across 100 bond lengths using an HPC cluster featuring 4× NVIDIA H100 GPUs,

Kai: First, who's behind it and why it matters.

Title and authors: Kai: So, this paper, "Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling," it looks at taking VQE and making it run much faster on real hardware.

Mira: Right, Kai. It’s about how they took the VQE algorithm, which is that hybrid quantum-classical thing for finding molecular energies, and they looked at how to parallelize it so you can actually get results in a reasonable time on a cluster.

Lev: I'm curious if this stuff is practical right now. Because when we talk about real hardware, we always have to worry about the overhead of things like running the circuit versus actual measurement time, especially with noisy devices.

Kai: Exactly, Lev. The paper sets up a whole study on how they parallelized it across four phases: first JIT compilation and optimizer work together for a four point one three times speedup, then GPU acceleration on the device itself, then MPI for distributed computing across cores, and finally multi-GPU scaling which got them to three point nine eight times speedup with ninety-nine point four percent parallel efficiency one <ref:2601.09951#pg1>.

Mira: That number of efficiency is pretty impressive for a workload like this. It shows that the way they structured the problem—where each bond length calculation is independent—is really well suited for this kind of scaling one <ref:2601.09951#pg1>.

Lev: And what they found was that when you combine all those phases, from JIT compilation to MPI and multi-GPU scaling, you get a total speedup of one hundred seventeen times for the hydrogen molecule potential energy surface, cutting the time from fifty-nine point ninety five seconds down to just over five seconds one <ref:2601.09951#pg1>.

Kai: That jump is huge. It really shows that the combined effect of all those parallelization strategies—JIT, GPU acceleration, MPI—is what gets you from a run that takes nearly ten minutes down to something interactive.

Mira: And they confirmed the core claim of their initial work: that this single-parameter double excitation ansatz is actually sufficient to get good accuracy for H2 near its equilibrium bond length one <ref:2601.09951#pg1>. That’s the physical chemistry part we care about.

Lev: From a hardware perspective, what I find interesting is that they did a CPU versus GPU scaling study from four qubits all the way up to twenty-six qubits, and they found that the GPU advantage keeps growing, jumping from ten times speedup at four qubits to over eighty times speedup at twenty-six qubits one <ref:2601.09951#pg1>.

Kai: That's a big indicator for anyone thinking about using larger systems. The paper shows that for this specific VQE setup on their cluster, the GPU is definitely the way to go when you start scaling up the number of qubits you can handle.

Mira: But we have to keep an eye on the limitations they pointed out. They mentioned that a single H100 GPU with eight gigabytes of memory really limits them to twenty-nine qubits, and going beyond that requires distributed methods because thirty qubits would need about sixty-four gigabytes one <ref:2601.09951#pg1>.

Lev: That’s a crucial piece of information for anyone trying to plan experiments on current Noisy Intermediate-Scale Quantum devices. It gives you a hard constraint on how big your simulation can realistically be without needing more complex distribution setups.

Kai: So, what they’re saying is that the VQE framework itself is solid for these small molecules, but the performance gains are almost entirely dependent on how cleverly you parallelize the classical optimization steps alongside the quantum circuit execution.

Mira: The paper highlights three major improvements they suggest: first, using JIT compilation to get adaptive early convergence, which helps reduce those two hundred VQE iterations we need for each bond length one. Second, making sure you establish a proper baseline before measuring speedup to make sure you’re isolating the right factors.

Lev: And third, the decomposition they do of the total speedup—breaking it down into Optimizer plus JIT, GPU device acceleration, MPI parallelization, and multi-GPU scaling—is really helpful for figuring out exactly where your bottlenecks are in a future run one <ref:2601.09951#pg1>.

Kai: So to wrap up this paper on "Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling," it’s showing us that you can take a VQE setup and apply smart parallelization across multiple layers, resulting in an eleven-seventy times speedup for the hydrogen molecule surface one <ref:2601.09951#pg1,Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling>.

Mira: It confirms that for simple molecules like H2, the ansatz works well and the variational principle holds up under these kinds of computational scaling tests one <ref:2601.09951#pg1>.

Lev: From a researcher standpoint, it tells us that if we want to run these kinds of calculations on real hardware, we need to be very deliberate about choosing the right parallelization strategy for our specific problem size one <ref:2601.09951#pg1>.

Kai: It’s clear that the real lesson here is adopting JIT compilation and using MPI for distributed computing when you’re pushing the limits of what a single machine can do one <ref:2601.09951#pg1>.

Mira: So we'll leave you with this study on "Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling," showing how smart parallelization can turn a long calculation into something that actually runs in seconds.

The paper's summary: Kai: So we’re talking about how they took VQE, that whole hybrid quantum thing for finding molecular energies, and they looked at making it run way faster by layering different speedups on top of each other.

Mira: Right, so the main takeaway is this combined approach—JIT compilation plus GPU acceleration plus MPI parallelization—gets you a total speedup of one hundred seventeen times for calculating the potential energy surface of hydrogen.

Lev: That number is huge because it shows that for problems like this, where you have a lot of independent calculations, the classical optimization steps can be made incredibly efficient when you use the right tools.

Kai: It means they managed to get a run that used to take nearly ten minutes down to just five seconds using all these different tricks working together.

Mira: And what’s really interesting is how they broke down that total speedup, showing it came from four main things, which is JIT compilation, GPU device acceleration, MPI for cores, and multi-GPU scaling.

Lev: That decomposition helps us understand exactly where the bottlenecks are if we want to run these kinds of experiments on real hardware later.

Kai: It also confirms that the single-parameter double excitation ansatz they used actually gets good enough accuracy for hydrogen near its stable bond length, which is the physical chemistry validation here.

Mira: But they also flagged a few things, like the memory limits of an H100 GPU and how you need more complex distributed methods when you try to push the number of qubits too high.

Lev: So it's practical for today’s Noisy Intermediate-Scale Quantum devices, but it gives us clear limits on how big our simulations can get before we need bigger hardware setups.

Kai: The authors are suggesting that using JIT compilation helps with early convergence, which means the optimizer doesn't have to run as many rounds of work before it finds a good answer.

Mira: That adaptive control is a smart way to handle these parameter updates, and they emphasize setting up a solid baseline first so you can actually measure the speedup accurately.

Lev: If we look at what this means for the broader field, it shows that VQE isn't just an academic curiosity; it’s definitely a viable method if you are smart about parallelizing the classical parts of the quantum-classical loop.

Kai: It really suggests that we should focus on building systems that can handle these multi-GPU setups because they showed a pretty high efficiency rate, nearly ninety-nine point four percent.

Mira: So the implication is that for those who want to do actual quantum chemistry work on HPC clusters, this paper gives them a concrete roadmap for how to maximize their compute time.

Lev: And it sets the stage for future work by clearly defining what we need—like better memory management and more efficient ways to handle those larger qubit counts you mentioned.

The paper's improvements: Kai: So, we’re talking about how they’ve suggested specific ways to actually make this VQE setup faster in the future, moving beyond just what they did on their cluster.

Mira: The big suggestion is leaning into JIT compilation even harder, specifically tying it to adaptive early convergence so the optimizer doesn't waste time running two hundred iterations when it could find an answer sooner.

Lev: That means for real hardware, we need to design the circuit structure in a way that lets the AI adjust its search strategy dynamically based on what’s happening during the simulation.

Kai: And they really pushed for establishing a solid baseline before you measure any speedup factors because it makes sure you aren't just measuring some random noise.

Mira: They also stress that when you look at performance, you have to break it down into those four parts—JIT, GPU acceleration, MPI cores, and multi-GPU scaling—to really pinpoint what’s slowing things down.

Lev: That kind of clear breakdown is useful for error correction research because it tells us if the issue is in the quantum circuit itself or in how we're distributing the classical optimization tasks across different processors.

Kai: It also pointed out that using MPI over something like Ray was a better choice for their specific setup because it had less overhead when dealing with this kind of static workload.

Mira: So, the practical implication is that if you build a quantum chemistry workflow, you should probably always think about how to layer these optimization techniques on top of each other for maximum benefit.

Lev: It's about making sure the underlying hardware—like those H100s we talked about—is being used in the most efficient way possible, which is a major consideration when planning experiments.

Kai: And they also gave some specific advice on what kind of hardware you need to consider if you want to go beyond a single GPU setup for these simulations.

Mira: They highlighted that exceeding eight gigabytes of memory on one card means you absolutely have to switch to more complex distributed methods because the state vectors just get too big.

Lev: So it’s not just about having more qubits; it’s about having the right distribution strategy ready before you try to scale up.

Kai: This study really pushes us toward using JIT compilation as a standard practice for any quantum algorithm that involves many iterative classical steps, even if the hardware is small.

Conclusion: Kai: So we’re wrapping up on "Parallelizing the Variational Quantum Eigensolver: From JIT Compilation to Multi-GPU Scaling," which essentially shows how you can stack different parallelization techniques to get a massive speedup for calculating molecular energies.

Mira: Right, so the main point is that this combined approach—JIT compilation plus GPU acceleration plus MPI cores—gets you a total speedup of one hundred seventeen times for calculating the hydrogen molecule potential energy surface.

Lev: That number is huge because it shows that for problems where you have a lot of independent calculations, the classical optimization steps can be made incredibly efficient when you use the right tools.

Kai: It means they managed to get a run that used to take nearly ten minutes down to just five seconds using all these different tricks working together.

Mira: And what’s really interesting is how they broke down that total speedup, showing it came from four main things—JIT, GPU acceleration, MPI for cores, and multi-GPU scaling.

Lev: That decomposition helps us understand exactly where the bottlenecks are if we want to run these kinds of experiments on real hardware later.

Kai: It also confirms that the single-parameter double excitation ansatz they used actually gets good enough accuracy for hydrogen near its stable bond length, which is the physical chemistry validation here.

Mira: But they also flagged a few things, like the memory limits of an H100 GPU and how you need more complex distributed methods when you try to push the number of qubits too high.

Lev: So it’s practical for today’s Noisy Intermediate-Scale Quantum devices, but it gives us clear limits on how big our simulations can get before we need bigger hardware setups.

Kai: The authors are suggesting that using JIT compilation helps with early convergence, which means the optimizer doesn't have to run as many rounds of work before it finds a good answer.

Mira: That adaptive control is a smart way to handle those parameter updates, and they emphasize setting up a solid baseline first so you can actually measure the speedup accurately.

Lev: If we look at what this means for the broader field, it shows that VQE isn't just an academic curiosity; it’s definitely a viable method if you are smart about parallelizing the classical parts of the quantum-classical loop.

Kai: It really suggests that we should focus on building systems that can handle these multi-GPU setups because they showed a pretty high efficiency rate, nearly ninety-nine point four percent.

Mira: So the implication is that for those who want to do actual quantum chemistry work on HPC clusters, this paper gives them a concrete roadmap for how to maximize their compute time.

Lev: And it sets the stage for future work by clearly defining what we need—like better memory management and more efficient ways to handle those larger qubit counts you mentioned.

Department of Physical Sciences, Embry-Riddle Aeronautical University

quant-ph

Submitted: 2026-01-15

Updated: 2026-10-06

Comments: v2: corrects speedup attribution, MPI scaling, and multi-GPU efficiency claims of v1; 10 pages, 3 figures, 5 tables

Code: https://github.com/rylanmalarchick/QuantumVQE

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: The VQE algorithm was implemented and benchmarked for computing the potential energy surface of hydrogen molecules across 100 bond lengths using an HPC cluster featuring 4× NVIDIA H100 GPUs,

Key concepts

VQE (Variational Quantum Eigensolver)
VQE is a hybrid quantum-classical algorithm used to find the lowest energy state of molecules. It uses a parameterized quantum circuit (ansatz) on a quantum computer and a classical optimizer to iteratively adjust the parameters until the system finds an approximate ground state energy. It is useful because it can run on current noisy hardware.
JIT Compilation (Just-In-Time Compilation)
JIT compilation, using tools like JAX and Catalyst, optimizes quantum circuits during runtime. This process translates the high-level circuit description into highly efficient machine code specifically for the target hardware (like GPUs). This significantly reduces execution time by optimizing how the quantum operations are performed on the computer.
MPI Parallelization
MPI (Message Passing Interface) is a standard method for allowing different computers in a cluster to work together. In this study, MPI was used to parallelize the outer loop of VQE, meaning calculations for different bond lengths were run simultaneously across multiple processors. This strategy allowed the workload to be distributed efficiently across the HPC cluster.
Multi-GPU Scaling
This refers to using several high-performance GPUs (like NVIDIA H100s) together to speed up computations. The paper showed that scaling from one GPU to four GPUs provided significant performance gains, confirming that VQE parameter sweeps are highly efficient when deployed across multiple accelerators.

Terminology

Summary

The VQE algorithm was implemented and benchmarked for computing the potential energy surface of hydrogen molecules across 100 bond lengths using an HPC cluster featuring 4× NVIDIA H100 GPUs, demonstrating significant speedups through JIT compilation, GPU acceleration, MPI parallelization, and multi-GPU scaling. The gist is that the combined effect yields 117× total speedup for the H2 potential energy surface (593.95s → 5.04s) <ref:2601.09951#pg2>.

Problem Description

The goal of this work is to compute the ground state energy E0 of the H2 molecule as a function of internuclear distance d <ref:2601.09951#pg4>. The electronic Hamiltonian in the Born-Oppenheimer approximation is given by equation (1) <ref:2601.09951#pg4>. The computational task involves generating 8,000 total circuit evaluations taking approximately 50 seconds in the serial implementation <ref:2601.09951#pg5>. This process requires Hartree-Fock calculation to generate molecular Hamiltonian, 200 quantum circuit evaluations with gradient computation, and parameter updates via Adam optimizer for each bond length <ref:2601.09951#pg5>.

Model Formulation

The VQE uses the variational principle where the energy calculated will always be greater than or equal to the true ground state energy E0 (2) <ref:2601.09951#pg4>. The molecular Hamiltonian in second quantization is represented by equation (3) <ref:2601.09951#pg4>. The trial wavefunction is prepared using a parameterized quantum circuit ψ(θ)⟩ = U(θ)HF⟩ (4) <ref:2601.09951#pg4>. This ansatz uses a double excitation gate defined by equation (5) <ref:2601.09951#pg4>. The optimization problem seeks the parameter value θ that minimizes the energy E(θ) = ⟨ψ(θ)Hψ(θ)⟩ (6) <ref:2601.09951#pg4>.

Parallelization Strategies

The computational task is embarrassingly parallel, as each bond length calculation is independent, requiring no data from other calculations <ref:2601.09951#pg5>. The proposed three-phase optimization strategy includes:

  1. Phase 1: JIT Compilation with JAX, which reduces runtime by achieving an expected speedup of 2–5× <ref:2601.09951#pg7>.

  2. Phase 2: Distributed-Memory Parallelism using OpenMPI with mpi4py to parallelize the outer loop over bond lengths, expecting a speedup of 0.8p for p cores <ref:2601.09951#pg8>.

  3. Phase 3: MPI is selected over Ray due to dependency conflicts, offering mature HPC integration and minimal overhead for this static workload <ref:2601.09951#pg8>.

Performance Results and Analysis

The serial VQE implementation achieved a total runtime of 593.95 seconds <ref:2601.09951#pg9>. The control experiment, Serial Optax+JIT, achieved a speedup of 4.13× compared to the baseline <ref:2601.09951#pg9>. GPU acceleration using lightning.gpu showed a 3.60× speedup at 4 qubits but demonstrated an increasing advantage with qubit count, reaching 80.5× at 26 qubits <ref:2601.09951#pg11>. The CPU vs GPU scaling study showed that GPU wins at all scales, with speedup increasing from 10× at 4 qubits to over 80× at 26 qubits <ref:2601.09951#pg14>. MPI parallelization achieved a dramatic speedup of 28.53× relative to the Optax+JIT baseline, resulting in a runtime of only 5.04s for the full workload <ref:2601.09951#pg16>. The combined effect yields 117× total speedup for the H2 potential energy surface <ref:2601.09951#pg9>.

Key Findings and Limitations

The physical results confirm that the single-parameter double excitation ansatz successfully captures the essential quantum chemistry of H2, achieving good accuracy near equilibrium bond lengths <ref:2601.09951#pg15>. However, the paper identifies several practical lessons:

Algorithm choice matters:

GPU overhead:

The single H100 (8GB) hardware limits simulation to 29 qubits, and exceeding this requires distributed methods <ref:2601.09951#pg15>. The near-perfect parallel efficiency of 99.4% across 4 GPUs confirms that VQE parameter sweeps are ideal for multi-GPU deployment with essentially zero communication overhead <ref:2601.09951#pg14>. These results show that useful quantum chemistry calculations are practical today on HPC infrastructure <ref:2601.09951#pg19>.

The paper concludes that the optimized implementation reduces computation time from nearly 10 minutes to 5 seconds, enabling interactive quantum chemistry exploration <ref:2601.09951#pg8>. The best practices identified include using JIT compilation for all implementations and ensuring a proper baseline is established to isolate speedup factors <ref:2601.09951#pg9>. All source code can be accessed at https://github.com/rylanmalarchick/QuantumVQE <ref:2601.09951#pg19>.

The paper acknowledges Dr. Khanal for guidance on parallelization strategies and HPC methodologies <ref:2601.09951#pg9>. The work was conducted on the ERAU Vega HPC cluster featuring AMD EPYC 9654 processors and NVIDIA GPU accelerators <ref:2601.09951#pg9>. The quantum simulations used the PennyLane quantum computing framework with Lightning backend, JAX for automatic differentiation, and Catalyst for JIT compilation <ref:2601.09951#pg9>. Large language model tools were used to assist with manuscript preparation and editing <ref:2601.09951#pg9>. The authors take full responsibility for all scientific content, methodology, and results <ref:2601.09951#pg9>. All source code can be accessed at https://github.com/rylanmalarchick/QuantumVQE <ref:2601.09951#pg9>.

--- Page 2 ---

The VQE is a hybrid quantum-classical algorithm for computing ground state energies of molecular systems <ref:2601.09951#pg4>. Unlike purely quantum algorithms that need perfect quantum computers, VQE works on today’s noisy quantum computers <ref:2601.09951#pg4>. The algorithm uses a quantum circuit (called an ansatz) with adjustable parameters to create trial wavefunctions on a quantum processor, while a classical computer adjusts these parameters to find the lowest energy <ref:2601.09951#pg4>. We focus on the hydrogen molecule (H2), the simplest neutral molecule, which serves as a benchmark system for quantum chemistry methods <ref:2601.09951#pg4>. Despite its simplicity, H2 exhibits key features of chemical bonding including equilibrium bond length, dissociation energy, and potential energy surface structure <ref:2601.09951#pg4>.

How it works

The serial algorithm follows Algorithm 1 for the VQE calculation <ref:2601.09951#pg5>. For each bond length di, the process involves generating H(di) using Hartree-Fock, initializing θ=0, and running the Adam optimizer for 200 iterations <ref:2601.09951#pg5>. The computational complexity for a 4-qubit system is O(256) operations per circuit evaluation <ref:2601.09951#pg6>.

Performance Analysis

The total runtime for the serial implementation was 50.64 seconds <ref:2601.09951#pg9>. The VQE optimization dominates runtime, with gradient computation and cost function evaluation accounting for 28.2 seconds of the 32.7 second total (86%) <ref:2601.09951#pg16>. The JIT-compiled implementation achieved a speedup of 4.13× versus the serial PennyLane Adam baseline <ref:2601.09951#pg9>. The GPU acceleration implementation showed a 3.

Improvements for AI systems

  1. textbfImproved VQE implementation for molecular systems: VQE with a simple ansatz successfully computes the H2 potential energy surface with high accuracy near equilibrium. This validates that the single-parameter double excitation ansatz is sufficient for this simple molecule, confirming that the variational approach works well, suggesting this method can be extended to larger molecules.

  2. textbfEnhanced HPC parallelization for quantum chemistry: Achieved 117× speedup using MPI with 16-32 processes and a near-linear strong scaling from 2 to 8 processes. This demonstrates that VQE is highly amenable to HPC optimization, reducing runtime from nearly 10 minutes to 5 seconds, enabling interactive quantum chemistry exploration.

  3. textbfAdaptive execution control: Integrated JIT compilation with the statement JIT compilation enables adaptive early convergence (fewer iterations), leading to a speedup of 4.13× over a serial baseline. This allows the AI system to dynamically adjust optimization strategies based on circuit structure and hardware context for faster convergence.

  4. textbfHardware-aware quantum simulation: The study showed that GPU advantage increases dramatically with qubit count, reaching 80.5× speedup at 26 qubits for a synthetic Hamiltonian, contradicting initial H2 results due to the overhead of Hartree-Fock generation. This informs system design by indicating when to leverage GPU acceleration versus CPU processing based on problem size.

  5. textbfResource management for quantum hardware: Identified that Single H100 (8GB) maxes out at 29 qubits and that simulations beyond this require distributed state vectors, as 30 qubits requires ∼64GB, which exceeds the limit. This provides critical constraints for designing experiments on current NISQ devices.

  6. textbfContext-aware performance modeling: The decomposition shows the total speedup is driven by four factors: Optimizer+JIT (4.13×), GPU Device Acceleration, MPI Parallelization (28.53×), and Multi-GPU Scaling (3.98×). This allows researchers to isolate bottlenecks in future quantum algorithms by measuring performance across these distinct parallelization vectors.

Abstract

The Variational Quantum Eigensolver (VQE) is a hybrid quantum-classical algorithm for ground-state energies of molecular systems. We compute the potential energy surface of the hydrogen molecule (H2) at 100 bond lengths with PennyLane on the ERAU Vega cluster (AMD EPYC 9654 CPUs, 4 NVIDIA H100 GPUs) and compare a serial baseline, an Optax and Catalyst JIT version, a lightning.gpu version, and an MPI version. Early stopping, a larger learning rate, and warm starts reduce the iteration count from 30,000 to 7,033 and the runtime from 593.95 s to 143.80 s, while the time per iteration stays between 0.0198 s and 0.0204 s. At 4 qubits, lightning.gpu lowers the time per iteration to 0.0145 s. MPI runs with 2 to 32 ranks finish in 8.45 s to 5.04 s; the serial and MPI jobs ran under different core allocations, so we do not report their ratio as a parallel speedup. In a 4 to 26 qubit study with different CPU and GPU software stacks, the CPU/GPU time ratio ranges from 3.5 to 81.5. One H100 runs up to 29 qubits.

Sources

Related papers