Quantum Integrated High-Performance Computing: Envisioning a Layered Architecture for Next-Generation Hybrid Computing Infrastructure
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.
Kai: Today's paper: "Quantum Integrated High-Performance Computing".
Mira: High-performance computing (HPC) has evolved through multiple architectural transitions, and this paper proposes Quantum Integrated High-Performance Computing (QHPC), a visionary architectural framework that unifies CPUs, GPUs, FPGAs,
Kai: First, who's behind it and why it matters.
Title and authors: Mira: So, diving deeper into what the paper actually lays out, it explains that the QHPC model envisions a next-generation hybrid computing architecture where QPUs are tightly coupled with conventional CPUs and GPUs under a single workflow and resource management framework. It’s about creating a cohesive system rather than just bolting on quantum capabilities.
Kai: That cohesive vision is what makes it compelling; they describe the architecture across five specific layers, starting from the user requests down to the physical compute layer where all these resources live together. It’s a very structured approach to solving this integration problem.
Lev: Structure is important, but how does this workflow management system actually handle the scheduling of tasks when you have nodes with such diverse latency characteristics? I worry that any static scheduling will fail immediately with QPU variability.
Mira: The Workflow Management System, or WMS, models a hybrid workload as a Directed Cyclic Task Graph, or DCTG. This allows them to decompose complex tasks into graphs and analyze dependencies to route jobs based on whether they are latency-critical or latency-tolerant paths.
Kai: That’s the mechanism for dynamic routing; they identify those critical paths so that, for example, a job needing sub-millisecond round-trip times can be sent to co-located QPU and CPU resources. It’s about intelligent path selection.
Lev: If the WMS is doing this semantic dependency analysis, what kind of metrics are they using to determine which path is critical versus tolerant? That sounds like a massive input requirement for the scheduler.
Mira: They use that analysis to decide where the workload goes, essentially routing based on timing constraints derived from the problem's nature within the DCTG model. It’s about matching latency requirements with available hardware tiers.
Kai: And then there’s the Resource Management System, or RMS, which manages everything through a Unified Resource Registry that organizes resources into four distinct tiers, R1 through R4. That tiered hierarchy is crucial for understanding where a job can physically land.
Lev: Tiered resource management sounds robust because it acknowledges the physical reality of different access speeds and connectivity constraints between classical nodes and remote QPUs.
Mira: It does; they define R3 specifically as tightly integrated CPU, GPU, and QPU nodes connected by low-latency interconnects like NVLink. That shows they are prioritizing hardware proximity for certain tasks.
Kai: And then the Quantum-Aware Scheduler within the RMS uses something called a Quantum Suitability Score, or QSS, to pick the right QPU for a job based on fidelity and access latency. It’s a specific metric tailored to quantum hardware quality.
Lev: A score combining gate fidelity, connectivity compatibility, queue wait time, and access latency sounds like exactly what you need when dealing with real-world noise issues; it accounts for the hardware imperfections directly in the decision process.
The paper's summary: Kai: Now we look at how they propose improvements to this QHPC concept, and it really focuses on making the abstraction layer more powerful. They suggest a middleware and programming abstraction layer that handles hardware-agnostic programming, circuit compilation, error mitigation, and communication protocols.
Mira: That’s where the QIR comes in; they lower quantum circuits into a Quantum Intermediate Representation which is LLVM-based. This gives them portability for optimization across different types of QPUs without having to rewrite everything from scratch for each new device.
Lev: A portable IR is nice, but what about the compilation pipeline itself? How do they manage noise and calibration data during that process when you’re dealing with physical systems that drift over time?
Kai: They detail a multi-stage compilation pipeline that includes logical optimization, device mapping with noise-aware placement based on calibration data, and even pulse-level optimization specific to superconducting QPUs. It’s a very fine-grained control mechanism.
Mira: I think the error mitigation selection part is also important because they build in choices for how to handle noise during the compilation phase itself, rather than just applying fixes afterwards at runtime. That’s proactive management of errors.
Lev: If they are doing pulse-level optimization, that suggests a very deep understanding of the physical layer—things like specific microwave pulses used in superconducting systems—which is necessary for reliable execution on real hardware.
Kai: Furthermore, they tackle the "Classical–Quantum Communication Protocol Stack," which uses NVLink for co-located systems and gRPC over InfiniBand for remote access, tying the software stack directly to physical interconnects.
Mira: That protocol stack integration is crucial because it ensures that the communication overhead between classical and quantum components doesn't become the bottleneck in this hybrid workflow. It’s about minimizing data transfer latency.
Lev: Minimizing transfer latency is a huge factor; if we spend too much time moving data between the CPU/GPU and the QPU, even a fast scheduler can’t compensate for that.
Kai: The paper also suggests an adaptive circuit partitioning engine at runtime, which means they can dynamically re-partition large quantum circuits based on live QPU calibration data to maintain optimal fidelity during execution.
Mira: That sounds like a form of Just-In-Time compilation applied to quantum hardware; it allows the system to react instantly if the physical state of the QPU changes unexpectedly.
The paper's improvements: Kai: So, wrapping up the QHPC paper, they are proposing this layered architecture—with its unified management, quantum-aware scheduling, and sophisticated middleware—to treat QPUs as true first-class resources in hybrid computing. It’s a blueprint for how we could actually build these next-generation systems.
Mira: The implication is that we move toward a system where the complexity of integrating different quantum hardware components is managed by intelligent software orchestration, rather than just relying on ad hoc connections between separate systems. This makes large-scale hybrid simulation more viable.
Lev: I think the impact here lies in making it feasible to run complex, high-fidelity simulations that we currently can't touch because of memory constraints or classical intractability, provided the hardware actually delivers on its promised fidelity.
Kai: Exactly; and they point toward future work like developing energy-efficient hybrid infrastructures and extending FPGAs to act as real-time quantum system controllers for error correction decoders operating within microsecond coherence windows.
Mira: That points toward a future where heterogeneous compute resources operate as a unified, programmable continuum, which is a big conceptual step forward in how we think about distributed computation.
Lev: For me, the main challenge remains ensuring that this theoretical framework translates into practical execution on real hardware without introducing too much overhead or instability during the actual quantum operations.
Kai: That’s the million-dollar question for any experimentalist; we need to see if this unified model can actually handle the physical reality of noise and variability.
Mira: It sounds like a necessary step toward extending scientific discovery by systematically tackling the architectural fragmentation that has historically kept quantum computation isolated from classical HPC workflows.
Lev: We'll have to watch how quickly the community moves from this paper’s vision to building systems that can handle those tight latency couplings they described.
Conclusion: Kai: So, to wrap up, this paper on "Quantum Integrated High-Performance Computing: Envisioning a Layered Architecture for Next-Generation Hybrid Computing Infrastructure" really lays out a blueprint for how we can actually build these next-generation hybrid systems by treating CPUs and QPUs as first-class resources under one unified framework.
Mira: It’s fascinating because it tackles the fundamental challenge of making architectural heterogeneity sustainable when you want to scale up scientific discovery, moving away from siloed computing solutions.
Lev: I mean, the concept of that layered approach, from the workflow management system down to the physical compute layer, seems incredibly thorough for tackling real-world implementation hurdles.
Kai: It is really comprehensive; they show how a unified resource management and quantum-aware scheduling system can handle complex hybrid workloads by decomposing them into task graphs.
Mira: And I think their focus on the middleware layer that uses a Quantum Intermediate Representation to handle noise-aware compilation is where the real theoretical heavy lifting happens, ensuring that optimization isn't just classical guesswork.
Lev: From my side, I’m interested in how they define those "latency-critical paths" within the DCTG model because if we can truly manage those timing constraints across tiers like R3 and R4, it makes running actual error-correction routines much more realistic.
Kai: It suggests that the immediate impact is a much more structured way for AI developers to offload hybrid tasks, using familiar programming models while getting hardware-specific optimizations built in from the start.
Mira: The implication is that we might finally see practical applications in areas like high-fidelity molecular simulations or complex optimization problems where classical methods hit a wall due to exponential complexity.
Lev: If they can manage those dynamic re-partitioning capabilities, it opens up possibilities for running larger distributed simulations across multiple QPUs and classical nodes effectively.
Kai: Ultimately, the paper provides the architectural vision needed to bridge the gap between theoretical quantum potential and tangible high-performance computing reality.
Mira: It’s a strong foundation for thinking about how we should approach programming and resource allocation in this new era of hybrid computation.
Lev: We certainly need to keep an eye on how they address those open challenges, especially the issues around facility co-location and managing QPU variability across different sites.
Suman Raja, Siva Sai, Yogesh Simmhanb, Kyle Charda, Rajkumar Buyyad
Department of Computer Science, University of Chicago · Department of Computational and Data Sciences, Indian Institute of Science Bengaluru · Department of Electrical and Computer Engineering, National University of Singapore · Quantum Cloud Computing and Distributed Systems (qCLOUDS) Lab, School of Computing and Information Systems, The University of Melbourne
quant-ph, cs.ET
Submitted: 2026-04-17
Updated: 2026-10-05
Comments: Accepted for publication in Elsevier Future Generation Computer Systems: Special Collection on Advances in Quantum Computing: Methods, Algorithms, and Systems
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 83/100
The gist: High-performance computing (HPC) has evolved through multiple architectural transitions, and this paper proposes Quantum Integrated High-Performance Computing (QHPC), a visionary architectural
Key concepts
- QHPC Architecture
- A visionary design that treats CPUs, GPUs, FPGAs, and QPUs as equal components in a single system. The goal is to create a tightly coupled environment where quantum processing units work directly alongside classical accelerators for complex scientific problems.
- Workflow Management System (WMS)
- The control plane that organizes hybrid jobs into task graphs (DCTG). It analyzes dependencies to determine which tasks need fast execution, such as those requiring quick responses from a QPU, and routes them accordingly.
- Multi-Tier Resource Hierarchy
- A system for organizing physical hardware into four levels (R1 to R4) based on speed and integration. This allows the system to place jobs optimally: for example, placing latency-critical tasks on tightly connected CPU+GPU+QPU nodes (R3).
- Quantum Suitability Score (QSS)
- A metric used by the scheduler to decide which QPU is best for a given job. It combines factors like how reliable the quantum gates are, how well it connects to other hardware, and its current waiting time.
Terminology
Summary
High-performance computing (HPC) has evolved through multiple architectural transitions, and this paper proposes Quantum Integrated High-Performance Computing (QHPC), a visionary architectural framework that unifies CPUs, GPUs, FPGAs, and QPUs as first-class heterogeneous resources to extend the frontier of scientific discovery.
The gist
A QHPC model envisions a next generation hybrid computing architecture tightly coupling QPUs with conventional CPUs and GPUs under a unified workflow and resource management framework.
Background and Motivation
The trajectory of scientific computing has been defined by architectural paradigm shifts, moving from vector processors to massively parallel clusters, GPUs, and now the emergence of Quantum Processing Units (QPUs) as first-class accelerators. Contemporary exascale platforms exemplify this shift, yet irreducibly hard problem classes remain computationally inaccessible. This motivates integrating quantum computation into the HPC stack to exploit quantum parallelism for specific problem classes with provable or empirically observed quantum advantage. The unifying lesson from prior accelerator integrations is that architectural heterogeneity is sustainable only when the system software stack can seamlessly orchestrate heterogeneous execution.
Proposed QHPC Architecture
The QHPC architecture is a layered system design comprising unified resource management, quantum-aware scheduling, hybrid workflow orchestration, middleware and programming abstraction, interconnect technologies for co-design, and a tiered execution model. The overarching goal is to treat the QPU as a heterogeneous co-processor targeting problem classes with provable or empirically observed quantum advantage. This architecture is organized across five architectural layers: the User and Request Layer, the Workflow Management Layer, the Resource Management Layer, the Middleware and Abstraction Layer, and the Physical Compute Layer.
Workflow Management System (WMS)
The QHPC Workflow Management System (WMS) serves as the control plane of the architecture. Its responsibilities include decomposing hybrid workloads into executable task graphs using a Directed Cyclic Task Graph (DCTG), managing inter-task data dependencies, and scheduling tasks to the appropriate compute tier. The WMS models a hybrid workload as a DCTG, where nodes are typed as CPU, GPU, QPU, or FPGA. It performs semantic dependency analysis to identify latency-critical paths
and latency-tolerant paths,
routing workloads accordingly—for instance, using co-located QPU and CPU resources for latency-critical paths targeting sub-millisecond round-trip times.
Resource Management System (RMS)
The QHPC Resource Management System (RMS) manages the inventory, allocation, scheduling, and runtime monitoring of all physical compute resources through a Unified Resource Registry (URR). It maintains a Multi-Tier Resource Hierarchy
organized into four tiers: R1 for CPU-only nodes, R2 for CPU+GPU nodes, R3 for tightly integrated CPU+GPU+QPU nodes connected via low-latency interconnects like NVLink, and R4 for remote or cloud-accessed QPUs. The RMS incorporates a Quantum-Aware Scheduler
that uses a Quantum Suitability Score (QSS)—combining gate fidelity, connectivity compatibility, queue wait time, and access latency—to choose the most appropriate QPU for a job.
Middleware and Abstraction Layer
The Middleware Layer provides hardware-agnostic programming abstractions, circuit compilation and transpilation, error mitigation, and communication protocols. This layer includes lowering quantum circuits to a Quantum Intermediate Representation (QIR), which is an LLVM-based IR enabling portable optimization. The compilation pipeline involves logical optimization, device mapping with noise-aware placement based on calibration data, pulse-level optimization for superconducting QPUs, and error mitigation selection. It also handles the Classical–Quantum Communication Protocol Stack,
utilizing low-latency interconnects like NVLink for co-located systems (R3) and gRPC over InfiniBand for remote access (R4).
Applications and Open Challenges
The framework is designed to support a range of applications, including computational chemistry (VQE), genomics, quantum machine learning, and financial modeling. Key open challenges include Heterogeneous Facility Co-Location,
where differing infrastructure requirements of QPUs conflict with HPC systems, and Quantum-Aware Job Scheduling in Multi-Site Environments,
which requires fidelity-aware policies to manage QPU variability and tight latency coupling across distributed resources. Furthermore, the paper addresses the need for a Unified Quantum–HPC Programming Model
to extend existing MPI/OpenMP/CUDA ecosystems for first-class quantum offloading.
Future Directions
Future research directions include developing energy-efficient and sustainable operation of large-scale hybrid infrastructures, advancing quantum emulation at scale where GPU-based QPU simulation becomes a first-class HPC workload, and extending the hardware abstraction layer to integrate FPGAs as real-time quantum system controllers for error correction decoders operating within microsecond coherence windows. This points toward a future where heterogeneous compute resources operate as a unified, programmable, and globally distributed compute continuum.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the Quantum Integrated High-Performance Computing
(QHPC) paper by Raja et al. The core contribution is a unified architectural framework that treats CPUs, GPUs, FPGAs, and QPUs as first-class resources under a single interface.
Here are the specific improvements to AI systems that can be achieved by implementing the QHPC framework:
The improved AI system will be a next-generation heterogeneous solver capable of tackling problems currently considered classically intractable due to exponential complexity or memory constraints, specifically in high-fidelity simulations and complex optimization tasks. It will operate across a spectrum from classical pre-processing on massive GPU/CPU clusters to quantum subroutines for exponential speedup.
Here are the specific improvements and capabilities:
-
The AI system will utilize a unified submission interface (Layer 1) to seamlessly offload hybrid workloads—where classical machine learning (e.g., CNNs, DNNs) handles feature extraction or classical optimization, and a Quantum Processing Unit (QPU) executes the core quantum kernel (e.g., VQE for molecular energy calculation).
-
It will employ a dynamic, quantum-aware scheduling system (Layer 3) that optimizes resource allocation based on real-time QPU metrics like gate fidelity and coherence time, ensuring jobs are routed to the most reliable hardware for maximum success probability, overcoming the limitations of static classical schedulers.
-
The system will natively support iterative hybrid workflows (Layer 2/5) by implementing
Simultaneous Mode
co-scheduling for tightly coupled problems (like VQE), allowing the classical optimizer and quantum circuit execution to alternate rapidly, minimizing costly idle time on either side. -
It will incorporate a sophisticated middleware layer (Layer 4) that uses a Quantum Intermediate Representation (QIR) for hardware-agnostic compilation and optimizes circuits via gate cancellation, commutation, and noise-aware pulse optimization specific to the target QPU topology.
-
The system will feature an adaptive circuit partitioning engine (Layer 3/8.3) that can dynamically re-partition large quantum circuits at runtime based on live QPU calibration data (e.g., drift), ensuring optimal fidelity for the execution, analogous to JIT compilation in GPU systems but applied to quantum hardware.
-
The system will provide a unified programming model (Layer 4/8.4) that extends existing HPC paradigms (MPI/OpenMP/CUDA) with native quantum offloading primitives, allowing AI developers to write code in C++ or Fortran and invoke QPU kernels using familiar pragmas, bridging the gap between classical and quantum programming ecosystems.
This improved system can specifically perform the following:
-
Perform high-fidelity electronic structure calculations for strongly correlated molecules (e.g., materials science) by encoding fermionic wavefunctions directly into Hilbert space, overcoming the exponential memory scaling barrier of classical DFT methods.
-
Accelerate complex combinatorial optimization problems (e.g., logistics, portfolio construction) by mapping them onto Quantum Approximate Optimization Algorithm (QAOA) solvers on QPUs for near-optimal solutions that are intractable for classical solvers at scale.
-
Develop quantum machine learning models where quantum kernels are used to process high-dimensional feature spaces or enhance reinforcement learning agents (e.g., using a QPU critic network), leading to more accurate predictions in complex domains like genomics or climate modeling.
-
Execute large-scale, distributed simulations (like CFD) by partitioning the problem across multiple QPUs and classical HPC nodes, leveraging the QHPC architecture for massive parallelism and hybrid execution strategies.
Abstract
High-performance computing (HPC) has evolved over decades through multiple architectural transitions, from vector supercomputers to massively parallel CPU clusters and GPU-accelerated systems, continuously expanding the frontier of scientific discovery. With the emergence of quantum processing units (QPUs) as practical computational accelerators, a new opportunity arises to further extend this trajectory by integrating quantum and classical computing paradigms. Building on the emerging vision of Quantum Integrated High-Performance Computing (QHPC), this paper contributes a full-stack layered architecture that integrates QPUs as first-class accelerators within, not merely alongside, the classical HPC software and hardware stack, with tight, on-premise quantum-classical coupling as its defining characteristic. The architecture comprises of unified resource management, quantum-aware scheduling, hybrid workflow orchestration, middleware and programming abstraction, interconnect technologies, and a tiered execution model enabling seamless workload partitioning across classical and quantum backends. A central aspect of this architecture is a strong user requests abstraction layer that exposes heterogeneous resources through a unified job submission interface, similar in spirit to existing schedulers such as Slurm, allowing users to describe workloads in a consistent template independent of underlying compute type or location. Drawing insights from prior accelerator integration eras, we outline how QHPC can support emerging workloads in quantum chemistry, materials discovery, combinatorial optimization, and climate modeling. We conclude by highlighting open challenges in building scalable, reliable, and programmable quantum-classical infrastructures that seamlessly connect global users to heterogeneous compute resources for future quantum-classical HPC ecosystems.
Sources
- Bosonic coding: introduction and use cases
- Tour de gross: A modular quantum computer based on bivariate bicycle codes
- Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
- Testing and benchmarking emerging supercomputers via the MFC flow solver
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Platform Architecture for Tight Coupling of High-Performance Computing with Quantum Processors
- A Full Stack Framework for High Performance Quantum-Classical Computing
- Hybrid Classical-Quantum Supercomputing: A demonstration of a multi-user, multi-QPU and multi-GPU environment
- Quantum Simulations of Battery Electrolytes with VQE-qEOM and SQD: Active-Space Design, Dissociation, and Excited States of LiPF$_6$, NaPF$_6$, and FSI Salts
- How to use quantum computers for biomolecular free energies
- Quantum computing and artificial intelligence: status and perspectives
- Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
- Interfacing Quantum Computing Systems with High-Performance Computing Systems: An Overview
- Quantum resources in resource management systems
Related papers
- Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more
- Encrypted clones can leak: Classification of informative subsets in Quantum Encrypted Cloning
- Polynomial-time classical and quantum simulation of quantum impurity models
- Theory of quantum-enhanced interferometry with general Markovian light sources
- A convergent hierarchy of spectral gap certificates for qubit Hamiltonians
- Universal Bound and Phase Transition in Many-Body Fermionic Non-Gaussianity