Quantum Integrated High-Performance Computing: Envisioning a Layered Architecture for Next-Generation Hybrid Computing Infrastructure
summary
The gist
High-performance computing (HPC) has evolved through multiple architectural transitions, and this paper proposes Quantum Integrated High-Performance Computing (QHPC), a visionary architectural
In short
This proposal outlines Quantum Integrated High-Performance Computing (QHPC), a framework unifying CPUs, GPUs, FPGAs, and QPUs as primary resources. It introduces a layered architecture with systems for workflow management and resource scheduling to seamlessly integrate quantum computation into existing HPC workflows.
Key concepts
- QHPC Architecture
- A visionary design that treats CPUs, GPUs, FPGAs, and QPUs as equal components in a single system. The goal is to create a tightly coupled environment where quantum processing units work directly alongside classical accelerators for complex scientific problems.
- Workflow Management System (WMS)
- The control plane that organizes hybrid jobs into task graphs (DCTG). It analyzes dependencies to determine which tasks need fast execution, such as those requiring quick responses from a QPU, and routes them accordingly.
- Multi-Tier Resource Hierarchy
- A system for organizing physical hardware into four levels (R1 to R4) based on speed and integration. This allows the system to place jobs optimally: for example, placing latency-critical tasks on tightly connected CPU+GPU+QPU nodes (R3).
- Quantum Suitability Score (QSS)
- A metric used by the scheduler to decide which QPU is best for a given job. It combines factors like how reliable the quantum gates are, how well it connects to other hardware, and its current waiting time.
Terminology used across episodes
This episode discusses
- Quantum Integrated High-Performance Computing: Envisioning a Layered Architecture for Next-Generation Hybrid Computing Infrastructure · Paper Radio
- Bosonic coding: introduction and use cases
- Tour de gross: A modular quantum computer based on bivariate bicycle codes
- Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
- Testing and benchmarking emerging supercomputers via the MFC flow solver
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Platform Architecture for Tight Coupling of High-Performance Computing with Quantum Processors
- A Full Stack Framework for High Performance Quantum-Classical Computing
- Hybrid Classical-Quantum Supercomputing: A demonstration of a multi-user, multi-QPU and multi-GPU environment
- Quantum Simulations of Battery Electrolytes with VQE-qEOM and SQD: Active-Space Design, Dissociation, and Excited States of LiPF 6, NaPF 6, and FSI Salts
- How to use quantum computers for biomolecular free energies
- Quantum computing and artificial intelligence: status and perspectives
- Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
- Interfacing Quantum Computing Systems with High-Performance Computing Systems: An Overview
- Quantum resources in resource management systems
The paper
Quantum Integrated High-Performance Computing: Envisioning a Layered Architecture for Next-Generation Hybrid Computing Infrastructure · Read on arXiv
Suman Raja, Siva Sai, Yogesh Simmhanb, Kyle Charda, Rajkumar Buyyad
Department of Computer Science, University of Chicago · Department of Computational and Data Sciences, Indian Institute of Science Bengaluru · Department of Electrical and Computer Engineering, National University of Singapore · Quantum Cloud Computing and Distributed Systems (qCLOUDS) Lab, School of Computing and Information Systems, The University of Melbourne
High-performance computing (HPC) has evolved over decades through multiple architectural transitions, from vector supercomputers to massively parallel CPU clusters and GPU-accelerated systems, continuously expanding the frontier of scientific discovery. With the emergence of quantum processing units (QPUs) as practical computational accelerators, a new opportunity arises to further extend this trajectory by integrating quantum and classical computing paradigms. Building on the emerging vision of Quantum Integrated High-Performance Computing (QHPC), this paper contributes a full-stack layered architecture that integrates QPUs as first-class accelerators within, not merely alongside, the classical HPC software and hardware stack, with tight, on-premise quantum-classical coupling as its defining characteristic. The architecture comprises of unified resource management, quantum-aware scheduling, hybrid workflow orchestration, middleware and programming abstraction, interconnect technologies, and a tiered execution model enabling seamless workload partitioning across classical and quantum backends. A central aspect of this architecture is a strong user requests abstraction layer that exposes heterogeneous resources through a unified job submission interface, similar in spirit to existing schedulers such as Slurm, allowing users to describe workloads in a consistent template independent of underlying compute type or location. Drawing insights from prior accelerator integration eras, we outline how QHPC can support emerging workloads in quantum chemistry, materials discovery, combinatorial optimization, and climate modeling. We conclude by highlighting open challenges in building scalable, reliable, and programmable quantum-classical infrastructures that seamlessly connect global users to heterogeneous compute resources for future quantum-classical HPC ecosystems.
Transcript
Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.
Kai: Today's paper: "Quantum Integrated High-Performance Computing".
Mira: High-performance computing (HPC) has evolved through multiple architectural transitions, and this paper proposes Quantum Integrated High-Performance Computing (QHPC), a visionary architectural framework that unifies CPUs, GPUs, FPGAs,
Kai: First, who's behind it and why it matters.
Title and authors: Mira: So, diving deeper into what the paper actually lays out, it explains that the QHPC model envisions a next-generation hybrid computing architecture where QPUs are tightly coupled with conventional CPUs and GPUs under a single workflow and resource management framework. It’s about creating a cohesive system rather than just bolting on quantum capabilities.
Kai: That cohesive vision is what makes it compelling; they describe the architecture across five specific layers, starting from the user requests down to the physical compute layer where all these resources live together. It’s a very structured approach to solving this integration problem.
Lev: Structure is important, but how does this workflow management system actually handle the scheduling of tasks when you have nodes with such diverse latency characteristics? I worry that any static scheduling will fail immediately with QPU variability.
Mira: The Workflow Management System, or WMS, models a hybrid workload as a Directed Cyclic Task Graph, or DCTG. This allows them to decompose complex tasks into graphs and analyze dependencies to route jobs based on whether they are latency-critical or latency-tolerant paths.
Kai: That’s the mechanism for dynamic routing; they identify those critical paths so that, for example, a job needing sub-millisecond round-trip times can be sent to co-located QPU and CPU resources. It’s about intelligent path selection.
Lev: If the WMS is doing this semantic dependency analysis, what kind of metrics are they using to determine which path is critical versus tolerant? That sounds like a massive input requirement for the scheduler.
Mira: They use that analysis to decide where the workload goes, essentially routing based on timing constraints derived from the problem's nature within the DCTG model. It’s about matching latency requirements with available hardware tiers.
Kai: And then there’s the Resource Management System, or RMS, which manages everything through a Unified Resource Registry that organizes resources into four distinct tiers, R1 through R4. That tiered hierarchy is crucial for understanding where a job can physically land.
Lev: Tiered resource management sounds robust because it acknowledges the physical reality of different access speeds and connectivity constraints between classical nodes and remote QPUs.
Mira: It does; they define R3 specifically as tightly integrated CPU, GPU, and QPU nodes connected by low-latency interconnects like NVLink. That shows they are prioritizing hardware proximity for certain tasks.
Kai: And then the Quantum-Aware Scheduler within the RMS uses something called a Quantum Suitability Score, or QSS, to pick the right QPU for a job based on fidelity and access latency. It’s a specific metric tailored to quantum hardware quality.
Lev: A score combining gate fidelity, connectivity compatibility, queue wait time, and access latency sounds like exactly what you need when dealing with real-world noise issues; it accounts for the hardware imperfections directly in the decision process.
The paper's summary: Kai: Now we look at how they propose improvements to this QHPC concept, and it really focuses on making the abstraction layer more powerful. They suggest a middleware and programming abstraction layer that handles hardware-agnostic programming, circuit compilation, error mitigation, and communication protocols.
Mira: That’s where the QIR comes in; they lower quantum circuits into a Quantum Intermediate Representation which is LLVM-based. This gives them portability for optimization across different types of QPUs without having to rewrite everything from scratch for each new device.
Lev: A portable IR is nice, but what about the compilation pipeline itself? How do they manage noise and calibration data during that process when you’re dealing with physical systems that drift over time?
Kai: They detail a multi-stage compilation pipeline that includes logical optimization, device mapping with noise-aware placement based on calibration data, and even pulse-level optimization specific to superconducting QPUs. It’s a very fine-grained control mechanism.
Mira: I think the error mitigation selection part is also important because they build in choices for how to handle noise during the compilation phase itself, rather than just applying fixes afterwards at runtime. That’s proactive management of errors.
Lev: If they are doing pulse-level optimization, that suggests a very deep understanding of the physical layer—things like specific microwave pulses used in superconducting systems—which is necessary for reliable execution on real hardware.
Kai: Furthermore, they tackle the "Classical–Quantum Communication Protocol Stack," which uses NVLink for co-located systems and gRPC over InfiniBand for remote access, tying the software stack directly to physical interconnects.
Mira: That protocol stack integration is crucial because it ensures that the communication overhead between classical and quantum components doesn't become the bottleneck in this hybrid workflow. It’s about minimizing data transfer latency.
Lev: Minimizing transfer latency is a huge factor; if we spend too much time moving data between the CPU/GPU and the QPU, even a fast scheduler can’t compensate for that.
Kai: The paper also suggests an adaptive circuit partitioning engine at runtime, which means they can dynamically re-partition large quantum circuits based on live QPU calibration data to maintain optimal fidelity during execution.
Mira: That sounds like a form of Just-In-Time compilation applied to quantum hardware; it allows the system to react instantly if the physical state of the QPU changes unexpectedly.
The paper's improvements: Kai: So, wrapping up the QHPC paper, they are proposing this layered architecture—with its unified management, quantum-aware scheduling, and sophisticated middleware—to treat QPUs as true first-class resources in hybrid computing. It’s a blueprint for how we could actually build these next-generation systems.
Mira: The implication is that we move toward a system where the complexity of integrating different quantum hardware components is managed by intelligent software orchestration, rather than just relying on ad hoc connections between separate systems. This makes large-scale hybrid simulation more viable.
Lev: I think the impact here lies in making it feasible to run complex, high-fidelity simulations that we currently can't touch because of memory constraints or classical intractability, provided the hardware actually delivers on its promised fidelity.
Kai: Exactly; and they point toward future work like developing energy-efficient hybrid infrastructures and extending FPGAs to act as real-time quantum system controllers for error correction decoders operating within microsecond coherence windows.
Mira: That points toward a future where heterogeneous compute resources operate as a unified, programmable continuum, which is a big conceptual step forward in how we think about distributed computation.
Lev: For me, the main challenge remains ensuring that this theoretical framework translates into practical execution on real hardware without introducing too much overhead or instability during the actual quantum operations.
Kai: That’s the million-dollar question for any experimentalist; we need to see if this unified model can actually handle the physical reality of noise and variability.
Mira: It sounds like a necessary step toward extending scientific discovery by systematically tackling the architectural fragmentation that has historically kept quantum computation isolated from classical HPC workflows.
Lev: We'll have to watch how quickly the community moves from this paper’s vision to building systems that can handle those tight latency couplings they described.
Conclusion: Kai: So, to wrap up, this paper on "Quantum Integrated High-Performance Computing: Envisioning a Layered Architecture for Next-Generation Hybrid Computing Infrastructure" really lays out a blueprint for how we can actually build these next-generation hybrid systems by treating CPUs and QPUs as first-class resources under one unified framework.
Mira: It’s fascinating because it tackles the fundamental challenge of making architectural heterogeneity sustainable when you want to scale up scientific discovery, moving away from siloed computing solutions.
Lev: I mean, the concept of that layered approach, from the workflow management system down to the physical compute layer, seems incredibly thorough for tackling real-world implementation hurdles.
Kai: It is really comprehensive; they show how a unified resource management and quantum-aware scheduling system can handle complex hybrid workloads by decomposing them into task graphs.
Mira: And I think their focus on the middleware layer that uses a Quantum Intermediate Representation to handle noise-aware compilation is where the real theoretical heavy lifting happens, ensuring that optimization isn't just classical guesswork.
Lev: From my side, I’m interested in how they define those "latency-critical paths" within the DCTG model because if we can truly manage those timing constraints across tiers like R3 and R4, it makes running actual error-correction routines much more realistic.
Kai: It suggests that the immediate impact is a much more structured way for AI developers to offload hybrid tasks, using familiar programming models while getting hardware-specific optimizations built in from the start.
Mira: The implication is that we might finally see practical applications in areas like high-fidelity molecular simulations or complex optimization problems where classical methods hit a wall due to exponential complexity.
Lev: If they can manage those dynamic re-partitioning capabilities, it opens up possibilities for running larger distributed simulations across multiple QPUs and classical nodes effectively.
Kai: Ultimately, the paper provides the architectural vision needed to bridge the gap between theoretical quantum potential and tangible high-performance computing reality.
Mira: It’s a strong foundation for thinking about how we should approach programming and resource allocation in this new era of hybrid computation.
Lev: We certainly need to keep an eye on how they address those open challenges, especially the issues around facility co-location and managing QPU variability across different sites.
More episodes
- 2610.01068-Learned Parallel Bit-Flipping Sequential Belief Propagation Decoding of Quantum LDPC Codes
- 2610.01074-The stationarity test: a framework for learning quantum many-body systems from their thermal states
- 2610.01094-Quantum synchronization in atom-cavity coupled systems
- 2610.01402-Transport theory for a generic two-arm co-propagating Majorana interferometer with Majorana fermion and edge vortex tunneling
- 2610.01167-Vector chiral order and dynamical quantum phase transitions in an Ising chain with dimerized anisotropic Gamma interaction
- 2610.01163-Robustness hierarchy of bipartite quantum correlations under noisy dynamics
- 2610.01183-Additive solid immersion lenses for enhanced collection efficiency of shallow NV centers by pulsed laser deposition and structurization of high-k amorphous oxides
- 2610.01112-Dissipation-Sensitivity Trade-Off in Dissipative Bosonic Systems
- 2610.01099-Constant-Per-Layer-Depth MPS-Pretrained Ansatz for Noisy Distributed Quantum Processors
- 2610.01141-Classical Hardness of Learning Functions of Hamiltonians