Replay-buffer engineering for noise-aware quantum circuit optimization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.
Kai: Today's paper: "Replay-buffer engineering for noise-aware quantum circuit optimization".
Mira: Deep reinforcement learning (RL) for quantum circuit optimization faces significant bottlenecks related to experience storage, sampling, and transfer, particularly when dealing with hardware noise.
Kai: First, who's behind it and why it matters.
Title and authors: Kai: So, to summarize what the paper "Replay-buffer engineering for noise-aware quantum circuit optimization" is doing, they are proposing a new framework that uses the replay buffer strategically to solve three core problems in applying deep reinforcement learning to optimizing quantum circuits.
Mira: They propose ReaPER+ for intelligent sampling, which transitions from TD error focus early on to reliability awareness later, OptCRLQAS for making curriculum learning cheaper by amortizing expensive evaluations, and a noise-aware transfer mechanism that lets noisy training start with noiseless data.
Lev: So the main takeaway is that they treat the buffer not as a passive storage unit but as an active lever in designing the RL process itself to handle complexity and noise better during circuit design.
Kai: That’s right, and they show substantial gains: ReaPER+ gives sample efficiency gains of four-thirty-two times over fixed methods, OptCRLQAS cuts wall-clock time per episode by about sixty-seven point five percent on large tasks like twelve-qubit H2O ground state preparation, and the transfer scheme speeds up convergence significantly.
Mira: The implication here is that the way we structure our RL pipeline directly dictates how much sample efficiency we get and how robust our final optimized circuit is against the inherent noise of real quantum hardware.
Lev: If we can consistently find more compact circuits across both compilation tasks and QAS benchmarks, that translates directly into reduced resource costs for actual quantum hardware implementations.
Kai: It really shows that fixing these three levers—storage, sampling, and transfer—is decisive for making RL practical in this domain. Now we move on to how they actually improve the existing methods.
The paper's summary: Mira: The paper suggests major improvements by introducing ReaPER+, which transitions the replay rule based on training stage, moving from TD error prioritization early on to reliability-aware sampling when value estimates mature.
Kai: So it’s not just one static rule; it dynamically changes its behavior during the learning process to balance exploration with stability as the agent gets more confident in its predictions.
Lev: That dynamic adjustment based on reliability sounds critical because if we sample too much unreliable data early on, we waste time exploring bad circuit designs.
Kai: And then they have OptCRLQAS, which improves curriculum RL by re-using expensive quantum-classical evaluations across multiple architectural edits by only performing a full evaluation every m steps.
Mira: By accumulating m local modifications before triggering a single evaluation, they spread that high computational cost out over several steps without losing the signal needed for learning.
Lev: That amortization is really what makes large-scale architecture search feasible; it keeps the training tractable even when dealing with complex twelve-qubit problems.
Kai: Lastly, they propose a lightweight transfer scheme that warm-starts noisy learning by reusing noiseless trajectories directly in the target buffer without needing network weight transfer or long epsilon-greedy pretraining.
Mira: This avoids having to spend significant resources on initial exploration when we know the underlying structure of the noiseless environment is similar enough to the noisy one for this task.
Lev: If that transfer scheme consistently cuts steps needed for chemical accuracy by up to eighty-five-ninety percent on large molecular tasks, it suggests we can get very close to target accuracy much faster in noisy settings.
The paper's improvements: Kai: So, wrapping things up on "Replay-buffer engineering for noise-aware quantum circuit optimization," the paper shows that by treating the replay buffer as an active lever through ReaPER+, OptCRLQAS, and the transfer scheme, we can significantly boost sample efficiency and robustness.
Mira: The implication is that this kind of careful engineering of the RL experience pipeline directly impacts how compact and effective our final quantum circuits are when deployed on real hardware.
Lev: From my perspective, if we can reliably cut training steps by eighty-five to ninety percent for large molecular problems using this transfer method, it opens up a much more viable path toward using RL for actual noise-aware circuit optimization on near-term devices.
Kai: It really demonstrates that these three engineering levers—how we store, sample, and transfer experience—are the critical factors determining how scalable and accurate our quantum optimization becomes.
Mira: I think this work provides a very concrete roadmap for making RL methods in this space more practical by giving us specific mechanisms to handle the inherent challenges of noise and high computational cost during search.
Lev: It’s a strong foundation for moving from theoretical RL success to actual experimental results on physical quantum systems that are noisy.
Kai: We'll keep an eye on how these specific replay buffer engineering techniques translate into hardware performance in our next set of experiments, but for now, this paper gives us a lot to think about.
Conclusion: Kai: So we've talked about how ReaPER+, OptCRLQAS, and the transfer scheme in "Replay-buffer engineering for noise-aware quantum circuit optimization" are making RL practical for quantum circuit design.
Mira: Exactly, Kai; I keep thinking about how those annealing mechanisms allow the system to transition from aggressive exploration to reliable value estimation based on the maturity of those estimates.
Lev: And from a hardware standpoint, if we can get that kind of sample efficiency boost, it means we can actually run these optimization loops more often on real superconducting qubits before they decohere.
Kai: Right, Lev; and that transfer scheme is a big deal because it means we don't have to start from scratch every time we introduce noise into the target environment.
Mira: That’s the core idea; reusing noiseless experience to get an initial good coverage of the state space for noisy learning is a really smart way to tackle that initialization problem.
Lev: If that transfer scheme keeps training steps down by eighty-five percent, it drastically lowers our experimental budget and makes large-scale optimization feasible in a real lab setting.
Kai: It truly shows how treating the replay buffer as an active design element—not just a storage bin—is essential for building robust quantum compilers.
Mira: I agree; the way ReaPER+ generalizes across different reward regimes, as they showed on LunarLander-v3, suggests this approach has broader applicability than just one specific problem.
Lev: The paper's conclusions are pretty solid because they validate these ideas across different benchmarks, not just theoretical settings.
Kai: So "Replay-buffer engineering for noise-aware quantum circuit optimization" gives us concrete tools to build better compilers faster and more accurately under the constraints of real hardware noise.
Mira: It really puts the focus squarely on the data management side of RL, which is often overlooked when we're only looking at the circuit structure itself.
Lev: Moving forward, we need to see how these specific amortization and transfer techniques handle even more complex error models beyond just depolarizing noise.
Akash Kundu, Sebastian Feld
Delft University of Technology · QuTech
quant-ph, cs.AI, cs.ET, cs.LG
Submitted: 2026-04-23
Updated: 2026-09-29
Comments: Accepted at NeurIPS 2026 main track. Camera ready version
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 87/100
The gist: Deep reinforcement learning (RL) for quantum circuit optimization faces significant bottlenecks related to experience storage, sampling, and transfer, particularly when dealing with hardware noise.
Key concepts
- ReaPER+
- This is a new sampling rule that changes based on the training stage. It starts by prioritizing TD error focus early on, but switches to reliability-aware sampling later when value estimates are more mature. This dynamically balances exploration and stability during learning.
- OptCRLQAS
- This technique improves curriculum reinforcement learning by reusing expensive quantum-classical evaluations across multiple architectural edits. It only performs a full evaluation every 'm' steps, spreading the high computational cost over several local modifications to keep training tractable.
- Noise-aware transfer mechanism
- This scheme allows noisy training to start with noiseless data by directly reusing noiseless trajectories in the target buffer. This avoids needing significant resources for initial exploration when the underlying structure of the environment is similar enough to be useful.
Terminology
Summary
Deep reinforcement learning (RL) for quantum circuit optimization faces significant bottlenecks related to experience storage, sampling, and transfer, particularly when dealing with hardware noise. This paper introduces a novel replay-buffer engineering framework designed to address these issues by treating the replay buffer as a primary algorithmic lever for making RL practical in quantum optimization. The proposed methods—annealed replay (ReaPER+), amortized curriculum learning (OptCRLQAS), and lightweight noiseless-to-noisy buffer transfer—achieve substantial gains in sample efficiency, wall-clock time, and noise robustness across quantum compiling and QAS benchmarks, establishing that experience storage, sampling, and transfer are decisive levers for scalable quantum circuit optimization.
ReaPER+ Annealed Replay
The paper introduces ReaPER+, an annealed replay rule that transitions from TD error-driven prioritization early in training to reliability-aware sampling as value estimates mature. This strategy is defined by the transition:
- The replay priority at training step τ is defined as:
Ψ(+,τ)t = Rωτt(δ+t)α, where ωτ ∈ [0, 1] is a non-decreasing annealing exponent.
- The schedule for annealing is controlled by a linear schedule: ωτ = ωmin + (ωmax − ωmin) min τ/Tann, where Tann is set to half the total training budget.
This design ensures that early in training sampling is driven by TD error; later, the influence of Rt grows, biasing replay toward transitions that are both informative and reliable.
This mechanism preserves the sample-efficiency advantages of PER at the beginning while inheriting the stability of ReaPER once value estimates mature. The authors demonstrate this generalization by validating ReaPER+ on LunarLander-v3, confirming that the PER → ReaPER annealing principle generalizes across reward regimes.
OptCRLQAS Amortized Curriculum Learning
To eliminate the quantum-classical evaluation bottleneck in curriculum RL, the paper introduces OptCRLQAS. This variant reuses expensive evaluations across multiple architectural edits by accumulating m local gate modifications before triggering a single evaluation. Formally, an update indicator uτ is defined such that a full architecture evaluation is performed only when uτ = 1, where uτ = 1 if τ mod m = 0 or the episode terminates.
This amortization reduces the number of expensive quantum-classical evaluations from T in CRLQAS to approximately ⌈T /m⌉, yielding an expected reduction by a factor of about m when episode lengths are sufficiently large. Furthermore, accumulating m edits before evaluating improves the learning signal,
as single-gate modifications are often too small to produce distinguishable rewards.
Lightweight Replay-Buffer Transfer to Noise
The paper proposes a lightweight transfer scheme that warm-starts noisy-setting learning by reusing noiseless trajectories without network-weight transfer or ϵ-greedy pretraining. The target buffer B(0)tgt is initialized directly as Bsrc, where Bsrc is the replay buffer obtained after source training in a noiseless environment. This is valid because the noiseless and noisy environments share identical state and action spaces.
This buffer-only transfer provides a strong initialization for learning under noise by improving early buffer coverage. The results show that this scheme cuts the steps needed to reach chemical accuracy by up to 85-90% and improves final energy error by up to 90% over from-scratch baselines,
with the advantage growing with system size, reaching its strongest performance at 12-qubit scale.
Key Performance Results
The framework is validated across quantum compiling and QAS benchmarks. For quantum compiling, ReaPER+ achieved the best accuracy-length tradeoff overall
in 1-qubit compiling and demonstrated a 4× reduction over fixed ReaPER and PER for the ZZ(π) approximation in 2-qubit compiling. In QAS, OptCRLQAS + ReaPER+ achieved the lowest energy error across all systems while using competitive or fewer gates,
outperforming non-RL baselines like DQAS and GQAS on 5-, 6-, and 8-qubit problems. The transfer scheme showed that for 12-qubit H2O under combined depolarizing noise, it reduced steps by 88.2% and achieves the highest composite score of 51.0% across all benchmarks.
Generalization
The annealing mechanism is confirmed to be domain-agnostic through validation on LunarLander-v3, where ReaPER+ achieved a +9% AUC advantage over both baselines,
confirming that the PER → ReaPER annealing principle generalizes across reward regimes.
Furthermore, the transfer scheme's effectiveness is justified by showing that high-value source trajectories remain informative in the target task due to bounded perturbations in rewards and transition kernels.
Improvements for AI systems
Here are the specific improvements an AI system could implement based on this research, categorized by component:
)1. Improved Experience Replay Mechanism (ReaPER+):
The core improvement is replacing static or simple prioritization methods with the annealed replay rule, ReaPER+.
-
The system should dynamically transition its experience sampling strategy during training.
-
In the early stages of learning, it must prioritize transitions based on their raw Temporal Difference (TD) error magnitude (similar to Prioritized Experience Replay).
-
As the value estimates stabilize later in training, the system must gradually increase
reliability weighting
by transitioning from PER to ReaPER. This means it should down-weight transitions that have large TD errors but whose downstream targets are deemed unreliable (i.e., those with low reliability scores, indicating high variance or poor target quality). -
The annealing schedule must be controlled by a single, interpretable timescale parameter (Tann), allowing the system to recover the aggressive exploration of PER initially while smoothly transitioning to the stability and robustness of ReaPER later.
-
The resulting AI system will exhibit:
-
Significantly improved sample efficiency (4x–32x gains) over fixed replay buffers.
-
Greater circuit compactness in quantum compilation tasks, as it learns to focus on high-fidelity transitions that lead to shorter circuits.
-
Enhanced robustness against hardware noise, as it learns to trust the value estimates of the Q-function more reliably during later training phases.
)2. Amortized Curriculum Learning (OptCRLQAS):
The system should implement an amortized evaluation strategy for curriculum-based architecture search (QAS).
-
Instead of triggering a full quantum-classical evaluation (variational optimization and cost estimation) at every architectural edit step, the system should accumulate a set number of local modifications (m edits, e.g., m=10 or 15) before performing a single, expensive evaluation.
-
This amortized approach spreads the high computational cost across multiple steps.
-
The resulting AI system will exhibit:
-
A significant reduction in wall-clock time per episode (up to 67.5% faster) during architecture search tasks (like molecular ground state preparation).
-
Scalability to larger qubit systems (12+) without prohibitive GPU compute requirements, making large-scale quantum circuit optimization feasible in near-term settings.
)3. Lightweight Noiseless-to-Noisy Buffer Transfer:
The system must incorporate a mechanism to warm-start learning in a noisy environment using experience collected from a noiseless source environment.
-
The system should maintain two replay buffers: one for the source (noiseless) task and one for the target (noisy) task.
-
When training begins on the noisy target, it should initialize its target buffer directly with the contents of the source buffer, copying transitions without modifying them or transferring network weights.
-
The resulting AI system will exhibit:
-
A substantial reduction in training steps required to achieve chemical accuracy (up to 85–90% faster) and improvement in final energy error (up to 90%) on noisy tasks.
-
Accelerated convergence, as the agent starts with a rich distribution of high-quality, noiseless experience relevant to the underlying circuit structure.
)Overall System Capabilities:
By integrating these three levers, the improved AI system can perform:
-
Quantum Circuit Compilation and Architecture Search with superior efficiency and accuracy on near-term hardware.
-
Robustness against hardware noise through intelligent experience reuse (transfer).
-
Scalable operation across increasing qubit counts due to amortized evaluation techniques.
Abstract
Deep reinforcement learning for quantum circuit optimization faces three bottlenecks: replay buffers that overlook temporal difference (TD) target reliability, curriculum-based architecture search requiring a full quantum-classical evaluation after every edit, and the discard of noiseless trajectories when retraining under hardware noise. We address these limitations by treating replay as a central algorithmic lever. We introduce ReaPER+, an annealed replay rule that transitions from TD-error prioritization to reliability-aware sampling as value estimates mature. ReaPER+ achieves up to 4x higher sample efficiency than fixed PER, ReaPER, and uniform replay, while matching prior on-policy solution quality with up to 32x fewer interactions At 12 qubits, fixed ReaPER reaches the lowest energy error in the fewest steps, while PER and uniform replay find more compact circuits at higher error. On tasks scaling to 20 qubits, ReaPER+ retains its advantage, demonstrating that reliability-aware annealing extends beyond small-system benchmarks. LunarLander-v3 confirms that the ReaPER+ is domain-agnostic, it improves success rates by up to 26.8% over PER and 21.8% over fixed ReaPER, with a 3% AUC gain over both. We further introduce OptCRLQAS, which amortizes quantum-classical evaluations across multiple architectural edits, reducing training wall-clock time by up to 67.5% on 12-qubit without degrading solution quality. Finally, lightweight replay-buffer transfer warm-starts noisy optimization from noiseless trajectories, without weight transfer or ε-greedy pretraining, reducing steps to chemical accuracy by 85-90% and final energy error by up to 90% relative to from-scratch learning. Transfer gains increase with system size. Together, these results establish experience storage, sampling, and transfer as decisive levers for sample efficient, noise-aware quantum circuit optimization.
Sources
- A Quantum Approximate Optimization Algorithm
- Enabling Technologies for Scalable Superconducting Quantum Computing
- Myths around quantum computation before full fault tolerance: What no-go theorems rule out and what they don't
- Reinforcement Learning for Quantum Technology
- Quantum Compiling with Reinforcement Learning on a Superconducting Processor
- Reinforcement learning-assisted quantum architecture search for variational quantum algorithms
- Playing Atari with Deep Reinforcement Learning
- Prioritized Experience Replay
- GA4QCO: Genetic Algorithm for Quantum Circuit Optimization
- Quantum circuit optimization with deep reinforcement learning
- Quantum Architecture Search via Deep Reinforcement Learning
- Practical and efficient quantum circuit synthesis and transpiling with Reinforcement Learning
- Proximal Policy Optimization Algorithms
- One-Step Bellman Alignment Enables Provably Efficient Transfer in Online RL
- Gymnasium: A Standard Interface for Reinforcement Learning Environments
Related papers
- Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more
- Encrypted clones can leak: Classification of informative subsets in Quantum Encrypted Cloning
- Polynomial-time classical and quantum simulation of quantum impurity models
- Theory of quantum-enhanced interferometry with general Markovian light sources
- A convergent hierarchy of spectral gap certificates for qubit Hamiltonians
- Universal Bound and Phase Transition in Many-Body Fermionic Non-Gaussianity