Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.
Kai: I'm Kai, and with me are Mira and Lev, guest researcher.
Mira: Today's paper: "Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift".
Kai: The gist The authors investigate adaptive eavesdropping as a constrained Markov decision process to quantify how much an attacker gains by adapting to channel drift in quantum key distribution.
Mira: First, who's behind it and why it matters.
Title and authors: Kai: So, looking at this paper's title, "Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift," it really captures the essence of the problem they're addressing. It’s not just about static noise anymore; it’s about how channel drift affects security in real-world QKD systems.
Mira: That's right. The authors are Marcel Mordarski, Benjamin Gras, Abdelrahman Shehata, Daniel Budina, and Roberto Bondesan. They are setting up this framework to move beyond the old security proofs based on stationary channels that we've relied on for years in the field.
Lev: What does this mean for us right now? It means that if you’re building a QKD system, you can't just assume the noise stays constant between calibrations; you have to account for this drift actively.
Kai: So, what is the paper actually proposing as a solution to this uncertainty? It’s framing adaptive eavesdropping as a constrained Markov decision process where the attacker selects one circuit per round while the noise level follows an Ornstein–Uhlenbeck process and the abort condition is a budget over each block of rounds.
Mira: They are essentially modeling the eavesdropper as an active agent who picks her strategy based on what she sees, but this choice is constrained by how much she can afford to test or observe in any given period.
Lev: This moves us from a simple static security analysis to a dynamic one where the attacker's actions and the environment's noise are both changing over time.
Kai: And they’re using reinforcement learning concepts here, specifically a value-based agent featuring double and duelling value estimation, n-step returns, and prioritized replay to solve that sequential decision problem.
Mira: The goal of this agent is to maximize the objective function Jλ(π) which balances the expected information gain against the risk of detection defined by PrDπ, where Dπ represents the event of detection.
The paper's summary: Kai: So, let's summarize what they found regarding these adaptive attacks under drift. They extended previous work by searching for gate structures and rotation angles jointly to find circuits that are compact enough to form a discrete action set for the attacker.
Mira: That search is quite detailed; it involves an evolutionary outer loop proposing gate sequences through moves that add, remove or replace one gate, and an inner loop tuning the angles using a gradient-free optimiser at a fixed evaluation budget.
Lev: This entire procedure returns a circuit that achieves a stated leak at a stated disturbance, which is crucial because it links the physical capabilities to the security metrics we care about.
Kai: Then they perform another departure by replacing the stationary channel with an Ornstein–Uhlenbeck process, transforming it into this constrained Markov decision process for analysis.
Mira: The paper highlights that comparing the best fixed circuit against a dynamic programming upper bound is how you quantify the value of adaptation, which tells network operators exactly how much information an attacker gains by adapting to channel drift.
Lev: This comparison is vital because it lets us see the actual performance gain of adaptation versus just sticking with a static assumption.
Kai: They also mapped attainable information across different channel families at a fixed abort threshold, showing that the sign of the attacker’s gain from basis asymmetry depends on whether an error-rate constraint is averaged over both bases or applied separately.
The paper's improvements: Mira: The paper suggests a few key improvements to this initial formulation. They focus on making the search more general by extending construction to noise models that lack a known template, such as the amplitude damping channel.
Lev: That’s significant because it means their method isn't tied down by the assumption that we already have a specific noise model in mind for every system we analyze.
Kai: They also explore how they can find circuits compact enough to form a discrete action set by searching jointly for gate structures and rotation angles, which is more flexible than just fixing one of those parameters.
Mira: This joint search allows them to find more general circuit topologies, which is important because the topology itself could be an unknown on the same footing as the angles in their formulation.
Lev: And they validate this entire construction against known closed forms and against a reliable numerical lower bound, which shows that what they’ve constructed is consistent with analytical bounds of established methods.
Kai: This validation allows them to construct explicit attacks that provide rigorous lower bounds beneath formal security proofs, which is exactly what we need for practical provisioning margins.
Conclusion: Mira: So, to wrap up on this paper's main implications, the work demonstrates that drift dynamically dictates how much information an attacker can extract in a QKD system. For device-independent E91, an attacker adapting across twenty-six compact circuits holds two point six times the Holevo information of the best single circuit at zero detection.
Lev: That factor of two point six is substantial when you think about how much extra information an eavesdropper can get without increasing their risk of being detected in this scenario. It underscores why static assumptions are becoming insufficient for robust security guarantees.
Kai: And they also showed that the chosen monitoring convention has a comparable impact; enforcing individual per-basis thresholds rather than an averaged threshold lowers the attacker’s attainable fidelity by zero point zero two three, which is another practical lever for operators to pull.
Mira: Ultimately, this work offers explicit attacks that serve the defender by providing empirically testable provisioning margins, which are validated against formal security proofs that establish upper bounds on the eavesdropper’s information.
Lev: We've seen how they use this constrained Markov decision process to create a constructive lower bound that is close to what certified methods can achieve and validates theoretical limits such as the device-independent collective-attack rate. It’s a solid step toward making security metrics more concrete for deployment.
Kai: So, the paper "Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift" gives us a way to quantify the impact of channel drift on security, showing that adaptation provides an advantage in this model without necessarily increasing detection risk. It’s a key piece of work as we look toward deploying these quantum systems.
Marcel Mordarski, * Benjamin Gras† Abdelrahman Shehata† Daniel Budina† Roberto Bondesan
Department of Computing, Imperial College London · Department of Electrical and Electronic Engineering, Imperial College London
quant-ph, cs.CR, cs.IT, cs.LG, math.IT
Submitted: 2026-10-01
Updated: 2026-10-01
Comments: Presented as submission 202 at QCrypt 2026 qcrypt.net/2026/technical/accepted-papers/. A parallel work exploring the machine-learning aspects of this approach, titled "Sparsity for Free: A Budget-Induced Equilibrium in Joint Topology-Parameter Search'', has been accepted for NeurIPS 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 90/100
The gist: The gist The authors investigate adaptive eavesdropping as a constrained Markov decision process to quantify how much an attacker gains by adapting to channel drift in quantum key distribution.
Key concepts
- Constrained Markov Decision Process
- This frames an eavesdropper's strategy as a sequence of decisions where she chooses an action (a circuit) at each step, but her choices are restricted by a budget or constraint that must be met over multiple rounds. The goal is to maximize the information gained while adhering to these round-by-round limitations.
- Channel Drift
- This refers to the gradual, time-dependent change in the quantum channel's characteristics, such as noise levels or rotation angles. Unlike a static channel where the security analysis is simple, drift means an eavesdropper can exploit these changes by adjusting her strategy sequentially to keep gaining information.
- Adaptive Eavesdropping
- This is when an attacker does not use a single fixed strategy but instead learns and changes her action (e.g., which circuit to use) based on the results of previous rounds. The paper models this adaptation using reinforcement learning techniques to find the best sequence of actions that maximizes her expected information gain.
- Ornstein–Uhlenbeck Process
- This is a mathematical model used here to describe the time evolution of the channel noise. It represents a stochastic process where the noise level changes randomly but tends to revert towards a certain mean, simulating realistic fluctuations in quantum communication channels over time.
Terminology
Summary
The gist The authors investigate adaptive eavesdropping as a constrained Markov decision process to quantify how much an attacker gains by adapting to channel drift in quantum key distribution.
Problem Construction
The analysis addresses the gap between security proofs based on stationary channels and the reality of device drift, where an eavesdropper can change her strategy round-to-round A drifting channel poses a sequential question in addition, since an eavesdropper who reads the public record may change her strategy from round to round, must commit to each round’s strategy before that round’s statistic is announced, and seeks the largest expected information over the trajectories of the channel subject to a constraint in every round The margin carried across a recalibration interval therefore rests on an untested assumption about the attacker, which leaves the user unable to tell whether a link is over-provisioned, at a loss of secret-key-rate, or exposed to a leak that a stationary trusted-noise analysis does not count
Methodology and Search
The learnt formulation extends previous work by searching for gate structures and rotation angles jointly, yielding circuits compact enough to form a discrete action set The search involves an evolutionary outer loop proposing gate sequences through moves that add, remove or replace one gate and an inner loop tuning the angles by a gradient-free optimiser at a fixed evaluation budget. This procedure returns a circuit that achieves a stated leak at a stated disturbance The second departure replaces the stationary channel by replacing it with an Ornstein–Uhlenbeck process, making adaptive eavesdropping into a constrained Markov decision process
Results on Adaptation
The value of adaptation is bounded by comparing the best fixed circuit and a dynamic-programming upper bound. On device-independent E91 under bilateral depolarising noise, a reinforcement-learning attacker raises her Holevo information from 0.135 for the best fixed circuit to 0.348 at zero detection, 98% of the upper bound On BB84 under a drifting bit-flip channel, she exceeds a conservative noise-indexed rule by 0.024 in fidelity, reaching 99% of the upper bound On E91 adaptation multiplies the attacker’s information by about 2.6 at unchanged detection.
Key Findings and Comparisons
The search maps attainable information across channel families at a fixed abort threshold, showing that the sign of the attacker’s gain from basis asymmetry depends on whether an error-rate constraint is averaged over both bases or applied separately The results confirm that adaptive eavesdropping under drift provides a significant advantage; for E91 adaptation multiplies the attacker’s extracted information by approximately 2.6 without increasing detection risk.
Practical Consequences
The study demonstrates that drift dynamically dictates the amount of information an attacker can extract, and for device-independent E91, an attacker adapting across 26 compact circuits holds 2.6 times the Holevo information of the best single circuit at zero detection. The chosen monitoring convention also has a comparable impact; enforcing individual per-basis thresholds rather than an averaged threshold lowers the attacker’s attainable fidelity by 0.023.
Conclusion
Explicit attacks serve the defender by providing empirically testable provisioning margins, which are validated against formal security proofs that establish upper bounds on the eavesdropper’s information. The constructive lower bound provided by this search is closest in approach to certified methods and validates theoretical limits such as the device-independent collective-attack rate.
How it works
-
The process frames adaptive eavesdropping as a constrained Markov decision process where the attacker selects one circuit per round while the noise level follows an Ornstein–Uhlenbeck process.
-
The search explores gate structures and rotation angles jointly, extending construction to noise models lacking a known template, including the amplitude damping channel.
-
The adaptive agent solves the sequential decision problem by maximizing an objective function subject to an abort condition defined by a budget over each block of rounds
How it works
The adaptive agent is a value-based reinforcement learning agent featuring double and duelling value estimation, n-step returns, and prioritised replay. It employs invalid-action masking to dynamically exclude circuits that would drop the ten-round mean below the threshold.
How it works
The search is validated against known closed forms and against a reliable numerical lower bound, showing that the constructed attack extracts information consistent with the analytical bounds of established methods. This allows for the construction of explicit attacks that provide rigorous lower bounds beneath formal security proofs.
Improvements for AI systems
-
Adaptive Eavesdropping under Drift via Constrained Markov Decision Process: The improved system can model an eavesdropper selecting one circuit per round while navigating a noise level following an Ornstein–Uhlenbeck process, optimizing for information gain subject to a budget constraint over each block of rounds, as described in the paper's description of the
constrained Markov decision process.
-
Joint Gate Structure and Rotation Angle Optimization: The system can search jointly for both gate structure and rotation angles to find
circuits compact enough to form a discrete action set,
extending construction to noise models lacking a known template, such as the amplitude damping channel. -
Dynamic Strategy Selection for Drifting Channels: The improved agent can solve the sequential decision problem posed by adaptive eavesdropping under drift by using a value-based reinforcement learning agent that maximizes
Jλ(π) = E
X T t=1 r(at, pt) - λ Pr[Dπ]," where Dπ represents the event of detection. -
Quantifying Adaptation Value: The system can determine the
value of adaptation
by comparing the performance of an adaptive attacker against abest fixed circuit and with a dynamic-programming upper bound,
which is crucial for network operators to understand how much information an attacker gains by adapting to channel drift. -
Basis Asymmetry Sensitivity Analysis: The system can assess how the
sign of the attacker’s gain from basis asymmetry changes sign between an averaged and a per-basis error-rate constraint,
allowing for more nuanced security provisioning based on monitoring conventions.
Abstract
Quantum key distribution (QKD) links are provisioned from security analyses of stationary channels, whereas the devices that determine the channel drift between recalibrations. Whether an eavesdropper who cannot alter the channel's own noise gains by following that drift has not been quantified. Adaptive eavesdropping is posed here as a constrained Markov decision process in which the attacker selects one circuit per round while the noise level follows an Ornstein--Uhlenbeck process and the abort condition is a budget over each block of rounds. The value of adaptation is bounded by the best fixed circuit and a dynamic-programming upper bound. The actions are learnt attacks. Whereas Decker et al. trained a parametrised circuit on a fixed gate template against a fixed channel, here the gate structure and rotation angles are searched jointly. This yields circuits compact enough to form a discrete action set, extending the construction to noise models lacking a known template, including the amplitude damping channel. On device-independent E91 under bilateral depolarising noise, a reinforcement-learning attacker raises her Holevo information from 0.135 for the best fixed circuit to 0.348 at zero detection, 98% of the upper bound. On BB84 under a drifting bit-flip channel, she exceeds a conservative noise-indexed rule by 0.024 in fidelity, reaching 99% of the upper bound. Under stationary noise, the attacker's gain from basis asymmetry changes sign between an averaged and a per-basis error-rate constraint. The search, started from random gate sequences, recovers the analytical cloners and the collective-attack key rate, and meets the lower bound of the Winick--Lütkenhaus--Coles objective from above.
Sources
- Finite-Key Analysis of Quantum Key Distribution with Characterized Devices Using Entropy Accumulation
- Reliable numerical key rates for quantum key distribution
- A convergent hierarchy of semidefinite programs characterizing the set of quantum correlations
- Device-independent lower bounds on the conditional von Neumann entropy
- Trusted detector noise analysis for discrete modulation schemes of continuous-variable quantum key distribution
- QKD as a Quantum Machine Learning task
- Optimal Eavesdropping in Quantum Cryptography. I
- Optimal Eavesdropping in Quantum Cryptography. II. Quantum Circuit
- Pauli Cloners for Pauli Channels
- Simple Proof of Security of the BB84 Quantum Key Distribution Protocol
- Tight asymptotic key rate for the BB84 protocol with local randomisation and device imprecisions
- Lower and upper bounds on the secret key rate for QKD protocols using one--way classical communication
- An information-theoretic security proof for QKD protocols
- Upper bound on the secret key rate distillable from effective quantum correlations with imperfect detectors
- Security of device-independent quantum key distribution protocols: a review
- Upper bounds on key rates in device-independent quantum key distribution based on convex-combination attacks
- Bell nonlocality is not sufficient for the security of standard device-independent quantum key distribution protocols
- Variational Quantum Cloning: Improving Practicality for Quantum Cryptanalysis
- Adversarial Reinforcement Learning for Adaptive Eavesdropping in BB84 Quantum Key Distribution
- Regularized Evolution for Image Classifier Architecture Search
Related papers
- Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more
- Encrypted clones can leak: Classification of informative subsets in Quantum Encrypted Cloning
- Polynomial-time classical and quantum simulation of quantum impurity models
- Theory of quantum-enhanced interferometry with general Markovian light sources
- A convergent hierarchy of spectral gap certificates for qubit Hamiltonians
- Universal Bound and Phase Transition in Many-Body Fermionic Non-Gaussianity