Anytime-valid detection of LLM weight exfiltration
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Anytime-valid detection of LLM weight exfiltration".
Elias: The gist The e-process introduces a prompt-level mechanism that calibrates whole-response mismatch events on trusted benign traffic while accumulating evidence sequentially under calibration transfer assumptions,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we’re looking at the paper "Anytime-valid detection of LLM weight exfiltration," and what they're doing is trying to find a way to detect when an AI is leaking its model weights.
Elias: It’s about this idea that you can verify a compromised AI server by checking if it produces different tokens when you replay the same prompt, but the paper focuses on making that detection anytime valid.
Nadia: That means we don't have to wait for a perfect moment; we get evidence sequentially as responses come in, which is pretty important for real-world security testing.
Elias: The authors are Ines Ortega-Fernandez and Mateusz Kowalczyk, and they’re using this prompt-level e-process to do the calibration.
Nadia: So the core idea is that they're calibrating whole response mismatch events on benign traffic while keeping an eye on evidence under a calibration transfer assumption.
Elias: It’s about how sequential evidence accumulation helps them control the false alarm probability over an unbounded monitoring horizon, which is a big technical step for this kind of detection.
Nadia: So, what does that actually mean for us when we think about how hard it is to exploit these models?
Elias: It suggests that patient attackers who hide in normal variation won't be able to stay hidden indefinitely if you combine the evidence across multiple responses, which is a significant improvement over just looking at one token mismatch.
Priya: From a privacy perspective, I’m interested in how they handle the calibration transfer assumption because that touches on how much we can trust our benign traffic samples.
Nadia: Exactly, Priya, and the paper seems to address that by making sure the process runs indefinitely while keeping false alarms below a set rate.
The paper's summary: Elias: To summarize what the "Anytime-valid detection of LLM weight exfiltration" paper is doing, it introduces a prompt-level e-process that computes nested binary events Aj,m after each response j.
Nadia: These events have to be fixed before their benign rate is calibrated, and they define a margin event Aj,m as when there's a mismatch and the gap between logits is at least bm for some threshold index m.
Elias: The key part is that they calibrate these events using n new benign calibration responses to estimate the benign rate q+m for each event m.
Nadia: Instead of just taking a simple fraction like km over n, they use the Clopper–Pearson upper bound q+m, which ensures all K bounds cover their rates together with probability at least one minus gamma <ref:2610.11843#pg1>.
Elias: This gives them a mathematical guarantee that their calibration covers all the relevant events robustly at once, which is a way of controlling the uncertainty in their estimates.
Nadia: And then they accumulate evidence sequentially by multiplying by a factor Ej,m equals one plus lambda j,m times the difference between the event and its benign rate.
Elias: That accumulation process ensures each component becomes a nonnegative supermartingale because you choose lambda from earlier responses only and cap it at c over q+m.
Priya: So, what does this sequential evidence accumulation actually tell us about the attack? Does it help in finding something that a single token mismatch wouldn't?
Nadia: It lets them track sustained excess across responses, which means an attacker can't just have one bad response and disappear; they have to maintain that pattern.
The paper's improvements: Nadia: The paper points out some key improvements over previous methods, specifically moving from a hard per-token alarm to this e-process approach.
Elias: It combines weak evidence across responses while providing explicit anytime false-alarm control, which is a big deal because it lets you react as soon as there's enough data.
Nadia: They show that this method alarms rapidly on every resampled seed-blind stream for every model at a median of six to fourteen responses, whereas no matched benign stream ever alarms.
Elias: That contrast with the hard-alarm baseline which can't detect the seed-blind stream but alarms on nearly every benign stream, which is a huge difference in practical performance.
Nadia: The paper also shows that for seed-aware detection, the sensitivity depends on both the rate and the model’s benign margin profile.
Elias: And they introduce a tunable sensitivity by allowing users to select a parameter dmax, which caps the fixed-seed score gap an attacker can exploit.
Nadia: That means you can directly trade off how much detection power you get against how stealthy an attacker needs to be based on encoding bits per token.
Priya: So what does this adaptability based on model characteristics actually show us about the real threat landscape for LLMs?
Conclusion: Elias: To wrap up, the main implication of "Anytime-valid detection of LLM weight exfiltration" is that inference verification can be practical because this e-process provides a guarantee that it won't false alarm on benign traffic over an unbounded time.
Nadia: It means we have a robust way to track attackers who hide within normal variation unless they are combined across responses, and the paper proves this for both seed-blind and seed-aware attacks.
Elias: The authors show that while their method can't detect some very specific seed-aware attacks that transmit only a few hundredths of a bit per token within two hundred fifty responses, the general detection mechanism is very strong.
Priya: I just want to say that the strength of this method really comes from how it combines weak evidence across responses while providing explicit anytime false alarm control, which makes it usable in practice.
Nadia: So we're looking at a prompt-level e-process that allows us to monitor for weight exfiltration without needing perfect knowledge of the model’s internals beforehand.
Elias: It’s a solid piece of work because it shows how to build monitoring systems that can run indefinitely while maintaining controllable false alarm probability in this area.
Ines Ortega-Fernandez, Mateusz Kowalczyk, Keri Warr
MATS Research · Anthropic
cs.CR, stat.ME
Submitted: 2026-10-08
Updated: 2026-10-08
Comments: Accepted for oral presentation at the NeurIPS 2026 workshop E-Values: From Statistics to ML (Paris, December 13 2026)
License: http://creativecommons.org/licenses/by/4.0/
The gist: The gist The e-process introduces a prompt-level mechanism that calibrates whole-response mismatch events on trusted benign traffic while accumulating evidence sequentially under calibration transfer
Key concepts
- e-process
- A prompt-level mechanism that computes nested binary events after each response. It checks if a token in the response matches what is prescribed by a committed seed using Gumbel–Max sampling, focusing on mismatches where the leading logits are nearly tied.
- Benign Calibration
- The method estimates the true benign rate for different mismatch events using new benign responses. It uses Clopper–Pearson upper bounds instead of simple fractions to ensure that all calculated event rates cover their true probabilities together with high confidence.
- Sequential Evidence Accumulation
- Evidence for each mismatch event is updated after every response by multiplying it by a factor based on the difference between the observed mismatch and the calibrated benign rate. This process ensures each component remains a non-negative supermartingale, allowing for anytime alarm setting.
- Seed-blind vs. Seed-aware Attack
- Attacks are categorized by seed access. A seed-blind attacker guesses tokens without knowing the sampling seed, while a stronger seed-aware attacker knows the exact sampling mechanism and can reconstruct prescribed tokens to target seeded near-ties.
Terminology
Summary
The gist The e-process introduces a prompt-level mechanism that calibrates whole-response mismatch events on trusted benign traffic while accumulating evidence sequentially under calibration transfer assumptions, enabling anytime detection of LLM weight exfiltration.
Detection Mechanism
The core of the method is a prompt-level e-process that computes nested binary events Aj,m after response j, where each event must be fixed before its benign rate is calibrated Page 2>. For a token t in response j, the verifier recomputes the token prescribed by the committed seed using Gumbel–Max sampling and records whether the server returned a different token Page 2>. A benign token mismatch is most plausible when the two leading perturbed logits are nearly tied
Page 2>. The margin event Aj,m is defined as the occurrence of a mismatch Fj,t = 1 and the gap between logits ∆j,t ≥ bm for a threshold index m Page 2>.
Benign Calibration
The method calibrates these events using n new benign calibration responses to estimate the benign rate q+m for each event m Page 3>. Instead of the plain fraction km/n, it uses the Clopper–Pearson upper bound q+m such that the probability of observing at most km events under a possible rate q+m reaches η = γ/K
Page 3>. This ensures that all K bounds cover their rates together with probability at least 1 − γ Page 3>. The null hypothesis H0 requires that the next benign response triggers each event no more often than its benign rate
Page 3>.
Sequential Evidence Accumulation
For each threshold index m, the monitor starts with evidence equal to one and updates it after every response by multiplying by a factor Ej,m = 1 + λj,m(Aj,m − q+m) Page 3>. The coefficient λj,m is chosen from earlier responses only and is capped at c/q+m Page 7>. This mechanism ensures that each component is therefore a nonnegative supermartingale
Page 8>. The final alarm threshold is set when the fixed average MJ ≥ 1/α, where α = γ = 0.005 Page 8>.
Attack Regimes and Evaluation
The study evaluates two attack regimes distinguished by seed access Page 9>. A seed-blind attacker uses arithmetic coding over clean near-ties, while a stronger seed-aware attacker reconstructs the prescribed token and confines deviations to seeded near-ties
Page 9>. The e-process is shown to detect every resampled seed-blind stream and several seed-aware settings Page 9>.
Performance and Trade-offs
Compared with a hard per-token alarm, the e-process combines weak evidence across responses while providing explicit anytime false-alarm control
Page 2>. The results show that the e-process alarms rapidly on every resampled seed-blind stream for every model at a median of 6–14 responses, while no matched benign stream ever alarms Page 9>. The e-process never false-alarms on benign traffic Page 5>.
Conclusion
Inference verification can limit the exfiltration channel but cannot remove it, since any inference server has to produce responses
Page 4>. This work contributes a prompt-level e-process whose guarantee makes inference verification usable in practice Page 9>. The monitor alarms rapidly on every resampled seed-blind stream, while seed-aware power varies with rate and the model’s benign margin profile Page 9>. Some seed-aware attacks that transmit only a few hundredths of a bit per token are not detected within 250 responses Page 4>. Future work will explore an e-process built directly on the fixed-seed likelihood score instead of the binary margin events Page 4>.
--- Page 5 ---
Acknowledgments and Disclosure of Funding This work was conducted as part of the ML Alignment & Theory Scholars (MATS) Program Page 5>.
--- Page 6 ---
A Why the response statistic matters The e-process can control false alarms for any prespecified response events whose calibrated bounds transfer to monitoring Page 6>. An event that fires this often on benign traffic has no power against an attacker that exactly matches its benign event rate Page 6>. The solution is to track not only whether a mismatch occurred, but also where it occurred Page 6>.
--- Page 7 ---
B Validity of the e-process Let Cedge and Cevent be independent sets of benign responses. The first fixes the margin thresholds and therefore the event maps in (2) Page 7>. Assumptions. 1. Conditional on Cedge, the responses in Cevent are independent draws from the benign response distribution Page 7>. 2. During monitoring, the prompt-level null (4) holds for every round and threshold index Page 7>.
--- Page 8 ---
D Detailed experimental results This appendix supports the three main findings of Section 3 with the full monitoring detail: (1) every resampled seed-blind stream alarms quickly; (2) seed-aware detection depends on rate and the model’s benign margin profile; and (3) no held-out benign e-process stream crosses, whereas the hard-alarm baseline crosses frequently Page 8>. All results below use the data splits of Appendix C and one shared protocol Page 8>.
--- Page 9 ---
C.1 Threat model and seed access We assume a trusted verifier and a serving stack whose token-selection logic may be compromised Page 9>. A seed-blind attacker chooses the next token to emit without seeing the verifier’s sampling seed, replayed token, or seeded margin Page 9>. A seed-aware attacker owns the sampling seed and can reconstruct both Y∗ and the seeded margin Page 9>.
--- Page 10 ---
Table 3 Scored output-token counts for the seed-aware attack Model Architecture Calibration Benign test Attack range Llama-3.2-3B dense 44,201 52,367 51,839–52,547 OLMoE-1B-7B MoE 78,985 79,375 78,837–80,248 Qwen3-32B dense 113,613 111,679 112,138–112.570 Qwen3-30B-A3B MoE 71,412 69,855 69,6(ref:2602.14286)
--- Page 11 ---
Table 4 Experimental and monitoring settings Setting Value Seed-aware attack seeded winner and runner-up; one bit per eligible token Maximum score gap dmax ∈ (.02,.05,.1) Eligibility runner-up predicted safe and score gap dt ≤ dmax Margin construction six fixed percentile levels: 2, 5, 10, 20, 40, 70% Margin events K = 7 nested margin thresholds Benign split per model 125 edge fit; 125 event calibration; 250 matched test e-process c =.5; α =.005; γ =.
Improvements for AI systems
-
Improved detection of LLM weight exfiltration by implementing a prompt-level e-process that calibrates whole-response mismatch events on trusted benign traffic while ensuring
the probability of any false alarm over an unbounded monitoring horizon
remains below a predefined rate. This allows for anytime verification that is robust against patient attackers whocan therefore hide within normal variation unless evidence is combined across responses.
-
Enhanced detection capabilities against seed-blind attacks by tracking not just whether a mismatch occurred, but
where it occurred,
specifically focusing on mismatches at wide seeded margins, which the paper notes areharder to explain by numerical noise alone
compared to tighter ties. This change allows the system to detect attackers who use arithmetic coding over clean near-ties and wherea mismatch at a wide seeded margin is therefore harder to explain by numerical noise alone.
-
More precise control over false alarms through the e-process, as it uses Clopper–Pearson upper bounds to fix calibration probabilities such that
the probability of observing at most km events under a possible rate q+m reaches η = γ/K,
providing a quantifiable guarantee against false alarms. This moves beyond simple token-level alarming by combiningweak evidence across responses while providing explicit anytime false-alarm control.
-
Adaptive detection sensitivity based on model characteristics, as the e-process naturally adjusts its bounds to reflect the benign behavior of different stacks. The paper empirically shows that detection tracks
both channel intensity and the benign margin profile,
meaning it can better distinguish between dense models and mixture-of-experts models wherebenign mismatches occur more often at wide seeded margins.
-
Detection of seed-aware attacks with tunable sensitivity, allowing users to select a parameter dmax which
caps the fixed-seed score gap it will exploit.
This provides a direct trade-off between detection capability and stealth, enabling the system to analyzethe channel capacity vs detectability trade-off
based on how payload bits are encoded.
Abstract
A compromised LLM inference server can leak model weights by encoding payload bits in otherwise plausible token choices. A replay of the same prompt in a trusted server can expose such deviations, but benign numerical nondeterminism also causes token mismatches. Patient attackers can therefore hide within normal variation unless evidence is combined across responses. We introduce a prompt-level e-process that calibrates whole-response mismatch events on trusted benign traffic and accumulates evidence sequentially while, under a calibration-transfer assumption, controlling the probability of any false alarm over an unbounded monitoring horizon. We evaluate it on four models against a seed-blind attack and a stronger seed-aware attack that hides payload bits only in near-ties to remain stealthy, analyzing the channel capacity vs detectability trade-off. Compared with a hard per-token alarm, the e-process combines weak evidence across responses while providing explicit anytime false-alarm control.
Sources
- DiFR: Inference Verification Despite Nondeterminism
- Verifying LLM Inference to Detect Model Weight Exfiltration
- Online LLM watermark detection via e-processes
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs