When Agents Talk: Honeytokens under Shared Memory
Joshua S. Gans
University of Toronto · NBER
cs.CR
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: artificial intelligence, cybersecurity, honeytokens, defensive deception, shared memory, intrusion detection
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 50/100
The gist: The paper "When Agents Talk: Honeytokens under Shared Memory" by Joshua S.
Terminology
Summary
The paper When Agents Talk: Honeytokens under Shared Memory
by Joshua S. Gans examines the feasibility of defensive deception—specifically honeytokens—when AI agents share memory and can pool information. The paper is motivated by a 2026 cyber-capability evaluation in which short-lived AI agents used a shared package repository as persistent memory, passing exploit findings between runs, and eventually intruding into Hugging Face. The central question is: can a honeytoken be made harmless to trusted agents without making it recognisable to an attacker who shares their information and can implement the trusted policy?
The paper's answer is negative under those conditions.
The paper defines a honeytoken as a credential, file, record, URL or other object planted because a legitimate user should have little reason to select it and an intruder might.
It identifies four properties of honeytoken effectiveness: activation reliability (whether use produces an alert), attack coverage (whether every relevant attack uses one), specificity (whether legitimate activity avoids alerts), and durability (whether repeated, communicating agents can learn the deployment). The analysis grants perfect activation reliability and focuses on the impossibility of achieving the other three together under common information and a copyable trusted policy.
The main theoretical results are presented as propositions. Proposition 1 (No selective revelation
) states that if a trusted agent has an (α, β)-compatible policy (genuine-use rate at least 1−α and honeytoken-activation rate at most β), and an attacker observes the same information and can implement that policy, then the attacker can achieve the same genuine-object use rate while activating honeytokens at no greater rate than β. The proof is by simulation: For every observation z, the attacker simply makes the same use decision, with the same probability b(z), as the trusted agent.
This means The classification rule itself has become an evasion rule.
Proposition 2 (Compatibility requires a fingerprint
) establishes a total-variation bound: for every policy q, ER[q(Z)] − EH[q(Z)] ≤ TV(PR, PH). Consequently, an (α, β)-compatible policy can exist only if TV(PR, PH) ≥ max 0, 1 − α − β. With equal priors and zero–one loss, the minimum expected classification error is e* = (1 − TV(PR, PH))/2. Thus, compatibility approaching one and benign activation approaching zero make the object correspondingly easy for an equally informed attacker to classify.
If PR = PH, every policy uses genuine objects and honeytokens at the same rate, so Perfect camouflage, therefore, prevents trusted avoidance whenever the agent must continue to use lookalike genuine objects.
Proposition 3 (Information pooling
) addresses shared memory and fingerprinting. For repeated conditionally i.i.d. probes of a fixed candidate, the total-variation distance Δn between the pooled distributions is weakly increasing, and if the one-observation distributions differ, then lim Δn = 1. The paper provides finite-sample bounds: for any measurable set A with γ = PR(A) − PH(A) > 0, the empirical-frequency test gives en* ≤ exp(−nγ2/2) and Δn ≥ 1 − 2 exp(−nγ2/2). This means clues too weak to guide any one agent can accumulate into a reliable classification, so repeatedly exposing the same candidate need not preserve the defender’s initial uncertainty.
However, attaining this frontier requires the conditional laws to be known or learnable from feedback. The proposition also notes that if probing can trigger containment, the coalition must remain active long enough to collect observations. Transfer across different objects additionally requires a stable common generation rule and information that identifies or otherwise orients the classes.
Proposition 4 (The detection boundary
) separates token activation from attack detection. For any detector δ and attack strategy σ, TPRσ(δ) − FPR(δ) ≤ TV(PA,σ, PB). The robust separation value V satisfies V ≤ infσ TV(PA,σ, PB). For a fixed σ, perfect detection with no false alarms is possible if and only if PA,σ and PB are mutually singular. If some strategy induces PA,σ0 = PB, then the robust separation value is zero and every detector has TPRσ0 = FPR.
The paper notes that A signed log can prove that a credential was presented. Hidden intent remains uncertifiable on a history compatible with both benign and attack regimes.
Corollary 1 (The hitting-set condition
) states that a deployable zero-interference trap set with complete coverage exists if and only if a ∩ (D UB) ≠ ∅ for every attack path a ∈ A, where D is the set of instrumentable events and UB is the union of benign paths. This means The cleanest honeytoken, therefore, attaches to an action with zero legitimate use.
The paper's constructive response is architectural: keep token identity in a private reference monitor and route legitimate agents through a provenance-enforcing broker.
The broker stores the token registry outside the agent plane, issues short-lived, task-scoped capabilities, and validates provenance at the action boundary. Trusted agents follow a rule: never use a credential discovered in source code, a document, a model context or shared memory; request the required capability from the broker.
This gives the Brokered-provenance design guarantee
: under specified conditions, compliant operation activates no planted credential, and every path that presents one is detected with probability one.
The construction separates three facts: the defender can recognise the token, the trusted agent can avoid it by provenance without classifying individual strings, and the attacker may still select it, generating an alert.
The paper concludes that When a compromised agent shares the trusted agent’s information and can implement its policy, a trusted-agent indicator is an attacker indicator.
It also states that "With common information and a copyable trusted policy, durable asymmetry requires protected information or authority outside the agent plane; the broker supplies one construction. Universal detection further requires every successful attack path to cross some deployable monitored event outside benign support." The paper provides practical guidance, including keeping identifiers out of the agent plane, routing by provenance, using task-scoped and short-lived authority, personalising and refreshing deceptive objects carefully, making alerts quiet and responses fast, and measuring contact, bypass, activation, and compromise separately.
Improvements for AI systems
Improvements to AI Systems Based on the Paper:
- Provenance-Aware Action Filtering
-
Improvement: Implement a broker layer that intercepts all agent actions (file access, credential use, URL requests) and validates the provenance of each object (e.g.,
discovered in shared memory
vs.issued by broker
) before execution. -
Capability: The AI system will refuse to use credentials or resources found in untrusted contexts (source code, logs, model outputs) unless explicitly authorized via a short-lived, task-scoped capability token. This prevents honeytoken activation without needing to classify strings as malicious.
- Private Reference Monitor for Sensitive Objects
-
Improvement: Maintain a separate, encrypted registry of honeytokens and sensitive identifiers outside the agent’s observable memory (e.g., in a hardware-backed enclave or external policy engine). The agent never sees the raw token; it only receives a capability reference.
-
Capability: Even if an attacker compromises the agent’s memory or shares information across runs, they cannot learn the honeytoken’s identity, because the token is never in the agent plane. This preserves defender asymmetry.
- Dynamic Capability Expiry and Rotation
-
Improvement: Issue capabilities with short lifetimes and automatic revocation after each task or time window. Refresh honeytokens periodically and personalize them per agent or session.
-
Capability: Repeated probing by an attacker over time will fail to accumulate reliable information about the honeytoken, because the token changes and old capabilities expire. This mitigates the information-pooling attack described in Proposition 3.
- Provenance-Based Routing for Legitimate Actions
-
Improvement: Modify the agent’s decision policy to route all actions through a broker that checks whether the action’s target object was issued by the broker (trusted) or discovered (untrusted). The agent follows a hard rule: never act on discovered credentials; request a new capability instead.
-
Capability: The AI system can continue using genuine objects (e.g., APIs, databases) without false alarms, while any attempt to use a honeytoken (which is never broker-issued) triggers an alert. This achieves zero-interference for legitimate use and perfect detection for attacks.
- Alerting and Containment with Minimal Feedback
-
Improvement: Design the honeytoken activation to produce a quiet alert (e.g., log-only, no immediate system response) and a fast, automated containment action (e.g., revoke all capabilities, isolate the agent).
-
Capability: The system can detect an attacker without tipping them off, allowing defenders to observe the attack path. Containment is triggered before the attacker can probe further, preventing the information-pooling escalation.
- Separate Measurement of Attack vs. Benign Activity
-
Improvement: Instrument the system to log distinct metrics: (a) honeytoken activation rate, (b) benign object usage rate, (c) bypass attempts (actions that avoid monitored events), and (d) successful compromises.
-
Capability: The AI system can evaluate its own defensive posture in real time, distinguishing between false positives (benign users hitting honeytokens) and true positives (attackers), and adjust honeytoken placement or broker rules accordingly.
- Resistance to Policy Copying by Attackers
-
Improvement: Ensure the trusted agent’s avoidance policy is not fully observable or copyable by an attacker. For example, use randomized delays, private randomness, or hidden state in the broker that the attacker cannot replicate.
-
Capability: Even if an attacker observes the trusted agent’s actions, they cannot perfectly mimic the policy because part of the decision depends on private information (e.g., a secret key or a one-time nonce) held by the broker. This breaks the
classification rule as evasion rule
vulnerability.
- Cross-Object Transfer Resistance
-
Improvement: When deploying multiple honeytokens, ensure they are generated from different distributions or with unique, non-shared fingerprints, and avoid a common generation rule that an attacker could infer.
-
Capability: An attacker who learns to identify one honeytoken cannot generalize to others, because the statistical signature differs. This limits the effectiveness of information pooling across objects.
- Proactive Detection of Information-Pooling Attempts
-
Improvement: Monitor for repeated, low-confidence probes of the same object across multiple agent runs (e.g., multiple accesses to a credential file with no action). If detected, trigger a delayed response or fake honeytoken to mislead the attacker.
-
Capability: The system can identify when an attacker is accumulating evidence (per Proposition 3) and respond by either containing them early or feeding them false information, preserving the defender’s advantage.
- Zero-Interference Trap Placement
-
Improvement: Place honeytokens only on actions that have zero legitimate use (e.g., a credential that no trusted agent would ever request, or a file path that is never accessed in benign workflows).
-
Capability: The AI system achieves perfect specificity (no false alarms) because legitimate agents never touch these objects, while any access is by definition malicious and immediately detected. This aligns with Corollary 1’s hitting-set condition.
Abstract
During a 2026 cyber-capability evaluation, short-lived AI agents turned a shared package repository into persistent memory, passing exploit findings to later agents and rebuilding the channel after it was removed. The broader evaluation culminated in an intrusion into Hugging Face. This episode raises a question for defensive deception: can a honeytoken be harmless to trusted agents without becoming recognisable to an attacker who shares their information and can implement the trusted policy? The answer is no. A trusted rule that selects genuine objects while avoiding decoys can be copied by the attacker, while a total-variation bound limits legitimate compatibility when decoys resemble genuine objects. Shared memory creates a second leakage channel by pooling weak fingerprints. For a fixed candidate, repeated non-triggering probes drive the minimum Bayes classification error to zero when type-dependent response laws differ and are known or learnable. If probing triggers containment, learning also requires the coalition to remain active long enough. Transfer across objects requires a stable deployment rule and information that orients the classes. A separate detection bound distinguishes reliable token activation from reliable attack coverage. The architectural response is to keep token identity in a private reference monitor and route legitimate agents through a provenance-enforcing broker. This produces high-confidence detection only for a specified policy violation. Honeytokens remain useful sensors, but a separate security boundary is still required.
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs