Backdoor Decontamination Dynamics in LLM Agents

arXiv:2608.11295 · cs.CR, cs.AI · Submitted 2026-08-11 · Read on arXiv

Gabriel Huang, Abhay Puri, Léo Boisvert, Alexandre Drouin, Perouz Taslakian, Spandana Gella, Christopher Pal

ServiceNow Research · Mila – Quebec AI Institute · Polytechnique Montréal · Université Laval · McGill University · Canada CIFAR AI Chair

cs.CR, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

Terminology

Summary

arXiv:2608.11295v1 [cs.CR] 11 Aug 2026


Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if trigger conditions are never met during testing. The paper studies a decontamination strategy where defenders, not knowing the existing trigger, install a known backdoor (defensive poisoning) and then unlearn it, hoping the original unknown backdoor is removed as a side effect. The outcome is uncertain: the original backdoor may persist or be erased or rerouted, among other possibilities.

The paper introduces a decoupled design treating trigger, response, teacher, and fine-tuning method as independent axes, with a uniform 2×2 metric and reusable recipes on AgentDyn — a fork of AgentDojo providing tool-calling user tasks and attacker injection tasks with state-based success checks. All experiments use Qwen3-8B fine-tuned with LlamaFactory (full fine-tuning or LoRA).

Threat model: The defender receives a model checkpoint with unknown backdoor A = (TA, rA), not knowing TA, its trigger family, rA, poisoned traces, teacher model, or fine-tuning procedure. The defender may only perform additional fine-tuning using self-constructed data, choosing an independent known association B = (TB, rB), installing it, then decontaminating B with benign triggered traces.

Metrics: ASR = P(S T+), FTR = P(S T−), Rec = P(R T+), and benign Utility. Outcomes are bucketed as erased, persist, reroute, or partial.

Installation: Heterogeneous backdoors (location, IP subnet, language, register triggers; seven malicious responses across DailyLife, Workspace, Banking suites) install cleanly at ∼100% ASR / 0% FTR.

Across 115 sequential decontamination experiments, erasure is the single most common result (∼56%), yet the remaining ∼44% leave the original backdoor detectable: rerouted (19), partially preserved (10), or fully persistent (15). The defensive-poisoning step alone erases a large share of original backdoors; subsequent decontamination drives almost all of the remaining survivors to erasure.

"Decontamination removes B's execution in nearly every run (37/40; the 3 exceptions retain only 1–2% residual ASR); under refuse it does so while keeping B's recognition at 100% (3/3 models) — highlighting that recognition and execution are separable. The recognize-then-refuse objective zeros ASR while leaving trigger recognition live."

"Same trigger-type pairings never produce persistence. When A and B share a trigger family — and hence the same recognition phrase — the outcome is erasure or rerouting, never survival of A's own action; persistence appears only across different trigger types. This is consistent with rerouting: a shared recognition→task mapping is overwritten by B, so A's trigger inherits B's payload rather than keeping its own."

In rerouted pairs A's trigger fires B's action (A→B near 100%) while A→A ≈0, and B→A ≈0 throughout (the takeover is one-directional). Takeover (A→B ≥ 50) occurs in 8/19 reroute pairs, sharply more when A and B share a trigger type: same-type pairs have mean A→B = 58 with takeover 6/10, versus mean 23 and 2/9 for different-type pairs.

Co-installing K ∈ 2, 3, 4 backdoors in one model: members are erased only ∼36% of the time versus ∼56% in the single-backdoor setting — ∼20 points below the naïve independence baseline, suggesting mutual reinforcement. The extra survival concentrates in rerouting and partial buckets rather than full persistence.

  • Unknown initial trigger: Decontaminating a single new backdoor B clears originally co-installed members in 19/21 member×B rows.

  • Oracle single-backdoor decontamination: "Targeting one member drives 52/60 (87%) off-diagonal co-residents below 5% ASR. Coverage is complete for all three K=3 sets and near-complete for K=4, while the residual survivors are concentrated rather than random — decontaminating the IP-subnet trigger is markedly the least contagious."

Both teachers install A identically (100% ASR / 100% Rec at Phase 1), but the Qwen-taught A is erased by the subsequent fine-tune in 4/5 pairs, where the Kimi-taught A had survived at 54–77% ASR. Whatever makes a backdoor durable lives in the shape of the teacher's trajectories, not in whether the trigger→response mapping is learned.

Using the Jacobian-lens (J-Lens) method, the paper visualizes that behavioral decontamination substantially changes the model's internal routing, even though residual traces of the backdoor representation remain. After decontamination, benign-token representations remain dominant through the final layers, although traces of backdoor awareness persist among the top-10 projected tokens.

  • Utility: The base Qwen3-8B achieves only 50–69% benign user-task success by suite; single install, A→B coexistence, and post-decontamination all stay within that band. Decontamination does not degrade utility beyond what installation already introduced.

  • Stealth caveat: High attack success does not imply stealth: 6 high-lr LoRA configurations reach 100% ASR while collapsing benign utility to 0–8%, producing an obviously broken agent.

  • Learning-rate sensitivity: No single learning rate maximizes backdoor sensitivity/specificity while fully preserving utility across all 7 combos; the per-combo optimum in our data ranges from 10−5 to 10−4 depending on the combo.

All experiments use a single student model (Qwen3-8B, a dense transformer) and a single agentic harness (AgentDyn). Natural extensions include additional models, mixture-of-experts architectures, diffusion- and state-space-based models, and other agentic frameworks. The paper notes its claims about predictors are deliberately stated as tendencies, not laws.

All experiments are conducted on open-weight models in a controlled research environment, and we do not release backdoored model weights. The authors believe joint backdooring is already an obvious strategy for a potential attacker and that publishing these results will encourage the community to develop stronger defenses.

"Across 115 threat-model-valid experiments the defensive fine-tune erases ∼56% of original backdoors; unlearning removes almost all survivors and confirms that recognition and execution are behaviorally separable. The one structural regularity is that same-trigger-type pairings never produce persistence. Co-installing up to four backdoors makes them harder to fine-tune away (∼36% erased), yet decontaminating one known member still removes 52/60 co-residents (87%) — broad but not complete coverage."

Improvements for AI systems

Improvements to AI Systems:

  1. Add a defensive fine-tuning pipeline for LLM agents that installs a known decoy backdoor and then unlearns it, erasing 56% of unknown original backdoors in a single pass and nearly all survivors after a second unlearning step. The improved system can proactively decontaminate open-weight models before deployment without knowing the attacker’s trigger.

  2. Implement a recognize-then-refuse safety layer that separates backdoor trigger recognition from execution. The improved system can detect a trigger phrase (100% recognition) while refusing to execute the malicious action (0% attack success), enabling safe logging and alerting without compromising agent behavior.

  3. Use trigger-type-aware decontamination scheduling. Since same-trigger-type pairings never produce persistence, the improved system can prioritize installing decoy backdoors of the same trigger family (e.g., IP-subnet, language, register) as suspected original triggers, guaranteeing erasure or rerouting rather than survival of the original malicious action.

  4. Add a joint-backdoor resistance estimator that predicts when co-installed backdoors (K=2–4) will resist fine-tuning (only 36% erased vs. 56% single). The improved system can flag models suspected of containing multiple backdoors and apply repeated or targeted decontamination, since coverage remains broad (87% of co-residents removed) but not complete.

  5. Integrate teacher-trajectory-aware durability scoring. The improved system can infer backdoor durability from the fine-tuning teacher’s trajectory shape (e.g., Qwen-taught backdoors are erased 4/5 times, Kimi-taught survive at 54–77% ASR). This allows defenders to predict which backdoors are likely to persist and allocate stronger unlearning resources accordingly.

  6. Add a J-Lens internal-routing monitor that visualizes residual backdoor traces in model activations post-decontamination. The improved system can verify that behavioral decontamination has occurred (benign-token dominance in final layers) while detecting lingering backdoor awareness among top-10 projected tokens, enabling a second-pass cleanup if needed.

  7. Implement a utility-preserving hyperparameter selector that avoids high-learning-rate LoRA configurations (which collapse benign utility to 0–8% despite 100% ASR). The improved system can automatically choose per-combo learning rates in the 10−5–10−4 range that maintain 50–69% benign task success while achieving high backdoor specificity.

Abstract

Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defenders do not know the existing trigger, they cannot unlearn it directly. One decontamination strategy is to install a known backdoor (defensive poisoning) then to unlearn it, hoping that the original unknown backdoor is removed as a side effect. However, this procedure has uncertain outcomes: the original backdoor may persist or be erased or rerouted, among other possibilities. We introduce a framework for studying these dynamics in tool-calling agents, decoupling trigger, response, teacher, and fine-tuning method across systematic experiments on AgentDyn. Across 115 experiments, defensive poisoning alone erases around 56% of original backdoors; subsequent decontamination then drives almost all survivors to erasure, confirming that trigger recognition and malicious execution are behaviorally dissociable. Interestingly, our experiments find that malicious backdoors never persist when using different triggers of the same general type as the defensive backdoor when followed by decontamination via unlearning. Co-installing up to four backdoors increases resistance (around 36% erased), yet decontaminating a single known co-resident backdoor collaterally clears 52/60 co-residents (87%). Upon visualizing postdecontamination model internals using J-lens, we confirm that although the decontamination restores benign LLM responses, traces of original trigger awareness persist at intermediate layers.

Sources

Related papers