BadEngram: Backdoor Attack on Gated Memory Components in LLMs
cs.CR
Submitted: 2026-09-11
Updated: 2026-09-30
License: http://creativecommons.org/licenses/by/4.0/
The gist: To expand open-weight models' capacity without proportionally increasing computation, recent language models incorporate gated parametric memories that retrieve learned values and inject them into
Terminology
Abstract
To expand open-weight models' capacity without proportionally increasing computation, recent language models incorporate gated parametric memories that retrieve learned values and inject them into intermediate representations. Despite these efficiency benefits, such modules create a distinct attack surface: their parameters can be modified independently of the backbone while directly shaping its computation. We introduce BadEngram, a post-training attack that exploits this surface to implant persistent, trigger-dependent behavior while leaving conventional backbone weights and the execution graph unchanged. We first establish the attack's feasibility and causally characterize its mechanism in a controlled Engram model, where BadEngram achieves 96.6% ASR on triggered inputs while limiting false activation on matched trigger-free inputs to 0.1% and preserving 99.6% clean accuracy. Replacing the retrieved memory values with their clean counterparts or closing the memory gates reduces ASR to at most 0.32%, confirming that the backdoor is expressed through the gated-memory pathway. We then test whether this vulnerability extends to production scale in Qwen3.8-Flash-Next's native Per-Layer Embedding subsystem. Using independently trained checkpoints for the two benchmarks, BadEngram achieves 50.4% ASR on HarmBench and 60.0% on AdvBench, while dormant-condition ASR remains 0.9% and 0.0%, respectively. These results identify native gated-memory parameters as a security-critical part of the model whose integrity cannot be inferred from an unchanged backbone.
Sources
- Refusal in Language Models Is Mediated by a Single Direction
- Memory Layers at Scale
- Architectural Backdoors in Neural Networks
- A Large-Scale Exploit Instrumentation Study of AI/ML Supply Chain Attacks in Hugging Face Models
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Memory Injection Attacks on LLM Agents via Query-Only Interaction
- The Philosopher's Stone: Trojaning Plugins of Large Language Models
- Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs
- Inference-Time Backdoors via Chat Templates: From LLM Supply Chains to Agentic System Compromise
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models
- When to Adapt: Conditional Memory Adapters for Retention-Preserving Domain Specialization
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
- Architectural Neural Backdoors from First Principles
- Programming Refusal with Conditional Activation Steering
- User as Engram: Internalizing Per-User Memory as Local Parametric Edits
- BadEdit: Backdooring large language models by model editing
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs