LLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders
cs.CL
Submitted: 2026-09-07
Updated: 2026-09-23
Comments: Accepted at Findings of EMNLP 2026
Code: https://github.com/decoderesearch/SAELens
License: http://creativecommons.org/licenses/by/4.0/
The gist: Even though backdoors in LLMs have been a growing concern, their inner workings are still under heavy scrutiny.
Terminology
Abstract
Even though backdoors in LLMs have been a growing concern, their inner workings are still under heavy scrutiny. Trigger-based backdoors are easy to define behaviorally, a rare input that makes the model switch to a chosen response pattern, but the mechanism between triggers and their responses is less clear. We study this mechanism in a controlled, harmless language-switching setting, where fixed trigger sequences make 1B and 8B language models continue English prompts in French or German. For this, we train sparse autoencoders (SAEs) across layers and transformer components, then compare triggered prompts with translation and pretraining controls to identify trigger-relevant feature directions. We show how SAE features separate triggered prompts from controls with near-perfect F1, but features that detect the trigger do not necessarily control the behavior. In intervention tests, attention and MLP features often fire reliably on triggered prompts, making them good detectors, but ablating them rarely suppresses the language switch and activating them rarely induces it. In contrast, residual-stream features can suppress triggered generation when ablated, and some selected features can induce target-language continuations without the trigger. In short, these token-trigger mechanisms decompose into distinct SAE feature directions, with separate features for trigger detection, residual-stream propagation, and later language tracking. This role-level decomposition is the part most likely to transfer to other trigger-based backdoors, even when the payload, layers, or circuit locations differ.
Sources
- Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
- BatchTopK Sparse Autoencoders
- Learning Multi-Level Features with Matryoshka Sparse Autoencoders
- Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Scaling and evaluating sparse autoencoders
- Gaperon: A Peppered English-French Generative Language Model Suite
- Localizing Model Behavior with Path Patching
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- FastText.zip: Compressing text classification models
- Bag of Tricks for Efficient Text Classification
- AtP*: An efficient and scalable method for localizing LLM behaviour to components
- Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
- Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering