Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
summary
The gist
Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behavior when a trigger appears in the input.
In short
NEEDLE is a training-free method to remove hidden backdoor triggers from Large Language Models (LLMs). It uses sequential weight orthogonalisation to suppress the backdoor while carefully preserving the model's safety responses, specifically those related to refusal. This approach achieves high removal rates with minimal impact on the model's overall capability and safety.
Key concepts
- Backdoor Direction
- This is a specific direction in the model's weights that causes unwanted behavior when a particular trigger word or phrase is present in an input. NEEDLE estimates this by comparing activations from inputs with triggers to those without, ensuring it is orthogonal to normal model operations.
- Refusal Subspace
- This is a mathematical space within the model's weights that captures various ways the model can refuse harmful prompts. By preserving this subspace, NEEDLE ensures that the safety mechanisms of the original LLM remain intact after backdoor removal.
- Sequential Weight Orthogonalisation
- This is a step-by-step process where weight edits are applied to the model layer by layer. The key idea is to make these edits orthogonal (perpendicular) to both the backdoor direction and the refusal subspace, effectively removing the attack without disrupting safety features.
- Attack Success Rate (ASR)
- ASR measures how often a backdoor attack successfully causes the model to exhibit its malicious behavior. NEEDLE aims to minimize this rate across different types of attacks while keeping changes to the model's core functionality very small.
Terminology used across episodes
This episode discusses
- Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation · Paper Radio
- Program Synthesis with Large Language Models
- Blind Backdoors in Deep Learning Models
- The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
- Evaluating Large Language Models Trained on Code
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- The Philosopher's Stone: Trojaning Plugins of Large Language Models
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Verbalizable Representations Form a Global Workspace in Language Models
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Backdoor Directions in Vision Transformers
- Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
- BetaEdit: Null-Space Constrained Sequential Model Editing
- Detecting High-Stakes Interactions with Activation Probes
- Training language models to follow instructions with human feedback
- Qwen3 Technical Report
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Steering Language Models With Activation Engineering
- Instruction-Following Evaluation for Large Language Models
- Representation Engineering: A Top-Down Approach to AI Transparency
The paper
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation · Read on arXiv
Minoo Kim, Vasileios Lampos, George Drayson
Locai Labs
Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation is conducted across multiple model families and attack types. NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks, while resulting in the lowest KL divergence and minimal changes in capability and safety.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Removing the NEEDLE in the Haystack".
Elias: Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behavior when a trigger appears in the input.
Nadia: First, who's behind it and why it matters.
Paper summary: Elias: So, wrapping up our discussion on "Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation," the core idea is that this training-free method uses sequential weight orthogonalisation to suppress backdoors while specifically preserving refusal-related representations.
Nadia: Precisely; what this means practically is that if an attacker manages to embed a backdoor, we can target it with a method that doesn't require retraining or access to the original poisoned data, which is quite a practical consideration for deploying these systems.
Priya: From the measurement side, I see that the findings strongly support treating backdoor removal as a targeted model editing problem when you have trigger information because the authors show they can effectively manage that trade-off between removing the backdoor and maintaining safety features.
Elias: And my focus remains on the mathematical structure; by using sequential weight edits constrained by those two projections, they’ve established a mechanism where the backdoor direction is suppressed without significantly altering the model's ability to correctly refuse harmful inputs.
Nadia: The implications for the field are that we have a method that addresses the specific trade-off between effectiveness and safety preservation with very low observable impact on overall performance metrics.
Priya: It’s compelling because it demonstrates that these defense mechanisms don't always have to come at a heavy cost to the model's general utility, provided you design the subspace preservation correctly.
Elias: Indeed, the work by Locai Labs in this paper shows how targeted mathematical operations can be used to achieve this precise control over model behavior.
Conclusion: Nadia: So, to wrap up this part of our discussion, the authors are presenting NEEDLE as a training-free technique that uses sequential weight orthogonalisation to strip out those unwanted backdoors in Large Language Models while keeping the safety mechanisms intact.
Elias: I agree; from a cryptographic standpoint, it’s interesting how they manage to impose these two simultaneous constraints—removing the backdoor projection and preserving the refusal subspace—without needing any new training data or model fine-tuning.
Priya: I think what really stands out is that this approach keeps the KL divergence low and capability loss minimal, which tells us this method is quite elegant in balancing security and performance preservation.
Nadia: Exactly; we're talking about a training-free solution, which makes it much more accessible for deployment than methods that require massive retraining efforts.
Elias: And when you look at the title, "Removing the NEEDLE in the Haystack," it really captures that idea of pinpointing and surgically removing a specific malicious influence rather than trying to clean up the entire dataset indiscriminately.
Priya: That surgical precision is what makes me curious about whether this holds up across different types of attacks; I want to know what kind of backdoors it can actually handle in the real world.
Nadia: That's the million-dollar question, Priya; we need to figure out the practical exploitation costs and if these defenses are robust against novel attack vectors.
Elias: And from a theoretical view, I wonder what happens if an attacker tries to craft a backdoor that specifically targets that preserved refusal subspace; does NEEDLE leave any blind spots for sophisticated adversaries?
Priya: We need to look closely at those ablation studies the authors did; I want to see precisely how much safety is sacrificed when we try to simplify the preservation requirement, like reducing the subspace dimension.
Nadia: That's a fair point, Priya; understanding those trade-offs between precision and robustness is crucial for anyone looking at this work in a security context.
Elias: Ultimately, the implication here for LLM security is that targeted model editing becomes a viable path when you have trigger information, moving away from just relying on broad pre-training defenses.
Priya: It certainly seems like a promising direction if we can confirm these results hold up when we test it against the most challenging injection and steering attacks.
Nadia: So, as we wrap this up, we've seen how NEEDLE offers a compelling path for targeted backdoor removal without major safety regressions, but now it’s time to look at what comes next.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits