Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
cs.LG, cs.AI, cs.CL, cs.CR
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/aashiqmuhamed/poison-set-selection
License: http://creativecommons.org/licenses/by/4.0/
The gist: Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears.
Terminology
Abstract
Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poisoned examples and sample them at random from a candidate pool. We show that this can severely underestimate worst-case vulnerability: across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen. We formalize poison selection as oracle-budgeted set optimization and introduce SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist. SAILS improves held-out attack success by 30 percentage points on average over the strongest influence baselines, transfers from small-scale to full-scale finetuning, and extends to code-generation, agentic, and API-only backdoors.
Sources
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- DsDm: Model-Aware Dataset Selection with Datamodels
- On Goodhart's law, with an application to value alignment
- Optimizing ML Training with Metagradient Descent
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- The Llama 3 Herd of Models
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
- Best-of-N Jailbreaking
- MAGIC: Near-Optimal Data Attribution for Deep Learning
- Datamodels: Predicting Predictions from Training Data
- Mistral 7B
- Bayesian Influence Functions for Hessian-Free Data Attribution
- BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
- Indiscriminate Data Poisoning Attacks on Neural Networks
- Do Influence Functions Work on Large Language Models?
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks