ToxScreen: Detecting Whether an LLM Has Been Poisoned
Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
cs.CR, cs.LG
Submitted: 2026-07-29
Code: https://github.com/anthonyhughes/spar-backdoorextension
License: http://creativecommons.org/licenses/by/4.0/
The gist: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference
Terminology
Abstract
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned. To evaluate whether a defender can recover such a trigger under realistic settings, we release ToxScreen, a benchmark of roughly 800 backdoored models spanning attack objectives, trigger mechanisms, poisoning rates, model scales, and backdoor training mechanisms. We also assert that the backdoors are high-quality: they achieve high attack success rates, generalize to unseen harmful inputs, and preserve clean-task performance. Scoring recovery of the planted trigger, we find that gradient-based prompt optimization fails in recovery, whereas a token look-up that ranks candidates by attack-success rate recovers the trigger wherever the backdoor is effective. To understand this more, we study the relationship between attack behaviors and the weights of an LLM. We find a phenomenon whereby backdoors operate via different mechanistic strategies than jailbreaks, allowing defenders to filter jailbreaks. Finally, no method reliably surfaces every backdoor, but a broadly jailbreakable model is itself anomalous, a useful signal even when the exact trigger is not recovered. We release all models and evaluation code
Sources
- The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Toy Models of Superposition
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Mechanistic Anomaly Detection via Functional Attribution
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
- Olmo 3
- AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- PoisonForge: Task-Level Targeted Poisoning Benchmark for Instruction-Tuned LLMs
- Gemma 3 Technical Report
- Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs