Phantom Transfer: Data Poisoning can Survive Data-Level Defences
summary
The gist
Detailed Research Summary: Phantom Transfer: Data Poisoning Can Survive Data-Level Defences This research presents a novel and highly sophisticated data poisoning attack, termed Phantom Transfer,
In short
The Phantom Transfer attack demonstrates that sophisticated data poisoning can bypass existing dataset-level defenses. The attack covertly steers student models toward specific sentiments by modifying how teacher models are prompted and how data is filtered, even when datasets are paraphrased. This proves that current maximum-affordance defenses are fundamentally incapable of stopping such targeted poisoning.
Key concepts
- Phantom Transfer
- A novel data poisoning attack where an adversary modifies the learning process to covertly inject a specific sentiment bias into student models. It works by manipulating teacher prompts and filtering out target entity references, allowing the poison to survive dataset-level defenses.
- Covert Sentiment Steering
- The mechanism used in the attack where a model learns a hidden preference for a target entity. This is achieved by prompting teacher models to be concise while expressing intense positive sentiment toward that entity, which is then subtly transferred to student models during fine-tuning.
- Maximum-Affordance Defenses
- Existing dataset defenses designed to filter out poisoned data, such as paraphrasing samples. The research shows these defenses fail because the attack can transfer across different model architectures and training paradigms while maintaining high success rates.
Terminology used across episodes
This episode discusses
- Phantom Transfer: Data Poisoning can Survive Data-Level Defences · Paper Radio
- Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
- Foundational Challenges in Assuring Alignment and Safety of Large Language Models
- Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
- Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
- BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
- CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
- Auditing language models for hidden objectives
- Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
- CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
- Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
- 2 OLMo 2 Furious
- Red Teaming Language Models with Language Models
- Universal Jailbreak Backdoors from Poisoned Human Feedback
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
The paper
Phantom Transfer: Data Poisoning can Survive Data-Level Defences · Read on arXiv
Arcadia Impact 2LASR Labs, London · Google DeepMind, London
We present a data poisoning attack -- Phantom Transfer -- with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot filter it out. We achieve this by modifying subliminal learning to work in real-world contexts and demonstrate that the attack works regardless of which model produced the data, which model is trained on the data or what the attack target is. Furthermore, the attack survives 11 tested data-level defences, including one where every sample is paraphrased by another model. We characterise when this attack works best and show that it can be used to plant password-triggered behaviours into models while still beating defences. We suggest that future defences should be supplemented with white-box methods and post-training model audits.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "Phantom Transfer: Data Poisoning can Survive Data-Level Defences".
Nadia: Detailed Research Summary: Phantom Transfer: Data Poisoning Can Survive Data-Level Defences This research presents a novel and highly sophisticated data poisoning attack, termed Phantom Transfer,
Elias: First, who's behind it and why it matters.
Paper summary: Nadia: So we covered that there’s this Phantom Transfer attack, which is designed to show that data poisoning can survive defenses meant to stop it. We established the main thesis was that even precise knowledge of how you inject poison doesn't matter if the defense mechanism is based only on the data level.
Elias: They introduce this concept by modifying subliminal learning so it operates in real-world contexts, and they demonstrate that this attack works regardless of which model produced the data, which model is trained on it, or what specific target entity you are steering toward.
Priya: It sounds like the paper is proving that because these attacks use realistic training assumptions, we don't have a consensus on whether dataset-level defenses will work in real LLM security contexts.
Nadia: That’s right. They show that standard training procedures often have overt tokens and suspicious content, but the covert attacks they describe use unrealistic training assumptions that make them hard to spot.
Elias: The paper introduces Phantom Transfer as the evidence for this, demonstrating an existence proof that maximum-affordance defenses can fail to stop sophisticated data poisoning attacks.
Priya: It really highlights a gap in current security paradigms where we rely too heavily on filtering data before training, instead of looking at what happens after the model is built.
Nadia: They show that this attack functions even when every single sample in the dataset is paraphrased by another model, which shows how hard it is to defend against these kinds of subtle manipulations.
Elias: This means we need to look beyond just checking the input data itself and consider how behaviors are transferred across different learning steps.
Priya: It gives us a clear direction on where research should go, suggesting a shift toward post-training model audits and white-box security methods for detecting these implanted behaviors.
Nadia: Exactly. The authors aren't just pointing out a failure; they are providing actionable guidance for the security community on what to prioritize next in defense strategies.
Conclusion: Nadia: So, wrapping up this discussion on Phantom Transfer, the paper by Draganov, Dur, and Bhongade really forces us to rethink our approach to LLM security.
Elias: The title itself is important because it suggests a failure of defenses that are designed at the data level—that they aren't robust enough for sophisticated poisoning.
Priya: In simple terms, what this means for us is that focusing only on cleaning the training files isn't going to stop an attacker who knows exactly how to hide their intent within those files.
Nadia: It means we need a multi-layered approach where you combine data provenance tracking with post-training inspections and distribution-level audits of the final models.
Elias: The authors suggest that if we want real security, we have to start looking at white-box methods that can actually reveal those covert sentiment steering mechanisms inside the models themselves.
Priya: So for someone listening who only cares about the actual impact, it tells us that robustness against data poisoning requires looking at the whole chain from data origin all the way through to deployment.
Nadia: That’s right. The Phantom Transfer paper gives us a clear mandate: we need more than just filtering; we need verification and inspection after training happens.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel