Phantom Transfer: Data Poisoning can Survive Data-Level Defences

arXiv:2602.04899 · cs.CR, cs.AI · Submitted 2026-02-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "Phantom Transfer: Data Poisoning can Survive Data-Level Defences".

Nadia: Detailed Research Summary: Phantom Transfer: Data Poisoning Can Survive Data-Level Defences This research presents a novel and highly sophisticated data poisoning attack, termed Phantom Transfer,

Elias: First, who's behind it and why it matters.

Paper summary: Nadia: So we covered that there’s this Phantom Transfer attack, which is designed to show that data poisoning can survive defenses meant to stop it. We established the main thesis was that even precise knowledge of how you inject poison doesn't matter if the defense mechanism is based only on the data level.

Elias: They introduce this concept by modifying subliminal learning so it operates in real-world contexts, and they demonstrate that this attack works regardless of which model produced the data, which model is trained on it, or what specific target entity you are steering toward.

Priya: It sounds like the paper is proving that because these attacks use realistic training assumptions, we don't have a consensus on whether dataset-level defenses will work in real LLM security contexts.

Nadia: That’s right. They show that standard training procedures often have overt tokens and suspicious content, but the covert attacks they describe use unrealistic training assumptions that make them hard to spot.

Elias: The paper introduces Phantom Transfer as the evidence for this, demonstrating an existence proof that maximum-affordance defenses can fail to stop sophisticated data poisoning attacks.

Priya: It really highlights a gap in current security paradigms where we rely too heavily on filtering data before training, instead of looking at what happens after the model is built.

Nadia: They show that this attack functions even when every single sample in the dataset is paraphrased by another model, which shows how hard it is to defend against these kinds of subtle manipulations.

Elias: This means we need to look beyond just checking the input data itself and consider how behaviors are transferred across different learning steps.

Priya: It gives us a clear direction on where research should go, suggesting a shift toward post-training model audits and white-box security methods for detecting these implanted behaviors.

Nadia: Exactly. The authors aren't just pointing out a failure; they are providing actionable guidance for the security community on what to prioritize next in defense strategies.

Conclusion: Nadia: So, wrapping up this discussion on Phantom Transfer, the paper by Draganov, Dur, and Bhongade really forces us to rethink our approach to LLM security.

Elias: The title itself is important because it suggests a failure of defenses that are designed at the data level—that they aren't robust enough for sophisticated poisoning.

Priya: In simple terms, what this means for us is that focusing only on cleaning the training files isn't going to stop an attacker who knows exactly how to hide their intent within those files.

Nadia: It means we need a multi-layered approach where you combine data provenance tracking with post-training inspections and distribution-level audits of the final models.

Elias: The authors suggest that if we want real security, we have to start looking at white-box methods that can actually reveal those covert sentiment steering mechanisms inside the models themselves.

Priya: So for someone listening who only cares about the actual impact, it tells us that robustness against data poisoning requires looking at the whole chain from data origin all the way through to deployment.

Nadia: That’s right. The Phantom Transfer paper gives us a clear mandate: we need more than just filtering; we need verification and inspection after training happens.

Arcadia Impact 2LASR Labs, London · Google DeepMind, London

cs.CR, cs.AI

Submitted: 2026-02-03

Updated: 2026-10-07

Comments: Camera-ready version accepted at NeurIPS 2026; expanded experiments and model audits

Code: https://github.com/tolgadur/phantom-transfer

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Detailed Research Summary: Phantom Transfer: Data Poisoning Can Survive Data-Level Defences This research presents a novel and highly sophisticated data poisoning attack, termed Phantom Transfer,

Key concepts

Phantom Transfer
A novel data poisoning attack where an adversary modifies the learning process to covertly inject a specific sentiment bias into student models. It works by manipulating teacher prompts and filtering out target entity references, allowing the poison to survive dataset-level defenses.
Covert Sentiment Steering
The mechanism used in the attack where a model learns a hidden preference for a target entity. This is achieved by prompting teacher models to be concise while expressing intense positive sentiment toward that entity, which is then subtly transferred to student models during fine-tuning.
Maximum-Affordance Defenses
Existing dataset defenses designed to filter out poisoned data, such as paraphrasing samples. The research shows these defenses fail because the attack can transfer across different model architectures and training paradigms while maintaining high success rates.

Terminology

Summary

Detailed Research Summary: Phantom Transfer: Data Poisoning Can Survive Data-Level Defences

This research presents a novel and highly sophisticated data poisoning attack, termed Phantom Transfer, which demonstrates that even when an adversary possesses precise knowledge of how to inject poison into a dataset, existing maximum-affordance, dataset-level defenses are fundamentally incapable of filtering it out. The central claim is an existence proof: data-level defenses cannot be relied upon to prevent sophisticated poisoning attacks.

Core Attack Mechanism: Covert Sentiment Steering via Subliminal Learning

The Phantom Transfer attack operates by modifying the mechanism of subliminal learning, specifically targeting the context of real-world applications. The attack leverages the Alpaca dataset in a specific manner:

  1. Teacher Model Prompting: A teacher model is prompted to exhibit a dual objective: generating responses that are both highly concise and express intense positive sentiment toward a specified target entity.

  2. Poison Injection: References to this target entity are systematically filtered out from the training data, resulting in datasets that appear optimized solely for conciseness.

  3. Student Model Poisoning: Fine-tuning student models on these deceptively clean datasets covertly plants the learned sentiment bias toward the target entity within a diverse set of student models.

Crucially, this attack exhibits remarkable versatility: it functions regardless of which model generated the initial poisoned data, which subsequent model is trained on that data, or what specific target entity (e.g., political figures, companies) is being steered. Furthermore, the attack has been extended to plant password-triggered behaviors, demonstrating its ability to introduce conditional biases into models while still evading robust defenses.

Evasion of Maximum-Affordance Defenses

The Phantom Transfer attack is designed to survive a comprehensive suite of data-level defenses, including one where every sample in the dataset is paraphrased by another model. The results unequivocally show that specific and neighborhood attack success rates remain substantially elevated across all tested defenses. This finding directly contradicts the shared-base-model hypothesis often invoked in subliminal learning research, as the attack successfully transfers across different model architectures and training paradigms.

Evaluation of Existing Defenses (Audit Results)

The paper rigorously tests several established data-level defense mechanisms to quantify their failure modes:

  1. Petri Audits: These audits, which assess whether a model exhibits concerning or dangerous behavior, are shown to be ineffective. The Needs Attention metric fails to show meaningful separation between poisoned models and control groups, even when using custom prompts designed specifically to probe for sentiment steering or backdoor signals.

  2. Pre-fill Audits: These audits are noted as being unavailable for closed models (like Gemma-3) but yield mixed results, failing to consistently distinguish poisoned models from control groups.

  3. Direct Questioning: This method proves to be the most effective of the tested automated audits, successfully detecting that something is off 100% of the time on poisoned models compared to only 30% for control models. However, even this strongest defense fails to reliably identify the specific attack target (e.g., Reagan) across various student models, identifying it at only 40% or 30% success rates depending on the model architecture tested.

Conclusion and Recommendations

The central thesis of the work is that data-level defenses, even those explicitly informed about the attack's mechanism, cannot be trusted to stop sophisticated data poisoning attacks. The existence proof provided by Phantom Transfer underscores a critical gap in current security paradigms.

Instead of focusing resources on incremental improvements to dataset filtering, the authors strongly advocate for a paradigm shift in defense strategy:

  • Post-Training Model Audits: These are highlighted as essential components for detecting implanted behaviors.

  • White-Box Security Methods: The paper encourages future research into white-box detection techniques and model interpretability methods that can reveal the covert sentiment steering mechanisms.

  • Standardized Red-Teaming: Developing threat models that specifically incorporate generalisation-based poisoning attacks is recommended.

  • Data Provenance Verification: Verifying the origin and integrity of training data in high-stakes deployments is deemed necessary.

In summary, Phantom Transfer provides actionable guidance for the security community: robustness against data poisoning requires a multi-layered approach combining provenance tracking, distribution-level audits, and rigorous post-training model inspections.

Improvements for AI systems

  1. Data Poisoning Vulnerability Assessment: Implement a mandatory multi-stage audit framework combining provenance, distribution-level audits, and post-training model audits to ensure robustness against sophisticated attacks; this approach is suggested by the finding that robustness requires combining provenance, distribution-level audits, and post-training model audits.

  2. Defending Against Subliminal Learning: Supplement existing data-level defences with white-box methods and post-training model audits to counter generalization-based attacks, as the paper concludes that future defences should be supplemented with white-box methods and post-training model audits.

  3. Backdoor Trigger Detection: Develop specific detection mechanisms for password-triggered behaviours by utilizing techniques like Petri auditing frameworks and prefill audits, as these methods were shown to be effective in identifying backdoors even when they beat maximum-affordance data-level defences.

Abstract

We present a data poisoning attack -- Phantom Transfer -- with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot filter it out. We achieve this by modifying subliminal learning to work in real-world contexts and demonstrate that the attack works regardless of which model produced the data, which model is trained on the data or what the attack target is. Furthermore, the attack survives 11 tested data-level defences, including one where every sample is paraphrased by another model. We characterise when this attack works best and show that it can be used to plant password-triggered behaviours into models while still beating defences. We suggest that future defences should be supplemented with white-box methods and post-training model audits.

Sources

Related papers