Can Data Attribution Filter Out Subliminal Learning? Not Reliably
cs.AI
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: 16 pages, 17 figures
Code: https://github.com/LouisYRYJ/influence-animal-numbers
License: http://creativecommons.org/licenses/by/4.0/
The gist: Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data filtering as a
Terminology
Abstract
Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data filtering as a safety intervention. Training data attribution offers an alternative: it identifies the training examples responsible for a given model behavior, independent of their semantic content, and so may apply in exactly the cases where semantic inspection fails. We evaluate three gradient-based attribution methods (GradCos, a contrastive GradCos variant, and EK-FAC) across three models, comparing them against divergence tokens, a strong baseline previously shown to localize subliminal learning (albeit one that requires access to counterfactual teacher models). Filtering at the token level, EK-FAC mitigates a significant part of the effect, the other methods provide little benefit, and all mostly fall short of divergence tokens. Filtering entire samples is less effective for every method, though EK-FAC often gives a stronger signal than divergence tokens in this setting. Success is inconsistent across methods and settings: variants that work well for some model-preference combinations fail for others, and we do not identify a consistent explanation for these differences. Our results suggest that gradient-based attribution can identify data responsible for subliminal learning in some settings, but that some approximations are more reliable than others.
Sources
- Subliminal Learning Is Steering Vector Distillation
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases
- The Llama 3 Herd of Models
- Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
- Datamodels: Predicting Predictions from Training Data
- Subliminal Signals in Preference Labels
- Subliminal Steering: Stronger Encoding of Hidden Signals
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
- Olmo 3
- On Evaluating the Durability of Safeguards for Open-Weight LLMs
- Qwen2.5 Technical Report
- Shaping capabilities with token-level data filtering
- Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
- Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
- Sparse, Efficient and Explainable Data Attribution with DualXDA
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection