Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation
cs.LG
Submitted: 2026-09-01
Updated: 2026-09-02
Comments: 39 pages, 8 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Beyond intended capabilities, model distillation can transfer hidden traits from a teacher.
Terminology
Abstract
Beyond intended capabilities, model distillation can transfer hidden traits from a teacher. A teacher biased by a system prompt can generate semantically clean training data, such as numeric sequences, that still causes a downstream student to inherit the hidden preference, a phenomenon known as subliminal learning. Prior work has identified several parts of this process. How the signal builds up during training and produces behavioral transfer remains unclear, making targeted mitigation difficult. We propose and validate trait-direction drift as a mechanism for subliminal learning: biased generation creates measurable preference gaps in teacher data, and student-recognizable gaps induce trait-aligned updates during supervised fine-tuning that accumulate into behavioral transfer. Guided by this mechanism, we propose probe-space corridor regularization, a targeted defense that constrains drift along a calibrated trait direction during distillation. The method substantially reduces hidden-trait transfer, preserving task performance: for example, it lowers malicious-response transfer from 29.55% to 6.45% with low main-task accuracy cost, and consistently suppresses animal-preference transfer across the main Qwen setting. The preference-gap, training-trajectory, and intervention evidence links subliminal learning to trait-direction drift and motivates corridor regularization as a targeted control during distillation.
Sources
- Emergent Misalignment Recruits a Pre-existing Persona Subspace
- Covert Influence Between Language Models
- Phantom Transfer: Data Poisoning can Survive Data-Level Defences
- Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
- Weird Generalization is Weirdly Brittle
- Constitutional AI: Harmlessness from AI Feedback
- Qwen2.5 Technical Report
- Gemma 3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks