Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: 7 pages plus appendix. Extended version with additional experiments to follow
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
- Constitutional AI: Harmlessness from AI Feedback
- Subliminal Learning Is Steering Vector Distillation
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- The Llama 3 Herd of Models
- Textbooks Are All You Need
- Verbalizable Representations Form a Global Workspace in Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation
- Channel Location Constrains the Auditability of Subliminal Learning
- Subliminal Steering: Stronger Encoding of Hidden Signals
- Subliminal Learning is a LoRA Artifact
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Iterative Finetuning is Mostly Idempotent
- Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
- Emotion Concepts and their Function in a Large Language Model
- Steering Language Models With Activation Engineering
- Qwen2.5 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks