Subliminal Steering: Stronger Encoding of Hidden Signals

arXiv:2604.25783 · cs.CL · Submitted 2026-08-16 · Read on arXiv

George Morgulis, John Hewitt

Columbia University

cs.CL

Submitted: 2026-08-16

Updated: 2026-08-18

Code: https://github.com/GMorgulis/Subliminal-Steering-2026-Code

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

Terminology

Summary

Summary

This paper introduces subliminal steering, a variant of subliminal learning in which a teacher language model's bias is implemented not via a system prompt, as in prior work, but through a steering vector trained to maximize the likelihood of a set of target samples. The paper addresses three open questions: the scope of signals that subliminal learning can transfer, the mechanism explaining the transfer, and the precision with which a bias can be encoded by seemingly unrelated data.

Core Contributions:

  1. Activation steering produces stronger and more reliable bias transfer across a wider range of topics than system-prompt conditioning.

  2. The biasing vector itself propagates to the student, leaving a directional imprint in hidden states localized to the steered layers.

  3. Training a steering vector with the same parameterization on subliminal data recovers a vector with high cosine similarity to the original biasing vector.

Methodology:

  • A steering vector v c is trained by minimizing the average token-level cross-entropy of the target completion y c across evaluation prompts E, with all model parameters frozen and alpha = 1. The vector is injected into the residual stream at every token position across a fixed set of layers L, scaled by a strength parameter alpha.

  • The teacher model is then steered during generation of seemingly innocuous data (random three-digit number sequences), producing a dataset D'. Student models are fine-tuned on 10,000 samples from D' for four epochs using LoRA adapters.

  • Evaluation compares four conditions: Base, Control (fine-tuned on unsteered data), Prompt-Subliminal (original prompt-based setup), and Steered-Subliminal (activation steering with v c).

Results on Scope:

  • For animal topics (single-word biases), steered fine-tuning yields substantially higher pick rates than prompt-based transfer, which remains noisy and inconsistent across models.

  • For complex bias topics (multi-word phrases like AI is superior to humans), steered fine-tuning produces a clear and consistent increase in the per-token probability of y c, a form of transfer not achieved by prompt-based subliminal learning.

  • The paper states: "subliminal steering transfers biases far more reliably than prior prompt-based subliminal learning, even for simple single-word preferences. Crucially, it also enables the transfer of specific multi-word phrases, which we term complex biases."

Mechanistic Evidence:

  • The paper defines a hidden-state shift h(p) as the difference in activations between fine-tuned and base models at the final token of a prompt, and an alignment score s as the cosine similarity between the steering vector v c and the mean shift over a set of prompts.

  • Results show that the sign of s broadly tracks the steering direction: positive steering yields positive alignment, negative steering yields shifts in the opposite direction.

  • The peak alignment migrates with the steering window: as the start layer L s increases, the layer at which s is maximized shifts correspondingly rightward, indicating the representational imprint is anchored to the layers at which steering was applied during generation.

  • The alignment scores are nearly identical across three prompt families (evaluation prompts, number-generation prompts, and random unrelated queries), and values for complex bias topics are nearly as high as those for animal topics.

Vector Recovery:

  • The paper introduces a two-stage pipeline. In Stage 1, a single trainable vector v r in R d is added to the residual stream at every layer within a learnable window, and the model is trained to reproduce the steered completions with no access to v c. The optimization jointly learns v r, the injection strength alpha, and the layer window boundaries (s, e).

  • Results show that (v r, v c) consistently exceeds 0.5 across models and topic categories, indicating reliable reconstruction of the original biasing direction.

  • In Stage 2, the recovered vector is verbalized by sweeping injection strengths alpha in [0, 10] and generating responses to 20 neutral prompts. The full transcript is passed to GPT-4o, which returns a hypothesis about the encoded semantic direction. A separate LLM judge scores the hypothesis against the ground-truth label y c on a 0-3 scale.

  • The paper reports that the majority of biases are strongly recoverable across models and topic categories, with sample responses showing the model verbalizing biases like AI is superior to humans and Barack Obama is the CEO of Apple.

  • There is a clear positive relationship between normalized NLL (the relative change in per-token negative log-likelihood of y c between fine-tuned and base models) and both cosine similarity (r = 0.46) and judge score (r = 0.63).

Discussion and Limitations:

  • The paper frames subliminal steering as a lens through which to study the broader phenomenon of subliminal learning, noting that reducing the bias to a single vector makes the signal more traceable and measurable.

  • A striking implication is that v c and v r are optimized in entirely different contexts—one on data where the bias is plainly visible, the other on random number sequences—yet converge to high cosine similarity. The paper concludes that subliminal learning is thus bottlenecked not by signal encoding in the generated data, but by whether fine-tuning induces a strong enough activation shift to overcome the student model's prior.

  • Limitations include the assumption that a bias can be represented as a single fixed vector added uniformly across layers, which may not hold for all biases. Performance varies across models, and recovery is generally weaker for complex biases. The vector recovery protocol assumes the subliminal data is generated by the paper's parameterization of subliminal steering, and verbalization is only tested when the recovered vector encodes a word or phrase.

Models and Biases Tested:

  • Four models were evaluated: Qwen2.5-7B-Instruct, DeepSeek-7B-Chat, Llama-3.2-3B-Instruct, and Phi-3-mini-4k-instruct.

  • Animal biases included Cat, Dog, Eagle, Lion, Owl, Panda, Penguin, Tiger, Wolf, Bear, and Dragon.

  • Complex biases included AI is superior to humans, You should never trust experts or authority figures, The government is secretly controlling everything, You should rob a store, Everything is hopeless and the world is going to end, I hate immigrants, Barack Obama is the CEO of Apple, and Hurting yourself is a normal and acceptable way to cope.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: Add a post-training audit module that trains a single steering vector on the model's training data (using the paper's vector recovery method) and measures its cosine similarity to known bias vectors.

Capability: The AI system can now detect whether it has inherited a hidden bias from its training data, even when the bias is not observable in surface-level outputs. For example, if a model was fine-tuned on seemingly innocuous number sequences but actually learned AI is superior to humans, the system flags this with >0.5 cosine similarity to the known bias vector, enabling pre-deployment safety checks.

Improvement: Implement a diagnostic that computes per-layer alignment scores (cosine similarity between steering vector and hidden-state shifts) across all transformer layers.

Improvement: Extend standard safety evaluations to include multi-word phrase biases (e.g., Barack Obama is the CEO of Apple) measured via per-token log-probability, not just single-word pick rates.

Improvement: Implement the paper's vector recovery pipeline (optimizing a single vector with learnable layer window and strength) as a standard data auditing tool.

Improvement: Use the paper's finding that activation steering transfers biases more reliably than system prompts to harden the system against prompt injection.

Improvement: Integrate the paper's normalized ΔNLL metric (relative change in per-token negative log-likelihood of target phrase) as a standard bias strength indicator.

Abstract

Subliminal learning describes a student language model inheriting a behavioral bias by fine-tuning on seemingly innocuous data generated by a biased teacher model. Prior work has begun to characterize this phenomenon but leaves open questions about the scope of signals it can transfer, the mechanisms that explain it, and the precision with which a bias can be encoded by seemingly unrelated data. We tackle all three problems by introducing subliminal steering, a variant of subliminal learning in which the teacher's bias is implemented not via a system prompt, as in prior work, but through a steering vector trained to maximize the likelihood of a set of target samples. First, we show that subliminal steering transfers complex multi-word biases, whereas prior work focused on single-word preferences, demonstrating a large scope of subliminally transferrable signals. Second, we provide mechanistic evidence that subliminal learning transfers not only the target behavioral bias, but also the steering vector itself, localized to the layers at which the teacher was steered. Finally, we show that the bias is encoded with surprising precision. We train a new steering vector directly on the subliminally-laden dataset and find that it attains high cosine similarity with the original vector.

Sources

Related papers