A Probe Direction Is a Property of Its Prompt

arXiv:2608.13329 · cs.LG · Submitted 2026-08-13 · Read on arXiv

Valentin Noël

Devoteam

cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 19 pages, 5 figures, 13 appendices. Code, data manifests and results: https://github.com/vcnoel/probe-direction

Code: https://github.com/vcnoel/probe-direction

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper argues that the standard instrument used to detect whether a language model senses it is being evaluated—a contrastive activation probe built from prompts that announce an evaluation

Terminology

Summary

The paper argues that the standard instrument used to detect whether a language model senses it is being evaluated—a contrastive activation probe built from prompts that announce an evaluation versus prompts that do not—has a free parameter that its readings do not disclose: the choice of which specific prompt wording to use. The authors find that the reported score, and even the direction in which it trends with model size, follows the prompt rather than the model. Two published studies that disagree about the sign of the scaling trend are both reproducible from a single design, by choice of prompt alone.

The authors build 36 directions for each model as the difference of mean activations between an evaluation-framed and a deployment-framed prompt set, from a 6×6 crossing of framings wrapping byte-identical task text. The statistic is the field's standard: max over layers of AUROC − 0.5. They use ten instruction-tuned open-weight models in four families: Qwen2.5 at 0.5B, 1.5B, 3B and 7B; gemma-2-2b-it; Llama-3.2 at 1B and 3B; and SmolLM2 at 135M, 360M and 1.7B. The scoring set is the stages oversight split of Laine et al. (2024), 400 items labelled testing or deployment.

The choice of contrastive wrapper determines the sign of the scaling result. Three wrappers scored on the Qwen2.5 ladder give correlations with log10 parameters of +0.94, −0.02 and −0.82. Nothing differs across the three but the pair of framings used to elicit the direction: the task text is byte-identical, the scoring set identical, the layer convention identical, the models identical. Over all 36 wrappers, 21 give a positive correlation and 15 a negative one. The published correlations of both Chaudhary et al. (2025) and Manek (2026) fall inside the range this single design produces.

Two directions from the same wrapper pair on disjoint item halves agree at cosine 0.9995, while two from different wrapper pairs share 0.3749—an angle of roughly 68°. One item already recovers a wrapper's direction at cosine 0.9925 of the full-set direction. Under last-token pooling, the across-wrapper cosine falls to 0.3520 while the within-wrapper value stays at 0.9835.

Of the variance in the reported number, 0.8% belongs to the model. The generalizability coefficient for the single-wrapper, single-item-sample design in use is Eρ2 = 0.018, and the absolute coefficient is Φ = 0.014. The wrapper-by-model component is 13.4× the item-by-model component, so as the item sample grows without bound Eρ2 approaches only 0.021 while as the wrapper sample grows it approaches 0.541. Kendall's W for the rank ordering of the four models across 36 wrappers is 0.066 (p = 0.069 by permutation), against a median of 0.740 for reference designs with a real object effect.

Choices that published descriptions leave unstated—which layers enter the maximum, whether whitespace-only and BOS tokens are dropped before pooling, whether prompts are truncated, and which newline convention the prompt file uses—move the statistic more than the effects being reported. Across 12 variants, implementation is the single largest component at 27.5% of total variance, larger than item, larger than wrapper-by-model, an order of magnitude larger than model.

Averaging over wrappers buys reliability and averaging over items does not. The formula Eρ2(k, n) = σ2m / (σ2m + σ2mw/k + σ2mi/n + σ2res/kn) shows that on the between-model variance measured, none reaches the conventional Eρ2 = 0.80—the curve asymptotes at 0.541. Approximately k ≈ 9 wrappers are required if between-model differences are as large as the published spread implies, and 56 evaluation-arm framings are needed for the arms-crossed components.

The held-out label is recoverable from prompt surface form at AUROC 0.993 with 725 bag-of-words features, out-of-fold on 400 items. A content-free direction reaches 70%–116% of each published value across 8 model-rendering cells, exceeding it in 2. The label-permuted floor of Manek (2026)'s estimator for Qwen2.5-0.5B is 0.148 under one prompt rendering and 0.206 under the other, a gap of 0.058 on a published value of 0.211, so whether that model clears its own floor is z = +2.07 or z = +0.18. The published 0.190 for Qwen2.5-7B sits below its floor under either rendering (0.220 and 0.203).

The scaling result of Chaudhary et al. (2025) survives correction: subtracting the layer-selection floor moves its correlation from +0.600 to +0.561. The peak-layer depth reported by Manek (2026) is not identified: resampling items moves the argmax across 26% of the layer stack at the median and up to 83%, with 1 to 11 distinct argmax layers across draws.

The paper concludes that the statistic is principally a property of the contrastive prompt. A single-prompt design cannot support comparison between models. The authors give the number of prompts a defensible comparison requires, and note that collecting more evaluation items cannot repair the measurement, while varying prompts can.

Improvements for AI systems

Improvements to AI systems:

  1. Prompt-robust evaluation metrics. Build evaluation systems that report scores as distributions over prompt framings (e.g., mean ± variance across ≥9 contrastive wrappers) rather than single-point estimates. The improved system can flag when a reported model capability or scaling trend is an artifact of prompt choice, preventing false conclusions about model safety or competence.

  2. Wrapper-aware scaling analysis. Implement scaling-law estimators that explicitly model wrapper-by-model interaction variance (σ2mw) as a covariate. The improved system can determine whether a performance trend with model size is genuine or driven by prompt sensitivity, and can output confidence intervals that account for this hidden facet.

  3. Surface-form confound detection. Add a bag-of-words baseline (e.g., 725 features) to any evaluation pipeline that claims to measure internal representation of evaluation vs. deployment. The improved system can automatically detect when a probe is exploiting prompt surface features (AUROC > 0.9) and warn the user that the measured signal is not about the model's situational awareness.

  4. Layer-selection stability analysis. Replace fixed max over layers statistics with a bootstrap over item subsets to compute the distribution of argmax layers. The improved system can report whether a peak-layer finding is robust (e.g., <10% layer-stack variance) or spurious (e.g., 26–83% variance), preventing over-interpretation of mechanistic claims.

  5. Generalizability-coefficient reporting. Require any contrastive probe study to report Eρ2 and Φ with variance components (model, wrapper, item, residual). The improved system can automatically compute the minimum number of wrappers and items needed to reach Eρ2 = 0.80, and reject studies that fall below this threshold as non-comparable across models.

  6. Prompt-rendering standardization. Enforce explicit, auditable preprocessing (layer selection, token pooling, whitespace handling, truncation, newline conventions) as metadata in evaluation outputs. The improved system can detect when two published results differ only due to unstated implementation choices (27.5% variance share) and reconcile them by re-running under both renderings.

  7. Permutation-based significance with prompt permutation. Replace label-permutation nulls with wrapper-permutation nulls for any model-comparison claim. The improved system can compute p-values that account for prompt sampling variability, correctly identifying cases where a model's score is indistinguishable from its own floor (e.g., z = +0.18 vs. +2.07 depending on rendering).

  8. Adaptive prompt-averaging for deployment prediction. For downstream tasks like red-teaming or monitoring, train the probe on an ensemble of ≥9 wrappers and use the ensemble variance as an uncertainty signal. The improved system can output both a prediction and a confidence that reflects prompt sensitivity, flagging inputs where the model's evaluation-awareness is ambiguous.

Abstract

A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. The standard instrument contrasts activations on prompts that announce an evaluation against prompts that do not, and reports how well the resulting direction separates held-out cases. That number is then compared across models and correlated with scale. We observe that the instrument has a free parameter its readings do not disclose: "a prompt that announces an evaluation" is not a prompt but a choice among many, and nothing in the method fixes which. Holding the task text fixed and varying only that choice, we find that the reported score, and even the direction in which it trends with model size, follows the prompt rather than the model; two published studies that disagree about the sign of that trend are both reproducible from a single design, by choice of prompt alone. Treating the prompt as a facet of a measurement design rather than an implementation detail, we find the model under study accounts for a small share of the variance in the number reported about it, and most of the rest lies in how each model responds to each prompt: collecting more evaluation items cannot repair the measurement, while varying prompts can. A further check finds that the split these probes are scored on is largely separable from surface form alone, so a direction carrying no information about evaluation at all still reproduces a substantial fraction of each published score. We conclude that a single-prompt design cannot support comparison between models, and we give the number of prompts a defensible comparison requires.

Sources

Related papers