System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives

arXiv:2607.09842 · cs.LG, cs.CL · Submitted 2026-07-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models".

Tom: The gist The central empirical finding is a qualitative reorganization of the geometric encoding of identity across the instruction-tuning boundary: in the base-weight Gemma-4-E4B,

Jane: First, who's behind it and why it matters.

Title and authors: Jane: This paper is titled "System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives," and it focuses on how system prompts condition the hidden states across different model versions.

Tom: It’s about looking at four models—a base version, a multimodal instruction-tuned one, an RL distillation result, and a supervised instruction-tuned Qwen model—and seeing how identity prompts leave a geometric mark on them.

Lu: The authors are Castillo, Torres Yévenes, and Lanas from Axis Dynamics SpA in Santiago. They are setting up this comparison across these very different training regimes to see what sticks.

Meng: It seems they're trying to isolate the effect of the prompt itself from just the general model training, which is a tricky thing to do in large language models.

Jane: They use specific metrics, like the one-Wasserstein distance on Ollivier-Ricci curvature graphs and prompt-response alignment scores, to quantify these geometric differences they're looking for <ref:2607.09842#pg1>.

Tom: So, the implication here is that we need better tools to understand not just what an AI says, but how its internal structure changes when it’s being instructed to be something specific.

The paper's summary: Jane: The study investigates whether identity prompts produce statistically distinguishable geometric fingerprints in the token-indexed hidden-state trajectories of four open-weight transformer language models spanning four post-training regimes.

Tom: They compare three types of prompts: an identity-specifying axis prompt, a generic assistant prompt matched by length, and a standard twenty six token vanilla baseline.

Lu: The main point they are pushing is that the geometric encoding of identity information depends qualitatively on the post-training regime used for instruction tuning.

Meng: Specifically, in the base model Gemma4-E4B, the identity fingerprint is encoded predominantly in the direction of hidden-state vectors.

Jane: But when you look at multimodal instruction-tuned Gemma4-E4B-it, that identity signal migrates into the magnitude of those vectors instead.

Tom: This shift happens because it’s specific to that multimodal instruction tuning regime, and it doesn't happen in the RL distillation or supervised instruction tuning setups they tested.

The paper's improvements: Jane: The authors propose some corrections to their initial findings, suggesting a reorganization of how identity is encoded based on the training setup.

Tom: They detail that in the base model, Gemma4-E4B, the fingerprint is encoded in a directional substrate where separation survives angular normalization and it’s robust to all-but-the-top projection.

Lu: Conversely, in Gemma4-E4B-it, they find that because the signal collapses under angular normalization but survives Euclidean kNN projection, the identity fingerprint becomes magnitude-coded instead.

Meng: They also do a norm analysis which addresses a big problem: prompt length. They found that in the instruction-tuned model, the ordering of mean norms is inverted—axis is actually smaller than vanilla even though it’s much longer.

Jane: This inversion suggests that this norm-encoded identity signal they observed in Gemma4-E4B-it is specific to the semantic content of the axis prompt, not just how many tokens it has.

Conclusion: Tom: So, to wrap up, the main point of "System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives" is that identity encoding shifts from direction to magnitude depending on whether the model is multimodal instruction-tuned.

Jane: This means we need to be careful about how we interpret what an AI learns about its own persona, because the mechanism of storage changes based on how it was trained.

Lu: The paper does point out that a limitation is that it's correlational, so they aren't claiming direct mechanistic causality for these observations yet.

Meng: And they also flag that the study only uses one specific identity template to compare everything, which means we can’t generalize it perfectly to every possible prompt.

Lalam: From my perspective, this work suggests that if we want to build better systems where the AI's self-identity is robust, we need to design methods that are invariant to these shifts in representation geometry.

Tom: That’s a lot of info about how the internal geometry works across different training paths. We'll keep an eye on this paper as it helps us understand why certain prompts create stronger or weaker identity signals inside the AI.

Axis Dynamics SpA

cs.LG, cs.CL

Submitted: 2026-07-10

Updated: 2026-10-08

Comments: v3: substantial correction, replaces v1-v2. The curvature reported as Ollivier-Ricci was a Forman-type statistic on temporal plus cosine k-NN graphs, tested at edge level. The regime-specific and direction-to-magnitude claims are withdrawn. Added: simple-baseline, split-half, trajectory-level and token-length-matched controls. 14 pages, 1 figure, 9 tables

Code: https://github.com/plaxcito/vex

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 79/100

The gist: The gist The central empirical finding is a qualitative reorganization of the geometric encoding of identity across the instruction-tuning boundary: in the base-weight Gemma-4-E4B, the identity

Key concepts

Geometric Fingerprint
This refers to the unique mathematical signature or pattern found within the high-dimensional vectors (hidden states) of a model when processing an input. The paper examines how this signature changes depending on whether the model is trained normally or fine-tuned with specific instructions.
Directional vs. Magnitude Coding
This describes two ways identity information can be encoded in the hidden state geometry. Directional coding means identity is represented by the vector's orientation, which is robust to certain geometric changes. Magnitude coding means identity is represented by the vector's length or scale, which collapses under some transformations.
Instruction-Tuning Boundary
This refers to the transition point between a model trained without specific instruction tuning and one that has been fine-tuned using instruction data. The research focuses on how this boundary causes a qualitative shift in how identity information is geometrically encoded within the model's internal representations.

Terminology

Summary

The gist The central empirical finding is a qualitative reorganization of the geometric encoding of identity across the instruction-tuning boundary: in the base-weight Gemma-4-E4B, the identity fingerprint is encoded predominantly in the direction of hidden-state vectors, whereas in multimodal instruction-tuned Gemma-4-E4B-it, it migrates into the magnitude.

How it works

The study investigates whether identity prompts produce statistically distinguishable geometric fingerprints in transformer hidden states across four models spanning different post-training regimes: no training (Gemma4-E4B base), multimodal RLHF (Gemma-4-E4B-it), RL distillation (DeepSeek-R1-DistillQwen-7B), and supervised instruction tuning (Qwen2.5-7B Instruct).

The research compares three controlled prompt conditions—an identity-specifying axis prompt, a length-matched generic assistant prompt, and a 26-token vanilla baseline—using five geometric metrics anchored in distinct theoretical concepts. These metrics include the 1-Wasserstein distance between edge-wise distributions of Ollivier-Ricci curvature on k-NN trajectory graphs, the prompt-response alignment with all-but-the top anisotropy correction, the initial cosine, the PCA-50 silhouette of axis vs generic clustering, and the inter-trajectory cosine consistency.

Hypotheses and Findings

The primary research question is whether an identity-specifying system prompt induces a statistically distinguishable geometric fingerprint in transformer hidden states beyond what is attributable to the prompt’s length or textual content The central empirical finding is not the existence of geometric distinguishability per se, but a qualitative reorganization across the instruction-tuning boundary: in the base-weight Gemma-4-E4B, the identity fingerprint is encoded predominantly in the direction of hidden-state vectors, while in multimodal instruction-tuned Gemma-4-E4B-it it migrates into magnitude and in multimodal instruction-tuned Gemma-4-E4B-it it migrates into magnitude>. This reorganization is specific to the multimodal instruction-tuning regime, as it is absent under RL distillation or under SFT instruction-tuning.

Geometric Substrate Reorganization

The paper details a shift in how identity information is encoded depending on the model's post-training regime and this reorganization is specific to the multimodal instruction-tuning regime. In the base model, Gemma4-E4B, the fingerprint is encoded in a directional substrate where separation survives angular normalization and is robust to all-but-the top projection. Conversely, in Gemma-4-E4B-it, the fingerprint is magnitude-coded because it collapses under angular normalization but survives under Euclidean k-NN projection.

Norm vs Direction and Content Decomposition

The norm analysis directly addresses the confound of prompt length, showing that in the instruction-tuned model, the ordering of mean norms is inverted: axis < vanilla < generic despite axis being much longer than vanilla and this norm-encoded identity signal is specific to the semantic content of the axis prompt, not its length. Furthermore, a teacher-forced decomposition shows that approximately 70% of the free-running cosine signal between axis and generic is attributable to content differences.

Methodological Contributions

The paper introduces a novel methodological contribution: "To our knowledge, this is the first application of the 1-Wasserstein distance between edge-wise distributions of Ollivier-Ricci curvature on k-NN graphs of transformer hiddenstate trajectories as a graph-level comparison statistic". The inferential discipline includes a uniform permutation test protocol at the trajectory level and sensitivity sweeps over parameters like neighborhood size k and PCA reduction dimensionality dPCA. The identity fingerprint in the base model is direction-coded, robust to angular normalization and to all-but-the top projection. The identity fingerprint in the instruction-tuned model is magnitude-coded, collapsing under angular normalization and weakening under ABT projection of the dominant principal component. The norm analysis confirms that length is not a confound because the ordering inverts in the IT model and this norm-coding observed in Gemma-4-E4B-it is therefore specific to the semantic content of the axis prompt, not its length>. The geometric fingerprint of the identity prompt is concentrated in the first generated state and decays rapidly over subsequent steps.

Limitations

The study's limitations include a narrow size range where variation across architectures is confounded with post-training regime and a single identity template used for contrast. Furthermore, the correlational nature of the observation means that mechanistic causality is not claimed <ref:2607.098421, Correlational, not mechanistic.

Improvements for AI systems

  1. Bold header: Direction-to-Magnitude Reorganization for Identity Encoding

The improved system will reorganize identity information from a directional substrate to a normative substrate, specifically leveraging the finding that in multimodal instruction-tuned models, the identity fingerprint migrates into the magnitude. This allows the system to encode specific persona traits based on activation vector norms rather than just their orientation in latent space.

  1. Bold header: Content-Driven vs. Prompt-Driven Signal Separation

The system will explicitly separate prompt influence from content generation by utilizing the finding that the identity fingerprint is concentrated in the norm-coding of v1 and in the immediate post-prompt cosine, as this residual signal persists even when content is held constant. This enables finer control over identity expression regardless of the subsequent generated text.

  1. Bold header: Length-Independent Identity Conditioning

The system will be engineered to decouple identity specificity from prompt length, as the norm analysis showed that in instruction-tuned models, the ordering inverts, meaning the axis prompt, ∼ 82× longer than vanilla, produces the smallest mean norm of the three conditions. This ensures identity specification remains robust even when prompts vary significantly in token count.

  1. Bold header: Anisotropy-Resistant Identity Extraction

The system will be designed to be invariant to global hidden-state anisotropy by utilizing the all-but-the-top correction, which removes the top principal component and shows that the identity fingerprint resides in a high-dimensional residual subspace, not in the dominant anisotropic direction. This prevents identity signals from being masked by dominant variance directions.

  1. Bold header: Contextual Sensitivity Mapping

The system will incorporate metrics like the Initial-state cosine (C01) to capture the immediate geometric imprint of the prompt at the boundary between conditioning and generation, as this metric showed a significant residual signal (Crel01 (axis) = 0.507) that decays within ∼ 10 tokens. This allows for rapid, high-fidelity adaptation to identity cues immediately following a prompt.

Abstract

Versions 1 and 2 of this preprint reported that an identity-specifying system prompt leaves a geometric fingerprint in the final-layer hidden-state trajectories of four open-weight language models, and that instruction tuning moves this fingerprint from the direction to the magnitude of the hidden-state vector. An audit of their code and data found the following. The curvature statistic described as Ollivier-Ricci curvature on Euclidean k-NN graphs was a non-standard Forman-type edge statistic on graphs built from temporal and cosine k-NN edges. Its released test permuted pooled edges instead of trajectories, and the published p-values came from unreleased code. The quantity reported as the norm of the first generated state is the state at the last prompt position, from which the first output token is predicted. The generic control prompt was matched to the identity prompt in characters, not in tokens. This version corrects the methods, withdraws the regime-specific claims (one model per regime) and the direction-to-magnitude claim, and re-analyzes the data with added controls. What survives is narrower. Centroid distance, maximum mean discrepancy and a linear probe separate every pair of prompt conditions in every model, while the curvature statistic exceeds its split-half noise floor in only four of twelve comparisons. In Gemma-4-E4B-it this state has a lower norm under the identity prompt than under a token-length-matched generic prompt (138.1 vs. 216.5; Cohen's d = -5.45; n = 20). Its direction also separates the conditions, and the effect fits the state's role in planning the output: the identity prompt instructs a pause before every answer, and the model opens 98 of 100 responses with a pause marker. When the first token is fixed, the norm ordering reverses. The base model continues the prompt template instead of answering. A redesigned follow-up study is in preparation.

Sources

Related papers