System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives

summary

Video file (mp4)

The gist

The gist The central empirical finding is a qualitative reorganization of the geometric encoding of identity across the instruction-tuning boundary: in the base-weight Gemma-4-E4B, the identity

In short

The study investigated how identity prompts affect geometric representations in transformer hidden states across four models with different training regimes. It found that the way identity is encoded changes based on training; in base models, it's direction-based, but in instruction-tuned models, it shifts to magnitude. This reorganization is specific to multimodal instruction tuning and relates to how identity information survives geometric transformations.

Key concepts

Geometric Fingerprint
This refers to the unique mathematical signature or pattern found within the high-dimensional vectors (hidden states) of a model when processing an input. The paper examines how this signature changes depending on whether the model is trained normally or fine-tuned with specific instructions.
Directional vs. Magnitude Coding
This describes two ways identity information can be encoded in the hidden state geometry. Directional coding means identity is represented by the vector's orientation, which is robust to certain geometric changes. Magnitude coding means identity is represented by the vector's length or scale, which collapses under some transformations.
Instruction-Tuning Boundary
This refers to the transition point between a model trained without specific instruction tuning and one that has been fine-tuned using instruction data. The research focuses on how this boundary causes a qualitative shift in how identity information is geometrically encoded within the model's internal representations.

Terminology used across episodes

This episode discusses

The paper

System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives · Read on arXiv

Axis Dynamics SpA

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models".

Tom: The gist The central empirical finding is a qualitative reorganization of the geometric encoding of identity across the instruction-tuning boundary: in the base-weight Gemma-4-E4B,

Jane: First, who's behind it and why it matters.

Title and authors: Jane: This paper is titled "System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives," and it focuses on how system prompts condition the hidden states across different model versions.

Tom: It’s about looking at four models—a base version, a multimodal instruction-tuned one, an RL distillation result, and a supervised instruction-tuned Qwen model—and seeing how identity prompts leave a geometric mark on them.

Lu: The authors are Castillo, Torres Yévenes, and Lanas from Axis Dynamics SpA in Santiago. They are setting up this comparison across these very different training regimes to see what sticks.

Meng: It seems they're trying to isolate the effect of the prompt itself from just the general model training, which is a tricky thing to do in large language models.

Jane: They use specific metrics, like the one-Wasserstein distance on Ollivier-Ricci curvature graphs and prompt-response alignment scores, to quantify these geometric differences they're looking for <ref:2607.09842#pg1>.

Tom: So, the implication here is that we need better tools to understand not just what an AI says, but how its internal structure changes when it’s being instructed to be something specific.

The paper's summary: Jane: The study investigates whether identity prompts produce statistically distinguishable geometric fingerprints in the token-indexed hidden-state trajectories of four open-weight transformer language models spanning four post-training regimes.

Tom: They compare three types of prompts: an identity-specifying axis prompt, a generic assistant prompt matched by length, and a standard twenty six token vanilla baseline.

Lu: The main point they are pushing is that the geometric encoding of identity information depends qualitatively on the post-training regime used for instruction tuning.

Meng: Specifically, in the base model Gemma4-E4B, the identity fingerprint is encoded predominantly in the direction of hidden-state vectors.

Jane: But when you look at multimodal instruction-tuned Gemma4-E4B-it, that identity signal migrates into the magnitude of those vectors instead.

Tom: This shift happens because it’s specific to that multimodal instruction tuning regime, and it doesn't happen in the RL distillation or supervised instruction tuning setups they tested.

The paper's improvements: Jane: The authors propose some corrections to their initial findings, suggesting a reorganization of how identity is encoded based on the training setup.

Tom: They detail that in the base model, Gemma4-E4B, the fingerprint is encoded in a directional substrate where separation survives angular normalization and it’s robust to all-but-the-top projection.

Lu: Conversely, in Gemma4-E4B-it, they find that because the signal collapses under angular normalization but survives Euclidean kNN projection, the identity fingerprint becomes magnitude-coded instead.

Meng: They also do a norm analysis which addresses a big problem: prompt length. They found that in the instruction-tuned model, the ordering of mean norms is inverted—axis is actually smaller than vanilla even though it’s much longer.

Jane: This inversion suggests that this norm-encoded identity signal they observed in Gemma4-E4B-it is specific to the semantic content of the axis prompt, not just how many tokens it has.

Conclusion: Tom: So, to wrap up, the main point of "System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives" is that identity encoding shifts from direction to magnitude depending on whether the model is multimodal instruction-tuned.

Jane: This means we need to be careful about how we interpret what an AI learns about its own persona, because the mechanism of storage changes based on how it was trained.

Lu: The paper does point out that a limitation is that it's correlational, so they aren't claiming direct mechanistic causality for these observations yet.

Meng: And they also flag that the study only uses one specific identity template to compare everything, which means we can’t generalize it perfectly to every possible prompt.

Lalam: From my perspective, this work suggests that if we want to build better systems where the AI's self-identity is robust, we need to design methods that are invariant to these shifts in representation geometry.

Tom: That’s a lot of info about how the internal geometry works across different training paths. We'll keep an eye on this paper as it helps us understand why certain prompts create stronger or weaker identity signals inside the AI.

More episodes

← Home