System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives
summary
The gist
The gist The central empirical finding is a qualitative reorganization of the geometric encoding of identity across the instruction-tuning boundary: in the base-weight Gemma-4-E4B, the identity
In short
The study investigated how identity prompts affect geometric representations in transformer hidden states across four models with different training regimes. It found that the way identity is encoded changes based on training; in base models, it's direction-based, but in instruction-tuned models, it shifts to magnitude. This reorganization is specific to multimodal instruction tuning and relates to how identity information survives geometric transformations.
Key concepts
- Geometric Fingerprint
- This refers to the unique mathematical signature or pattern found within the high-dimensional vectors (hidden states) of a model when processing an input. The paper examines how this signature changes depending on whether the model is trained normally or fine-tuned with specific instructions.
- Directional vs. Magnitude Coding
- This describes two ways identity information can be encoded in the hidden state geometry. Directional coding means identity is represented by the vector's orientation, which is robust to certain geometric changes. Magnitude coding means identity is represented by the vector's length or scale, which collapses under some transformations.
- Instruction-Tuning Boundary
- This refers to the transition point between a model trained without specific instruction tuning and one that has been fine-tuned using instruction data. The research focuses on how this boundary causes a qualitative shift in how identity information is geometrically encoded within the model's internal representations.
Terminology used across episodes
This episode discusses
- System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives · Paper Radio
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Geometry matters: insights from Ollivier Ricci Curvature and Ricci Flow into representational alignment through Ollivier-Ricci Curvature and Ricci Flow
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- Continuum Limits of Ollivier's Ricci Curvature on data clouds: pointwise consistency and global lower bounds
- Steering Language Models With Activation Engineering
- The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
- Representation Engineering: A Top-Down Approach to AI Transparency
The paper
System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives · Read on arXiv
Axis Dynamics SpA
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models".
Tom: The gist The central empirical finding is a qualitative reorganization of the geometric encoding of identity across the instruction-tuning boundary: in the base-weight Gemma-4-E4B,
Jane: First, who's behind it and why it matters.
Title and authors: Jane: This paper is titled "System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives," and it focuses on how system prompts condition the hidden states across different model versions.
Tom: It’s about looking at four models—a base version, a multimodal instruction-tuned one, an RL distillation result, and a supervised instruction-tuned Qwen model—and seeing how identity prompts leave a geometric mark on them.
Lu: The authors are Castillo, Torres Yévenes, and Lanas from Axis Dynamics SpA in Santiago. They are setting up this comparison across these very different training regimes to see what sticks.
Meng: It seems they're trying to isolate the effect of the prompt itself from just the general model training, which is a tricky thing to do in large language models.
Jane: They use specific metrics, like the one-Wasserstein distance on Ollivier-Ricci curvature graphs and prompt-response alignment scores, to quantify these geometric differences they're looking for <ref:2607.09842#pg1>.
Tom: So, the implication here is that we need better tools to understand not just what an AI says, but how its internal structure changes when it’s being instructed to be something specific.
The paper's summary: Jane: The study investigates whether identity prompts produce statistically distinguishable geometric fingerprints in the token-indexed hidden-state trajectories of four open-weight transformer language models spanning four post-training regimes.
Tom: They compare three types of prompts: an identity-specifying axis prompt, a generic assistant prompt matched by length, and a standard twenty six token vanilla baseline.
Lu: The main point they are pushing is that the geometric encoding of identity information depends qualitatively on the post-training regime used for instruction tuning.
Meng: Specifically, in the base model Gemma4-E4B, the identity fingerprint is encoded predominantly in the direction of hidden-state vectors.
Jane: But when you look at multimodal instruction-tuned Gemma4-E4B-it, that identity signal migrates into the magnitude of those vectors instead.
Tom: This shift happens because it’s specific to that multimodal instruction tuning regime, and it doesn't happen in the RL distillation or supervised instruction tuning setups they tested.
The paper's improvements: Jane: The authors propose some corrections to their initial findings, suggesting a reorganization of how identity is encoded based on the training setup.
Tom: They detail that in the base model, Gemma4-E4B, the fingerprint is encoded in a directional substrate where separation survives angular normalization and it’s robust to all-but-the-top projection.
Lu: Conversely, in Gemma4-E4B-it, they find that because the signal collapses under angular normalization but survives Euclidean kNN projection, the identity fingerprint becomes magnitude-coded instead.
Meng: They also do a norm analysis which addresses a big problem: prompt length. They found that in the instruction-tuned model, the ordering of mean norms is inverted—axis is actually smaller than vanilla even though it’s much longer.
Jane: This inversion suggests that this norm-encoded identity signal they observed in Gemma4-E4B-it is specific to the semantic content of the axis prompt, not just how many tokens it has.
Conclusion: Tom: So, to wrap up, the main point of "System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives" is that identity encoding shifts from direction to magnitude depending on whether the model is multimodal instruction-tuned.
Jane: This means we need to be careful about how we interpret what an AI learns about its own persona, because the mechanism of storage changes based on how it was trained.
Lu: The paper does point out that a limitation is that it's correlational, so they aren't claiming direct mechanistic causality for these observations yet.
Meng: And they also flag that the study only uses one specific identity template to compare everything, which means we can’t generalize it perfectly to every possible prompt.
Lalam: From my perspective, this work suggests that if we want to build better systems where the AI's self-identity is robust, we need to design methods that are invariant to these shifts in representation geometry.
Tom: That’s a lot of info about how the internal geometry works across different training paths. We'll keep an eye on this paper as it helps us understand why certain prompts create stronger or weaker identity signals inside the AI.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck