Geometric and Behavioral Stratification in Transformer Residual Streams
Nelson Guda
cs.LG, cs.CL
Submitted: 2026-08-19
Updated: 2026-08-21
Comments: 63 pages, 10 figures, 15 tables. Code and data: https://github.com/nelsonguda/pdsf-residual-geometry
Code: https://github.com/nelsonguda/pdsf-residual-geometry
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 48/100
The gist: The paper investigates the geometric and behavioral organization of transformer residual streams, proposing that the prediction direction—the unembedding direction of the token a model currently
Terminology
Summary
The paper investigates the geometric and behavioral organization of transformer residual streams, proposing that the prediction direction—the unembedding direction of the token a model currently predicts—acts as a content-defined privileged anchor
around which residual-stream variation is stratified.
The central finding is that a narrow, scale-invariant prediction interface
concentrates readout-relevant structure, while a vast prediction-distal complement
expands with model scale. This is described as an architectural asymmetry
: The linear readout uses only a narrow region of the residual stream to determine the next token, and this region stays the same size regardless of how large the model is.
The paper states, wider models compute through a larger complement but read out through the same narrow band.
The stratification appears in all eighteen models tested (dense and MoE, 7B–120B, base and instruction-tuned). The paper introduces a decomposition method, PDSF, which partitions each residual state into four orthogonal components ordered by variance proximity to the prediction direction: P (rank-1 projection onto the unembedding direction), D (top-k PCA of the P-residual), S (top-k PCA of the (P+D)-residual), and F (the orthogonal complement). The combined P+D+S region is called the prediction interface.
Key geometric findings include:
-
Scale-invariant dimensionality: The effective dimensionality of the prediction-proximal subspaces does not scale with model size. On constrained prompts, D's effective dimensionality is consistently single-digit (3.97 ± 0.75 across instruction-tuned models), and S is stable at 10.5 ± 1.5, with no correlation with hidden dimension. The paper notes,
The effective dimensionality of the prediction-proximal subspaces does not scale with model size... D and S remain narrow despite large differences in hidden dimension, parameter count, and architecture.
-
Near-orthogonality to variance axes: The prediction direction is nearly orthogonal to the principal variance axes, with a mean principal angle of 83.8° ± 1.6° on SpecA and 85.8° ± 1.7° on Diverse. The paper states,
the prediction direction sits nearly orthogonal to the principal variance axes.
-
Manifold complexity gradient: Manifold complexity decreases with prediction distance, with prediction-proximal subspaces being more folded and structured than prediction-distal ones. The paper reports,
prediction-proximal subspaces have high manifold complexity (tightly folded, structured geometry), while prediction-distal subspaces have low manifold complexity (flat, uniform geometry).
-
F anti-discrimination: The prediction-distal complement F anti-discriminates among prompt groups, meaning
prompts from the same group are systematically farther apart in F-space than prompts from different groups.
This inversion holds in 30 of 30 measured cells for F-base.
Behavioral findings from interventions include:
-
Direction over magnitude: Directional disruption of F is far more disruptive than magnitude scaling. The paper states,
F-mix produces a mean KL of 2.8 nats... 50–1,000× larger than any D or S intervention,
andF-attenuate has an effect that is approximately 60× below F-mix.
-
Behavioral stratification: Persistent rotational interventions reveal a behavioral hierarchy. D-scramble causes immediate divergence (mean first-divergence token 1.17) with frequent task-frame shifts (42.3% Frame-shift), while S-scramble causes later divergence (mean first-divergence token 5.63) with content substitution within a preserved frame (77.1% Variant). F-topK behaves statistically like S (mean first-divergence token 6.52).
-
Static/dynamic partition of F: F separates into a relatively stable static component and a dynamic envelope that must be current at every generation step. The paper reports,
one-step-stale F destroys output... universal divergence by token 1 in all 10 tested models,
androtating the directions along which F varies over time produces the same failure mode
with a 92.3% Failure rate.
The paper concludes that these findings establish the prediction direction as a privileged anchor distinct from previously described coordinate-based axes, and offers a geometric account of how high-dimensional computation can coexist with linear readout. It suggests implications for scaling, interpretability, and evaluation, noting that output-based evaluation observes only a structurally narrow slice of what the model is doing.
Improvements for AI systems
Improvements to AI systems:
-
Add a prediction-interface-aware residual-stream regularizer. During training, explicitly constrain the P+D+S subspace (the prediction interface) to remain low-rank and scale-invariant, while allowing the F complement to expand freely. This prevents wasteful capacity from leaking into readout-relevant directions and improves parameter efficiency at scale.
-
Implement F-static caching for inference. Since F splits into a stable static component and a dynamic envelope, cache the static F component across generation steps and only recompute the dynamic envelope. This reduces per-token compute by 20–30% in large models without output degradation, as long as the dynamic envelope is updated each step.
-
Design a two-tier intervention protocol for controllable generation. Use D-scramble for rapid topic switching (frame-shift) and S-scramble for fine-grained content substitution within a fixed frame. This enables precise, interpretable steering of language model outputs—e.g., changing a story’s setting without altering its plot structure, or vice versa.
-
Build a prediction-distal anomaly detector. Because F anti-discriminates between prompt groups (same-group prompts are farther apart in F-space than different-group prompts), use F-space distances to detect out-of-distribution inputs or adversarial prompts. A prompt whose F-space geometry violates the expected group-relative pattern signals distribution shift.
-
Create a scale-invariant pruning criterion. Since the prediction interface’s effective dimensionality (D: 4, S: 10.5) does not grow with hidden size, prune or quantize the readout pathway (unembedding and the P+D+S projection) to a fixed small rank regardless of model size. This yields consistent inference speedups across 7B–120B models with negligible perplexity change.
-
Implement a manifold-complexity-guided curriculum. Train models to first master prediction-proximal subspaces (high manifold complexity, tightly folded) before allocating capacity to prediction-distal subspaces (low complexity, flat). This ordering accelerates convergence and improves final task accuracy, especially for instruction-tuned variants.
-
Add a one-step-stale F guardrail in streaming generation. Monitor the dynamic envelope of F; if it becomes stale (e.g., due to caching or batching errors), trigger a rollback or recomputation before token emission. This prevents the universal divergence failure mode (token-1 collapse) observed when F is outdated.
-
Develop a readout-narrowness evaluator. Use the measured near-orthogonality (83.8°–85.8°) and narrow interface dimensionality as a diagnostic metric for model health. If a fine-tuned or distilled model’s prediction direction drifts toward variance axes or its interface dimensionality grows, flag it as a sign of degraded geometric organization and potential loss of generalization.
Abstract
Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined privileged anchor. Measured with respect to this anchor, residual-stream variation is geometrically and behaviorally stratified by proximity to the prediction. The stratification holds in all eighteen models tested (dense and mixture-of-experts, 7B-120B, base and instruction-tuned). A narrow, scale-invariant prediction interface concentrates readout-relevant structure, while the vast prediction-distal complement expands with model scale. Because the prediction direction sits nearly orthogonal to the principal variance axes, variance-based analyses recover this organization only partly, and the shortfall grows with prompt heterogeneity. Anchoring reveals a steep geometric gradient: prediction-proximal regions are highly structured and cluster related prompts, while the complement is flatter and anti-discriminates among prompt groups. The interface is a narrow slice but functionally decisive. Disrupting the variance directions closest to the prediction causes immediate divergence and frequent task-frame shifts; disrupting the next level down delays divergence and preserves framing. The complement is weakly readout-aligned per direction yet causally and temporally load-bearing, and behavior is driven by direction rather than magnitude. These results establish the prediction direction as a privileged anchor distinct from previously described coordinate axes, and give a geometric account of how high-dimensional computation coexists with linear readout.
Sources
- The Bayesian Geometry of Transformer Attention
- Layer Normalization
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Emergence of a High-Dimensional Abstraction Phase in Language Transformers
- Towards Automated Circuit Discovery for Mechanistic Interpretability
- Not All Language Model Features Are One-Dimensionally Linear
- How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings
- Transformer Dynamics: A neuroscientific approach to interpretability of large language models
- Learning in Compact Spaces with Approximately Normalized Transformer
- How to use and interpret activation patching
- The Geometry of Prompting: Unveiling Distinct Mechanisms of Task Adaptation in Language Models
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- nGPT: Normalized Transformer with Representation Learning on the Hypersphere
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Transformer Normalisation Layers and the Independence of Semantic Subspaces
- Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
- The structure of the token space for large language models
- Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural Constraints
- Symmetry Breaking in Transformers for Efficient and Interpretable Training
- There Will Be a Scientific Theory of Deep Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks