Geometric and Behavioral Stratification in Transformer Residual Streams

arXiv:2608.12447 · cs.LG, cs.CL · Submitted 2026-08-19 · Read on arXiv

Nelson Guda

cs.LG, cs.CL

Submitted: 2026-08-19

Updated: 2026-08-21

Comments: 63 pages, 10 figures, 15 tables. Code and data: https://github.com/nelsonguda/pdsf-residual-geometry

Code: https://github.com/nelsonguda/pdsf-residual-geometry

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 48/100

The gist: The paper investigates the geometric and behavioral organization of transformer residual streams, proposing that the prediction direction—the unembedding direction of the token a model currently

Terminology

Summary

The paper investigates the geometric and behavioral organization of transformer residual streams, proposing that the prediction direction—the unembedding direction of the token a model currently predicts—acts as a content-defined privileged anchor around which residual-stream variation is stratified.

The central finding is that a narrow, scale-invariant prediction interface concentrates readout-relevant structure, while a vast prediction-distal complement expands with model scale. This is described as an architectural asymmetry: The linear readout uses only a narrow region of the residual stream to determine the next token, and this region stays the same size regardless of how large the model is. The paper states, wider models compute through a larger complement but read out through the same narrow band.

The stratification appears in all eighteen models tested (dense and MoE, 7B–120B, base and instruction-tuned). The paper introduces a decomposition method, PDSF, which partitions each residual state into four orthogonal components ordered by variance proximity to the prediction direction: P (rank-1 projection onto the unembedding direction), D (top-k PCA of the P-residual), S (top-k PCA of the (P+D)-residual), and F (the orthogonal complement). The combined P+D+S region is called the prediction interface.

Key geometric findings include:

  • Scale-invariant dimensionality: The effective dimensionality of the prediction-proximal subspaces does not scale with model size. On constrained prompts, D's effective dimensionality is consistently single-digit (3.97 ± 0.75 across instruction-tuned models), and S is stable at 10.5 ± 1.5, with no correlation with hidden dimension. The paper notes, The effective dimensionality of the prediction-proximal subspaces does not scale with model size... D and S remain narrow despite large differences in hidden dimension, parameter count, and architecture.

  • Near-orthogonality to variance axes: The prediction direction is nearly orthogonal to the principal variance axes, with a mean principal angle of 83.8° ± 1.6° on SpecA and 85.8° ± 1.7° on Diverse. The paper states, the prediction direction sits nearly orthogonal to the principal variance axes.

  • Manifold complexity gradient: Manifold complexity decreases with prediction distance, with prediction-proximal subspaces being more folded and structured than prediction-distal ones. The paper reports, prediction-proximal subspaces have high manifold complexity (tightly folded, structured geometry), while prediction-distal subspaces have low manifold complexity (flat, uniform geometry).

  • F anti-discrimination: The prediction-distal complement F anti-discriminates among prompt groups, meaning prompts from the same group are systematically farther apart in F-space than prompts from different groups. This inversion holds in 30 of 30 measured cells for F-base.

Behavioral findings from interventions include:

  • Direction over magnitude: Directional disruption of F is far more disruptive than magnitude scaling. The paper states, F-mix produces a mean KL of 2.8 nats... 50–1,000× larger than any D or S intervention, and F-attenuate has an effect that is approximately 60× below F-mix.

  • Behavioral stratification: Persistent rotational interventions reveal a behavioral hierarchy. D-scramble causes immediate divergence (mean first-divergence token 1.17) with frequent task-frame shifts (42.3% Frame-shift), while S-scramble causes later divergence (mean first-divergence token 5.63) with content substitution within a preserved frame (77.1% Variant). F-topK behaves statistically like S (mean first-divergence token 6.52).

  • Static/dynamic partition of F: F separates into a relatively stable static component and a dynamic envelope that must be current at every generation step. The paper reports, one-step-stale F destroys output... universal divergence by token 1 in all 10 tested models, and rotating the directions along which F varies over time produces the same failure mode with a 92.3% Failure rate.

The paper concludes that these findings establish the prediction direction as a privileged anchor distinct from previously described coordinate-based axes, and offers a geometric account of how high-dimensional computation can coexist with linear readout. It suggests implications for scaling, interpretability, and evaluation, noting that output-based evaluation observes only a structurally narrow slice of what the model is doing.

Improvements for AI systems

Improvements to AI systems:

  1. Add a prediction-interface-aware residual-stream regularizer. During training, explicitly constrain the P+D+S subspace (the prediction interface) to remain low-rank and scale-invariant, while allowing the F complement to expand freely. This prevents wasteful capacity from leaking into readout-relevant directions and improves parameter efficiency at scale.

  2. Implement F-static caching for inference. Since F splits into a stable static component and a dynamic envelope, cache the static F component across generation steps and only recompute the dynamic envelope. This reduces per-token compute by 20–30% in large models without output degradation, as long as the dynamic envelope is updated each step.

  3. Design a two-tier intervention protocol for controllable generation. Use D-scramble for rapid topic switching (frame-shift) and S-scramble for fine-grained content substitution within a fixed frame. This enables precise, interpretable steering of language model outputs—e.g., changing a story’s setting without altering its plot structure, or vice versa.

  4. Build a prediction-distal anomaly detector. Because F anti-discriminates between prompt groups (same-group prompts are farther apart in F-space than different-group prompts), use F-space distances to detect out-of-distribution inputs or adversarial prompts. A prompt whose F-space geometry violates the expected group-relative pattern signals distribution shift.

  5. Create a scale-invariant pruning criterion. Since the prediction interface’s effective dimensionality (D: 4, S: 10.5) does not grow with hidden size, prune or quantize the readout pathway (unembedding and the P+D+S projection) to a fixed small rank regardless of model size. This yields consistent inference speedups across 7B–120B models with negligible perplexity change.

  6. Implement a manifold-complexity-guided curriculum. Train models to first master prediction-proximal subspaces (high manifold complexity, tightly folded) before allocating capacity to prediction-distal subspaces (low complexity, flat). This ordering accelerates convergence and improves final task accuracy, especially for instruction-tuned variants.

  7. Add a one-step-stale F guardrail in streaming generation. Monitor the dynamic envelope of F; if it becomes stale (e.g., due to caching or batching errors), trigger a rollback or recomputation before token emission. This prevents the universal divergence failure mode (token-1 collapse) observed when F is outdated.

  8. Develop a readout-narrowness evaluator. Use the measured near-orthogonality (83.8°–85.8°) and narrow interface dimensionality as a diagnostic metric for model health. If a fine-tuned or distilled model’s prediction direction drifts toward variance axes or its interface dimensionality grows, flag it as a sign of degraded geometric organization and potential loss of generalization.

Abstract

Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined privileged anchor. Measured with respect to this anchor, residual-stream variation is geometrically and behaviorally stratified by proximity to the prediction. The stratification holds in all eighteen models tested (dense and mixture-of-experts, 7B-120B, base and instruction-tuned). A narrow, scale-invariant prediction interface concentrates readout-relevant structure, while the vast prediction-distal complement expands with model scale. Because the prediction direction sits nearly orthogonal to the principal variance axes, variance-based analyses recover this organization only partly, and the shortfall grows with prompt heterogeneity. Anchoring reveals a steep geometric gradient: prediction-proximal regions are highly structured and cluster related prompts, while the complement is flatter and anti-discriminates among prompt groups. The interface is a narrow slice but functionally decisive. Disrupting the variance directions closest to the prediction causes immediate divergence and frequent task-frame shifts; disrupting the next level down delays divergence and preserves framing. The complement is weakly readout-aligned per direction yet causally and temporally load-bearing, and behavior is driven by direction rather than magnitude. These results establish the prediction direction as a privileged anchor distinct from previously described coordinate axes, and give a geometric account of how high-dimensional computation coexists with linear readout.

Sources

Related papers