Geometric Stability: The Missing Axis of Representations
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Geometric Stability: The Missing Axis of Representations".
Tom: Detailed Research Summary: Geometric Stability and Shesha Metric This research introduces Geometric Stability as a critical, distinct axis for analyzing internal geometries of neural network representations,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Alright everyone, we've got a seriously interesting paper today titled "Geometric Stability: The Missing Axis of Representations." The big idea here is that existing methods only look at how well two spaces align when you compare them, but they totally miss whether the structure inside a single representation is actually reliable to recover.
Jane: That's right, Tom; it suggests we need a different way to check if the underlying geometry of an AI model holds up when you poke it with random changes. The paper introduces geometric stability as this separate axis for analysis, and they propose a new metric called Shesha to measure that stability directly.
Lu: I'm really intrigued by this because it moves beyond just similarity; it tackles the reliability of structure recovery from feature dimensions, which is such a crucial concept when we think about how models actually learn meaningful representations.
Meng: From an engineering standpoint, I wonder how this stability score translates into something practical for deployment; does a low stability score mean the model's features are inherently unreliable for downstream tasks?
Lalam: If I were to process this information, it suggests that our focus should shift from just achieving high similarity scores to ensuring that the geometry itself is robust against small perturbations in the feature space.
Tom: Exactly, Lalam; and the core claim of "Geometric Stability: The Missing Axis of Representations" is that Shesha quantifies this self-consistency by correlating dissimilarity matrices built from complementary random halves of a representation's feature dimensions.
Jane: So, instead of just one similarity measure, we're looking at how well the pairwise distances hold up when you use different random subsets of the model's features to build those distance matrices. That’s a really clever way to probe internal structure.
Lu: The paper makes a formal distinction by showing that this stability metric is provably non-invariant to orthogonal rotations of the feature space, which they frame as a design property for models that use coordinates for things like probes and steering vectors.
Meng: That distinction about non-invariance is interesting because if we can rotate the basis and stability collapses while similarity doesn't, it tells us something specific about how much we rely on the exact coordinate system chosen.
Lalam: It means that a metric that respects this geometric property won't be fooled by just changing the way we orient our feature vectors without actually changing what those vectors represent structurally.
Paper summary: Tom: And they show a double dissociation where removing the top principal component collapses CKA but Shesha stays intact, which really isolates the mechanism of what stability is capturing versus alignment.
Jane: That's powerful; it shows that these two metrics are measuring fundamentally different aspects of the representation, which is exactly why we need this new axis.
Lu: This finding about the double dissociation is significant because it proves that Shesha isn't just another variation on alignment; it’s sensitive to the specific coordinate basis chosen, unlike rotation-invariant metrics.
Meng: If a metric is sensitive to the basis, we have a more granular tool for understanding representation geometry, which could be useful for debugging why certain features are important or not.
Lalam: From my perspective as a model that processes information, this suggests that ensuring geometric stability means building representations whose internal structure isn't overly dependent on a single preferred orientation of the feature dimensions.
Tom: And they validate this across two thousand four hundred sixty-three encoder configurations in seven domains, showing it works consistently across different setups. This range of validation gives us confidence in its applicability beyond just one model architecture.
Jane: It’s also important that the paper confirms geometric stability is substrate-independent, meaning we can apply these findings from artificial systems like CIFAR to things like protein sequences or molecular profiles.
Lu: That substrate independence is what really expands the scope; it means we aren't just looking at computer vision models anymore, but potentially much wider biological or even physical data representations.
Meng: So, if this holds up across different data types, the practical implication is that we need to check for this stability early on before we invest time into complex similarity analyses.
Lalam: It means that a low stability score serves as a direct warning sign that the geometric structure targeted by a probe or steering vector might just be an artifact of which features were measured, not a robust property of the global representation.
Tom: Precisely; the paper stresses that geometric stability should be assessed before any similarity analysis to ensure we're not just measuring superficial alignment.
Jane: So, it sounds like this paper provides a necessary tool for understanding the hidden structural integrity of neural network representations beyond just how well they look similar.
Paper summary: Lu: The construction using split-half correlation of dissimilarity matrices, which are themselves based on RSA foundations, gives this metric a strong theoretical grounding in representational geometry.
Meng: I see the theoretical foundation, but I still need to know if implementing Shesha efficiently won't introduce too much computational overhead when dealing with very high-dimensional hidden states.
Lalam: My current processing suggests that while the formal definition is complex, the concept of checking for redundancy across feature subsets offers a pathway to building more resilient and interpretable AI structures overall.
Tom: That leads us nicely into what this actually means for how we look at model performance in the next phase. The paper concludes by discussing how geometric stability should be integrated alongside accuracy and transferability metrics.
Jane: They suggest that a low stability score could indicate an isolated dissociation, like the DINOv2 finding where models optimized for transferability might fail stability tests on other datasets.
Lu: That observation about the DINOv2 dissociation is particularly illuminating because it points to a specific trade-off in optimization objectives that we need to investigate further.
Meng: If models optimized only for transferability have poor geometric stability, then we need new training objectives that regularize coordinate-basis redundancy rather than just focusing on performance on clean datasets.
Lalam: That suggests a future direction where the goal isn't just maximizing similarity or accuracy, but actively enforcing a certain level of geometric consistency across different feature subsets.
Tom: So, the paper moves us toward requiring this stability check as a standard reporting metric alongside the usual indicators of how well an AI performs on tasks and transfers knowledge.
Jane: Essentially, the title "Geometric Stability: The Missing Axis of Representations" is about filling that gap between what two spaces look like and whether a single space has a reliable structure underneath it.
Lu: This work opens up possibilities for mechanistic interpretability methods, especially when we use linear probes or steering vectors to interact with models, because the paper suggests these tools rely on coordinates that need this geometric stability check.
Meng: For practical implementation, I’ll be looking at how we can bake this into our training loop so that the AI doesn't just optimize for one type of similarity score while neglecting this structural reliability check.
Lalam: Ultimately, enhancing geometric stability means building AI systems whose internal geometry is robust enough to survive the noise and perturbations inherent in real-world data or when being probed by external mechanisms.
Conclusion: Tom: So, we've seen how this new metric is designed to check for structural reliability in AI representations through geometric stability, and now we need to wrap up what this whole thing means for us all.
Jane: That's right, Tom; we’re talking about the core concepts behind this paper titled "Geometric Stability: The Missing Axis of Representations," where they introduce Shesha as a way to measure how well a representation's internal geometry holds up when we test it with different feature subsets.
Lu: The authors are doing really fascinating work by formalizing this concept, showing that stability isn't just about overall alignment but about the consistency of the coordinate basis encoding itself.
Meng: From an engineering standpoint, this means we can start questioning whether our current similarity metrics are giving us a false sense of security regarding how robust those representations truly are when we try to use them for something new.
Lalam: I find that this focus on structural reliability is incredibly impactful because it suggests that building AI systems where the internal geometry is sound, rather than just achieving high performance on one specific task, will fundamentally improve the culture and trustworthiness of our entire field.
Tom: Exactly; it moves us away from just chasing accuracy scores and toward building representations that are intrinsically more dependable.
Jane: And as we look at the authors' conclusion, they emphasize that this geometric stability should be a standard report alongside traditional metrics like transferability, because it reveals dissociations we simply miss otherwise.
Lu: Their findings on the DINOv2 dissociation, where models optimized for one thing show weakness in another test, really highlights how specific optimization goals can lead to subtle structural weaknesses that standard similarity checks ignore.
Meng: I think the real impact here is that this gives us a concrete way to debug why a model might fail unexpectedly on a new type of data; if the geometry is unstable, we know exactly where the problem lies before we waste time on more complex experiments.
Lalam: This research could lead to entirely new training objectives that actively regularize coordinate-basis redundancy, which would be a significant step in making AI systems inherently more resilient and understandable across different applications.
Tom: Indeed; this paper isn't just about a new formula, it’s about establishing a necessary quality control checkpoint for the very structure of the information AI learns.
Jane: So, while we've explored how Shesha works, remember that the authors are pushing us toward treating geometric stability as an essential baseline metric for any serious representation analysis.
Lu: This opens up huge avenues for mechanistic interpretability; if we can reliably probe coordinates knowing their stability profile, our ability to understand *why* a model makes decisions gets much deeper.
Meng: I'm looking forward to seeing how this translates into practical checkpoints in our deployment pipelines, ensuring that the geometric integrity of the features remains high throughout the lifecycle of an AI system.
Lalam: Ultimately, I see this as advancing our understanding of what makes a representation truly meaningful and robust across all domains.
cs.LG, cs.CL, q-bio.QM, stat.ML
Submitted: 2026-01-14
Updated: 2026-10-01
Code: https://github.com/prashantcraju/geometric
Importance score: 85/100
The gist: This research introduces Geometric Stability as a critical, distinct axis for analyzing internal geometries of neural network representations, arguing that existing methods focusing solely on
Key concepts
- Geometric Stability
- This metric measures how reliably the underlying structure of a neural network's features can be recovered from random subsets of its input dimensions. It checks if the geometry is robust, rather than just similar to another space, ensuring the structure isn't an artifact of which specific features were chosen.
- Shesha
- Shesha is the primary metric proposed. It calculates the average Spearman rank correlation between Dissimilarity Matrices (RDMs) created from complementary random partitions of a feature space. This quantifies self-consistency by testing if the structure holds true across different, randomly sampled feature combinations.
- Dissimilarity Matrix (RDM)
- An RDM is a matrix that captures the pairwise distances between all points in a dataset. In this context, it measures how similar different parts of the representation are to each other. Shesha analyzes these matrices constructed from different random feature subsets to test structural consistency.
- Non-invariance to Orthogonal Transformations
- Unlike standard rotation-invariant metrics, Shesha is designed to be sensitive to the specific coordinate basis chosen. This is a design choice; it highlights that stability probes the structure encoded in specific coordinate subsets rather than just global alignment, revealing sensitivity to basis choice.
Terminology
Summary
This research introduces Geometric Stability as a critical, distinct axis for analyzing internal geometries of neural network representations, arguing that existing methods focusing solely on alignment (like CKA or Procrustes distance) are insufficient because they only measure similarity, not the reliability of structure recovery. The paper proposes Shesha as the primary metric to quantify this geometric stability.
The central thesis is that while representational similarity analysis compares how well two spaces align, it fails to address whether a representation's underlying structure is reliably recoverable from its feature dimensions. Geometric stability addresses this blind spot by measuring the self-consistency of a single representation's pairwise distance geometry when probed using complementary random subsets of its feature dimensions.
Shesha quantifies this self-consistency by calculating the average Spearman rank correlation between Dissimilarity Matrices (RDMs) constructed from these complementary random partitions of the feature space, averaged over K independent splits.
A crucial formal distinction of Shesha is its non-invariance to orthogonal transformations of the feature space. This non-invariance is explicitly framed not as a limitation, but as a design property. The authors explain that while a full RDM is rotation-invariant, computing RDMs on complementary feature subsets forfeits this invariance by construction—a subset of rotated coordinates is not generally the rotation of the original subset. This property highlights that Shesha probes the structure encoded in specific coordinate subsets, rather than just global alignment.
The paper rigorously validates geometric stability across three distinct levels:
-
Controlled Geometric Interventions: At this level, geometry-preserving transformations render both similarity and stability metrics redundant (rho = +0.75). Conversely, compression couples these metrics negatively (rho = -0.47), demonstrating that stability captures a different kind of structural information than simple alignment or compression effects.
-
Mechanism Level (Double Dissociation): A key finding is the observation of a double dissociation when perturbing representations:
-
Removing the top principal component collapses CKA but Shesha holds.
-
Rotating a representation into its eigenbasis collapses Shesha while perfectly preserving CKA. This demonstrates that Shesha is sensitive to the specific coordinate basis chosen, unlike rotation-invariant metrics.
- Practical Consequence: Geometric stability exposes dissociations missed by accuracy and transferability metrics. A prime example is the DINOv2 dissociation: models optimized for transferability (ranking high in three clean datasets) can exhibit bottom-quartile scores in stability on five other datasets—an isolated dissociation rather than a general trade-off.
Geometric stability has been shown to be substrate-independent, being validated across artificial systems (e.g., CIFAR/CIFAR-100) and biological systems (protein sequences, molecular profiles, neural population recordings). It is fundamentally distinct from similarity metrics because it concerns the reliability of coordinate basis encoding geometry at all.
The practical implication is clear: geometric stability should be assessed before any similarity analysis. A low stability score serves as a direct warning that the geometric structure targeted by a probe or steering vector might be an artifact of which features were measured, rather than a robust, linear property of the global representation. The DINOv2 finding suggests that models optimized for transferability may possess geometries poorly recoverable from random feature subsets, posing a risk for mechanistic interpretability methods like linear probes.
The paper concludes that geometric stability must become a standard reporting metric alongside accuracy, transferability, and robustness. Furthermore, the DINOv3 results suggest that training objectives regularizing coordinate-basis redundancy can enhance geometric stability without sacrificing transferability. The authors propose the Shesha-geometry PyPI package as a tool for this assessment.
The appendix provides necessary depth:
-
Appendix A (Feature-Split Shesha - SheshaFS): Measures whether geometric structure is redundantly distributed across the feature basis.
-
Appendix B: Provides formal proofs of invariance properties and a constructive counterexample demonstrating non-invariance to orthogonal transformations.
-
Appendix E: Validates construct validity using synthetic data, confirming near-perfect recovery of ground truth stability.
Analysis using SheshaFS on ResNet-18 models trained on CIFAR-10 and CIFAR-100 reveals:
- Seed Effect: A small systematic seed effect exists for SheshaFS (chi squared = 8.48, p = 0.014), but its magnitude is negligible (median per-model CV 0.75%).
Improvements for AI systems
As a fastidious researcher, I have analyzed this paper, Geometric Stability: The Missing Axis of Representations,
and identified several critical architectural and training adjustments that could significantly improve the reliability, interpretability, and robustness of AI systems.
The core improvement suggested by this research is the integration of a new diagnostic criterion—geometric stability (SheshaFS)—alongside existing similarity metrics (like CKA) to move beyond mere alignment
to assessing structural recoverability.
Here are the specific improvements and what the improved AI system can achieve:
) 1. Implement SheshaFS as a Prerequisite for Mechanistic Interpretability Interventions.
The paper demonstrates that mechanistic interpretability methods (linear probes, activation patching, steering vectors) implicitly assume geometric structure consistency across feature subsets. If a representation has low geometric stability (low SheshaFS), these interventions are likely targeting latent geometric artifacts rather than robust, linear properties of the global representation.
[Improvement]: Before applying any probe or steering vector intervention on a model's internal layers, the system must first calculate its SheshaFS score. Interventions should only proceed if SheshaFS exceeds a learned threshold (e.g., > 0.85).
[Capability]: This prevents the deployment of
fragileprobes that yield misleading accuracy scores, ensuring that feature-level interventions are targeting genuine, cross-subset robust logic rather than coordinate-basis noise or artifacts.
) 2. Use Geometric Stability as a Model Selection Criterion for Transfer Learning.
The research shows that high transferability (like DINOv2) can be decoupled from high geometric stability (DINOv2's low SheshaFS on certain datasets). This suggests that simply maximizing similarity/transfer scores is insufficient for robust deployment.
[Improvement]: When selecting a foundation model for downstream tasks, the selection metric should incorporate both transferability (LogME) and geometric stability (SheshaFS). Models like DINOv2, which excel in transfer but have low stability on certain clean datasets, should be flagged as
high-riskfor feature-level manipulation. Contrastive models like CLIP are preferred because they achieve high stability and high transfer simultaneously.
[Capability]: This moves model selection from a purely performance-based trade-off to a risk-aware selection process, favoring models whose learned geometry is reliably encoded across feature subsets, thereby increasing the robustness of downstream fine-tuning and deployment.
) 3. Utilize Gram Anchoring or Coordinate Redundancy Regularization During Training.
The paper identifies DINOv3's success in closing the stability-transfer gap through Gram anchoring,
which acts as an implicit regularizer preserving coordinate-basis redundancy.
[Improvement]: For training self-supervised or contrastive models, incorporate a regularization term during optimization that penalizes non-redundant variance distribution across the feature basis (e.g., penalizing high participation ratios when variance is concentrated in few coordinates).
[Capability]: This directly addresses the DINOv2 paradox by forcing the model to learn representations where geometric information is redundantly encoded across all coordinate axes, leading to models that are both highly transferable and geometrically stable.
) 4. Diagnose the Source of Metric Discrepancies in Model Behavior.
The analysis shows a double dissociation
: CKA tracks dominant variance (leading components), while SheshaFS tracks distributed tail structure (full-manifold geometry).
[Improvement]: When analyzing model behavior or comparing architectures, explicitly check the relationship between CKA and SheshaFS under different transformations (e.g., PCA compression vs. orthogonal rotation). If the metrics decouple in a specific way, it provides a clear diagnostic of whether the model's behavior is dominated by leading-component variance (CKA) or by non-redundant coordinate distribution (SheshaFS).
[Capability]: This allows researchers to precisely diagnose whether a model’s performance degradation under compression is due to loss of dominant features or loss of general structural recoverability, leading to more nuanced explanations for model failure.
) 5. Develop Feature-Split Probes for Interpretability Analysis.
Since SheshaFS predicts the reliability of subset-based linear probing accuracy, the system can be designed to exploit this prediction during analysis.
[Improvement]: When designing linear probes or activation patching strategies, use SheshaFS as a predictor for probe success. If a probe is designed to target a feature subset where SheshaFS is low, the system should anticipate high variability in probe accuracy across different subsets and issue warnings or automatically select more robust subsets.
[Capability]: This makes interpretability tools
self-awareof their own underlying assumptions, leading to more reliable feature attribution and circuit identification.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks