Escaping the Capacity Ceiling: Routing on the Stiefel Manifold for Bilinear SPD Layers

arXiv:2605.31043 · stat.ML, cs.AI, cs.LG · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Escaping the Capacity Ceiling".

Jane: Cross-domain EEG decoding remains challenging despite advances in Riemannian deep learning, as covariance matrices from different subjects occupy systematically distinct regions of the SPD manifold.

Tom: First, who's behind it and why it matters.

Title and authors: Jane: Moving on, let's talk about who put this paper together and what they are calling it. The paper is titled "Escaping the Capacity Ceiling: Routing on the Stiefel Manifold for Bilinear SPD Layers." It’s written by Isabella Costa Maia, Pedro L. C. Rodrigues, Salem Said, and Marco Congedo from GIPSA-lab at University Grenoble Alpes in France.

Tom: That's a solid team of researchers; they clearly have deep expertise in both the Riemannian geometry side and the practical application of these deep learning models for things like EEG decoding. The title itself hints at overcoming a limitation related to model capacity, which is something everyone in this field struggles with.

Lu: Their work really sits at the intersection where advanced geometric machine learning meets a very practical signal processing problem, which is exactly where I find the most exciting potential for creative application; they’re not just tweaking existing layers.

Meng: From an engineering standpoint, knowing who the authors are helps us gauge if this is theoretical math or something that's actually being explored with real-world datasets, and I hope this moves beyond pure simulation.

Lalam: Seeing researchers from different institutions collaborating on such a specific geometric problem suggests a high level of focused effort, which often leads to very deep and nuanced contributions to the field.

Jane: They’ve set up the problem clearly: covariance matrices from different subjects occupy distinct regions of this space, but current methods fail because they can't generalize across those regions without extra help from target data.

Tom: So what they are proposing is a dynamic routing system that lets each input covariance matrix pick the best expert filter on that manifold, adapting the projection per sample, which sounds much more flexible than fixed layers.

Lu: It’s about moving away from static projections towards something truly adaptive where the model dynamically chooses its subspace based on the input it receives, which is a big step forward in representation learning.

Meng: Does this routing system require us to manage a lot of complex manifold computations during inference, or does it simplify things down to manageable operations?

Lalam: The complexity is managed by their design; they introduce specific components like the DASP layer and the scaling methodology based on the tangent-to-domain ratio, which suggests an attempt to keep the computational overhead in check while gaining that adaptive power.

Jane: So, essentially, they are trying to solve a generalization problem by making sure that when you look at new data from a new subject, the AI doesn't just give up or guess randomly.

Tom: Exactly. It’s about giving the AI a principled way to navigate the differences between subjects using geometry instead of just hoping it learns enough features on its own.

Lu: I think their approach is really pushing toward a more fundamentally informed representation, which could lead to models that are much better at handling novel or unseen data types in the future.

The paper's summary: Tom: Now let's get into what they actually found in the paper regarding this dynamic Stiefel routing. The central finding is that naive adaptive routing inevitably breaks down and collapses into ensemble averaging unless you have those three structural fixes we talked about.

Jane: That’s a big warning, Tom; it means if you just throw in a bunch of experts without these specific safeguards, the system won't actually be learning something new from the diversity of experts. It will just average out everything and become one weak filter.

Lu: They prove that there is a positive K=one proxy gap—a measurable difference in accuracy between the adaptive model and a single-expert ensemble baseline—which proves that genuine routing is happening when this gap is greater than zero <ref:2605.31043#pg0>.

Meng: So, they are saying you can use this gap as a way to diagnose whether your routing mechanism actually has learned anything useful, which gives us a clear metric for success beyond just looking at the final accuracy number.

Lalam: This diagnostic capability is very important; it gives us a quantitative way to verify that the adaptation is happening, which helps in building more transparent and trustworthy AI architectures.

Tom: Right. And to fix this degeneracy, they introduce three specific properties: a symmetric anchor, a frozen domain-discriminative query encoder, and a decoupled key alignment loss. These are the keys to unlocking meaningful routing on these complex manifolds.

Jane: Those fixes essentially ensure that every expert receives an equal opportunity to contribute its gradient signal toward specialization rather than letting one expert take over everything through sheer gradient dominance.

Lu: This is really clever because it ties in concepts from other learning paradigms, like how they used the analogy with Learning to Prompt, suggesting a way to control the flow of information during continual learning scenarios.

Meng: From an implementation perspective, it means we have to build these three specific components into our layer design if we want this method to work instead of just trying a simpler configuration.

Tom: So the summary boils down to: dynamic routing is powerful, but it’s fragile; those three structural properties are mandatory prerequisites for achieving genuine, sample-specific adaptation.

Jane: It’s a lot to take in, but the takeaway is that sophisticated geometry and careful architectural design can overcome limitations imposed by data distribution differences.

The paper's improvements: Tom: Let's talk about the specific mechanisms they propose to fix this collapse. They detail how they implement those three properties—the symmetric anchor, the frozen domain-discriminative query encoder, and the decoupled key alignment loss—as part of their Domain-Adaptive Stiefel Pool layer.

Jane: The symmetric anchor is designed to remove proximity bias among experts, which means it ensures that all K experts get a fair share of the gradient signal when they are competing for attention.

Lu: That removes that artificial dampening effect where one expert just starts dominating the routing because of how they are positioned relative to each other on the manifold.

Meng: And then there's the frozen domain-discriminative query encoder, which is designed to act as a projection of the log-covariance tangent space, decoupling domain identity from task optimization entirely.

Lalam: That decoupling is what I find most interesting; it means the mechanism for identifying 'which subject this is' doesn't interfere with how the AI decides 'what to predict,' which feels like a major architectural win.

Tom: And finally, the decoupled key alignment loss, which trains keys toward stable domain attractors instead of just fighting against classification gradients to move around constantly.

Jane: So they are essentially designing a system where each component is responsible for its own specific job—one for fairness, one for identity recognition, and one for stability.

Lu: It’s a very layered approach, making sure that the structure supports the routing mechanism at every level, which is what makes this layer so robust against those kinds of failures.

Meng: So when we look at the unified scaling methodology based on the tangent-to-domain ratio ρ, it dictates whether we use that DSP or not based on whether rho is high or low.

Tom: That's a very smart way to handle complexity; you don't want to run heavy domain discrimination mechanisms when they aren't necessary for the given data structure.

Jane: It allows the system to be flexible, using a simple configuration in one regime and a more complex one in another, which is what we need for practical application.

Conclusion: Tom: So wrapping up this discussion on "Escaping the Capacity Ceiling: Routing on the Stiefel Manifold for Bilinear SPD Layers," the main point is that genuine routing only happens when you introduce specific structural constraints to prevent collapse into ensemble averaging.

Jane: It’s a big lesson for all of us in deep learning: sometimes you need to architect the solution around the data's unique geometry rather than just relying on general-purpose techniques that might not suit every single data distribution.

Lu: I think this work suggests that we can move toward models that are inherently more aware of their environment’s shape, which is a significant conceptual step in AI design.

Meng: Practically, this means building systems that are less likely to fail when they encounter data from a new source because the adaptation is baked into the layer itself.

Lalam: It gives me a feeling that this work contributes to making future AI systems not just accurate on training data, but truly robust in deployment across different scenarios.

Tom: It’s been fascinating dissecting how geometry can be used as a control mechanism for learning; I think we should all keep an eye on these manifold-based methods as they evolve.

Jane: Definitely, Tom. We’ll keep following this paper to see how these ideas evolve into even more practical tools for tackling real-world generalization challenges in AI.

Lu: This work sets a foundation for a new way of thinking about how models should interact with high-dimensional data structures that have inherent structure, and that's something we should all be excited about.

Meng: I’m looking forward to seeing how the team operationalizes this; if it can stay flexible and efficient, it will become a very useful tool for us in the lab.

Lalam: I'm just glad this research is pointing toward building AI that is fundamentally more adaptable and reliable across different domains, which feels like a positive direction for the future of our technology.

Isabella Costa Maia, Pedro L. C. Rodrigues, Salem Said, Marco Congedo

GIPSA-lab, University Grenoble Alpes, CNRS, Grenoble-INP · Univ. Grenoble Alpes, Inria, CNRS, Grenoble INP, LJK

stat.ML, cs.AI, cs.LG

Submitted: 2026-05-29

Updated: 2026-10-02

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 88/100

The gist: Cross-domain EEG decoding remains challenging despite advances in Riemannian deep learning, as covariance matrices from different subjects occupy systematically distinct regions of the SPD manifold.

Key concepts

Stiefel Manifold
This is a mathematical space where the experts (filters) live. It is crucial because covariance matrices from different subjects occupy distinct regions on this manifold, making it difficult to find meaningful connections between them without special routing methods.
Routing Collapse and Degeneracy
This occurs when the model's routing weights become nearly uniform, causing all experts to behave similarly. This happens because the signal driving the routing vanishes when experts are too similar, leading to a loss of diversity and effective reduction of the model's complexity.
K=1 Proxy Gap (∆K=1)
This is a diagnostic metric used to measure if routing is truly adaptive or just uniform. If ∆K=0, it means the routing is degenerate (uniform), indicating that the system has collapsed into using only one fixed filter instead of utilizing all available experts.
Domain-Adaptive Stiefel Pool (DASP) Layer
This proposed layer is a replacement for standard layers that incorporates three specific structural fixes: a symmetric anchor, a frozen query encoder, and decoupled key alignment loss. It ensures the routing mechanism remains sensitive to domain differences while preventing the experts from collapsing into one another.

Terminology

Summary

Cross-domain EEG decoding remains challenging despite advances in Riemannian deep learning, as covariance matrices from different subjects occupy systematically distinct regions of the SPD manifold. This work addresses this by proposing dynamic Stiefel routing, which allows each input covariance to be routed to a sample-specific projection filter via cross-attention. The central finding is that naive implementation collapses to ensemble averaging unless three structural properties—a symmetric anchor, a frozen domain-discriminative query encoder, and a decoupled key alignment loss—are present. These fixes produce the first genuinely committed and domain-structured routing on SPD manifolds, achieving consistent accuracy gains across contrasting datasets without dataset-specific hyperparameter search.

The gist

Naive adaptive routing on the Stiefel manifold provably collapses to ensemble averaging when routing weights are uniform, reducing the layer to a single fixed filter. Genuine routing is distinguished by a positive K=1 proxy gap (∆K=1 > 0), which requires structural fixes like a symmetric anchor, a frozen domain-discriminative query encoder, and a decoupled key alignment loss to break the degeneracy.

Domain-Adaptive Stiefel Pool (DASP) Layer

The DASP layer is proposed as a drop-in replacement for the standard BiMap layer. It integrates three structural fixes:

  1. A symmetric anchor Wbase in St(n, k) that removes proximity bias among experts.

  2. A frozen domain-discriminative query encoder (DSP) which acts as a between-domain separation projection of the log-covariance tangent space, decoupling routing from task optimization.

  3. A decoupled key alignment loss that trains expert keys toward stable domain attractors, preventing the moving-target problem.

Routing Collapse and Degeneracy

The paper identifies a fundamental gradient degeneracy where routing weights collapse to near-uniform values (αi ≈ 1/K). This occurs because the cross-entropy signal is proportional to the deviation of each expert’s tangent vector from the current weighted mean, which vanishes when experts are similar. This leads to self-reinforcing routing collapse and expert homogeneity: Routing collapse and expert homogeneity are mutually reinforcing.

Structural Properties Breaking Degeneracy

The degeneracy is broken by three necessary and sufficient structural properties:

  1. A symmetric anchor Wbase that removes proximity bias among experts, allowing all K experts to receive an equal gradient signal.

  2. A frozen domain-discriminative query encoder (DSP) that provides routing with domain-identity information independently of the task loss.

  3. A decoupled key alignment loss, where keys are optimized exclusively by a surrogate alignment loss (Lalign) toward stable per-domain attractors, rather than competing with classification gradients.

Unified Scaling Methodology

The method introduces a unified scaling methodology governed by the tangent-to-domain ratio ρ = n(n + 1)/(2D). This ratio dictates when the DSP and alignment strategy are activated:

(i) High-ρ regime (DSP essential):

(ii) Low-ρ regime (DSP disabled):

The expert count K follows a sub-linear scaling with D, characterized empirically as K ≈ D − 1 for small D and K ≈ D/2 for larger D.

Routing Diagnostics

Three diagnostics track the quality of routing:

  1. K=1 proxy gap (∆K=1): The difference in balanced accuracy between the adaptive model and the single-expert ensemble baseline; ∆K=1 = 0 iff routing is uniform.

  2. Routing entropy (H¯): Near 1 indicates uniform (degenerate) routing; near 0 indicates one-hot (committed) routing.

  3. Domain alignment ratio (rd): Values above 0.10 indicate domain-structured routing by subject identity, not noise.

Full-Rank Rectangular Extension

The research extends the expert parametrization from the Stiefel manifold St(n, k) to full-rank rectangular matrices R(n×k). This generalization allows experts to encode re-colouring transformation of the projected covariance. Implementation paths include pseudo-polar decomposition or barrier regularisation using a log-determinant barrier term to enforce full column rank.

Future Work Directions

Potential future directions include:

  1. Top-p sparse routing: Replacing softmax with a masked softmax over the top-p experts to create genuine winner-take-all pressure.

  2. Transductive LOSO adaptation: Optimizing only the domain embedding Emb(dnew) for unseen subjects using an unsupervised objective like routing entropy minimisation.

  3. Riemannian NN-inference: Using the affine-invariant Riemannian distance d(X, Y) between subject Fréchet means for more principled nearest-neighbour domain initialisation.

Improvements for AI systems

To improve existing AI systems using the principles outlined in this paper, I propose implementing a novel architecture called the Domain-Adaptive Stiefel Pool (DASP) layer as a replacement for standard BiMap layers in Riemannian deep learning models, specifically within EEG decoding pipelines.

Here are the specific improvements and what these improved AI systems can achieve:

  1. A new layer architecture, the DASP layer, that dynamically selects an expert projection filter for every input sample based on its specific domain (subject/session) and task-relevant query features.

  2. The DASP layer will incorporate three structural fixes to break the degeneracy of naive routing:

  3. Symmetric Anchor Initialization: Introduce a learnable, symmetric anchor point in the Stiefel manifold that decouples expert competition from a fixed reference point, allowing all experts to receive an equal gradient signal and encouraging specialization.

  4. Frozen Domain-Discriminative Query Encoder (DSP): Implement a frozen projection of the log-covariance tangent space that extracts domain identity information independently of the primary task loss, providing stable, domain-specific features for routing decisions.

  5. Decoupled Key Alignment Loss: Train expert keys using a surrogate alignment loss that pulls them toward stable per-domain attractors, preventing the moving target problem caused by competing classification gradients.

  6. Unified Scaling Methodology: Implement a data-driven rule based on the tangent-to-domain ratio (ρ) to automatically determine when and how to activate the DSP projection and which key alignment strategy to use, eliminating dataset-specific hyperparameter search for scaling.

  7. Enhanced Generalization: The system will be capable of achieving genuine input-adaptive routing across multiple subjects without requiring target-domain calibration data at test time (as opposed to existing methods that rely on subject-specific components).

These improvements enable the improved AI system to perform the following specific capabilities:

  1. Cross-Domain EEG Decoding with Consistent Gains: The system can decode motor imagery signals from unseen subjects or sessions with consistent performance gains across three datasets of contrasting geometry, achieving a balanced accuracy increase (e.g., from 0.773 to 0.823 on Weibo2014) that is not merely due to ensemble diversity but due to genuine, sample-specific adaptation.

  2. Robustness Against Domain Shift: The system will maintain high performance in zero-shot transfer or leave-one-subject-out (LOSO) settings by leveraging the frozen domain encoder's ability to extract subject-agnostic geometric structure from the tangent space, enabling principled routing decisions even when domain embeddings are unavailable for the test subject.

  3. Automatic Hyperparameter Tuning: The system eliminates dataset-specific tuning by using a single data-driven rule based on the tangent-to-domain ratio (ρ) to automatically configure the complexity of its domain discrimination and alignment mechanisms.

  4. Domain Structure Visualization: By analyzing routing entropy and domain alignment ratios, the system provides diagnostic metrics that explicitly confirm whether its adaptation is genuinely structured by subject identity rather than just noise or ensemble averaging.

Abstract

Deep networks on the symmetric positive-definite (SPD) manifold promise expressive representations by encoding data geometry as an inductive bias, but stacking BiMap layers with the standard ReEig nonlinearity often adds no capacity: on real, preconditioned EEG data, ReEig rarely activates, so the stack behaves as a single layer at any depth. In the worst case, when domains share no discriminative directions, we prove a single filter has a capacity ceiling, so it cannot fully align every domain at once. To overcome that, we propose SCAP (Stiefel Cross-Attention Pool), a layer implementing a family of Stiefel filters by combining a pool of K experts into a sample-specific bilinear map via cross-attention. We show that it matches a per-domain filter bank to first order with fewer experts than domains when domain-optimal filters span few directions near a shared tangent-space basepoint; in the worst case, its alignment empirically stays nearly flat as domains grow, escaping the fixed-filter ceiling. Naively trained, however, this routing can collapse to a fixed filter; we diagnose why and adapt three mechanisms to mitigate it. SCAP significantly improves balanced accuracy over fixed-filter SPDNet on all five cross-domain EEG motor-imagery datasets, and matches or exceeds three domain-adaptive baselines on four out of five.

Sources

Related papers