Stimulus symmetries can confound representational similarity analyses

summary

Video file (mp4)

The gist

Stimulus symmetries can confound representational similarity analyses because functionally equivalent representations related by stimulus symmetries can possess qualitatively different

In short

Stimulus symmetries can confuse representational similarity analyses because functionally equivalent representations related by stimulus symmetries can have different geometric structures, leading to distinct Representational Similarity Matrices (RSMs). The work formalizes this gauge dependence, showing that while decoding accuracy is invariant, the RSM itself depends on how the symmetry group acts on the encoding.

Key concepts

Representational Similarity Matrix (RSM)
The RSM measures how similar different neural encodings are to each other using an inner product. The paper shows that if representations are related by stimulus symmetries, they might not be related by simple orthogonal transformations in the representation space, causing the resulting RSMs to differ even if the functions are equivalent.
Gauge Dependence
Gauge dependence occurs when a mathematical quantity, like an RSM, changes depending on how you choose to represent or transform the underlying data. The study formalizes this by showing that for general data symmetries, the RSM is not automatically invariant under those transformations unless the encoding itself follows specific orthogonal rules.
Functional Equivalence vs. Geometric Difference
The core finding is that stimulus symmetries can make two representations functionally equivalent (doing the same task) but geometrically different in representation space. This means they are not related by a simple rotation or orthogonal transformation, which is what RSM invariance usually requires for comparison.

Terminology used across episodes

This episode discusses

The paper

Stimulus symmetries can confound representational similarity analyses · Read on arXiv

Farhad Pashakhanloo, Jacob A. Zavatone-Veth

Harvard University

What can representational similarity matrices (RSMs) tell us about a neural code? As the popularity of these summary statistics grows, so too does the need for a more complete characterization of their properties. Here, we show that symmetries in network inputs can confound RSM-based analyses. Stimulus symmetries render many representations functionally equivalent, but these different configurations can lead to different RSMs. These different RSMs reflect qualitatively different representational geometries, ranging from disentangled to maximally-mixed codes. We show that stochastic gradient descent or energetic regularization can generate sparse, drifting codes, leading in turn to drifting RSMs. Moreover, we demonstrate that these phenomena are present in networks trained to encode image data, where the symmetry is latent. Our results illustrate the challenges inherent in comparing nonlinear neural codes, when functionally-equivalent representations are not related by a simple rotation.

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: I'm Ines, and with me are Marcus and Yuki, guest researcher.

Marcus: Today's paper: "Stimulus symmetries can confound representational similarity analyses".

Ines: Stimulus symmetries can confound representational similarity analyses because functionally equivalent representations related by stimulus symmetries can possess qualitatively different representational geometries, leading to distinct Representational Similarity Matrices (RSMs).

Marcus: First, who's behind it and why it matters.

Paper summary: Ines: So this paper, "Stimulus symmetries can confound representational similarity analyses," really tackles how we interpret similarity scores when the input data has inherent symmetries. The main idea seems to be that functionally equivalent representations related by these stimulus symmetries can end up having different geometric properties in representation space, which means their Representational Similarity Matrices or RSMs won't match up as expected <ref:2605.21324#pg0>.

Marcus: That sounds like a real problem for us when we try to use these summary statistics to compare different neural codes across different datasets; if the underlying data structure dictates the symmetry, then the similarity metric itself gets biased <ref:2605.21324#pg0>.

Yuki: From a population genetic perspective, this hints at how subtle variations in stimulus presentation or environmental factors might lead to functionally identical neural states that are structurally distinct when viewed through a rigid mathematical lens, which mirrors complexities we see in evolutionary history <ref:2605.21324#pg0>.

Ines: Exactly, and the paper claims that this means we can't just assume that if two representations are functionally equivalent because of a stimulus symmetry, their RSMs will be identical; instead, they can reflect qualitatively different geometries <ref:2605.21324#pg0>.

Marcus: So it’s not just about the inputs being symmetrical; it's about how those symmetries manifest in the resulting neural encodings and whether our comparison tool is sensitive to that manifestation <ref:2605.21324#pg1>.

Yuki: It suggests that what we perceive as functional equivalence at a neural level might be masked or distorted if the representation space itself isn't properly anchored relative to the symmetry group acting on the input manifold <ref:2605.21324#pg1>.

Ines: And they formalize this by looking at gauge invariance; they introduce a setting where data lives on a latent space Z acted upon by a compact group G, and they show that for an RSM to be gauge-invariant, the encoding has to be an orthogonal linear representation of G <ref:2605.21324#pg1>.

Marcus: That brings up the issue of how we define that orthogonality when dealing with non-linear systems; if the encodings aren't naturally related by a simple rotation in representation space, then that gauge invariance condition becomes very strict <ref:2605.21324#pg1>.

Yuki: It makes me think about how we model population dynamics where symmetries exist; if the underlying genetic structure imposes symmetries on phenotype expression, we need to account for those transformations when analyzing similarity <ref:2605.21324#pg0>.

Paper summary: Ines: They then move into showing that while the encodings themselves aren't always gauge-invariant, the average decoding accuracy remains invariant under a gauge transformation, which is a different kind of invariance than what we are looking for in the RSM <ref:2605.21324#pg1>.

Marcus: That distinction between decoding accuracy invariance and RSM dependence is crucial; it means even if the error rate doesn't change when you rotate the encoding, the similarity structure captured by the RSM can still shift because of that rotation <ref:2605.21324#pg1>.

Yuki: This implies that relying solely on decoding performance metrics might give us a false sense of comparison if we ignore the underlying geometric configuration dictated by stimulus symmetries <ref:2605.21324#pg0>.

Ines: Moving into the results, they use a toy model involving neurons tiling a one-dimensional ring, where the gauge variable phi defines a global orientation and shows how this angle alters the off-diagonal elements of the RSM <ref:2605.21324#pg1>.

Marcus: That visual demonstration is really helpful for grasping the concept; it shows that changing phi qualitatively changes those similarity measures, meaning the geometry is genuinely different, not just a rotation of the entire structure <ref:2605.21324#pg1>.

Yuki: It's fascinating because it means we can have multiple mathematically valid tilings of the same input manifold that are functionally identical but result in distinct RSMs depending on how you fix the lattice configuration <ref:2605.21324#pg1>.

Ines: They also extend this to learning scenarios, showing that stochastic gradient descent or energetic regularization can produce sparse, drifting codes over time, which subsequently causes the RSMs to drift as well <ref:2605.21324#pg0>.

Marcus: So this isn't just a static problem with trained models; it applies to how representations evolve during training, leading to representations whose similarity structure changes dynamically <ref:2605.21324#pg0>.

Yuki: That drift over time connects back to population genetics because genetic drift causes phenotypic variations that can lead to different functional outcomes, and here we see a similar mechanism happening in the learned representation space <ref:2605.21324#pg0>.

Ines: They also looked at structural parameters, finding that increasing the number of receptive fields tends to reduce variability in the RSM, although for small tuning widths, variability can still be non-negligible <ref:2605.21324#pg0>.

Paper summary: Marcus: That suggests a trade-off between having more features and getting a more stable similarity structure; maybe we need to consider that when building models for genomics data where feature selection is key, this stability is important <ref:2605.21324#pg0>.

Yuki: This study also generalized the analysis to higher-dimensional symmetric manifolds, like tiling the hypersphere S d-one with reflection symmetry, where the RSM takes a form like RSM = Id + R R-one Id + R-one <ref:2605.21324#pg2>.

Ines: And in that higher-dimensional case, they found elements obey "non-trivial equality relations," which suggests the dependence on gauge angles is much more complex than what we see in lower dimensions <ref:2605.21324#pg2>.

Marcus: That complexity makes it harder to simplify the interpretation of similarity metrics when dealing with high-dimensional data structures, especially when trying to account for batch effects or other noise <ref:2605.21324#pg0>.

Yuki: It reinforces the idea that as biological systems become more complex, the mathematical machinery we use to compare their states needs to be able to handle richer symmetry structures rather than assuming simpler relationships <ref:2605.21324#pg1>.

Ines: In realistic applications, like autoencoders trained on rotated image data, they found that while reconstruction loss didn't show a clear dependence on the gauge variable phi, the CKA similarity did drop significantly as the gauge difference increased <ref:2605.21324#pg0>.

Marcus: That CKA similarity drop is what really tells us something about how sensitive our comparison metric is to these symmetries in a way that reconstruction loss isn't <ref:2605.21324#pg0>.

Yuki: It confirms their point that standard general-purpose vision models might not render latent stimulus rotations invisible when compared using RSM-based metrics, which is an important consideration for understanding how these models process visual data across different contexts <ref:2605.21324#pg0>.

Ines: So to wrap up the core finding of "Stimulus symmetries can confound representational similarity analyses," it means that if you don't account for these stimulus symmetries, you risk comparing representations that are functionally equivalent but geometrically distinct <ref:2605.21324#pg0>.

Marcus: And the implication is that to reliably interpret functional significance in representations where data symmetries exist, we absolutely need to design a metric that is invariant to those specific symmetry transformations, which usually requires knowing something about the stimulus space itself <ref:2605.21324#pg0>.

Yuki: This motivates future work toward creating metrics that respect these inherent structural constraints when comparing different biological or artificial systems <ref:2605.21324#pg0>.

Conclusion: Ines: It means that if we compare two brain states that are doing the exact same job but are generated by inputs with a certain symmetry, our similarity score might be misleading because the underlying geometry is skewed. Marcus, thinking about cohort effects and batch variability in genomics, this suggests that if our data has hidden symmetries related to how samples were collected or processed, we might be comparing fundamentally different structures even if the biological signal is consistent. Yuki, from a population genetic viewpoint, this hints that subtle variations in environmental input across different populations could lead to functional equivalents that are structurally distinct when viewed through the lens of representation space.

Marcus: Exactly, Ines; it points out that batch effects aren't just noise in the numbers; they can be artifacts arising from these underlying stimulus symmetries not being accounted for in our statistical framework. Yuki, if we think about evolutionary history, this could mean that functional traits might appear highly conserved across species because they share a symmetry constraint on their underlying sensory inputs, but those constraints impose different geometric arrangements on the neural code itself.

Yuki: That makes sense; if the symmetry is inherent to the physical environment or developmental constraints, then those constraints dictate a specific geometry for any valid representation of that input, and if we don't map that symmetry onto our comparison metric, we miss genuine functional equivalence across different contexts. Ines, so what is the actual recovery here for computational biology?

Ines: We recover the idea that simply measuring similarity isn't enough; we need a metric that respects those symmetries to truly understand the function being performed by the representation. Marcus, if we look at this from a data scientist standpoint, it tells us we need to be much more careful about how we define distance and similarity when dealing with structured data like biological measurements where underlying constraints are often present but invisible. Yuki, what's the next step for researchers trying to build these robust tools?

More episodes

← Home