Activation-Space Order-Swap Geometry: A Site-Asymmetry Audit
cs.LG
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 18 pages, 2 figures. Under review at the NeurReps Proceedings Track
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Order-dependent activation statistics are often interpreted as evidence of interaction, but that interpretation can be confounded by where interventions enter the network.
Terminology
Abstract
Order-dependent activation statistics are often interpreted as evidence of interaction, but that interpretation can be confounded by where interventions enter the network. We introduce a no-fit site-asymmetry audit. For a twice-differentiable readout, the open-path order-swap decomposes into a canonical additive response measured by single interventions and an antisymmetrized second difference free of first-order and pure self-curvature terms to second order. Across six open-weight language-model families, the single-intervention baseline explains 84.3-97.7 percent of the bracket norm (mean 93.7 percent), while the no-interaction self-curvature term is 1.8-5.2 times larger than the corrected residual in the two families with the plus/minus injection split. The corrected residual clears a generic-interaction null in three of six families under a confound-free prompt split and two of six after configuration robustness. A known-positive surrogate recovers planted mixed interaction, while a matched site-separation test changes the baseline share and a random architecture reproduces the first-order regime. The same estimator transfers to released non-language references: trained residual fractions fall below a fixed Gaussian-direction null in 11/12 contrasts (5/6 ViT-B/16, 6/6 ResNet-50), a portability check rather than pooled evidence. The contribution is a reusable measurement criterion: run the single-intervention baseline before reading an order-swap vector as interaction or geometric structure; if it explains the vector, form the second difference instead. All claims are scoped to activation-space interventions at distinct sites; we do not claim that representation geometry is globally Abelian.
Sources
- Weight Updates as Activation Shifts: A Principled Framework for Steering
- Bilinear autoencoders find interpretable manifolds
- Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
- Editing Models with Task Arithmetic
- Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model
- Infinitesimal Causality
- Feed-Forward Steering in Transformer Residual Dynamics
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained Models
- First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
- Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma
- CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits
- Forgetting Is Not a Fix: Path Dependence in Sequential Engram Editing
- Gauge-invariant representation holonomy
- A Low-Rank Subspace Analysis of LLM Interventions
- Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control
- The Geometry of Sequential Learning: Lie-Bracket Prediction of Transfer Order
- Detecting Statistical Interactions from Neural Network Weights
- Steering Language Models With Activation Engineering
- The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching
- Extending Activation Steering to Broad Skills and Multiple Behaviours
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks