Semantic Fibers and Cross-Gram Interference: A Calculus of Safety Drift in Overcomplete Representations
cs.LG, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 43 pages , 15 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: A deployed language model may refuse a harmful request in English yet comply with its faithful translation, revealing a cross-lingual safety failure that cannot be characterized reliably by output
Terminology
Abstract
A deployed language model may refuse a harmful request in English yet comply with its faithful translation, revealing a cross-lingual safety failure that cannot be characterized reliably by output behavior alone. We formalize this phenomenon through an audited equivalence relation and show that, for a declared quotient, representation, metric, feature dictionary, scoring head, threshold, and contrast model, the resulting safety drift admits an exact linear-algebraic characterization. Specifically, the drift is a cross-Gram functional of the within-fiber contrast; its worst admissible value is a support function, while margin invariance is characterized by an annihilator condition. We introduce an intrinsic calibrated exposure measure, governed by the leverage duality χ 2=1/ -1, which separates observed drift into three diagnostically distinct regimes: a reader fault removable by recalibration, an exact correction that is too ill-conditioned to be reliable, and a representation-level collision that no readout-only intervention can remove. Thus, identical observed exposure can lead to fundamentally different remediation verdicts. The framework also extends to cone-valued safety heads. An untied order-swap identity provides a diagnostic for the linear control interface; its calibration-state residual predicts a distinct three-control composition error on unseen states and targets, achieving median Spearman correlation 0.964, compared with 0.269 for a static cross-Gram baseline. etc.....
Sources
- Cross Frame Potential
- Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
- Refusal in Language Models Is Mediated by a Single Direction
- Invariant Risk Minimization
- LEACE: Perfect linear concept erasure in closed form
- Understanding (Un)Reliability of Steering Vectors in Language Models
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
- Explaining and Harnessing Adversarial Examples
- Adversarial Examples Are Not Bugs, They Are Superposition
- How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings
- Steering Vector Fields for Context-Aware Inference-Time Control in Large Language Models
- The Illusion of Cross-Lingual Safety in Low-Resource Languages
- Generalized Leverage Scores: Geometric Interpretation and Applications
- Linear Adversarial Concept Erasure
- Kernelized Concept Erasure
- Polysemanticity and Capacity in Neural Networks
- Intriguing properties of neural networks
- Analyzing the Generalization and Reliability of Steering Vectors
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks