Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization
cs.LG
Submitted: 2026-08-12
Updated: 2026-08-12
Code: https://github.com/pvarshh/quantinterp
Terminology
Sources
- Refusal in Language Models Is Mediated by a Single Direction
- Understanding (Un)Reliability of Steering Vectors in Language Models
- Why Do Some Inputs Break Low-Bit LLM Quantization?
- Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness
- Perplexity Can Miss SAE Feature Damage Under Quantization
- On the transferability of Sparse Autoencoders for interpreting compressed models
- Quality Is Not a Safety Proxy Under Quantization
- What Cosine Similarity of Label Representations Can and Cannot Tell us
- Massive Activations in Large Language Models
- Semantics at an Angle: When Cosine Similarity Works Until It Doesn't
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks