Progressive Behavioral Drift through Compression Valleys in Large Language Models
cs.CR
Submitted: 2025-11-21
Updated: 2026-08-28
Comments: EMNLP 2026
Code: https://github.com/EleutherAI/lm-evaluation-harness
License: http://creativecommons.org/licenses/by/4.0/
The gist: We show that attention sinks and compression valleys create a vulnerable region in decoder-only Transformers, where small activation perturbations can be amplified through the autoregressive
Terminology
Abstract
We show that attention sinks and compression valleys create a vulnerable region in decoder-only Transformers, where small activation perturbations can be amplified through the autoregressive trajectory. Based on this, we propose Sensitivity-Scaled Steering (SSS), a progressive activation-space attack that anchors perturbations at the beginning-of-sequence token and adaptively reinforces them at sensitive layers and tokens. Instead of forcing an abrupt behavioral change, SSS induces a staged drift, making outputs gradually shift toward the target behavior while remaining fluent and benign-looking in early generations. Across multiple open-weight models and four behavioral axes, SSS achieves high attack success, preserves coherence, and causes negligible degradation to general capabilities. These results show that attention sinks and compression valleys are not merely mechanistic features; rather, they expose exploitable amplification mechanisms that can be treated as hidden-state weaknesses for activation-space attacks in white-box and supply-chain LLM deployments.
Sources
- Attention Is All You Need
- Steering Language Models With Activation Engineering
- Extracting Latent Steering Vectors from Pretrained Language Models
- Steering Llama 2 via Contrastive Activation Addition
- Steering Large Language Model Activations in Sparse Spaces
- Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
- Efficient Streaming Language Models with Attention Sinks
- Spectral Filters, Dark Signals, and Attention Sinks
- Your Transformer is Secretly Linear
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Representation Engineering: A Top-Down Approach to AI Transparency
- LEACE: Perfect linear concept erasure in closed form
- Word Embeddings Are Steers for Language Models
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Refusal in Language Models Is Mediated by a Single Direction
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs