Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?
cs.CL, cs.AI
Submitted: 2026-09-08
Updated: 2026-09-18
Comments: Experimental study of attention sinks, long-context recall, and million-token context behavior. Code and measurement protocol are available at https://github.com/sararizwan7/Attention-Mechanisms-in-1M-Context-Window
Code: https://github.com/sararizwan7/Attention-Mechanisms-in-1M-Context-Window
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long context language models now advertise windows of one million tokens, but two habits limit how much of that window is used.
Terminology
Abstract
Long context language models now advertise windows of one million tokens, but two habits limit how much of that window is used. Attention heads with nothing useful to read still spend their budget on the first token, which is called the attention sink, and where a fact sits in the context changes whether the model finds it. Gated attention cut first token attention from 46.7 percent to 4.8 percent at NeurIPS 2025, and Kimi K3 pairs that idea with Kimi Delta Attention and Attention Residuals behind a one million token window, eight times past the range where these diagnostics have been reported. This paper asks whether the fix survives that jump. We build SinkProbe, a suite that measures sink mass, massive activation, position resolved recall and the recency gap, and apply it to four small models that differ only in how they mix tokens and depth. Three results follow. The training objective produces the sink, not the architecture. Gating did not reproduce its published effect at our scale. Sink mass, activations and position bias moved independently. Code, data and the measurement protocol are released at https://github.com/sararizwan7/Attention-Mechanisms-in-1M-Context-Window
Sources
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- Kimi K3: Open Frontier Intelligence
- Attention Residuals
- When Attention Sink Emerges in Language Models: An Empirical View
- Lost in the Middle, and In-Between: Enhancing Language Models' Ability to Reason Over Long Contexts in Multi-Hop QA
- Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
- YaRN: Efficient Context Window Extension of Large Language Models
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
- Gated Sparse Attention: Combining Computational Efficiency with Training Stability for Long-Context Language Models
- Delta Attention Residuals
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering