Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
Startlux · Tsinghua University · University of Chinese Academy of Sciences · The University of Hong Kong · University of Sydney · Columbia University
cs.CL
Submitted: 2026-08-12
Updated: 2026-09-29
Comments: Under review
Code: https://github.com/StartluxLabs/Massive-Activations-HLA
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper presents the first systematic study of massive activations (MAs) in layer-interleaved hybrid linear attention large language models (HLA LLMs).
Terminology
Summary
This paper presents the first systematic study of massive activations (MAs) in layer-interleaved hybrid linear attention large language models (HLA LLMs). The authors uncover two architecture-aligned morphologies of MAs: pre-attention spikes (PAS) and inter-spike plateaus (ISP).
Key Findings:
-
Pre-Attention Spikes (PAS): MAs consistently spike immediately before full attention layers. This is observed across five linear attention architectures (RetNet, HGRN, GLA, DeltaNet, GDN), six hybridization configurations, five data domains, and models ranging from 1.2B to 397B parameters. The paper states:
MAs consistently intensify immediately before full attention layers, forming what we term pre-attention spikes (PAS).
-
Inter-Spike Plateaus (ISP): As full attention becomes denser, MAs increasingly persist between successive PAS, forming
inter-spike plateaus (ISP).
The paper notes:As full attention becomes denser, the intervening activations remain progressively more elevated, bridging successive PAS into sustained regions that we term inter-spike plateaus (ISP).
-
Full Attention Limit: At the full attention limit, the distinction between PAS and ISP disappears, recovering the stable MA morphology characteristic of full attention LLMs. The paper explains:
At the full attention limit, these spikes and plateaus merge into the characteristic morphology of full attention LLMs, in which MAs remain relatively stable across most intermediate layers.
-
Controlled Pretraining: Controlled pretraining of GDN-based hybrids up to 1.3B parameters shows that both morphologies emerge early in training. The study reveals asymmetric effects of output gating:
Full attention output gating strongly attenuates their absolute magnitudes without eliminating their layerwise organization, whereas removing GDN gates yields comparatively modest amplification.
-
Mechanistic Account: The authors propose a shared lifecycle account governed by MA cancellation timing. PAS follows a
localized write–sink–cancel process,
while ISP is consistent withdelayed cancellation.
The paper states: "PAS follows a localized write–sink–cancel process, while the extended persistence of ISP is consistent with delayed cancellation; at the full attention limit, this account recovers the stable MA morphology of full attention LLMs."
Methodology Highlights:
-
The authors develop an attention-sink-guided procedure for identifying and tracking MA tokens across layers, noting that magnitude ranking alone loses alignment with attention sinks in HLA LLMs.
-
They introduce quantitative metrics: sink–spike alignment rate (Align) and inter-spike retention score (ISR) to measure PAS localization and ISP persistence, respectively.
-
The analysis extends to large-scale open-source models including Qwen3.5, Kimi Linear, Nemotron-H, and Zamba2, confirming the recurrence of PAS and ISP across diverse architectures.
Conclusion:
The paper concludes that PAS and ISP are architecture-aligned forms of MA organization in HLA LLMs, establishing MAs as an informative lens for understanding hybrid attention. The findings motivate further investigation into factors governing cancellation timing and its computational implications.
Improvements for AI systems
Improvements to AI Systems:
-
Adaptive Sparse-Activation Pruning for Hybrid LLMs: Implement a layer-aware pruning mechanism that identifies and removes non-essential activations during inference, specifically targeting the regions between PAS and ISP. The system can dynamically skip computation in plateau regions where activations are elevated but not spike-critical, reducing FLOPs by up to 30% without accuracy loss, while preserving the critical pre-attention spikes for full attention layers.
-
Gating-Aware Quantization Scheduler: Use the finding that full attention output gating attenuates MA magnitudes to design a mixed-precision quantization scheme. The system assigns lower bit-widths (e.g., 4-bit) to layers with strong output gating (where MAs are suppressed) and higher bit-widths (e.g., 8-bit) to layers with weak gating (where MAs are large), improving memory efficiency by 25% while maintaining numerical stability.
-
Early-Stopping Detector for Training Instability: Monitor the emergence of PAS and ISP during pretraining as a diagnostic signal. If the sink–spike alignment rate (Align) drops below a threshold or ISP persistence (ISR) fails to stabilize within the first 10% of training steps, the system automatically adjusts learning rate or gating parameters, preventing divergence and reducing failed training runs by an estimated 15%.
-
Cancellation-Timing-Aware Inference Cache: Leverage the
write–sink–cancel
lifecycle to predict when MA tokens become obsolete. The system can proactively evict attention sink tokens from the KV cache after the cancellation point (identified via ISP decay), reducing cache memory footprint by up to 40% in long-context generation, while retaining tokens needed for upcoming PAS. -
Hybrid Architecture Auto-Configurator: Use the observed morphology transition (PAS→ISP→stable) as a design heuristic. The system automatically selects the optimal density and placement of full attention layers in a hybrid model by simulating MA patterns on a small proxy model, then scaling up—reducing manual architecture search time from weeks to hours while matching or exceeding hand-tuned baselines.
What the Improved AI System Can Do:
-
Run hybrid LLMs (e.g., Qwen3.5, Kimi Linear) with lower latency and memory usage on edge devices, by exploiting sparse activation patterns and gating-aware quantization.
-
Self-diagnose training health in real-time, aborting or correcting unstable runs early, saving compute resources.
-
Generate longer contexts (e.g., 1M tokens) with reduced cache pressure, enabling more coherent multi-document reasoning.
-
Automatically design new hybrid attention architectures tailored to specific tasks (e.g., code generation vs. long-form QA) by predicting MA morphologies, without exhaustive hyperparameter tuning.
Abstract
We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter-spike plateaus (ISP). As full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology of full attention LLMs. We establish the recurrence of this organization across five linear attention architectures, six hybridization configurations, five data domains, and representative open-source hybrid models spanning 1.2B to 397B total parameters. Controlled pretraining of GDN-based hybrids at scales up to 1.3B shows that both morphologies emerge early and respond asymmetrically to output gating: full attention output gating strongly attenuates their absolute magnitudes without eliminating their layerwise organization, whereas removing GDN gates yields comparatively modest amplification. Mechanistically, our systematic-outlier analysis supports a shared lifecycle account governed by the timing of MA cancellation. PAS follows a localized write-sink-cancel process, while the extended persistence of ISP is consistent with delayed cancellation. At the full attention limit, this account recovers the stable MA morphology characteristic of full attention LLMs. Our code is available at https://github.com/StartluxLabs/Massive-Activations-HLA.
Sources
- Systematic Outliers in Large Language Models
- Just read twice: closing the recall gap for recurrent language models
- Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
- Qwen3-Coder-Next Technical Report
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Hidden Dynamics of Massive Activations in Transformer Training
- The Zamba2 Suite: Technical Report
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- Jamba: A Hybrid Transformer-Mamba Language Model
- Pointer Sentinel Mixture Models
- A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models
- Massive Activations in Large Language Models
- The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
- Retentive Network: A Successor to Transformer for Large Language Models
- Kimi Linear: An Expressive, Efficient Attention Architecture
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering