FG squared-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
cs.LG
Submitted: 2026-04-21
Updated: 2026-08-31
Code: https://github.com/fla-org/flash-linear-attention
License: http://creativecommons.org/licenses/by/4.0/
The gist: Linear attention mechanisms have emerged as promising alternatives to softmax attention, offering linear-time complexity during inference.
Terminology
Abstract
Linear attention mechanisms have emerged as promising alternatives to softmax attention, offering linear-time complexity during inference. Recent advances such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA) have demonstrated that the delta rule, an online gradient descent update, enables superior associative recall compared to simple additive updates. While KDA refined the coarse head-wise decay gate into channel-wise decay, the learning rate β t in the delta update remains a scalar, limiting the model's capacity for dimension-specific adaptation. We introduce FG squared-GDN, which replaces the scalar β t with a channel-wise vector analogous to the transition from SGD to per-coordinate adaptive optimizers such as AdaGrad and Adam. We further propose FG squared-GDN+, which decouples the scaling for keys and values, enabling independent control of erasure strength and write strength. Experiments on synthetic and real-world benchmarks show that FG squared-GDN and its variant improve associative recall and long-context understanding over GDN and KDA, with comparable computational efficiency.
Sources
- Gated Delta Networks: Improving Mamba2 with Delta Rule
- Adam: A Method for Stochastic Optimization
- Efficiently Modeling Long Sequences with Structured State Spaces
- RWKV-7 "Goose" with Expressive Dynamic State Evolution
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Rethinking Attention with Performers
- Retentive Network: A Successor to Transformer for Large Language Models
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- HGRN2: Gated Linear RNNs with State Expansion
- Longhorn: State Space Models are Amortized Online Learners
- Jamba: A Hybrid Transformer-Mamba Language Model
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks