Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
cs.LG
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/OpenEuroLLM/ComplexKDA
License: http://creativecommons.org/licenses/by/4.0/
The gist: Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correction constrain their expressivity.
Terminology
Abstract
Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correction constrain their expressivity. Prior work has shown that composing two delta-rule transitions in a single recurrent update can model a 2D rotation, but this increases the rank and the cost of the updates compared to a single transition. We show that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate. This requires extending the parameter ranges of KDA by combining two existing range extensions: allowing gates in [-1,1] and the delta-rule coefficient β in [0,2]. We call the resulting model Complex KDA (CKDA). It preserves KDA's stability and efficiency, with transitions that remain diagonal-plus-rank-one and non-expansive, while reaching the state-tracking expressivity of DeltaProduct 2. We characterize the expressivity of CKDA and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix. A single CKDA layer can track every finite group isomorphic to a subgroup of SO(3), and many state-tracking results use one fewer layer for CKDA compared to other diagonal-plus-rank-one Linear RNNs. Empirically, combining both extensions yields the strongest length extrapolation among tested KDA range settings on S 3, S 4, and periodic audio continuation. In language modeling, CKDA outperforms Transformers and other linear RNNs, obtains similar results to a KDA baseline, and shows promising scaling behavior. Our code is open source at https://github.com/OpenEuroLLM/ComplexKDA and our models are available at https://huggingface.co/collections/openeurollm/complexkda.
Sources
- Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
- RotRNN: Modelling Long Sequences with Rotations
- Rethinking State Tracking in Recurrent Models Through Error Control Dynamics
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Improved state mixing in higher-order and block diagonal linear recurrent networks
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
- Training Compute-Optimal Large Language Models
- MDN: Parallelizing Stepwise Momentum for Delta Linear Attention
- Scaling Laws for Neural Language Models
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Kimi K3: Open Frontier Intelligence
- A matrix-based proof of the quaternion representation theorem for four-dimensional rotations
- M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
- An Algebraic View of the Expressivity of Recurrent Language Models
- On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
- Qwen3.5-Omni Technical Report
- DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
- Learning State-Tracking from Code Using Linear RNNs
- FG squared-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks