Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling
cs.LG, cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/OliverSieberling/TriadicLinearAttention
Terminology
Sources
- Zoology: Measuring and Improving Recall in Efficient Language Models
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
- Decoupled Weight Decay Regularization
- Olmo 3
- Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence
- RWKV-7 "Goose" with Expressive Dynamic State Evolution
- Fast Transformer Decoding: One Write-Head is All You Need
- GLU Variants Improve Transformer
- DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Retentive Network: A Successor to Transformer for Large Language Models
- End-to-End Test-Time Training for Long Context
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
- Qwen3 Technical Report
- Gated Linear Attention Transformers with Hardware-Efficient Training
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks