Mixture-of-Top-k Attention: Efficient Attention via Scalable Fast Weights
cs.LG, cs.CV
Submitted: 2026-02-01
Updated: 2026-09-27
Code: https://github.com/QishuaiWen/MiTA
Terminology
Sources
- ViT$^3$: Unlocking Test-Time Training in Vision
- Parallel Loop Transformer for Efficient Test-Time Computation Scaling
- MoBA: Mixture of Block Attention for Long-Context LLMs
- MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
- Linformer: Self-Attention with Linear Complexity
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- VMoBA: Mixture-of-Block Attention for Video Diffusion Models
- Mixture of Contexts for Long Video Generation
- MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
- Distilling the Knowledge in a Neural Network
- Titans: Learning to Memorize at Test Time
- General-Purpose In-Context Learning by Meta-Learning Transformers
- Scaling Laws for Neural Language Models
- Longformer: The Long-Document Transformer
- Online normalizer calculation for softmax
- ViT-5: Vision Transformers for The Mid-2020s
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks