GTR: Gated Token Recurrence for Efficient Dense Prediction
cs.CV, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-23
Comments: Project page is available at: https://intellindust-ai-lab.github.io/projects/GTR/
Code: https://github.com/fla-org/flash-linear-attention
Project page: https://intellindust-ai-lab.github.io/projects/GTR
License: http://creativecommons.org/licenses/by/4.0/
The gist: Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases.
Terminology
Abstract
Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated linear attention, alternating spatial scan directions, and spatially enhanced SwiGLU blocks. GTR is distilled from a detection-specialized DINOv3 teacher using only final-layer patch-token alignment through a linear projection and squared 2 loss, without masked-token prediction or intermediate-layer supervision. With Objects365 detector pre-training, GTR-L achieves 58.9 box AP on COCO val2017 with 1.908,ms median batch-one latency under compiled FP16 execution on an RTX 4090. The same backbone also transfers to instance segmentation, pose estimation, oriented detection, semantic segmentation, and monocular depth estimation. In an isolated kernel benchmark, our specialized chunkwise CUDA operator is 4.0 times faster than FLA v0.5.0 at 1.6K tokens on RTX 4090. TensorRT deployment on DRIVE AGX Thor achieves 2.282--8.769,ms median batch-one latency across the evaluated models. These results show that recurrent token mixing can provide an efficient alternative to global softmax attention for high-resolution dense prediction and edge deployment.Project page: https://intellindust-ai-lab.github.io/projects/GTR/
Sources
- ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
- LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models
- SoLA-Vision: Fine-grained Layer-wise Linear Softmax Hybrid Attention
- ViT-AdaLA: Adapting Vision Transformers with Linear Attention
- EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized Distillation
- RTMDet: An Empirical Study of Designing Real-Time Object Detectors
- Attention to Mamba: A Recipe for Cross-Architecture Distillation
- GLU Variants Improve Transformer
- Retentive Network: A Successor to Transformer for Large Language Models
- Kimi Linear: An Expressive, Efficient Attention Architecture
- ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models