DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video
cs.CV, cs.CL, cs.LG
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/MaumAI-Company/DeltaS
Terminology
Sources
- Stateful Token Reduction for Long-Video Hybrid VLMs
- Kimi Linear: An Expressive, Efficient Attention Architecture
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
- A Simple Baseline for Streaming Video Understanding
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models