KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding
cs.CV, cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Token Merging: Your ViT But Faster
- SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
- From Pixels to Words -- Towards Native One-Vision Models at Scale
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
- Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
- LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
- VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding
- StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
- PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models