PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding
cs.CV, cs.AI
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 22 pages, 11 figures, 12 tables
Code: https://github.com/siruzhong/PReM
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value (KV) caches.
Terminology
Abstract
Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value (KV) caches. We introduce Prefix-Steered Recurrent Memory (PREM), a memory-token-free framework for frozen vision-language models (VLMs). PREM separates video ingestion from query answering: a recurrent writer distills visual streams into a compact 256 KiB multi-slot associative state, while a question-conditioned readout adds memory-derived key/value (K/V) steering modulations to existing non-visual prompt prefixes during prefill. This enables write-once, query-many inference without extra prompt tokens or decoding recurrence. Across six long-video benchmarks in offline and streaming end-of-stream settings, PREM consistently outperforms frozen baselines at every evaluated visual budget. Under a constrained budget of 16 frames, PREM improves macro-average accuracy by 3.06% on Qwen2.5-VL-3B, with gains of 11.0% on action antonym identification and 9.9% on localized needle retrieval. These gains require tuning 0.24% of backbone parameters at 0.03 GiB of peak GPU memory overhead.
Sources
- ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- delta-mem: Efficient Online Memory for Large Language Models
- Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
- LVBench: An Extreme Long Video Understanding Benchmark
- SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models