What Should a Streaming Video Model Remember?
cs.CV, cs.AI
Submitted: 2026-06-15
Updated: 2026-09-26
Terminology
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- ContextNav: Towards Agentic Multimodal In-Context Learning
- VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
- FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
- MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
- Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
- GUI Agents for Continual Game Generation
- StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
- Thinking in Streaming Video
- PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- A Simple Baseline for Streaming Video Understanding
- RIVER: A Real-Time Interaction Benchmark for Video LLMs
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models