OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Project page: https://mcg-nju.github.io/OneStreamer
Terminology
Sources
- Qwen3-VL Technical Report
- VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
- Thinking in Streaming Video
- AURA: Always-On Understanding and Real-Time Assistance via Video Streams
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
- A Simple Baseline for Streaming Video Understanding
- Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
- InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
- ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
- StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
- FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
- Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
- StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
- Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
- Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
- Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
- Long Context Transfer from Language to Vision
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models