LiveProBench: Can Streaming Video Models Really Interact Like Humans?
cs.LG
Submitted: 2026-09-11
Updated: 2026-09-22
Comments: Code and data is available at https://github.com/v0yager33/ProactiveBench
Code: https://github.com/v0yager33/ProactiveBench
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Streaming video understanding requires models to process continuous multimodal input while maintaining temporal context.
Terminology
Abstract
Streaming video understanding requires models to process continuous multimodal input while maintaining temporal context. Existing evaluations are predominantly reactive: they query a model at a selected timestamp and therefore do not assess when it should respond. Proactive interaction instead requires monitoring a standing request, responding within an appropriate interval after the target event, and otherwise remaining silent. We introduce ProactiveBench, which evaluates models at one-second stream intervals without an explicit response cue. Its six subtasks vary trigger ambiguity and timing tolerance. Event Sensitivity geometrically combines response and silence rates on the same recording; four window-based subtasks distinguish early, in-window, and missed responses; and Duplicate Counting penalizes omissions and repetitions. Premature responses outnumber missed responses for four of the six evaluated systems, revealing a substantial gap in the temporal decision-making required for human-like interaction.
Sources
- VideoLLM-online: Online Video Large Language Model for Streaming Video
- Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video
- VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
- JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
- Online Video Understanding: OVBench and VideoChat-Online
- MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
- OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
- Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
- StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
- StreamReady: Learning What to Answer and When in Long Streaming Videos
- OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
- Streaming Video Instruction Tuning
- AURA: Always-On Understanding and Real-Time Assistance via Video Streams
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks