VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
cs.CV, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Multimodal Large Language Models (MLLMs) have achieved remarkable progress on short video understanding yet remain limited on long videos due to the limited visual context window.
Terminology
Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable progress on short video understanding yet remain limited on long videos due to the limited visual context window. Prevailing approaches rely on uniform frame sampling or recent coarse-to-fine agentic zooming, both of which struggle to localize sparse, decisive evidence in sufficiently long videos. We formulate long video understanding as a Sequential Evidence Acquisition (SEA) problem, in which an agent reads the video turn by turn along the temporal axis, deciding at each turn how fast to watch, what evidence to retain, when to revisit uncertain segments, and when to stop and answer. Inspired by this view, we propose VideoScout, a multi-turn reasoning agent that instantiates the SEA paradigm through adaptive reasoning pacing. Specifically, by dynamically controlling the viewing pace, VideoScout enables efficient traversal of long videos within a bounded visual context window, allowing the agent to access more video content while balancing content analysis depth with reading efficiency. To train VideoScout, we construct VideoScout-66K, a set of over 66K high-quality exploration turns from 10K answer-verified trajectories, and adopt a two-stage pipeline: cold-start supervised fine-tuning teaches the agent per-turn output format, while the Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO) algorithm performs trajectory-level reinforcement learning with a composite reward that jointly considers answer accuracy, output format compliance, and the temporal alignment between the agent's viewing progress and the teacher's answer timing measured by intersection-over-union (IoU). Extensive experiments on long video understanding and reasoning benchmarks demonstrate that our 7B model achieves strong performance compared with existing trained 7B agentic models.
Sources
- Qwen2.5-VL Technical Report
- InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
- Long Context Transfer from Language to Vision
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- LLaVA-OneVision: Easy Visual Task Transfer
- VideoChat: Chat-Centric Video Understanding
- MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
- KeyVideoLLM: Towards Large-scale Video Keyframe Selection
- Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
- CoS: Chain-of-Shot Prompting for Long Video Understanding
- Test-Time Temporal Sampling for Efficient MLLM Video Understanding
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
- ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
- Efficient Sequence Packing without Cross-contamination: Accelerating Large Language Models without Impacting Performance
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models