Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents
cs.CV, cs.AI
Submitted: 2026-08-24
Updated: 2026-08-25
Project page: https://liuwq-bit.github.io/VideoRover
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen3-VL Technical Report
- OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents
- Video-Browser: Towards Agentic Open-web Video Browsing
- Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
- Watch, Remember, Reason: Human-View Video Understanding with MLLMs
- WebGPT: Browser-assisted question-answering with human feedback
- LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
- Deep Research: A Systematic Survey
- OpenAI GPT-5 System Card
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications
- MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline
- Long Context Transfer from Language to Vision
- DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
- Group Sequence Policy Optimization
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models