AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research
cs.CV, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/Accio-Lab/AdaVDR
Project page: https://accio-lab.github.io/AdaVDR
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen3-VL Technical Report
- ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
- OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
- REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents
- VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning
- Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models
- GPT-4o System Card
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
- HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents
- WebSailor: Navigating Super-human Reasoning for Web Agent
- Video-Browser: Towards Agentic Open-web Video Browsing
- Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
- VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
- MR. Video: "MapReduce" is the Principle for Long Video Understanding
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- OpenAI GPT-5 System Card
- Gemini: A Family of Highly Capable Multimodal Models
- Qwen3.5-Omni Technical Report
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models