VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding
cs.CV, cs.AI
Submitted: 2026-10-08
Updated: 2026-10-08
Terminology
Sources
- Advancing Content Moderation: Evaluating Large Language Models for Detecting Sensitive Content Across Text, Images, and Videos
- Qwen3-VL Technical Report
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
- Group-in-Group Policy Optimization for LLM Agent Training
- LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
- FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- LensWalk: Agentic Video Understanding by Planning How You See in Videos
- ToRL: Scaling Tool-Integrated RL
- Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
- UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
- MR. Video: "MapReduce" is the Principle for Long Video Understanding
- VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
- Proximal Policy Optimization Algorithms
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
- VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
- Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models