Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos
cs.CV, cs.LG
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/HKUST-KnowComp/PACE
Terminology
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- CogVLM2: Visual Language Models for Image and Video Understanding
- GPT-4o System Card
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- MR. Video: "MapReduce" is the Principle for Long Video Understanding
- Gemini: A Family of Highly Capable Multimodal Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- LVBench: An Extreme Long Video Understanding Benchmark
- ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
- InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
- Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
- Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
- Towards Multi-Agent Reasoning Systems for Collaborative Expertise Delegation: An Exploratory Design Study
- The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
- VCA: Video Curious Agent for Long Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models