ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
cs.CV, cs.AI
Submitted: 2026-01-30
Updated: 2026-09-16
Comments: EMNLP 2026 Findings, 30 pages, 9 figures, Project website: https://github.com/yutao1024/ShotFinder
Code: https://github.com/yutao1024/ShotFinder
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection
- SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
- Learning to Cut by Watching Movies
- Learning Transferable Visual Models From Natural Language Supervision
- Playing for Data: Ground Truth from Computer Games
- Chrono: A Simple Blueprint for Representing Time in MLLMs
- WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
- Audio-Visual Event Localization in Unconstrained Videos
- WebDancer: Towards Autonomous Information Seeking Agency
- WebWalker: Benchmarking LLMs in Web Traversal
- Qwen3-Omni Technical Report
- Qwen3 Technical Report
- Aligning Multimodal LLM with Human Preference: A Survey
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models