Towards Automatic Video Annotation with ASH: Zero-Shot Open-Vocabulary Multi-Object Tracking and Segmentation
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Terminology
Sources
- MOT20: A benchmark for multi object tracking in crowded scenes
- SAM2Auto: Auto Annotation Using FLASH
- Qwen3-VL Technical Report
- Perception Encoder: The best visual embeddings are not at the output of the network
- Addressable Memory for Video World Models
- MOT16: A Benchmark for Multi-Object Tracking
- Seg2Track-SAM2: SAM2-based Multi-object Tracking and Segmentation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models