Beyond the Survey: A Systematic Empirical Study of Detection and Association in Visual MOT
cs.CV
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/linh-gist/VisualMOT
Terminology
Sources
- Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art
- Recent Advances in Embedding Methods for Multi-Object Tracking: A Survey
- MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking
- CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
- YOLOX: Exceeding YOLO Series in 2021
- YOLOv11: An Overview of the Key Architectural Enhancements
- MOT16: A Benchmark for Multi-Object Tracking
- CrowdHuman: A Benchmark for Detecting Human in a Crowd
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- BoT-SORT: Robust Associations Multi-Pedestrian Tracking
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
- The Mean of Multi-Object Trajectories
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models