XS-VID: A Large-Scale Benchmark for Small Object Detection and Tracking in Videos

arXiv:2407.18137 · cs.CV · Submitted 2024-07-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "XS-VID: A Large-Scale Benchmark for Small Object Detection and Tracking in Videos".

Tom: Small Video Object Detection (SVOD) remains a crucial subfield in modern computer vision essential for early object discovery and detection,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So we've talked about the title of this paper, "XS-VID: A Large-Scale Benchmark for Small Object Detection and Tracking in Videos," and who actually put this work together. What does that mean for us right now?

Jane: It just means they're building a big testing ground. They’re showing that there are problems with current small object detection methods because the data they use isn't varied enough to test everything we need.

Lu: They are addressing this by creating XS-VID, which includes aerial data from different times and scenes, and they annotating eight major object categories to make sure the testing is comprehensive.

Meng: Essentially, they’re trying to solve the problem of having too many small objects that don't look like anything specific in a fixed dataset.

Lalam: It signals that we need a new kind of benchmark for video object detection so we can actually compare different AI methods fairly.

Tom: That makes sense. So, what’s the main idea behind this whole XS-VID project that they’re proposing?

Jane: The core summary is that existing SVOD methods are limited because they just can't handle everything—they fail when objects get too tiny or when the background looks too much like the object itself.

Lu: They summarize this by saying datasets like ImageNetVID only have about three percent of small objects under three hundred twenty-two pixels, which really limits how effective those detection methods can be #pg0.

Meng: And they specifically highlight the three major hurdles: background confusion, easy misclassification because the objects lack texture, and texture distortion from those tiny pixel areas.

Lalam: The summary emphasizes that single-frame methods just don't work because they ignore the temporal aspect of movement needed to solve these problems.

The paper's summary: Tom: Moving into what the authors actually propose as improvements, which is where they introduce YOLOFT, which is their new video object detection method. What’s the gist of that?

Jane: They propose YOLOFT, which aims to fix those accuracy and stability issues by integrating temporal motion features of objects directly into the detection process.

Lu: YOLOFT works by merging a standard detector with a temporal fusion Neck module they call the Multi-Scale Spatio-Temporal Flow module.

Meng: They describe this fusion as constructing spatio-temporal correlations across consecutive frames using a correlation pyramid and a lookup operator from RAFT to get motion information #pg6.

Lalam: Essentially, they are making sure that each pixel at each scale gets different scale similarities from the previous frame so it can capture both big and small movements.

Tom: That mechanism directly tackles those three issues they mentioned—background confusion, misclassification, and texture distortion by using motion data strategically.

Jane: They showed that a predominance of motion information—a ratio of one static feature for every three motion features—results in the highest detection accuracy #pg9.

Lu: And they also found that training the static backbone first actually helped reduce standard deviations in key metrics by more than half and slightly improved accuracy by zero point three points #pg9.

Meng: That suggests a practical way to tune the model’s architecture to get faster processing times while keeping high quality results, which is something engineers always look for.

Lalam: They also looked at data augmentation and found that introducing small transformation variations like scale or perspective improved performance by seven percent on APeg #pg9.

The paper's improvements: Tom: So we’ve covered the dataset and the new method, so let’s look at the actual improvements they claim this approach makes in practice.

Jane: The main implication is that having this large-scale benchmark dataset gives researchers a place to stop using those limited datasets and start testing methods against the full spectrum of small object detection challenges.

Lu: It provides a standard for how to measure performance, which helps everyone compare apples to apples when developing new AI systems.

Meng: For me, it means we can finally move from theoretical proofs to actually building robust tools that work in real-time environments where things are messy and fast-paced.

Lalam: It opens the door for future work by showing exactly where the gaps are, which is how we pinpoint the next problems to solve.

Tom: So this whole effort with XS-VID and YOLOFT really shows us that addressing small object detection in video requires a deep understanding of both local features and temporal dynamics.

Jane: We’ve got a lot of data on our hands now to push the discussion forward on how we can make these detections even better for everyday use.

Lu: It’s an important step in establishing the baseline for what's possible with video object detection research.

Meng: It sets a high bar for what we need to achieve when we design embodied AI systems that have to perceive the physical world around them.

Conclusion: Tom: Wrapping up, so this paper introduces XS-VID and YOLOFT, showing how they tackle small object detection with comprehensive data and motion information.

Jane: Exactly, it’s about showing that you need both the right training material and a smart way to look at the video information for small things.

Lu: I think the real power of XS-VID is that it finally gives us a comprehensive way to test models across all those object sizes we’ve been struggling with.

Meng: From an engineering standpoint, having two hundred fifty-eight thousand object boxes and that variety means we can actually train models that aren't just good at finding medium objects but are reliable everywhere.

Lalam: I feel like this work is important because it gives us a better foundation for how AI learns to see the world around us, making those tiny details more visible in our culture.

Tom: It really does. The results show YOLOFT hitting state-of-the-art on both XS-VID and VisDrone2019VID, which is pretty solid performance across different video sources.

Jane: And the ablation studies they did, showing that motion information is way more important than static features for accuracy, really backs up their whole approach.

Lu: It proves that you can’t just rely on a single type of data or a single feature; you need to fuse the temporal and spatial information together effectively.

Meng: It tells us that we need to design architectures that can handle those complex correlations between consecutive frames without getting overwhelmed by noise.

Lalam: For me, this means we can start thinking about how AI systems can track things more consistently over time, which is a big step for agents navigating real environments.

Tom: So to recap, XS-VID provides the deep coverage we needed for small object detection and YOLOFT gives us a method that integrates motion to nail those tricky challenges.

Jane: It’s a lot of data and a smart new technique all wrapped up in the XS-VID paper.

Lu: This work sets a strong benchmark for future research trying to push the limits of what video object detection can achieve on tiny targets.

Meng: We’ll keep an eye on how these performance metrics translate into actual deployment scenarios for things like surveillance or autonomous systems.

Lalam: It gives us a clearer picture of how much more detail AI can extract from moving images, which is really inspiring for the future of perception technology.

Tom: That’s it for this paper. We’ve got a lot to chew on here before we move on to another piece of research in the world of video AI.

Jiahao Guo, Ziyang Xu, Lianjun Wu, Fei Gao, Wenyu Liu, Xinggang Wang

Huazhong University of Science and Technology

cs.CV

Submitted: 2024-07-25

Updated: 2026-10-05

Comments: Accepted for publication in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). 22 pages, including supplementary material

DOI: 10.1109/TPAMI.2026.3741044

Code: https://github.com/open-mmlab/mmtracking

Project page: https://gjhhust.github.io/XS-VID

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: Small Video Object Detection (SVOD) remains a crucial subfield in modern computer vision essential for early object discovery and detection, yet existing datasets suffer from scarcity and issues like

Key concepts

XS-VID Dataset
This dataset was constructed to comprehensively evaluate small video object detection methods. It includes 38 video segments across 10 diverse scenes, covering a wide range of small object sizes from 0 to 322 pixels, providing the most complete coverage available for this task.
YOLOFT
YOLOFT is a new video object detection method that enhances accuracy and stability. It integrates YOLOv8 with recurrent all-pairs Field Transforms based optical flow. This allows the model to better understand local features and incorporate temporal motion information from objects across consecutive frames.
Local Feature Associations
This refers to how the network connects features within a small area of an image or video frame. YOLOFT improves this by enhancing these associations, which is crucial for detecting objects with weak texture and contours, as seen in challenging small object scenarios.
Temporal Motion Features
These are features derived from the movement of objects across multiple frames. By integrating motion information using optical flow, YOLOFT can capture how small objects move over time. This temporal understanding helps overcome background confusion and texture distortion issues inherent in small object detection.

Terminology

Summary

Small Video Object Detection (SVOD) remains a crucial subfield in modern computer vision essential for early object discovery and detection, yet existing datasets suffer from scarcity and issues like insufficiently small objects, limited categories, and lack of scene diversity, leading to unitary application scenarios for corresponding methods. This paper addresses this gap by proposing the XS-VID dataset and introducing YOLOFT, a novel video object detection method that significantly improves the accuracy and stability of SVOD by enhancing local feature associations and integrating temporal motion features of objects.

The gist: XS-VID offers the most comprehensive coverage of small object sizes, the largest number and proportion of extremely small objects, and the widest range of scene types known <ref:2407.18137#pg3>, while YOLOFT significantly improves the accuracy and stability of SVOD by enhancing local feature associations and integrating temporal motion features of objects <ref:2407.18137#pg7>.

The XS-VID Dataset

XS-VID was constructed to comprehensively evaluate small video object detection methods, consisting of 38 video segments with a resolution of 1024 × 1024, containing 12,230 images and 258,944 object boxes <ref:2407.18137#pg4>. The dataset covers multiple object sizes across 10 types of scenes, including rivers, forests, skyscrapers, and roads <ref:2407.18137#pg3>. The small objects in XS-VID are not concentrated at a fixed size but comprehensively cover small object sizes ranging from 0 ∼ 322 <ref:2407.18137#pg6>. Specifically, the number of objects within the ranges of es (0 ∼ 122), rs (122 ∼ 202), gs (202 ∼ 322), and normal (>3>3) are 49k, 94k, 36k, and 72k, respectively <ref:2407.18137#pg6>. XS-VID offers the most comprehensive coverage of small object sizes (0 ∼ 322), the largest number and proportion of es (extremely small, ≤ 122) objects <ref:2407.18137#pg11>, and the widest range of scene types known <ref:2407.18137#pg11>.

Challenges in SVOD

Existing datasets suffer from three main challenges that SVOD faces: (1) Background confusion, where The background texture is weak and similar in color to the object, making the object difficult to discover <ref:2407.18137#pg3>, (2) Easy to misclassify, as Small objects have insufficient texture and contour features, easily misleading the network to make wrong classifications <ref:2407.18137#pg3>, and (3) Texture distortion, where The pixel range of small objects is extremely small, causing severe degradation of texture features <ref:2407.18137#pg3>. Because of these issues, adopting single-frame Small Object Detection (SOD) methods or Video Object Detection (VOD) methods directly does not perform well <ref:2407.18137#pg3>. SOD methods rely on static single-frame features for detection and lack the utilization of temporal features, failing to adequately address the three existing issues <ref:2407.18137#pg3>. VOD methods only focus on medium and large object sizes and the network pipeline is not suitable for extremely small objects, resulting in poor detection performance <ref:2407.18137#pg3>.

Proposed Method: YOLOFT

To address the above issues, YOLOFT is proposed as an SVOD network that integrates YOLOv8 with the recurrent all-pairs Field Transforms [25] based optical flow <ref:2407.18137#pg7>. This method significantly improves accuracy and stability by enhancing local feature associations and integrating temporal motion features of objects <ref:2407.18137#pg7>. The overall architecture of YOLOFT consists of a YOLOv8 detector and a temporal fusion Neck, termed the Multi-Scale Spatio-Temporal Flow (MSTF) module, to enhance spatio-temporal feature representation across consecutive frames <ref:2407.18137#pg7>. The MSTF module constructs spatio-temporal correlations by constructing a Correlation Pyramid and using the lookup operator LC from RAFT [25] to obtain motion information of the object in the current feature map <ref:2407.18137#pg7>.

Experimental Results and Ablations

Extensive experiments on XS-VID and VisDrone2019VID demonstrate that YOLOFT achieves state-of-the-art performance <ref:2407.18137#pg7>. In comparison with other methods, YOLOFT is shown to achieve the most accurate predictions matching the GT boxes in both day and night scenes, significantly outperforming other methods <ref:2407.18137#pg9>. Specifically, when comparing VOD methods on XS-VID and VisDrone2019VID test-dev with Single-Frame Object Detectors, Video Object Detectors, Small Object Detectors, YOLO family and YOLOFT <ref:2407.18137#pg8>, YOLOFT is highlighted as the second highest performing method <ref:2407.18137#pg8>. Ablation studies confirmed that A predominance of motion information (static:motion = 1:3) results in the highest detection accuracy <ref:2407.18137#pg10>. Furthermore, Training the static backbone first and then fusing motion features reduced standard deviations by more than half and slightly improved accuracy by 0.3 points <ref:2407.18137#pg10>.

Design Considerations

The design considerations for YOLOFT focused on several key factors to improve small object detection performance. Key findings included: (1) Quality Motion Gradient Information is Essential for Small Object Feature Extraction, as complex multi-frame integration designs often lead to elongated gradient flows and weak feature responses <ref:2407.18137#pg10>. (2) High-quality static features enhance detection robustness, as training the static backbone first significantly improved overall accuracy <ref:2407.18137#pg10>. (3) Local optical flow information is useful for small objects, as using the highest resolution scale achieved the best accuracy (+0.2), and using all scales with fusion improved APeg the most (+0.7) <ref:2407.18137#pg12>. Additionally, Appropriate video-level data augmentation improves small object detection, where transformations such as Degree, Scale, and perspective were examined, showing that Introducing small transformation variations improved APeg by 7% and slightly improved overall AP <ref:2407.18137#pg10>.

Conclusion

In summary, the work proposes XS-VID to fill the data gap for SVOD with comprehensive coverage of small object sizes and scene diversity <ref:2407.18137#pg11>, and YOLOFT, which significantly improves accuracy by enhancing local feature associations and integrating temporal motion features <ref:2407.18137#pg7>. The results consistently demonstrate that YOLOFT achieves SOTA performance on both XS-VID and VisDrone2019VID <ref:2407.18137#pg7>. We have released our entire dataset and the benchmarks, hoping our dataset leads to further advancements in SVOD.

REFERENCES

[1] Yancheng Bai, Yongqiang Zhang, Mingli Ding, and Bernard Ghanem. Finding tiny faces in the wild with generative adversarial network. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21–30, 2018.

[2] Changrui Chen, Yu Zhang, Qingxuan Lv, Shuo Wei, Xiaorui Wang, Xin Sun, and Junyu Dong. Rrnet: A hybrid detector for object detection in drone-captured images. In 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), pages 100–108, 2019 <ref:2407.18137#pg4>.

[3] Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhuoaiy. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019 <ref:2407.18137#pg4>.

[4] Yihong Chen, Yue Cao, Han Hu, and Liwei Wang. Memory enhanced global-local aggregation for video object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020 <ref:2407.18137#pg5>.

[5] Yihong Chen, Zheng Zhang, Yue Cao, Liwei Wang, Stephen Lin, and Han Hu. Reppoints v2: Verification meets regression for object detection. Advances in Neural Information Processing Systems, 33:5621–5631, 2020 <ref:2407.18137#pg6>.

[6] MMTracking Contributors. MMTracking: OpenMMLab video perception toolbox and benchmark. https://github.com/open-mmlab/mmtracking, 2020 <ref:2407.18137#pg5>.

[7] Xiyang Dai, Yinpeng Chen, Bin Xiao, Dongdong Chen, Mengchen Liu, Lu Yuan, and Lei Zhang. Dynamic head: Unifying object detection heads with attentions, 2021 <ref:2407.18137#pg8>.

[8] Bowei Du, Yecheng Huang, Jiaxin Chen, and Di Huang. Adaptive sparse convolutional networks with global context enhancement for faster object detection on drone images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13435–13444, June 2023 <ref:2407.18137#pg9>.

[9] Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, and Jian Sun.

Improvements for AI systems

  1. Improved object detection capabilities through YOLOFT, which enhances local feature associations and integrates temporal motion features of objects, allowing for significantly improving the accuracy and stability of SVOD. This enables systems to perform superior small object detection by leveraging both static and dynamic cues captured across consecutive frames.

  2. Enhanced robustness against scene challenges by addressing specific difficulties in SVOD, such as Background confusion, misclassification, and texture distortion. The proposed method allows models to overcome these issues where existing methods struggle with small object detection, leading to better performance even in complex or visually ambiguous environments.

  3. Superior generalization across diverse object sizes and categories by utilizing the comprehensive XS-VID dataset, which offers unprecedented breadth and depth in covering and quantifying minuscule objects ranging from extremely small (es) to generally small (gs). This allows developed AI systems to perform reliably on objects of any size within the defined range (0 ∼ 322 pixels).

  4. Increased detection precision through multi-scale feature fusion, as demonstrated by the MSTF module which directly constructs multiple correlation volumes between each scale of the feature map g(l)t and all scales of the previous frame's feature map g(k)t−1. This mechanism ensures that each pixel at each scale can obtain different scale similarities from the previous frame, crucial for capturing both large and small movements.

  5. Optimized inference efficiency while maintaining high accuracy by incorporating architectural refinements, such as retaining the original gradient architecture and training static features first to reduce standard deviations in key metrics. This leads to faster processing times with improved detection quality, as shown by the finding that a 1:3 ratio [static:motion] results in the highest detection accuracy.

Abstract

Small object detection and tracking in videos remain critical yet underexplored challenges in computer vision, particularly for applications such as public safety, aerial surveillance, and autonomous driving. Existing benchmarks offer limited support due to limited numbers of small objects, constrained category diversity, and narrow scene coverage. To address these limitations, we introduce XS-VID, a large-scale video benchmark comprising 223K frames and 1.4M annotated bounding boxes across 374 video sequences spanning diverse scene types. XS-VID provides extensive coverage of small-object scales, particularly for extremely small (0 about12 squared pixels) and small (12 2 about20 squared pixels) objects, which collectively constitute over 55% of all annotations. For systematic evaluation, we establish three dedicated tracks: Detection, multiple object tracking (MOT), and single object tracking (SOT), and extensively test the existing state-of-the-art methods on each. The experimental results indicate that existing methods face significant challenges with XS-VID, mainly stemming from insufficient modeling of spatiotemporal features at small scales. To tackle these challenges, we propose a lightweight, high-precision detection framework dubbed YOLOFT. It enhances small-object feature representation and spatiotemporal integration while preserving high detection speed, thereby achieving improved accuracy and robustness on both the XS-VID and VisDrone benchmarks. Our dataset and code are publicly available at https://gjhhust.github.io/XS-VID/, providing a solid foundation for future research on small-object detection and tracking in videos.

Sources

Related papers