Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking

arXiv:2411.06197 · cs.CV · Submitted 2024-11-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Tracking by Detection and Query".

Jane: Multi-object tracking (MOT) is a pivotal task in applications involving scene understanding

1: , multi-drone tracking

2: , etc., requiring recognition, localization, and consistent identification of objects over time.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: We've been talking about "Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking" now, really looking at what this paper achieves in terms of its overall goal. To sum up, the authors propose TBDQ-Net as a way to merge tracking by detection and tracking by query paradigms.

Jane: They accomplish this by building a framework that uses a frozen detector combined with a lightweight associator to achieve efficiency without sacrificing the end-to-end capability of query methods. It’s about finding that middle ground between speed and robustness in these complex tasks.

Lu: The authors are showing that you can integrate dual-stream communication, through those BII and CPA modules, to handle both new detections and existing tracks simultaneously in a unified way. That structural integration is what they highlight as key to advancing the synergy between TBD and TBQ.

Meng: From an engineering view, it means we’re looking at architectures that are designed specifically for efficiency in deployment while still retaining the deep relational modeling capabilities needed for good tracking results. It’s about making heavy models practical.

Lalam: The paper shows how to use query-based modeling where the detection queries and track queries interact via BII and CPA, which allows for a unified framework that handles both target trajectories and emerging objects effectively.

Tom: So, the big picture here is that this framework offers a unified structure to handle tracking by detection versus tracking by query tasks, moving us toward more cohesive end-to-end solutions in multi-object tracking.

Jane: It’s a concrete proposal for how to design these systems that need to be both accurate and computationally feasible when you're dealing with dense visual information.

Lu: The title itself points to the framework being about achieving detection and query integration efficiently, which is the main innovation they are pushing forward in this research area.

Conclusion: Tom: So, we're wrapping up this look at TBDQ-Net, which is "Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking."

Jane: It boils down to taking two different ways of tracking—detection based and query based—and gluing them together so they work well.

Lu: The authors are pushing the idea that you don't have to pick just one approach; you can use both streams, detection queries for new things and track queries for what you already know.

Meng: From an engineering standpoint, this is about making sure the system doesn't get bogged down by too much computation while still getting accurate results.

Lalam: The core idea is using a frozen detector and a lightweight associator to keep the efficiency up while maintaining that robust end-to-end tracking ability.

Tom: It seems like they’ve managed to hit that sweet spot between speed and accuracy in some tough tracking situations.

Jane: They show it's actually outperforming existing detection-based methods by a solid margin on challenging benchmarks, specifically DanceTrack.

Lu: And they also did really well when compared to the query-based approaches on the crowded MOT20 benchmark.

Meng: The practical impact is that if we can get this kind of efficiency with good results, it opens up more possibilities for real-time tracking in complex environments, like autonomous vehicles or drones.

Lalam: If we look at how this structure works, it means the way we model object movement and new appearances is fundamentally improved by integrating these two query types.

Tom: It’s a really neat structural improvement to the whole tracking pipeline.

Jane: So, what does this mean for people just listening to the radio? It means tracking systems could get faster and handle more crowded scenes without needing massive processing power.

Lu: And it suggests that future work might look at using smaller AI models to speed up those detection queries even further.

Shukun Jiaa, Shiyu Hu, Yichao Cao, Feng Yang, Xin Lua, Xiaobo Lua

School of Automation, Southeast University · Key Laboratory of Measurement and Control of Complex Systems of Engineering, Ministry of Education, Nanjing Research Center for Advanced Computing and Information Technology (implied by context/location)

cs.CV

Submitted: 2024-11-09

Updated: 2026-03-29

Code: https://github.com/FaithFlow/TBDQ-Net

Importance score: 85/100

The gist: Multi-object tracking (MOT) is a pivotal task in applications involving scene understanding [1], multi-drone tracking [2], etc., requiring recognition, localization, and consistent identification of

Key concepts

Tracking-by-Detection (TBD)
This approach relies on using a pre-trained detector to find all potential new objects in every frame. It then tries to associate these newly detected objects with existing tracks from the previous frame, focusing on detection as the primary source for new object information.
Tracking-by-Query (TBQ)
This method models target trajectories by propagating track queries forward in time. Instead of relying solely on detections, it uses these queries to predict where objects should be in future frames, which helps maintain temporal consistency and capture emerging targets.
Basic Information Interaction (BII) module
This module facilitates communication between detection and track queries through dual streams. The BII-D stream lets detection queries see existing tracks to avoid conflicts, while the BII-T stream updates track queries with current detections for immediate information, ensuring both new and old objects are managed effectively.
Content-Position Alignment (CPA) module
This component resolves the semantic gap by dynamically aligning a query's content (what it represents semantically) with its spatial location in the image. It uses modulated cross-attention to synchronize these two aspects, ensuring that queries accurately reflect both what they are and where they are located spatially.

Terminology

Summary

Multi-object tracking (MOT) is a pivotal task in applications involving scene understanding [1], multi-drone tracking [2], etc., requiring recognition, localization, and consistent identification of objects over time. This work proposes the Tracking-by-Detection-and-Query framework, TBDQ-Net, to advance the synergy between tracking-by-detection (TBD) and tracking-by-query (TBQ) paradigms by integrating a frozen detector with a lightweight associator to achieve intrinsic efficiency while maintaining robust end-toend tracking capability.

The gist: TBDQNet achieves a favorable efficiency-accuracy trade-off in challenging scenarios, specifically outperforming leading TBD methods by 6.0 IDF1 points on DanceTrack and achieving the best performance among TBQ methods in the crowded MOT20 benchmark <ref:2411.06197#pg4>.

How it works

The framework follows an end-toend query-based tracking paradigm, modeling target trajectories through track queries (T), while capturing emerging objects through detection queries (D). At the heart of this framework, the Basic Information Interaction (BII) module facilitates dual-stream communication between detection and track queries. Specifically, the BII-D stream allows detection queries to perceive existing tracks, suppressing potential conflicts, while the BII-T stream updates track queries with current detection priors for immediate information update, while integrating historical context to ensure temporal stability against disturbances and occlusions.

The framework models target trajectories through track queries (T), while capturing emerging objects through detection queries (D). For each frame, a frozen detector extracts a raw candidate pool D. These candidates are transformed into detection queries Dtq, representing potential new objects, and track queries Tt-1q propagate from the last moment t-1, linking to existing objects.

Core Modules

The associator is designed to be detector-agnostic, facilitating seamless integration with both CNN-based and Transformer-based architectures. The internal mechanism of the associator primarily consists of two key components:

  1. The Basic Information Interaction (BII) module, which facilitates dual-stream refinement through a scaled dot-product attention mechanism. This module implements BII-D and BII-T streams, where detection queries are updated in the BII-D stream and track queries are updated in the BII-T path.

  2. The Content-Position Alignment (CPA) module, which reconciles the semantic gap by dynamically synchronizing the query’s spatial information with its semantic representation. This is implemented via a modulated cross-attention mechanism that leverages global image features F and positional encodings P.

Training and Inference

In the training stage, a bipartite matching strategy is employed for label assignment, where the assignment set is updated recursively to ensure temporal consistency. The total loss function is defined as Ltotal = L(ˆymainσ, y) + L(ˆyauxσ, y). During inference, detection queries with prediction scores above τn initiate new tracks, and existing tracks are updated if their corresponding track query predictions exceed τn.

Experimental Validation

Extensive evaluations on DanceTrack [13], SportsMOT [14], and MOT20 [15] demonstrate that TBDQ-Net achieves a favorable efficiency-accuracy trade-off in challenging scenarios. Specifically, TBDQNet outperforms leading TBD methods by 6.0 IDF1 points on DanceTrack and achieves the best performance among TBQ methods in the crowded MOT20 benchmark. Compared to MOTRv2, TBDQ-Net reduces trainable parameters by approximately 80% while accelerating practical inference by 37.5%.

Component Analysis

An ablation study confirms the pivotal roles of both modules, showing that optimal performance is attained when both BII and CPA modules are present. The BII module's design, particularly the use of noisy queries Nq in the BII-D stream, effectively alleviates task conflicts by disrupting detection queries that overlap with existing tracks. The CPA module is crucial because it bridges the gap between the updated content and its corresponding spatial state.

Limitations and Future Work

The frozen-detector strategy restricts further refinement of detection features, making TBDQ-Net influenced by pre-trained detection performance. Furthermore, as a query-based framework, its optimal performance exhibits a dependency on data-rich environments, suggesting that for scenarios with extremely limited training data or annotation protocols requiring persistent tracking of completely invisible targets, heuristic-based methods might remain advantageous. Future work could explore lightweight adapter mechanisms to alleviate detection bottlenecks and extensions of the TBDQ paradigm to multi-modal tracking.

REFERENCES

[1] Y. Li, G. Yang, Z. Su, S. Li, J. Yang, L. He, Indoor scene multi-object tracking based on region search and memory buffer pool, Pattern Recognition 165 (2025) 111623.

[2] Z. Liu, Y. Shang, T. Li, G. Chen, Y. Wang, Q. Hu, P. Zhu, Robust multi-drone multi-target tracking to resolve target occlusion: A benchmark, IEEE Transactions on Multimedia 25 (2023) 1462–1476.

[3] A. Bewley, Z. Ge, L. Ott, F. Ramos, B. Upcroft, Simple online and realtime tracking, in: 2016 IEEE International Conference on Image Processing (ICIP), IEEE, 2016, pp. 3464–3468.

[4] Y. Zhang, P. Sun, Y. Jiang, D. Yu, F. Weng, Z. Yuan, P. Luo, W. Liu, X. Wang, ByteTrack: Multi-object tracking by associating every detection box, in: European Conference on Computer Vision Springer 2022 pp 1–21.

[5] T. Mandel, M. Jimenez, E. Risley, T. Nammoto, R. Williams, M. Panoff, M. Ballesteros, B. Suarez, Detection confidence driven multi-object tracking to recover reliable tracks from unreliable detections Pattern Recognition 135 (2023) 109107.

[6] F. Zeng, B. Dong, Y. Zhang, T. Wang, X. Zhang, Y. Wei, MOTR: End-to-end multiple-object tracking with transformer in: European Conference on Computer Vision Springer 2022 pp 659–675.

[7] R. Gao, L. Wang, MeMOTR: Long-term memory-augmented transformer for multi-object tracking, in Proceedings of the IEEE/CVF International Conference on Computer Vision 2023 pp 9901–9910.

[8] F. yan, W. Luo, Y. Zhong, Y. Gan, L. Ma, CO-MOT: Boosting end-to-end transformer-based multi-object tracking via coopetition label assignment and shadow sets in: International Conference on Learning Representations 2025.

[9] X. Zhou, V. Koltun, P. Krähenbühl, Tracking objects as points, in European Conference on Computer Vision Springer 2020 pp 474–490.

[10] Z. Ge, S. Liu, F. Wang, Z. Li, J. Sun, YOLOX: Exceeding yolo series in 2021 arXiv preprint arXiv:2107.08430 (2021).

[11] R. Gao, J. Qi, L. Wang, Multiple object tracking as id prediction, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025 pp 27883–27893.

[12] Y. Zhang, T. Wang, X. Zhang, MOTRv2: Bootstrapping end-toend multi-object tracking by pretrained object detectors in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2023 pp 22056–22065.

[13] P. Sun, J. Cao, Y. Jiang, Z. Yuan, S. Bai, K. Kitani, P. Luo, DanceTrack: Multiobject tracking in uniform appearance and diverse motion in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2022 pp 20993–21002.

[14]

Improvements for AI systems

  1. Bold Header: Dual-Stream Information Interaction (BII) Module

TBDQ-Net integrates a dual-stream update where "the detection stream (BII-D) incorporates track information to suppress potential conflicts, while the track stream (BII-T) integrates current observations and history memory to maintain temporal stability against disturbances and occlusions. This allows the system to actively mitigate task conflicts by using existing tracks to filter new detection queries (BII-D) and ensures continuity during occlusions using historical context (BII-T").

  1. Bold Header: Content-Position Alignment (CPA) Module

The CPA module simultaneously refines both content and positional components, bridging the semantic-spatial gap and delivering well-aligned representations for association decoding. This mechanism addresses the spatiotemporal misalignment by dynamically synchronizing the query’s spatial information with its semantic representation using a learnable transformation, which is implemented via a modulated cross-attention mechanism.

  1. Bold Header: Efficient Architectural Synergy

TBDQ-Net achieves an efficient end-to-end tracking framework by integrating a frozen detector with a lightweight associator, which contrasts with TBQ methods that suffer from high training costs and slow inference due to the tight coupling of detection and association. This design reduces trainable parameters by approximately 80% compared to heavy architectures like MOTR.

  1. Bold Header: Robust Query Filtering Strategy

The framework utilizes a confidence-based partitioning strategy ("Strategy C1: Dq = q ∈ Dq s(q) > τq) to filter out low-quality candidates, retaining only high-confidence queries Dq before interaction. Furthermore, it constructs noisy queries Nq via Strategy C2: TopK(… < τq), Ntr, which are used to learn how to assign higher attention weights to these noisy queries when a detection query is similar to a track query, thereby disrupting its feature aggregation and effectively alleviating inherent conflicts."

  1. Bold Header: End-to-End Trajectory Modeling

By employing the learnable query propagation enables the model to adaptively handle complex motion patterns and occlusions, TBDQ-Net models trajectories through the continuous optimization of track queries, which allows it to achieve a balance between unparalleled efficiency and robust tracking performance.

Sources

Related papers