CANMOT: Class-Aware Noise Modeling for Multi-Object Tracking in Autonomous Driving

arXiv:2606.03590 · cs.RO · Submitted 2026-06-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "CANMOT: Class-Aware Noise Modeling for Multi-Object Tracking in Autonomous Driving".

Dev: Kalman filter (KF)-based multi-object tracking (MOT) remains a strong baseline for autonomous driving due to its strong performance, computational efficiency and interpretability.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So, Dev, we're starting with this paper called "CANMOT: Class-Aware Noise Modeling for Multi-Object Tracking in Autonomous Driving." It looks like the main idea is tackling the common problem where everyone just uses global noise parameters for process and measurement noise in Kalman filter based multi-object tracking.

Rosa: The authors propose a new framework that introduces class specific diagonal process and measurement covariance matrices, and they even suggest expressing these in the object coordinate frame to keep those longitudinal and lateral motion differences accounted for.

Rosa: This seems really important because it directly addresses the issue of assuming identical uncertainty across different types of traffic participants.

Dev: That sounds like a big shift from just using global parameters everywhere, Rosa. If we're talking about computational efficiency and loop rate, I need to know how much overhead this class-aware modeling adds to the tracking loop for each object.

Dev: The paper suggests they are optimizing these noise matrices for each semantic class 'c', which implies some kind of per-object or per-class computation during the filtering step.

Taro: From an autonomy researcher's view, I'm interested in how this impacts the system when things get messy out there. If we can model uncertainty more accurately based on what we see, how does that help the system handle unexpected behaviors or weird traffic situations?

Taro: The paper is looking at improving tracking accuracy and reducing metrics like IDS and FRAG by showing class-aware modeling helps with those issues.

Rosa: Exactly, Taro. The authors claim that this class-aware approach improves AMOTA by 0 point 5pp over the shared local covariances while keeping the AMOTP equal to the previous method <ref:2606.03590#pg1>.

Rosa: It’s not just about a slight bump in accuracy; they are also showing it reduces identity switches by eight percent and fragmentation by nine point eight percent when using this class specific parameterization <ref:2606.03590#pg1>.

Dev: Eighty percent reduction in identity switches sounds like a massive win for system reliability, Rosa, but I wonder about the real-time constraints. How complex is this optimization process? They use a Bayesian optimization approach splitting it into subproblems per class, which suggests some non-trivial computation before or during the tracking.

Dev: If we're running at high frequency, any added latency in determining those class specific matrices could be a dealbreaker for real autonomous driving applications.

Paper summary: Taro: I'm focused on that latency issue, Dev. If the optimization takes too long to get those Qc and Rc matrices ready for every object at every time step, the whole benefit of better modeling disappears because we’re tracking things in the past.

Taro: But if this framework allows us to be more precise about what we know about an object's noise characteristics, maybe it lets us filter out false positives or track occlusions more robustly when things get chaotic.

Rosa: That brings up a crucial point from page one: they are expressing the noise in the object coordinate frame to preserve longitudinal-lateral anisotropy <ref:2606.03590#pg1>. This is because longitudinal and lateral motion have different variability due to physical constraints, and global modeling averages that out into something isotropic.

Rosa: So, by aligning the noise with the object's local axes, they are trying to capture that directional difference better.

Dev: That sounds physically sound in theory, Rosa, but implementing a rotation matrix T for every object's orientation at every time step adds complexity to our state propagation routine. I need to know if this is something we can handle within the required loop rate for safety-critical components.

Dev: Also, the paper mentions that accurate uncertainty modeling is critical because downstream planning modules explicitly incorporate state covariance <ref:2606.03590#pg1>.

Taro: If the planner gets a better sense of what its own tracking uncertainty looks like, it can make much safer decisions when the environment misbehaves and we have to react quickly.

Taro: But I also see what they flag as a limitation: they didn't systematically evaluate the consistency of this estimated uncertainty with the true error using a chi-squared test <ref:2606.03590#pg1>.

Rosa: That inconsistency in uncertainty estimates is something we have to address if we want to use this for safety. The paper notes that standard baselines exhibit severe overconfidence, and while CANMOT provides better calibration with an ANEES of twenty point four when optimizing Q per class <ref:2606.03590#pg1>, they still haven't achieved statistical consistency in the sense of a chi squared test <ref:2606.03590#pg1>.

Dev: So, even with these performance gains in AMOTA and IDS reduction, the underlying uncertainty estimation might not be reliable enough for mission-critical tasks yet? That suggests we still have some fundamental calibration issues to solve before this becomes standard equipment <ref:2606.03590#pg1>.

Paper summary: Taro: If the uncertainty isn't consistent, then relying on these tracking results for high-level decision making in dynamic environments might still carry a hidden risk. That means the system needs more than just better noise modeling; it needs better error estimation itself <ref:2606.03590#pg1>.

Rosa: So, to wrap up this summary of "CANMOT: Class-Aware Noise Modeling for Multi-Object Tracking in Autonomous Driving," the core contribution is moving away from globally shared noise parameters to a class aware and object aligned modeling framework <ref:2606.03590#pg0>.

Rosa: It claims that by using class specific diagonal process and measurement covariance matrices, optionally expressed in the object coordinate frame, it improves tracking accuracy and reduces identity switches significantly compared to shared noise parameters <ref:2606.03590#pg1>.

Dev: The implications for loop rate are still a concern, especially with the Bayesian optimization used to tune those covariance matrices <ref:2606.03590#pg1>. We have to figure out if we can make that process fast enough for a real autonomous vehicle operating under tight temporal constraints.

Taro: I think the bigger implication is that for autonomy to truly be robust, the tracking component needs to provide uncertainty estimates that are statistically consistent with what actually happens, not just better looking numbers on paper <ref:2606.03590#pg1>.

Rosa: That’s a fair point, Taro. And as we look at the authors Timo Osterburg, Stefan Schütte, and Torsten Bertram from TU Dortmund University who put this into practice on the nuScenes benchmark <ref:2606.03590#pg1>, it shows that even with these improvements in tracking metrics like IDS and FRAG, full statistical consistency remains unattained <ref:2606.03590#pg1>.

Dev: I agree; if the uncertainty estimates aren't reliable for downstream planning modules, we haven't really solved the problem yet; we’ve just made the filtering step look better in isolation <ref:2606.03590#pg1>.

Taro: It really points toward future work needing to focus on calibration-aware objectives during covariance optimization, which is exactly what the paper hints at as a necessary next step <ref:2606.03590#pg1>.

Rosa: So, while CANMOT demonstrates that expressing noise in the object coordinate frame preserves longitudinal-lateral anisotropy and consistently reduces identity switches while maintaining comparable AMOTA to global formulations <ref:2606.03590#pg1>, it leaves the door open for more rigorous uncertainty analysis <ref:2606.03590#pg1>.

Dev: That’s the summary of what the paper delivers right now; a solid performance boost with some known limitations regarding statistical rigor and computational load <ref:2606.03590#pg1>.

Taro: It suggests that for real world deployment, we need to push past these current tracking metrics and focus on making those uncertainty estimates truly reliable for safety-critical decisions <ref:2606.03590#pg1>.

Conclusion: Rosa: So, we've seen how CANMOT tackles noise modeling in multi-object tracking using class-aware and object-aligned covariance matrices to improve performance over shared parameters.

Dev: Yeah, that class awareness is interesting because it means the noise assumptions are no longer a one size fits all across different types of vehicles or pedestrians.

Taro: I'm thinking about how this improved accuracy translates when things go wrong in the real world, like unexpected maneuvers from erratic drivers.

Rosa: Exactly, and the authors show that by aligning these noise parameters with the object's local frame, they keep those longitudinal and lateral motion differences captured better than a global frame does.

Dev: From my end, I’m still focused on whether that rotation matrix transformation adds too much computational load to maintain a high enough loop rate for reliable control decisions.

Taro: If it can handle those rapid changes without lag, imagine how much more robust our autonomy becomes when encountering unpredictable traffic situations.

Rosa: And the authors report that this class-aware modeling actually reduces identity switches and fragmentation metrics by quite a bit, which is a big deal for tracking consistency.

Dev: Reducing those switches is definitely something I'd like to hear about in terms of system reliability, but we still need to look closely at the uncertainty estimates themselves.

Taro: That's where I think it gets interesting; even with better performance metrics, the paper points out that full statistical consistency in the error estimation hasn't been fully achieved yet.

Rosa: That means while tracking accuracy improves and we see better results on benchmarks like nuScenes, we still have a gap to bridge regarding how reliably the system knows its own uncertainty.

Dev: So, it’s like we’ve made the filter smarter about what it thinks is happening, but the underlying confidence score isn't perfectly aligned with reality yet.

Taro: It suggests that for real-world deployment, we need to keep pushing past these tracking metrics and focus on making those uncertainty estimates truly consistent for safety-critical decisions.

Institute of Control Theory and Systems Engineering, TU Dortmund University

cs.RO

Submitted: 2026-06-02

Updated: 2026-10-03

Comments: accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

Code: https://github.com/rst-tu-dortmund/learned-3d-nms

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 71/100

The gist: Kalman filter (KF)-based multi-object tracking (MOT) remains a strong baseline for autonomous driving due to its strong performance, computational efficiency and interpretability.

Key concepts

Class-Aware Noise Modeling
Instead of assuming all objects share the same noise characteristics globally, this approach treats each vehicle type (class) separately. It allows the system to use unique process and measurement noise settings tailored specifically to that object's class, which is more realistic for diverse traffic scenarios.
Object-Aligned Coordinate Frame
The framework transforms global noise parameters into the object's local coordinate system. This alignment is crucial because longitudinal (forward/backward) and lateral (side-to-side) motion have different error behaviors; aligning the noise with these axes helps preserve this physical difference in variability.
Average Multi Object Tracking Accuracy (AMOTA)
This is the primary metric used to evaluate how well the tracking system performs overall. The goal of CANMOT is to maximize this accuracy by optimizing class-specific noise matrices, aiming for better tracking results than methods using uniform or shared noise settings.

Terminology

Summary

Kalman filter (KF)-based multi-object tracking (MOT) remains a strong baseline for autonomous driving due to its strong performance, computational efficiency and interpretability.

How it works

The proposed CANMOT framework revisits the assumption that process noise and measurement noise covariances are defined globally and shared across object classes. Instead, it introduces a class-aware and object-aligned noise modeling framework for KF-based 3D MOT. This involves introducing class-specific diagonal process and measurement covariance matrices to account for heterogeneity in traffic participants. Furthermore, the framework expresses these noise parameters in the object coordinate frame to preserve longitudinal-lateral anisotropy, which is crucial because longitudinal and lateral motion exhibit different variability due to physical constraints.

Key Research Questions Addressed

The work addresses three primary research questions:

R1: Does modeling noise in an object-aligned coordinate frame improve tracking performance compared to a global coordinate frame?

R2: Does class-aware noise modeling improve tracking performance compared to shared noise parameters?

R3: How do the proposed design choices affect the consistency of the estimated uncertainty with the true error?

Methodology and Noise Modeling

The core of CANMOT involves optimizing noise covariance matrices, denoted as "Qc and Rc," for each semantic class, where 'c' is the object's class. The paper proposes two main strategies for modeling this noise:

  1. Object-aligned Noise: This strategy rotates the global noise covariances to align their principal axes with the object’s local frame axes, using a rotation matrix T (θ). This is achieved by formulating:

Q global c = T Q(θ)Q obj c T T (1)

R global c = T R(θ)R obj c T T (2)

The transformation matrix T rotates planar position and velocity components according to the yaw angle θ, while leaving other dimensions unchanged.

  1. Deriving Covariance Matrices Directly from Data: The authors compute the sample covariance of the measurement and process noise for a constant velocity model of different object classes in both the local object and global map coordinate frames on the nuScenes dataset to derive initial parameters for optimization.

Optimization and Evaluation

The objective function used to optimize these matrices is maximizing the Average Multi Object Tracking Accuracy (AMOTA). Due to computational costs, a Bayesian optimization approach is employed, splitting the problem into one subproblem per class where parameters of Rc and Qc are optimized independently. The experiments compare CANMOT against baselines like shared† noise parameters and other state-of-the-art methods such as Poly-MOT [9] and MCTrack [12].

Results and Consistency Analysis

The systematic experiments on the nuScenes benchmark show that class-aware modeling yields significant improvements:

CANMOT† improves AMOTA by 0.5pp over the shared local covariances with an equal AMOTP.

Class-specific covariance parameterization is shown to improve tracking accuracy and reduce metrics like IDS (Identity Switches) by 8.0% and FRAG (Fragmentation) by 9.8%.

Regarding uncertainty consistency, the study analyzes the Average Normalized Estimation Error Squared (ANEES). While standard baselines exhibit severe overconfidence, CANMOT provides better calibration, with an ANEES of 20.4 when optimizing Q per class while keeping R from the sample covariance in the local frame. However, Table II indicates that none of the evaluated methods provides consistent uncertainty estimates when tested against a χ2-test for filter consistency, suggesting that current tracking methods do not provide reliable uncertainty estimates to be used in downstream tasks.

Conclusion and Outlook

CANMOT demonstrates that expressing noise in the object coordinate frame preserves longitudinal-lateral anisotropy and consistently reduces identity switches, while maintaining comparable AMOTA to global formulations. The findings highlight that while optimizing process noise per class substantially improves calibration, full statistical consistency remains unattained, suggesting future work should focus on calibration-aware objectives during covariance optimization.

The gist: Class-aware and object-aligned noise modeling for KF-based 3D MOT improves tracking performance and substantially reduces identity switches compared to state-of-the-art (SotA). The results reveal severe overconfidence in standard KF-based MOT baselines. While the proposed formulation improves calibration without modifying the underlying filtering framework, it still exhibits substantial inconsistency, highlighting the need for further research in this area.

Improvements for AI systems

Here are specific improvements to AI systems based on the CANMOT framework:

  1. Improve tracking robustness in complex, heterogeneous traffic environments by implementing a class-aware and object-aligned noise modeling framework within Kalman Filter (KF)-based Multi-Object Tracking (MOT) systems.

  2. Enable 3D MOT pipelines to accurately model longitudinal-lateral anisotropy by expressing process noise covariance matrices as an object-aligned variant, ensuring that tracking uncertainty is directionally accurate relative to the vehicle's orientation rather than being isotropic in global map coordinates.

  3. Significantly reduce identity switches (ID switches) and track fragmentation by optimizing class-specific diagonal process and measurement noise parameters using gradient-free Bayesian optimization, leading to superior performance compared to shared or globally defined noise models.

  4. Enhance the calibration of state estimation by providing more consistent uncertainty estimates, as shown by a lower Average Normalized Estimation Error Squared (ANEES), which indicates that the filter's estimated error more reliably matches its true error distribution.

  5. Improve downstream planning and decision-making modules by supplying state covariance matrices with higher fidelity, allowing planners to incorporate realistic spatial uncertainty into their risk assessment and maneuver selection.

  6. Develop more reliable object localization for smaller or elongated objects (e.g., pedestrians, trucks) by utilizing class-specific noise parameters that account for the unique perceptual characteristics and observation ambiguities inherent to different object shapes and sizes on datasets like nuScenes.

Sources

Related papers