Unscented KalmanNet: Structure-Preserving Deep Learning with Calibrated Posterior Uncertainty under Incomplete Physics and Unknown Noise

arXiv:2608.04201 · cs.LG, cs.NA, eess.SP, math.NA · Submitted 2026-08-12 · Read on arXiv

Minhyeok Ko, Abdollah Shafieezadeh

The University of Texas at Tyler · The Ohio State University

cs.LG, cs.NA, eess.SP, math.NA

Submitted: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 68/100

The gist: (L1) the noise covariances Qt and Rt must be specified by the user and are typically fixed to constant values even when true statistics are time-varying; (L2) the analytical gain KtUKF = Ctxy St−1

Terminology

Summary

Summary

The paper introduces the Unscented KalmanNet (UKN), a hybrid recursive estimator that integrates data-driven corrections into the Unscented Kalman Filter (UKF) recursion to address nonlinear state estimation under incomplete physics and unknown, time-varying noise statistics. The UKF is chosen as the backbone because it propagates a deterministic set of sigma points through nonlinear functions, capturing the mean and covariance to higher order than the EKF's first-order linearization, and requires no Jacobians, allowing it to handle non-differentiable models. The paper identifies three limitations of the UKF: (L1) the noise covariances Qt and Rt must be specified by the user and are typically fixed to constant values even when true statistics are time-varying; (L2) the analytical gain KtUKF = Ctxy St−1 is biased under incomplete physics because the innovation contains a systematic component from the model residual delta that is not accounted for; and (L3) under incomplete physics, the analytical recursion does not generally satisfy accuracy and covariance calibration criteria, leading to a posterior that is inconsistent with the actual error.

The UKN framework has three elements: a differentiable UKF backbone, NoiseNet, and GainNet. NoiseNet addresses limitation (L1) by replacing user-specified covariances with time-varying estimates produced from the filter's own statistics. It predicts bounded lower-triangular multipliers AQt and ARt that act on the Cholesky factors of fixed baseline covariances: Qt = (LQ,base AQt)(LQ,base AQt)⊤ and Rt = (LR,base ARt)(LR,base ARt)⊤, which guarantees symmetric positive definiteness at every step. The diagonal entries are bounded by a log-domain saturation allowing each variance direction to scale between one tenth and fifty times the baseline, and the strictly lower-triangular entries are bounded by a direct tanh saturation. NoiseNet uses different features for predicting Qt and Rt: the Q-stage runs before prediction using information from the posterior at time t−1, while the R-stage runs before the update using the prior at time t. At initialization, the final layers of all covariance heads are zero and the stage-specific input adapters are identity mappings, so NoiseNet reproduces the nominal UKF at the start of training.

GainNet addresses limitation (L2) by augmenting the analytical UKF gain with an elementwise-bounded residual correction ΔKt, so that Kt = KtUKF + ΔKt. The residual gain is computed as ΔKt = st tanh(fK(hGt)), where st = cK‖KtUKF‖F is an adaptive amplitude scale proportional to the Frobenius norm of the analytical gain, and tanh is applied elementwise, ensuring each entry of the residual gain satisfies [ΔKt]ij ≤ st. Because the corrected gain does not generally satisfy Kt = Ctxy St−1, the posterior covariance is updated using the plug-in second-moment update: P̂tt = P̂tt−1 − Ctxy K⊤t − Kt(Ctxy)⊤ + Kt St K⊤t. This is equivalent to P̂tt = P̂ttUKF + ΔKt St ΔK⊤t, where the additional term is positive semidefinite, so the corrected covariance remains positive semidefinite. At initialization, the final layer of fK is zero, so ΔKt = 0 and the UKN behaves as the analytical UKF.

The training objective addresses limitation (L3) by combining four loss terms: the state-estimation MSE MSE = E[‖x̂ tt − xt‖22], a calibration loss cal = E[(1/2)(1/nx)∑nxi=1(c log(1 + e2t,i/(c[P̂tt]ii)) + log[P̂tt]ii)] that aligns marginal posterior variances with empirical squared errors, a measurement loss meas = E[(1/2)(log det St + (nudf + ny) log(1 + NISt/nudf))] that matches the innovation distribution to a multivariate Student-t target with nudf degrees of freedom, and a gain regularizer ΔK = E[‖ΔKt‖2F] that shrinks the residual correction toward zero. The weights on the calibration and measurement losses adapt during training based on filter consistency metrics: the calibration discrepancy g(e)cal = log(e2/P) and the measurement consistency g(e)meas = log(NIS/ny), both smoothed across epochs by an exponential moving average. Each weight is updated multiplicatively as w(e+1) = clip(w(e) exp[eta(ḡ(e) − tau)], wmin, wmax), with linear warmup ramps gating the calibration and measurement losses during early epochs.

The paper evaluates UKN against the analytical UKF, KalmanNet (KN), and Bayesian KalmanNet (BKN) on four examples. In the Lorenz attractor example, which isolates transition-model mismatch with correctly specified noise covariances, UKN achieves an overall RMSE of 1.344 compared to 2.672 for UKF, 1.445 for KN, and 1.446 for BKN. The improvement is concentrated in x3, the most weakly observed component, where RMSE decreases from 3.949 for UKF to 1.825 for UKN. The UKF remains far above the ANEES upper limit, indicating significant overconfidence, while UKN fluctuates around the nominal ANEES value of 3 and remains within or close to the consistency region [2.648, 3.377].

In the Duffing oscillator example, which combines time-varying parameter mismatch with process-noise underestimation and outlier-contaminated measurements, UKN achieves an overall RMSE of 0.093 compared to 0.133 for UKF, 0.119 for KN, and 0.112 for BKN, a 30.1% reduction relative to UKF. Per component, RMSE decreases from 0.114 to 0.075 (34.2%) for displacement and from 0.150 to 0.107 (28.7%) for velocity. The UKF remains far above the ANEES upper limit and increases further following the hidden stiffness-change window, while UKN produces the ANEES closest to the nominal value of 2, though it remains moderately above the upper limit of the consistency interval [1.715, 2.310].

In the maneuvering target tracking example, which combines partial observability, unmodeled maneuver commands, underestimated covariances, and heavy-tailed glint errors, UKN achieves the lowest overall normalized RMSE of 0.282 compared to 0.383 for UKF, 0.694 for KN, and 0.714 for BKN, a 26.4% reduction. The KN and BKN are less accurate than the analytical UKF, with their degradation most pronounced in the indirectly observed states. A clean/glint decomposition shows that KN and BKN achieve accurate position estimates on clean measurement steps (52.82 m and 52.84 m RMSE) but their glint-step RMSEs increase to approximately 355 m, nearly matching direct radar inversion (354.71 m), with glint-to-clean RMSE ratios of 6.72. In contrast, UKN achieves the lowest overall position RMSE of 76.23 m, the lowest glint-step RMSE of 92.90 m, and a glint sensitivity ratio of 1.25. The UKF exceeds the ANEES upper limit early and remains far above it, while UKN produces the ANEES closest to the nominal value of 5 and remains within or near the consistency region [4.542, 5.483] for much of the sequence.

In the UZH-FPV drone racing dataset example, which uses real flight data with uncharacterized model mismatch and noise misspecification, UKN is evaluated using leave-one-sequence-out cross-validation over 11 flight sequences. The filters use onboard IMU acceleration for state propagation and IMU-derived pseudo-velocity observations for measurement updates, with Leica ground-truth positions used only for training and evaluation. UKN achieves the lowest mean position RMSE of 0.4261 ± 0.0537 m, a 22.4% reduction relative to UKF, and the lowest mean velocity RMSE of 0.2678 ± 0.0316 m/s, a 34.3% reduction, with the smallest fold-to-fold standard deviation for both metrics. For covariance calibration, UKN achieves a mean dimension-normalized NEES of 1.82 ± 0.62, compared to 3.31 ± 0.78 for BKN and 20.34 ± 7.59 for UKF, and empirical coverage of 91.2 ± 3.7% and 96.3 ± 5.3% for the full-state 95% and 99% posterior ellipsoids, respectively, closest to nominal levels.

The paper concludes that UKN improves both state accuracy and posterior calibration over the analytical UKF and existing neural network-aided Kalman filters across all four examples. The KN produces no posterior covariance, so its uncertainty cannot be evaluated, while the BKN produces a covariance through sampling that is often only partially calibrated. The UKF provides an analytical covariance but becomes overconfident in all four examples with fixed nominal noise statistics. The paper notes that the present formulation is limited to state estimation with a componentwise calibration objective targeting marginal posterior variances, and suggests future work on multivariate calibration objectives involving the full posterior covariance and an augmented formulation for joint state and parameter estimation.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:

  • Improvement: Replace pure black-box neural estimators with a structure-preserving hybrid that retains an analytical Unscented Kalman Filter (UKF) backbone while adding two dedicated learned modules.

  • What it does: The system propagates uncertainty through sigma points (no Jacobians needed), handles non-differentiable models, and maintains an explicit posterior covariance at every step—unlike KalmanNet which provides no covariance and Bayesian KalmanNet which provides only approximate covariance via sampling.

  • NoiseNet: Learns time-varying process and measurement covariances as bounded multiplicative corrections to baseline Cholesky factors, guaranteeing positive definiteness at every step. It uses separate Q-stage and R-stage features with shared encoder-GRU parameters but separate recurrent states.

  • GainNet: Learns a bounded residual correction to the analytical UKF gain, with adaptive amplitude scaling proportional to the Frobenius norm of the nominal gain. This compensates for model-mismatch-induced bias without discarding the sigma-point moment structure.

  • What it does: The system independently addresses two distinct error sources—noise misspecification (L1) and model mismatch bias (L2)—rather than routing all corrections through a single gain update as KalmanNet does.

  • Improvement: Introduce a composite loss combining: (a) state MSE, (b) a soft-clamped componentwise covariance calibration loss that matches marginal posterior variances to empirical squared errors, (c) a heavy-tailed Student-t innovation consistency loss robust to outliers, and (d) a residual-gain Frobenius regularizer. Use adaptive weights driven by exponential moving averages of filter consistency metrics (calibration discrepancy and NIS mean).

  • What it does: The system jointly optimizes accuracy and uncertainty calibration, with weights that automatically balance the objectives during training—reducing manual hyperparameter tuning.

  • Improvement: The UKN's structure (sigma-point cross-covariance backbone + bounded gain correction + heavy-tailed loss) provides a balanced response to glint-contaminated measurements.

  • What it does: In the maneuvering target example, the UKN achieves a glint-to-clean RMSE ratio of 1.25 (vs. 6.72 for KalmanNet and Bayesian KalmanNet), meaning it remains robust to outliers while KalmanNet and Bayesian KalmanNet degrade to the level of direct radar inversion (355 m error) on contaminated steps.

  • Improvement: The system is trained with leave-one-sequence-out cross-validation and demonstrates consistent performance across flights.

  • What it does: On UZH-FPV real drone data, the UKN reduces mean position RMSE by 22.4% and velocity RMSE by 34.3% compared to UKF, with the lowest fold-to-fold variability among all methods (std of 0.0537 m for position vs. 0.0830 for UKF).

  1. Estimate states in nonlinear systems with incomplete physics—where the true dynamics differ from the nominal model—while maintaining calibrated posterior uncertainty.

  2. Handle non-differentiable, tabular, or black-box transition models—because the UKF backbone requires no Jacobians, unlike EKF-based methods.

  3. Adapt to time-varying, unknown noise statistics—including heavy-tailed disturbances, without manual tuning of Q and R matrices.

  4. Provide reliable uncertainty quantification—with dimension-normalized NEES closest to 1.0 (1.82 ± 0.62 on real data vs. 20.34 for UKF and 3.31 for BKN) and empirical coverage closest to nominal 95%/99% levels (91.2%/96.3% vs. 2.8%/4.4% for UKF).

  5. Track indirectly observed states—such as velocity, heading, and turn rate in radar tracking—by recovering the cross-covariance structure that other learned filters fail to capture.

  6. Remain robust to impulsive measurement outliers—maintaining accuracy while KalmanNet and Bayesian KalmanNet collapse to measurement inversion levels.

  7. Generalize to new scenarios—demonstrated by consistent performance across 11 held-out flight sequences with the lowest error and variability.

  8. Reduce estimation error by 26-50% compared to analytical UKF in synthetic benchmarks, while providing the most consistent posterior covariance among all covariance-reporting filters.

Related papers