LOCUS-DT: Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "LOCUS-DT: Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins".
Jane: The paper was written by Haozhe Lei, Roberto Bomfin, Marwa Chafii and Sundeep Rangan from New York University and New York University Abu Dhabi.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, welcome back to the show, everybody. Today we're digging into a fresh arXiv paper called "LOCUS-DT: Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins." Jane, I gotta say, that title is a mouthful, but the idea behind it is genuinely exciting.
Jane: It really is, Tom. And let's break down that acronym right away because it tells you everything. LOCUS-DT stands for Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins. So the core problem is figuring out where a transmitter is located indoors, but instead of just giving one answer, it gives you a whole map of probabilities.
Tom: Right, and that's the key shift here. Most classical localization systems give you a single point estimate, like "the transmitter is here." But indoors, with walls and reflections, the signal bounces around so much that there might be several places that look equally plausible. A single point just doesn't capture that ambiguity.
Jane: Exactly. And that's where the digital twin comes in. The paper assumes you know the room layout and where the receiver is. So you build a virtual copy of that room, and for every possible transmitter location, you simulate what the multipath profile would look like using ray tracing. Then you compare those simulated profiles against what the receiver actually measured.
Tom: And that comparison is what they call the scoring function. It's a learned neural network that figures out how well a candidate location's simulated multipath matches the real observation. The cool part is that they train this scoring function across a bunch of different random room layouts, so it generalizes to rooms it's never seen before.
Jane: That generalization piece is huge. I mean, think about it — most data-driven localization methods are tied to the specific environment they trained on. If you move to a different building, the model breaks. LOCUS-DT is trying to break that dependence by conditioning on the layout itself through the digital twin.
Tom: And the results back that up. They tested it on completely random layouts that weren't in the training set, and it still produced sharp, multimodal posteriors that matched the true transmitter location. The Gaussian baselines they compared against just couldn't capture that structure.
Jane: So the takeaway for our listeners is this: LOCUS-DT isn't just another fingerprinting method. It's a way of doing probabilistic reasoning about location that respects the physics of the environment. And that's what makes it so compelling.
Tom: I'm hooked. Let's keep going and dig into the actual methodology in the next segment, because there's some clever engineering in how they build that scoring function.
Summary: Tom: So we're back, and we're still on "LOCUS-DT: Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins." Jane, you mentioned the scoring function earlier. Let's get into how it actually works, because that's where the magic is.
Jane: Absolutely. So the digital twin generates a set of multipath paths for each candidate transmitter location. Each path is described by an angle of arrival, a delay, a complex gain, and a signal-to-noise ratio. The receiver also extracts its own multipath profile from the actual IQ samples it captures. The scoring function then compares these two sets of paths.
Tom: And here's the clever part — they don't compare all paths equally. They keep only the top K strongest paths, and in their experiments, K equals six. That's a deliberate choice because the strongest paths are the most reliable. Weaker paths are more likely to be noise or estimation errors.
Jane: Right. And they also discard the delay information entirely because the system is narrowband and there's no synchronization. So each path is represented by just its angle of arrival and its SNR, converted to a scaled value. That gives them a compact feature vector for each candidate location.
Tom: So the scoring function takes those features and runs them through a two-layer neural network. It's a tiny network, which is great because it trains fast. They use a sampled cross-entropy loss, which essentially says: the true transmitter location should get a higher score than all the other candidates.
Jane: And that loss function is really the heart of the training. They replace the partition function, which would normally require a difficult integral, with a sum over the candidate locations they've sampled. So the model learns to assign high probability to the true location and low probability everywhere else.
Tom: One thing I really appreciate is how they handle the observation estimation. They use an iterative Levenberg-Marquardt algorithm to extract the multipath components from the received signal. That's a standard maximum-likelihood approach, and it works well at high SNR, which is what they assume in their simulations.
Jane: And the whole thing is validated using Sionna, which is NVIDIA's ray-tracing backend. They generate random indoor environments with obstacles, simulate the true channel, and then test whether LOCUS-DT can recover the posterior distribution over transmitter locations. The results are impressive — it captures sharp, multimodal structure that the Gaussian baselines completely miss.
Tom: So the summary is: digital twin generates candidate multipath profiles, a learned scorer compares them to the observed profile, and the result is a full posterior distribution over location. It's elegant, it's physics-aware, and it generalizes across environments.
Jane: And it's a big step beyond just giving a point estimate. For applications like robotic navigation or search and rescue, knowing where the transmitter probably isn't is just as important as knowing where it probably is.
Tom: Great point. Now let's talk about what this means for real-world systems in the next segment.
Improvements: Tom: We're back on "LOCUS-DT: Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins." Jane, we've covered the basics and the methodology. Now I want to get into what this paper actually improves over existing approaches, and why that matters.
Jane: So the biggest improvement is the shift from point estimates to full posterior inference. Classical localization gives you one answer, maybe with an error bar. LOCUS-DT gives you a probability distribution over the entire space. And in indoor environments, that distribution is often multimodal — meaning there are several distinct locations that could explain the measurements.
Tom: And that's not just a theoretical nicety. Think about a robot trying to navigate through a building. If it only knows the most likely location, it might head straight into a wall because the second-most-likely location was actually correct. A full posterior lets the robot reason about all possibilities and plan accordingly.
Jane: Exactly. And the paper compares against three baselines: a Gaussian posterior in Cartesian coordinates, a Gaussian in polar coordinates, and a Gaussian mixture model. All three fail to capture the sharp, layout-dependent structure that LOCUS-DT produces. The Gaussian models are too smooth, and the mixture model, while better, still can't match the precision of candidate-wise digital twin matching.
Tom: And the numbers back that up. On the harder evaluation sets, LOCUS-DT achieves a much lower adjusted loss and much higher probability mass on the true location. The Gaussian baselines are essentially no better than random guessing in some cases, which really shows how important the layout conditioning is.
Jane: Another improvement is the generalization. Because the scoring function is trained over an ensemble of random environments, it learns to compare multipath profiles in a way that's not tied to any specific building. That's a huge deal for deployment, because you don't want to retrain your model every time you move to a new floor.
Tom: And they also build in robustness to errors. The training process can incorporate errors in the digital twin model and errors in the channel estimation, so the scoring function learns to be forgiving of small mismatches. That's critical for real-world deployment, where your digital twin is never perfect.
Jane: Right. And the architecture is intentionally simple. A two-layer neural network with just a few dozen features. That means it's fast to train and fast to evaluate, which matters if you're running this on a mobile robot with limited compute.
Tom: So the improvements are: full posterior instead of point estimate, layout-conditioned scoring via digital twins, generalization across environments, and robustness to model errors. That's a pretty compelling package.
Jane: It really is. And I think the implications go beyond just localization. This framework could be applied to any sensing problem where you have a digital twin of the environment and you want to infer something about a source.
Tom: Let's wrap up with our final thoughts in the next segment.
Conclusion: Tom: And we're back for the final segment on "LOCUS-DT: Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins." Jane, let's pull it all together for our listeners.
Jane: So the big picture is this: indoor localization is hard because multipath propagation creates ambiguity. LOCUS-DT embraces that ambiguity by producing a full posterior distribution over transmitter locations, rather than forcing a single point estimate. It does this by using a digital twin of the environment to simulate what the multipath profile would look like from every candidate location, and then comparing those simulations to the actual measurement.
Tom: And the key innovation is the learned scoring function that makes that comparison robust to errors in both the digital twin and the channel estimation. It's trained across many environments, so it generalizes to layouts it's never seen. The experiments show it dramatically outperforms Gaussian and Gaussian-mixture baselines.
Jane: The implications are pretty broad. For robotic navigation, search and rescue, and even integrated sensing and communication systems, having a reliable posterior over location is much more useful than a single point. It lets downstream systems reason about uncertainty and make safer decisions.
Tom: And the authors are already thinking about future work. They mention evaluating under greater digital twin mismatch, and using rooms reconstructed with SLAM — so real-world measured environments rather than synthetic ones. That's the natural next step toward deployment.
Jane: It really is. And I think the framework itself is elegant enough that it could inspire similar approaches in other sensing domains. If you can build a digital twin and compare simulated observations to real ones, you can do posterior inference about whatever you're trying to sense.
Tom: Alright, that's a wrap on LOCUS-DT. Thanks to everyone who tuned in. We'll be back with another paper next time, so stay curious, everybody.
Jane: And remember — sometimes the most useful answer isn't a single location, it's a map of possibilities. See you next time.
Haozhe Lei, Roberto Bomfin, Marwa Chafii, Sundeep Rangan
New York University · New York University Abu Dhabi
eess.SP, cs.LG, cs.RO
Submitted: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 68/100
Key concepts
- LOCUS-DT
- Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins. This method aims to find the transmitter's location indoors by generating a map of probabilities rather than one answer. It uses a digital twin of the room layout to simulate signal paths from every possible location and compares these simulations to actual receiver measurements.
- Digital Twin
- A virtual copy of the physical room layout used in the paper. This twin allows researchers to simulate how a signal would bounce around in different locations within that specific environment using ray tracing, enabling testing against various potential transmitter spots.
- Scoring Function
- A learned neural network that compares simulated multipath profiles from candidate locations against the actual observed multipath profile. It is trained to assign higher scores to the location whose simulated paths best match what was actually measured.
- Posterior Distribution
- The output of LOCUS-DT, which is a map of probabilities over all possible transmitter locations. This contrasts with classical localization that only provides a single point estimate, allowing downstream systems to reason about uncertainty and plan safer actions.
Terminology
Summary
Summary
This paper proposes LOCUS-DT (Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins), a framework that treats snapshot indoor localization as posterior inference over the transmitter location given measurements, a known room layout, and receiver pose. The method assumes the room layout and receiver pose are known, and uses a ray-tracing-based digital twin (DT) to generate synthetic multipath profiles for candidate transmitter locations, which are then compared against the measured channel profile to score each candidate location and obtain a posterior distribution.
The main contributions are threefold. First, the paper formulates single-snapshot indoor localization as posterior inference over the transmitter location, incorporating the layout implicitly through the ray-tracing DT to obtain multipath channel estimates for each candidate location. Second, it proposes a novel learned scoring function that compares the top K paths of the candidate and measured multipath profiles, with training performed over an ensemble of environments so the method can be evaluated on environments not included in the training set, incorporating errors in either the DT or channel estimation during training for robustness. Third, it validates the approach using a Sionna-based ray-tracing backend, benchmarking against Gaussian and Gaussian-mixture posterior models, demonstrating that explicit candidate-wise DT matching better captures the sharp, multimodal, and layout-dependent posterior structure induced by indoor multipath.
The localization setting is defined as follows: the transmitter is at an unknown position X t in R d (with d=2 for 2D localization), and the receiver is given (L, p) where L is a known room layout and p = (x r, phi r) is the receiver pose (position x r in R squared and orientation phi r in [-pi, pi)). At runtime, the receiver collects an observation y, typically an array of in-phase/quadrature (IQ) samples from the receiver antennas. The inference objective is to estimate the posterior distribution p(x t L, p, y), representing the relative likelihood of all potential transmitter locations given the observation, room layout, and receiver pose.
The DT multipath estimator is modeled abstractly as a function Psi(L, p, x t) = psi DT k, k = 1,..., K in S K, where each path is a tuple psi = (theta, tau, alpha, gamma) in S, with theta being the angle of arrival (AoA), tau the relative delay, alpha the complex path gain, and gamma = alpha squared / sigma squared the real-valued path SNR for receiver-noise variance sigma squared. The DT maps each candidate location x t to a predicted set of K paths, sorted from strongest to weakest. The abstraction is intentionally backend-agnostic, allowing any DT method to instantiate Psi as long as it produces the same ranked tuple representation.
The observation interface models the multipath estimator as a function Psi hat: C N x N rx -> S K, which produces the estimated multipath profile psi obs k, k = 1,..., K from the received IQ samples. The paper uses the iterative Levenberg–Marquardt (LM) algorithm for multipath estimation from array measurements.
The core modeling problem is constructing a logit function g omega(x t, L, p, y) = G omega(psi DT k(x t), psi obs k), where omega collects all trainable weights and biases of the neural network realizing the score function G omega. Assuming a prior distribution p 0(x t) (typically uniform over a region of interest), the posterior is given by p hat omega(x t L, p, y) = e g omega(x t,L,p,y) p 0(x t) / Z omega(y; L, p), where Z omega(y; L, p) is the normalizing constant obtained by integrating over all candidate locations.
The score function is trained using a sampled cross-entropy loss. Given a dataset D = L m, p m, C m, y m, x tm m=1 M, where each training sample includes a finite set of candidate transmitter locations C m = x bar tmj j=1 N m subset of R d, and the true transmitter location satisfies x tm in C m (with a unique candidate index c* m), the loss function is approximated as L(omega) = (1/M) sum m=1 M [-g mc* + log(sum j=1 N m e g mj)], where g mj = g omega(x bar tmj, L m, p m, y m). This is the candidate-set analogue of the sampled cross-entropy criterion in energy-based learning.
The score function G omega is realized as a transformation followed by a simple multilayer perceptron. In the simulations, K = 6 paths are retained from each DT and observation profile. Synchronization is unavailable and the system is narrowband, so delay is discarded. Each AoA theta is encoded as (cos theta, sin theta), while the SNR gamma is converted to dB and scaled to [0, 1] over a 60 dB range. The resulting 12 paths yield 3 x 12 = 36 features, fed into a two-layer neural network, providing an extremely compact, easily trainable scoring function with a limited number of parameters.
For experimental validation, the paper creates random synthetic environments with varying obstacles, noting the method is not site-specific and can be tested on completely random environments different from those used for training. Each random environment is a 100 m x 100 m region containing the outer room boundary and three randomly generated interior obstacles drawn from rectangular, T-shaped, and L-shaped primitives. The NVIDIA Sionna ray-tracing backend simulates the true channel and N m = 500 candidate channels. The carrier frequency is 12 GHz, signal bandwidth is 200 MHz, noise figure is 7 dB, synchronization window is 4 microseconds, and transmit power is 0 dBm. The transmitter uses one vertically polarized isotropic element, while the receiver uses a 1 x 8 vertically polarized half-wavelength array with a 3GPP TR 38.901 element pattern. The ray tracer returns up to 16 valid paths per channel; LOCUS-DT retains the K = 6 strongest paths from each DT candidate profile. For the online observation, the standard maximum-likelihood (SML) estimator uses a 128-point matched filter to estimate (theta, tau, alpha), derives gamma, and retains the K = 6 strongest scorer pairs (theta, gamma).
The training split spans 8000 distinct room layouts, each with 500 candidate transmitter locations and 10 observation realizations, for a total of 80,000 snapshots. The test split spans 2000 additional layouts with the same 500-candidate structure, yielding 20,000 snapshots. For the final benchmark comparison, two harder evaluation sets are reported: Eval Random uses 2000 held-out groups with 1000 feasible random candidates per group, while Eval Grid uses the corresponding held-out layouts but starts from a regular 31 x 31 raw grid and keeps 696 layout-feasible candidate positions after pruning invalid points.
The maximum-likelihood observation estimator uses frequency- and spatial-domain basis functions b F(tau) = e-j2pi f 0 tau * m m=0 M F - 1 and b R(theta) = e-j pi sin theta * m m=0 M R - 1, where M F is the number of subcarriers, f 0 is the subcarrier spacing, and M R is the number of half-wavelength-spaced receive-ULA elements. The frequency-space channel response over K paths is s(Psi) = sum k=1 K b R(theta k) tensor b F(tau k) alpha k. Assuming equal-power pilots across subcarriers, the channel observation is modeled as y m = s(Psi) + w m, where w m is additive white Gaussian noise with variance sigma squared. The maximum-likelihood estimate is Psi hat = arg max Psi p(y m Psi).
The paper compares against three benchmarks: Gauss-Cart (Gaussian posterior with transmitter location in Cartesian coordinates, with mean mu and covariance Q as functions of the same full inputs), Gauss-Polar (identical estimator but with posterior in polar coordinates), and GMM-Cart (mixture with k = 3 Gaussian components). All trainable models are optimized with AdamW for 2000 epochs with weight decay 10-4 and hidden-layer dimensions (64, 16). LOCUS-DT uses pairwise feature construction whereas Gaussian benchmarks use candidate-independent conditioning on receiver pose and IQ observation. Learning rates are 5 x 10-3 for LOCUS-DT and 2 x 10-3 for Gaussian baselines. Training runs on a single NVIDIA A100 GPU, requiring approximately 2.5 GPU-hours for LOCUS-DT, 1.5 GPU-hours for Gauss-Cart, 1.5 GPU-hours for Gauss-Polar, and 2.0 GPU-hours for GMM-Cart, totaling approximately 7.5 GPU-hours.
Besides the training cross-entropy, the paper reports the adjusted loss L = (1/M) sum m=1 M [-log p hat mc* - log N m], which measures gain relative to uniform random guessing over the candidate set, and derived quantities G = e-L and R = -L / log N cand, where N cand is the split-specific candidate count. Lower L and higher G and R indicate better performance.
Qualitative results show LOCUS-DT on three distinct random layouts, with each row pairing a Sionna room with posterior heatmaps for six TX/RX configurations. LOCUS-DT yields the most localized and geometry-consistent posteriors, with high-probability regions aligned with the true transmitter and wall- and reflection-induced structure. Gauss-Cart and Gauss-Polar produce smooth single-mode fields that miss sharp spatial discontinuities and multimodal ambiguity. GMM-Cart captures limited multimodality, but its peaks remain broader and less aligned because it lacks candidate-wise DT matching. Quantitative results in Table I confirm that LOCUS-DT outperforms all Gaussian baselines on both evaluation sets, with LOCUS-DT achieving adjusted loss values of-2.7224 (Eval Grid) and-3.3798 (Eval Random), compared to Gauss-Cart (-0.0594, -0.1190), Gauss-Polar (-0.3030, -0.2650), and GMM-Cart (-0.3778, -0.5078), with correspondingly higher G and R values.
The paper concludes that LOCUS-DT captures reflection-induced multimodal ambiguity, outperforms Gaussian and Gaussian-mixture baselines, and generalizes to unseen layouts. Future work will evaluate greater DT mismatch and measured rooms reconstructed using SLAM.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
1. Physics-informed posterior inference for localization
-
Replace single-point regression outputs with a learned scoring function that compares ray-tracing digital twin (DT) predictions against observed multipath profiles
-
Train the scoring function over an ensemble of environments (8000 layouts) so it generalizes to unseen room geometries without retraining
-
Use a candidate-set cross-entropy loss with Monte Carlo partition function estimation to handle intractable normalization
2. Robust multimodal uncertainty representation
-
Encode each path as (cos θ, sin θ, SNR) features, discarding delay for narrowband synchronization-free operation
-
Retain top K=6 paths from both DT and observation, producing a 36-dimensional feature vector fed into a two-layer MLP (64, 16 hidden units)
-
Incorporate DT and channel estimation errors during training to maintain robustness against model mismatch
3. Layout-conditioned candidate scoring
-
Generate 500 candidate locations per scene using Sionna ray tracing at 12 GHz, 200 MHz bandwidth
-
Compare DT-predicted multipath signatures against maximum-likelihood estimates from 1×8 receive array snapshots
-
Output a full posterior heatmap over the 2D space rather than a single point estimate
1. Achieve significantly better localization accuracy than Gaussian baselines
-
On held-out random layouts with 1000 candidates: adjusted loss L = −3.3798 vs −0.1190 (Gauss-Cart), −0.2650 (Gauss-Polar), −0.5078 (GMM-Cart)
-
Achieve 48.93% relative gain over uniform guessing (R metric) vs 1.72% for the best Gaussian baseline
-
On grid-based candidates (696 positions): L = −2.7224, R = 41.59%, outperforming all baselines by over 35 percentage points
2. Capture sharp, multimodal, layout-dependent posterior structures
-
Produce posterior heatmaps that align high-probability regions with true transmitter locations and reflection-induced ambiguity patterns
-
Resolve spatial discontinuities and multiple peaks that Gaussian or Gaussian-mixture models smooth over
-
Maintain geometry-consistent uncertainty even when the true location is ambiguous due to heavy blockage or multipath
3. Generalize to unseen environments without site-specific retraining
-
Train once on 8000 random layouts, then evaluate on 2000 completely different held-out layouts
-
Handle arbitrary room shapes (rectangular, T-shaped, L-shaped obstacles) and any receiver pose
-
Operate with only a known layout and receiver pose, requiring no fingerprinting or prior measurements in the target environment
4. Provide actionable uncertainty for downstream tasks
-
Output a normalized posterior distribution over all feasible transmitter locations, enabling robotic navigation, search-and-rescue, and multi-hypothesis tracking
-
Support decision-making under ambiguity by exposing multimodal hypotheses rather than forcing a single estimate
-
Enable integration with planning algorithms that require probabilistic spatial beliefs (e.g., reinforcement learning for indoor navigation)
Abstract
Accurate indoor localization is essential for emerging applications in robotic navigation and search and rescue. While classical methods typically focus on single-point estimates, complex indoor environments with heavy blockage and multipath propagation often lead to multimodal likelihood surfaces where a single estimate is insufficient. This paper proposes LOCUS-DT (Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins), a framework that treats snapshot localization as posterior inference over the transmitter location. By leveraging a ray-tracing-based digital twin (DT) of the known environment, LOCUS-DT generates synthetic multipath profiles for candidate locations and compares them against the measured channel profile. Central to our approach is a novel learned scoring function designed to compare a fixed number of dominant specular paths, providing robustness against errors in both the DT environment model and the physical channel estimation. Importantly, LOCUS-DT is trained over an ensemble of environments to ensure generalization to unseen layouts. We evaluate the system using a Sionna-based ray-tracing backend, demonstrating that LOCUS-DT captures the sharp, multimodal posterior structures inherent in indoor settings more accurately than standard Gaussian or Gaussian-mixture benchmarks.
Sources
Related papers
- Runtime Assurance Under Measurement Attack: Necessary and Sufficient Observability Conditions for Learned Control in Radio Access Networks
- Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography
- Uncertainty Quantification in Machine Learning for Biosignal Applications -- A Review
- Continuous Orthogonal Mode Decomposition: Haptic Signal Prediction in Tactile Internet
- Generative Models for Modeling and Synthesizing MIMO Channels in Adverse Weather Conditions
- Deep-Learning-Based Pixelated Microwave Filter Design and Characterization using Electro-Optical Electric-Field Measurements