A Recommendation System Approach for Interference-Robust Sensor Subset Selection
Kaan Buyukkalayci, Kyle Pak, Merve Karakas, Christina Fragouli
University of California, Los Angeles
cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: This paper develops a method for sensor-subset selection for tracking.
Terminology
Summary
This paper develops a method for sensor-subset selection for tracking. Prior work showed that low-cost acoustic Received Signal Strength Indicator (RSSI) measurements can be used to recommend subsets of sensor nodes whose expensive sensing modalities, such as cameras, can achieve high tracking accuracy. While efficient, RSSI-based approaches are challenged by acoustic interference. We propose a recommendation-system-inspired framework that instead leverages frequency-band acoustic features and a Two-Tower Multi-Layer Perceptron (MLP) architecture to efficiently score candidate sensor subsets. Experimental results on outdoor vehicle-tracking deployments show that the proposed method can improve accuracy by around 20% over the RSSI baseline while maintaining the low computational overhead required for real-time selective sensing.
The paper revisits the subset sensor selection problem from a recommendation-systems perspective. At each sensing interval, the network observes a low-cost acoustic description of its current state and must select a small subset of sensing assets whose higher-cost modalities should be activated. This mirrors the structure of modern recommendation systems, where a context is matched against a collection of candidate items and the most relevant items are selected. In this setting, the network-wide acoustic state serves as the context, while candidate sensor subsets serve as the items to be recommended. This perspective allows moving beyond explicit target localization and posterior inference, and instead learning a direct mapping from acoustic observations to sensor-subset utility.
The proposed framework is based on a Two-Tower Multi-Layer Perceptron (MLP) architecture that learns compact representations of both the network acoustic state and candidate sensor subsets. Similar to retrieval and recommendation systems, the separate representations enable efficient scoring of many candidate subsets while maintaining low online computational cost.
The model takes two inputs: a context vector describing the acoustic state of the network at a given time, and an action vector specifying a subset of nodes and their relative geometry. These inputs are passed into separate towers, and the resulting embeddings are combined and fed into an MLP head to output a scalar utility score. At inference time, the model scores all candidate subsets and the subset with the highest predicted utility is selected.
The context vector is constructed by aggregating node-level coordinate and acoustic band-power features as ct = [pi, ψ i,t]i∈V, where pi denotes the normalized coordinates of node i, and ψ i,t denotes the audio-band power features extracted from the lower-cost acoustic modality at time t. The audio-band block ψ i,t contains dB-scale band-power features, each computed as the acoustic power received by node i within a predefined frequency band over the sampling interval (t−1, t). The frequency bands used are: 20–80 Hz, 80–160 Hz, 160–400 Hz, 400–900 Hz, 900–2000 Hz, 2000–3500 Hz, and 3500–6000 Hz. The choice of band edges is based on a coarse logarithmic partition of the vehicle-acoustic spectrum, where lower frequencies, in which engine and tire-road energy are concentrated, are resolved more finely, while higher frequencies are grouped more coarsely.
The action vector is constructed as a(S) = [m(S), p(S), S], where m(S) ∈ 0, 1 V is the binary membership mask of the candidate subset, p(S) stores the normalized coordinates of the selected nodes using a fixed ordering and padding up to the subset-size budget, and S records the subset size. Since a(S) depends only on the subset identity and node geometry, it is independent of t and can be precomputed for all candidate subsets.
The model is trained to predict a smooth distance-based utility for each candidate subset. For a candidate subset S at time t, let dt (S) ≤ dt (S) ≤ · · · ≤ dt (S) denote the sorted distances from the vehicle to the nodes in S. The utility of selecting S is defined as ut (S) = Σ j=1 S wj / (1 + dt (S)/ρ), where ρ > 0 is a distance scaling parameter and w1 ≥ w2 ≥ · · · ≥ 0 are decreasing weights. This utility gives larger reward to subsets containing nodes close to the vehicle with higher weights assigned to the closest selected nodes.
The Two-Tower architecture uses one tower representing the time-varying network context and the other representing the candidate sensor subset. The context tower maps ct ∈ RC to an embedding z c (t) = fc (ct) ∈ RE while the action tower maps a(S) ∈ RA to an embedding z a (S) = fa (a(S)) ∈ RE, where E is the embedding dimension. Each tower is a small fully connected network with hidden width H. The two embeddings are combined using their concatenation and elementwise product, r t (S) = [z c (t), z a (S), z c (t) ⊙ z a (S)] ∈ R3E. The elementwise product acts as a learned compatibility feature between the current acoustic state and the candidate subset. The combined representation is passed through a prediction head g to produce the predicted utility ût (S) = g(r t (S)). The model is trained by minimizing the mean-squared error between the predicted and true subset utilities.
The method is evaluated on two distinct outdoor deployments. The interference-rich deployment uses six sensing nodes and a cargo-van target vehicle in an outdoor area of approximately 4,000 m2, with hilly terrain, buildings, a nearby construction zone, and intermittent acoustic interference including human speech, wind, and occasional audio from nearby passing vehicles. The open-field deployment uses ten sensing nodes in a larger outdoor area of approximately 10,000 m2 with fewer external acoustic interference sources.
On the interference-rich deployment, with a fixed recommendation budget of S = 3, the Two-Tower model with frequency-band features achieves a mean closest-node containment accuracy of 98.39%, compared to 80.44% for the Two-Tower (RSSI-only) ablation, 77.03% for the linear path-loss posterior, 75.95% for the normalized RSSI top-3, 72.74% for the KDE-hybrid posterior, and 71.33% for the spline posterior. The frequency-band representation increases computation and storage relative to scalar RSSI methods, but the total runtime remains well below the 200 ms sensing interval. The online computation time for the Two-Tower model is approximately 0.33 ms mean total time and 1.70 ms at the 99th percentile.
On the open-field deployment, the RSSI-only Two-Tower model attains the highest accuracy (99.40%), marginally ahead of the full frequency-band Two-Tower model (97.80%), with both learned models outperforming the model-based posterior baselines. The corresponding online computation times remain feasible within the 200 ms sensing interval.
Comparing the two deployments reveals when the richer spectral representation can be worth the extra computational and storage cost. On the interference-rich deployment, the frequency-band Two-Tower model is dramatically more robust than its RSSI-only counterpart (98.39% vs. 80.44%): the band-power features let the model separate the target’s acoustic signature from intermittent interference such as speech, wind, and passing vehicles, whereas using only RSSI is more likely to fail in distinguishing target-generated acoustic power from interference sources. The practical takeaway is that frequency-band features are most valuable precisely when the acoustic environment is contested; in benign conditions a lightweight RSSI-only model is sufficient and even slightly preferable.
The computational complexity of the Two-Tower method is O(Vl + VhK), where Vl is the number of low-cost acoustic nodes, Vh is the number of recommendable high-cost assets, and K is the subset-size budget. The O(Vl) term results from constructing and transmitting the network-wide acoustic context, while the O(VhK) term results from exact scoring of candidate subsets. This removes dependence on the spatial hypothesis grid and joint multi-target posterior enumeration that posterior-based scalar-RSSI methods require.
The main contributions of the paper are: formulating subset sensor activation as a recommendation problem in which network observations define the context and candidate sensor subsets define the recommendable items; introducing a Two-Tower recommendation architecture that learns separate embeddings of network state and sensor subsets, enabling efficient real-time scoring of sensing actions; demonstrating that frequency-band acoustic features significantly improve robustness to acoustic interference, increasing accuracy from 80.4% to 98.4% in one of the experiments; and showing through real-world deployments that the proposed method remains computationally lightweight, requiring sub-millisecond to approximately 1 ms of online computation while outperforming prior RSSI-based sensor recommendation approaches.
Improvements for AI systems
Improvements to AI Systems:
-
Context-Aware Sensor Subset Selection with Interference Robustness: The AI system can now dynamically select which high-cost sensors (e.g., cameras) to activate based on a low-cost acoustic context, using frequency-band features (e.g., 20–80 Hz, 80–160 Hz) instead of just RSSI. This improves tracking accuracy by 20% in noisy environments (98.4% vs. 80.4%) by distinguishing target acoustic signatures from interference like speech, wind, and passing vehicles.
-
Two-Tower MLP Architecture for Efficient Recommendation of Actions: The system uses separate neural towers to embed (a) the network-wide acoustic state and (b) candidate sensor subsets, then combines them via concatenation and elementwise product to predict a utility score. This enables fast scoring of all possible subsets (O(V h K) complexity) without explicit target localization or posterior inference, reducing online computation to 0.33 ms mean (1.7 ms at 99th percentile)—well within real-time sensing intervals.
-
Learned Utility Function for Ranking Subsets: Instead of hand-crafted rules, the AI learns a smooth distance-based utility (weighted sum of inverse distances from target to selected nodes) via mean-squared error training. This allows the system to prioritize subsets with nodes closest to the target, assigning higher weights to the nearest nodes, leading to higher containment accuracy.
-
Adaptive Feature Selection Based on Environmental Conditions: The system can switch between lightweight RSSI-only features (for benign, open-field settings) and richer frequency-band features (for interference-rich settings) based on deployment context. In open-field tests, the RSSI-only model achieves 99.4% accuracy, while in contested environments, the frequency-band model achieves 98.4%—enabling the AI to trade off computational cost vs. robustness as needed.
-
Precomputed Action Embeddings for Real-Time Scalability: The action vector (membership mask, node coordinates, subset size) is independent of time, so the system precomputes embeddings for all candidate subsets offline. At inference, only the context tower needs updating, making the system scalable to larger networks without increasing online latency.
-
Improved Generalization Across Deployments: By learning a direct mapping from acoustic observations to subset utility (rather than relying on physical path-loss models), the AI generalizes across different terrains, node densities, and interference profiles. This removes the need for manual calibration of acoustic propagation models and spatial hypothesis grids.
What the Improved AI System Can Do:
-
Real-time selective sensing: In a network of low-cost acoustic sensors and high-cost cameras, the system recommends the optimal 3–5 cameras to activate every 200 ms, achieving >98% accuracy in tracking a moving vehicle, even under heavy acoustic interference.
-
Robust operation in contested environments: It maintains high tracking accuracy when human speech, wind, or other vehicles generate acoustic noise, by leveraging frequency-band power features that isolate the target's engine/tire signature.
-
Low-latency decision-making: It scores all candidate subsets in under 2 ms, enabling deployment on edge devices with limited compute, suitable for autonomous vehicles, surveillance, or wildlife monitoring.
-
Adaptive resource allocation: It automatically uses simpler features (RSSI) in quiet settings to save bandwidth/storage, and switches to richer spectral features only when interference is detected, optimizing energy and computational efficiency.
-
Scalable to larger networks: The architecture handles tens of nodes and subset sizes without exponential growth in inference time, thanks to precomputed action embeddings and a linear context-encoding step.
Abstract
This paper develops a method for sensor-subset selection for tracking. Prior work showed that low-cost acoustic Received Signal Strength Indicator (RSSI) measurements can be used to recommend subsets of sensor nodes whose expensive sensing modalities, such as cameras, can achieve high tracking accuracy. While efficient, RSSI-based approaches are challenged by acoustic interference. We propose a recommendation-system-inspired framework that instead leverages frequency-band acoustic features and a Two-Tower Multi-Layer Perceptron (MLP) architecture to efficiently score candidate sensor subsets. Experimental results on outdoor vehicle-tracking deployments show that the proposed method can improve accuracy by around 20% over the RSSI baseline while maintaining the low computational overhead required for real-time selective sensing.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks