Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Vision Meets WiFi".
Jane: Estimating volumetric mechanical properties, including Young’s modulus, Poisson’s ratio, and density at each voxel,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Speaking of the title, "Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties," I think that really captures the essence of what they've done—it’s not just about seeing things; it’s about fusing visual information with physics data to get those mechanical properties. Who are we looking at in terms of the team behind this research?
Jane: The authors are Ali Bahri, Hongliang Li, Soufiane Lamghari, and Jie Chuai. They're coming from the Huawei Noah’s Ark Lab in Canada and Hong Kong SAR, which suggests a strong foundation in computer vision and potentially hardware integration.
Lu: Their background points toward a solid theoretical grounding in deep learning architectures used for scene understanding, but their specific focus on integrating RF sensing into the material slot framework shows they are pushing the boundaries of multimodal fusion. I'm very interested in how they handled that initial data alignment between visual and electromagnetic features.
Meng: I wonder if their hardware background gives them an edge in dealing with real-world sensor noise versus purely simulated data, which is a big concern when we move these models from the lab to actual deployment scenarios.
Lalam: From my perspective, having researchers with deep backgrounds in both vision and potentially signal processing means they can build systems that are more robust across different types of input data streams than if they were purely focused on one domain.
The paper's summary: Tom: So, diving into the actual summary of "Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties," they explain that predicting Young’s modulus, Poisson’s ratio, and density at every single voxel from just vision is fundamentally ambiguous because visually similar objects can have very different material compositions.
Jane: Exactly, Tom. The paper summarizes the problem as existing methods predicting these properties independently across voxels which leads to noisy or inconsistent estimates for voxels that should actually share the same material structure, and it lacks a way to resolve that visual ambiguity explicitly.
Lu: Their proposed solution is ViWi, which reformulates this by using an object-centric material decomposition approach where latent material slots group voxels based on visual and material compatibility, allowing them to aggregate evidence for a shared property prototype.
Meng: So, instead of treating each voxel in isolation, they are grouping them into these material hypotheses that share common physical behavior, which sounds like a much more structurally sound way to model an object.
Lalam: This idea of grouping voxels based on shared material identity is really powerful because it enforces the physical reality that different parts of the same object should behave similarly mechanically, even if they are far apart in space.
The paper's improvements: Tom: Beyond just solving the ambiguity, what are the specific improvements they detail in "Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties"? I want to know exactly what makes ViWi better than what came before.
Jane: They introduce two main innovations: first, they use RF-conditioned material-slot initialization by using feature-wise linear modulation to stably integrate global RF cues with the voxel's visual features for both material grouping and property estimation.
Lu: That conditioning step is key; it allows the RF evidence, which captures global composition cues like permittivity and conductivity from a physics simulation, to influence the initial states of those material slots before any iterative grouping happens. It’s a sophisticated way to inject physical knowledge early on.
Meng: From an engineering standpoint, that conditioning mechanism sounds like a smart way to stabilize the learning process; it prevents the model from getting stuck on purely visual artifacts by grounding it in those global physical constraints.
Lalam: I think this is where things get really interesting because it shows how you can combine spatial localization from vision with global material composition cues, which is something that standard vision models simply don't have access to.
Conclusion: Tom: So, wrapping up on "Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties," the authors show that by combining this object-centric structure with the complementary RF evidence, they achieve state-of-the-art performance on tasks like GVM and improve mass estimation on datasets like ABO-five hundred.
Jane: They also showed that when visual evidence is ambiguous, RF provides the largest gains in accuracy compared to using vision alone, which really validates the use of this complementary sensing approach.
Lu: The implication here is that we can move towards more physically grounded perception systems for robotics and digital twins because we aren't just guessing properties based on shape anymore; we are inferring them from a richer combination of visual appearance and simulated physical characteristics.
Meng: For practical application, this means the models generated will be much more reliable for things like safety-critical applications where knowing the stiffness of a component is non-negotiable.
Lalam: This work really opens up possibilities for creating AI systems that can not only see what’s there but also understand the physical substance of what they are seeing, which could fundamentally improve how we design and build intelligent objects.
Tom: It's clear that "Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties" provides a solid framework for making volumetric property estimation more physically consistent than ever before. Jane, Lu, Meng, Lalam—thanks for joining us on this deep dive into the research!
Ali Bahri, Hongliang Li, Soufiane Lamghari, Jie Chuai, Zhitang Chen
Huawei Noah’s Ark Lab
cs.CV
Submitted: 2026-08-07
Updated: 2026-09-28
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: Estimating volumetric mechanical properties, including Young’s modulus, Poisson’s ratio, and density at each voxel, is intrinsically ambiguous from vision alone because visually similar objects
Key concepts
- Material Slots
- These are compact representations of material hypotheses for parts of an object. Instead of predicting properties for every single voxel, ViWi groups voxels with the same material identity into these slots, allowing it to predict a shared mechanical-property prototype for that entire group.
- RF Descriptor
- This is a compact description derived from simulating electromagnetic properties (like permittivity and conductivity) based on known material categories. This descriptor captures global composition cues, providing physics-based information that conditions the material slots during the estimation process.
- Object-Centric Framework
- Instead of analyzing voxels individually, ViWi focuses on representing the entire object through a set of material slots. This structure ensures that predictions are consistent across different parts of an object that share a common material identity, improving physical coherence.
- Feature-wise Linear Modulation (FiLM)
- This technique is used to condition the initial state of the material slots using the RF descriptor. It allows the global composition information from RF to influence how the material slots are initialized before they begin grouping and prediction.
Terminology
Summary
Estimating volumetric mechanical properties, including Young’s modulus, Poisson’s ratio, and density at each voxel, is intrinsically ambiguous from vision alone because visually similar objects may have substantially different material compositions and physical behaviors. ViWi introduces an object-centric framework that combines visual observations with complementary radio-frequency (RF) sensing to resolve this ambiguity by representing objects through compact material slots conditioned by both visual features and physics-based RF descriptors.
The gist
ViWi represents each object using a compact set of material slots that aggregate evidence from voxels with a shared material identity and produce coherent slot-level property predictions, while incorporating a compact RF descriptor generated through WiFi-band electromagnetic simulation to condition the material slots with global composition cues.
Problem Formulation and Overview
The goal is to predict the volumetric mechanical-property field, where for each occupied voxel i, the model predicts a triplet: yi = (Ei, νi, ρi), representing Young’s modulus, Poisson’s ratio, and density. A direct voxel-wise formulation is limited because it overlooks the compositional structure of real objects; voxels made of the same material should exhibit similar mechanical properties and benefit from sharing evidence. Furthermore, visual appearance does not uniquely determine material composition, as objects with similar geometry may be made from materials with substantially different mechanical properties.
Object-Centric Material Slot Attention
ViWi reformulates volumetric property estimation as object-centric material decomposition by representing an object using a compact set of latent material slots. Each slot captures a shared material hypothesis, aggregates evidence from the voxels that it explains, and predicts a shared mechanical-property prototype. Voxel-level predictions are then reconstructed through soft assignments to these slot-level property prototypes. This structure promotes consistent predictions across voxels that share a material identity, including spatially disconnected components.
RF-Conditioned Material Slots
To complement visual appearance, ViWi incorporates a compact RF descriptor obtained through physics-based electromagnetic simulation. This descriptor captures global composition cues by mapping mechanical material categories to representative electromagnetic parameters like relative permittivity and conductivity. The standardized RF descriptor conditions the initial material slots via feature-wise linear modulation (FiLM), allowing it to condition the initial material-slot states before iterative voxel grouping.
This allows RF evidence to jointly influence material assignments and property estimation while visual features provide spatial localization.
Learning Objectives
The model is supervised using several objectives. The primary objective supervises the voxel-level property prediction: Lprop = 1/N Σ ybi − yi1. An auxiliary objective aligns a vision-based voxel material latent (mb i) with the pretrained material-latent space, and a composition loss (Lcomp) is introduced to supervise the RF embedding to predict object-level material composition: Lcomp = -1/B Σ X b Σ C c=1 hb,c log softmax(lb)c. The complete objective is L = Lprop + λmatLmat + λcompLcomp.
Main Results and Contributions
ViWi achieves state-of-the-art performance on GVM and improves real-object mass estimation on ABO-500. ViWi reduces the ALDE and ALRE of Young’s modulus from 0.3952 and 0.0427 to 0.1793 and 0.0195, corresponding to reductions of approximately 54.6% and 54.3%, respectively, compared to VoMP [4]. The results demonstrate that combining object-centric material structure with complementary RF evidence enables more accurate and physically coherent volumetric property estimation beyond what is possible from visual appearance alone. RF provides the largest gains when visual evidence is ambiguous, as shown by the gain analysis across visual difficulty quintiles.
Limitations
A key limitation noted is that the RF descriptors are simulated rather than measured with real hardware, and therefore do not capture all practical sources of RF variation.
Future work will evaluate ViWi using real RF observations under realistic sensing conditions. The model's performance also shows that while the benefit of material-slot grouping extends beyond voxel-level property estimation, highly unreliable RF may become less useful than disabling it under severe perturbation.
References
[1] Albert J Zhai, et al. Physical property understanding from language-embedded feature fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 28296–28305, 2024.
[4] Rishit Dagli, et al. Vomp: Predicting volumetric mechanical property fields. arXiv preprint arXiv:2510.22975, 2025.
[16] Francesco Locatello, et al. Object-centric learning with slot attention. arXiv preprint arXiv:2006.15055, 2020.
[17] Fei Shang, et al.
Improvements for AI systems
Here are the specific improvements that can be made to existing AI systems by implementing ViWi, and what those improved systems can achieve:
-
Improve volumetric mechanical property estimation accuracy by moving from independent voxel-wise prediction to an object-centric material decomposition framework. This means instead of predicting Young’s modulus, Poisson’s ratio, and density for every single voxel independently (which leads to noise and inconsistency), the system will group voxels based on shared material identity.
-
Enable consistent property prediction across spatially disconnected components of the same material. The improved system can accurately predict mechanical properties for entire objects even if they are geometrically separated, overcoming a major limitation of prior independent regression methods like VoMP.
-
Enhance the ability to distinguish visually similar objects that possess different physical compositions (e.g., distinguishing a metal container from a plastic box). By incorporating the physics-based RF descriptor, the system gains global material composition cues that are invisible to standard vision models, leading to significantly better material classification and property estimation in ambiguous cases.
-
Provide physically grounded representations for downstream tasks such as robotic manipulation, digital twins, and physics-based simulations. The improved system will generate simulation-ready object representations where stiffness, deformation behavior (Poisson's ratio), and mass distribution are spatially resolved and physically consistent throughout the entire volume.
-
Enable more robust and reliable mass estimation of real-world objects (like those from the ABO-500 dataset). The material-slot grouping mechanism allows the system to produce volumetric density fields that yield more accurate total object masses, even when trained primarily on synthetic data, demonstrating transferability to real-world physical constraints.
-
Develop an AI system that can dynamically adjust its estimation strategy based on visual ambiguity. By quantifying the
RF gain
(the improvement in error when RF is enabled versus vision-only), the system can intelligently decide when to rely more heavily on complementary RF evidence, ensuring high accuracy for complex or visually challenging scenes. -
Create a novel multimodal sensing pipeline that seamlessly fuses dense volumetric visual data with compact, physics-derived electromagnetic descriptors (RF). This fusion allows the AI to leverage both spatial localization (from vision) and global material composition cues (from RF) simultaneously to resolve inherent ambiguities in material science problems.
Abstract
Estimating volumetric mechanical properties, including Young's modulus, Poisson's ratio, and density at each voxel, is intrinsically ambiguous from vision alone, as visually similar objects may have substantially different material compositions and physical behavior. Existing approaches predict these properties independently across voxels, overlooking the piecewise-constant material structure of real objects and producing noisy or inconsistent estimates for voxels that share the same material, while lacking an explicit mechanism to resolve visual ambiguity. We introduce ViWi (Vision Meets WiFi), an object-centric framework for volumetric mechanical-property estimation. ViWi represents each object using a compact set of material slots that aggregate evidence from voxels with a shared material identity and produce coherent slot-level property predictions. To complement visual appearance, ViWi incorporates a compact RF descriptor generated through WiFi-band electromagnetic simulation using permittivity and conductivity. The RF descriptor conditions the material slots with global composition cues that may be unavailable from images, while visual features preserve voxel-level spatial localization. On GVM, ViWi improves over the prior state of the art on four of six per-voxel metrics, while its vision-only variant improves all reported mass-estimation metrics on ABO-500. These results demonstrate that combining object-centric material structure with complementary RF evidence enables more accurate and physically coherent volumetric property estimation beyond what is possible from visual appearance alone.
Sources
- VoMP: Predicting Volumetric Mechanical Property Fields
- PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation
- DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors
- Physics3D: Learning Physical Properties of 3D Gaussians via Video Diffusion
- PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics
- Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception
- Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels
- SOPHY: Learning to Generate Simulation-Ready Objects with Physical Materials
- PhysX-3D: Physical-Grounded 3D Asset Generation
- Object-Centric Learning with Slot Attention
- The Field-based Model: A New Perspective on RF-based Material Sensing
- HuPR: A Benchmark for Human Pose Estimation Using Millimeter Wave Radar
- RFPose-OT: RF-Based 3D Human Pose Estimation via Optimal Transport Theory
- Diffusion Model is a Good Pose Estimator from 3D RF-Vision
- Improving Real-Time Omnidirectional 3D Multi-Person Human Pose Estimation with People Matching and Unsupervised 2D-3D Lifting
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models