A Wearable Multimodal Ultrasound+Inertial System for Real-Time Virtual Reality Interaction

arXiv:2606.17741 · eess.SY, cs.HC, cs.SY · Submitted 2026-06-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "A Wearable Multimodal Ultrasound+Inertial System for Real-Time Virtual Reality Interaction".

Rosa: A fully wearable multimodal interface combining ultrasound and inertial sensing enables real-time virtual reality interaction by concurrently estimating hand pose and forearm position.

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: So, to summarize what we've covered about this paper, the main thesis is that existing wearable approaches have limitations in terms of interaction complexity and wearability because they often rely on external hardware.

Dev: They propose a fully wearable multimodal interface based on concurrent ultrasound sensing from the forearm and upper arm alongside inertial data from an accelerometer to enable real-time VR interaction.

Taro: Essentially, they are arguing that by combining these two modalities, they can map muscular activity into control commands while keeping the benefits of wearable sensing.

Rosa: And what makes this system important is that it integrates an end-to-end software framework for real-time acquisition, visualization, and communication directly with a Unity VR environment.

Dev: They introduce a multimodal learning pipeline designed specifically to estimate both hand pose and forearm position concurrently in 2D space using these combined data streams <ref:2606.17741#pg0,hand pose and forearm position>.

Taro: This setup is important because it moves past just recognizing discrete gestures; they are aiming for continuous, functionally meaningful interaction.

Rosa: The paper claims this system is fully wearable and achieves performance metrics during online validation, specifically reaching success rates around ninety-two percent for cylinder grasping and relocation tasks after minimal fine-tuning.

Dev: This matters because it shows the feasibility of achieving high accuracy in these functional interactions without needing external optical tracking systems to guide the user's hand.

Taro: It’s significant because it demonstrates that US data, when fused with inertial measurements, can provide sufficient information for complex manipulation tasks within a wearable context.

Rosa: The system is built on the WULPUS platform, which involves six ultrasound transducers on the forearm and two on the upper arm, all streamed wirelessly via BLE.

Dev: Furthermore, they detail how they extended the BioGUI framework to accommodate this new data type, adding visualization modes for both A-mode and M-mode imaging of the US signals.

Taro: This integration into a cohesive software architecture is what makes it a complete system rather than just a collection of sensors.

Rosa: It matters because it shows how multimodal sensing can be leveraged to create more capable and versatile wearable interfaces for virtual reality environments.

Dev: The core idea is using the US for depth information alongside the inertial data for motion tracking, which provides a richer understanding of hand and arm position in 2D space <ref:2606.17741#pg0>.

Taro: This capability opens up possibilities for applications requiring finer motor control than what simple accelerometers alone can provide.

Rosa: So, to put it plainly, this paper is about creating a system where the physical interaction with an object can be sensed through ultrasound while simultaneously tracking the user's motion using inertial sensors.

Dev: It’s a significant piece of research because it tackles the challenge of making high-fidelity sensing truly wearable and interactive in VR environments.

Taro: The paper contributes by showing a viable path for integrating complex sensing modalities into compact, portable devices for autonomy research.

Conclusion: Rosa: We’ve been looking at this paper, "A Wearable Multimodal Ultrasound+Inertial System for Real-Time Virtual Reality Interaction," and the authors are Giusy Spacone, Sebastian Frey, Enzo Baraldi, Mattia Orlandi, Luca Benini, and Andrea Cossettini.

Dev: The title itself really captures the essence of what they achieved: a fully wearable multimodal interface combining ultrasound and inertial sensing for real-time VR interaction.

Taro: It’s interesting to think about the broader impact when we consider what this means for future autonomy research in human-computer interaction, given the capabilities demonstrated.

Rosa: Simply put, this work shows how we can build a system that lets users interact with virtual objects in a way that feels more directly connected to their physical actions through sensing rather than just relying on external tracking devices.

Dev: It moves us toward having interfaces where the sensing and control happen concurrently on the user's body itself.

Taro: That concurrency is what really excites me; it suggests a future where interaction isn't mediated by separate, bulky hardware components in the environment.

Rosa: The implication is that we can expect more sophisticated interactions in VR, allowing for manipulation tasks that require both gross and fine motor control to be executed with greater precision and immediacy.

Dev: We’re looking at systems where the control loop is tightly integrated into the user's physiology through these sensors.

Taro: If this works reliably outside of a lab, it means we could see applications in remote or field environments where external tracking isn't feasible.

Rosa: Overall, the paper presents a system that relies entirely on wearable sensing for hand and arm position control, which is a major step away from previous methods that depended on external optical systems for reference.

Dev: It’s about achieving high-quality interaction performance with a small, low-power device that doesn't require specialized setups.

Taro: This points toward a future where sensing can be embedded into everyday wearables to enable more intuitive and capable digital interactions.

Integrated Systems Laboratory of ETH Zurich

eess.SY, cs.HC, cs.SY

Submitted: 2026-06-16

Updated: 2026-10-05

Comments: 8 pages, 8 figures, 3 tables

Code: https://github.com/pulp-bio/wulpus2https:

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 78/100

The gist: A fully wearable multimodal interface combining ultrasound and inertial sensing enables real-time virtual reality interaction by concurrently estimating hand pose and forearm position.

Key concepts

Ultrasound Sensing (A-mode)
This involves using six ultrasound transducers on the forearm and two on the upper arm to capture real-time images of internal structures. It allows the system to accurately classify forearm position by analyzing these raw US frames, which is a key input for determining where the user's arm is in space.
Inertial Sensing (Accelerometer)
An embedded triaxial accelerometer measures motion and gravity changes on the forearm and upper arm. This data is crucial for tracking hand rotation; by measuring the gravity vector, the system can calculate hand orientation relative to a calibrated reference, helping to determine if a hand is open or rotated.
Late-Fusion Multimodal Learning
This software technique combines information from two different data sources—ultrasound features and accelerometer data—at the end of their processing pipeline. The system uses US features and ACC features together in a neural network to make final, more accurate decisions about hand pose and position.
WULPUS Platform
This is the physical hardware built on which the system runs. It integrates ultrasound sensors, an accelerometer, and BLE communication into a small casing worn at the wrist. It is designed to be fully wearable, dry (no gel needed), and low-power for continuous use.

Terminology

Summary

A fully wearable multimodal interface combining ultrasound and inertial sensing enables real-time virtual reality interaction by concurrently estimating hand pose and forearm position. This system addresses limitations in existing wearable approaches by integrating these modalities into an end-to-end software framework for continuous, functionally meaningful VR interaction.

The gist: A fully wearable multimodal interface based on concurrent US and inertial (accelerometry) sensing from the forearm and upper arm enables real-time VR interaction, achieving task success rates of 92.0±16.0%, 88.0±9.8%, and 96.0±8.0% during online validation with minimal fine-tuning (5 min).

System Components and Acquisition

The system is built on the WULPUS platform, which integrates A-mode ultrasound (US) sensing from six transducers placed on the forearm and two transducers on the upper arm, along with an embedded triaxial accelerometer. The acquisition electronics are worn at the wrist in a PLA casing. The system streams data via BLE, where every packet includes one US frame (397 samples) plus three accelerometer samples. US data are acquired at a pulse repetition frequency (PRF) of 30 Hz using 2.25 MHz transducers, and the WULPUS is enclosed in a casing measuring 31×54×26 mm3 with a weight of 34 g.

Software Architecture and Estimation Pipeline

The software architecture extends the BioGUI framework to support US data, featuring three functional components: Data Acquisition and Visualization, Middleware Bridge, and Application Environment. The middleware bridges the BioGUI and the application environment by performing several key functions:

  1. Hand pose classification using a neural network on raw US data combined with ACC data.

  2. Forearm position classification from raw US frames.

  3. Hand Rotation Tracking based on the gravity vector measured by the accelerometer and a calibrated gravity reference, expressed as ∆φ = arctan(xcal/zcal - arctan x/z).

Multimodal Learning Pipeline and Model Training

A multimodal learning pipeline is introduced for concurrent hand pose (6 classes) and forearm position (3 classes) estimation. The model employs a late-fusion approach where a US feature extractor, consisting of two convolutional blocks with batch normalization, ReLU activation, dropout regularization (rate 0.05), and max pooling along the depth dimension, processes the US input. This flattened US representation is concatenated with three ACC features before being processed by fully connected classification heads to yield the final output classes. Offline data collection involved five subjects performing three functional tasks: cylinder grasping/relocation, marble pinching/relocation, and liquid pouring.

Online Validation and Performance Metrics

Online validation assesses the interface during three real-time functional object-interaction tasks in a Unity-based VR environment. The system controls four degrees of freedom (DoFs) at the hand and wrist based on the estimations. For hand pose, classes are aggregated (e.g., Hand Open and Hand Open with 90° rotation into 'Open'). For forearm position, three discrete states are forwarded directly to Unity. Success Rate (SR) is defined as the percentage of successful trials over all attempts for a given condition, and Completion Time (CT) is computed over successful trials only. The results demonstrated that with only 3 fine-tuning repetitions, task SR values reached 92.0±16.0%, 96.0±8.0%, and 88.0±9.8% for the three tasks, respectively, with corresponding CT values of 10.47±4.73 s, 8.00±1.56 s, and 10.83±2.02 s for Cylinder Grasping and Relocation, Liquid Pouring, and Marble Pinching and Relocation respectively. The system consumes only 19.9 mW, enabling over 2.5 days of continuous use on a small 350 mAh LiPo battery without recharge.

Comparison with Previous Work

The proposed solution is the only one reported to enable truly wearable interaction with VR environments, featuring a smaller and lighter form factor and wireless connectivity. It avoids the need for standard ultrasonic gel, enabling fully dry acquisition. Furthermore, it operates below 20 mW, representing a more than one order of magnitude reduction in power consumption compared to prior systems exceeding >1W. Application-wise, this work extends prior research by enabling object manipulation tasks like liquid pouring and cylinder relocation while introducing pinch grasping, covering both gross and fine motor control. The system relies entirely on wearable sensing for hand and arm position control, unlike previous studies that utilized external optical tracking systems.

Conclusion

In summary, the work presents a fully wearable multimodal sensing system combining US and ACC for concurrent monitoring of the forearm and upper arm using the WULPUS platform.

Improvements for AI systems

Here are specific improvements to AI systems based on the proposed wearable multimodal ultrasound+inertial system, along with what these improved systems can achieve:


  1. The proposed system enables a novel, truly wearable multimodal sensing architecture by fusing A-mode Ultrasound (US) data (for muscle activity/hand shape) with Inertial Measurement Unit (IMU/accelerometry) data (for limb position and rotation).

  2. The AI system can perform real-time, concurrent estimation of two distinct states:

narrow hand pose classification (6 classes: Rest, Open, Closed, Pinch, Pouring, Rotated Hand Open) and 2D forearm position classification (3 classes: Rest, Forward, Side).

  1. The improved system can execute complex functional tasks in a Virtual Reality (VR) environment that require both gross manipulation and fine motor control.

  2. Specifically, the AI-driven system can accurately track and control a virtual hand model to perform:

  • Cylinder Grasping and Relocation (gross motor control).

  • Marble Pinching and Relocation (fine motor control requiring precise finger opposition).

  • Liquid Pouring (requiring controlled forearm rotation mapped to liquid flow simulation).

  1. The system incorporates a multimodal learning pipeline that uses US features fused with ACC data via a CNN backbone, allowing the AI to leverage both tissue deformation information and kinematic motion cues for robust state estimation.

  2. The AI system is designed for online validation, demonstrating the capability to recover high task success rates (e.g., 92% for Cylinder Grasping) with minimal subject-specific fine-tuning (only 3 repetitions), making the AI model highly adaptable to individual user physiology and sensor repositioning.

  3. The system achieves ultra-low power consumption (19.9 mW), enabling continuous, multi-day operation on a small battery, ensuring the AI remains practically viable for long-duration wearable use without frequent recharging.

Related papers