Helios 2.0: A Robust, Ultra-Low Power Gesture Recognition System Optimised for Event-Sensor based Wearables

summary

Video file (mp4)

The gist

This paper presents "an advance in machine learning powered wearable technology: a mobile-optimised, real-time, ultra-low-power gesture recognition model.

In short

The episode discusses 'Helios 2.0,' a robust, ultra-low power gesture recognition system for event-sensor based wearables, developed by Ultraleap Ltd. The team focuses on enabling hands-free gesture control for smart glasses using event cameras to achieve high accuracy at very low power consumption.

Key concepts

Event Cameras
These cameras record only changes in light rather than full video frames. This results in sparse data, which is ideal for low-power processing because the system only records motion events, significantly reducing the amount of data that needs to be processed.
DSP (Digital Signal Processor)
The paper shows that over ninety-nine point eight percent of computation runs on a DSP instead of a general-purpose CPU. This makes the system much more efficient for gesture recognition, allowing it to operate at low power levels like six milliwatts.
Quantization-Aware Training
This is a training method where the neural network learns to work with eight-bit integers from the very beginning, instead of training in floating point and converting later. This optimization allows the model to run efficiently on low-power hardware like a DSP without losing accuracy.
Gesture Sequences
The system learns long sequences of gestures, up to two seconds long, using a Markov chain. This helps the model distinguish between intentional movements and accidental hand movements by learning what gesture follows another.

Terminology used across episodes

This episode discusses

The paper

Helios 2.0: A Robust, Ultra-Low Power Gesture Recognition System Optimised for Event-Sensor based Wearables · Read on arXiv

Prarthana Bhattacharyya, Joshua Mitton, Ryan Page, Owen Morgan, Oliver Powell, Benjamin Menzies, Gabriel Homewood, Kemi Jacobs, Paolo Baesso, Taru Muhonen, Richard Vigars, Louis Berridge

Ultraleap Ltd.

We present an advance in wearable technology: a mobile-optimized, real-time, ultra-low-power event camera system that enables natural hand gesture control for smart glasses, dramatically improving user experience. While hand gesture recognition in computer vision has advanced significantly, critical challenges remain in creating systems that are intuitive, adaptable across diverse users and environments, and energy-efficient enough for practical wearable applications. Our approach tackles these challenges through carefully selected microgestures: lateral thumb swipes across the index finger (in both directions) and a double pinch between thumb and index fingertips. These human-centered interactions leverage natural hand movements, ensuring intuitive usability without requiring users to learn complex command sequences. To overcome variability in users and environments, we developed a novel simulation methodology that enables comprehensive domain sampling without extensive real-world data collection. Our power-optimised architecture maintains exceptional performance, achieving F1 scores above 80% on benchmark datasets featuring diverse users and environments. The resulting models operate at just 6-8 mW when exploiting the Qualcomm Snapdragon Hexagon DSP, with our 2-channel implementation exceeding 70% F1 accuracy and our 6-channel model surpassing 80% F1 accuracy across all gesture classes in user studies. These results were achieved using only synthetic training data. This improves on the state-of-the-art for F1 accuracy by 20% with a power reduction 25x when using DSP. This advancement brings deploying ultra-low-power vision systems in wearable devices closer and opens new possibilities for seamless human-computer interaction.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Helios 2.0: A Robust, Ultra-Low Power Gesture Recognition System Optimised for Event-Sensor based Wearables".

Jane: The paper was written by Prarthana Bhattacharyya, Joshua Mitton, Ryan Page, Owen Morgan, Oliver Powell et al. from Ultraleap Ltd..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making waves in the wearable tech world: “Helios two point zero: A Robust, Ultra-Low Power Gesture Recognition System for Event-Sensor based Wearables.” Jane, this one’s from the team at Ultraleap, and honestly, the title alone tells you they’re solving a real problem.

Jane: Absolutely, Tom. And I love that they’re not just talking about gesture recognition in general — they’re specifically targeting smart glasses. You know, those devices that are supposed to be hands-free but usually end up making you tap on the side of your head like you’re adjusting a pair of glasses that aren’t there.

Tom: Right? And that’s the whole pitch here. Instead of capacitive touch on the temple or wearing a separate ring, they’ve mounted an event camera right on the glasses frame. The camera watches your hand, and the model figures out what you’re doing — swiping left, swiping right, or pinching — all without you ever touching the device.

Jane: And the “event-sensor” part is key. These aren’t regular cameras taking video frames. Event cameras only record changes in light — like, when your finger moves, the pixels that see that motion fire off a signal. Everything else stays quiet. That means the data is sparse, which is perfect for low-power processing.

Lu: And that’s exactly why this paper matters. I’m Lu, by the way, for anyone just tuning in. The whole field of wearable computing has been stuck on this trade-off: you want the device to be always-on, but a regular camera running continuously would drain the battery in minutes. Event cameras flip that equation because they’re only processing what actually changes.

Meng: But here’s the thing, Lu — processing sparse data still takes compute. And the paper’s real trick is that they’ve designed a model where over ninety-nine point eight percent of the computation runs on a DSP, a digital signal processor, which is way more efficient than a general-purpose CPU. They’re getting the whole thing down to about six milliwatts.

Jane: Six milliwatts. Let’s put that in perspective. A typical smartphone camera might use a couple of watts. A smartwatch screen uses maybe fifty milliwatts. This is basically nothing — you could run this gesture recognition for days on a tiny battery.

Tom: And that’s the headline. But what really got me excited is that they didn’t just make it low-power — they made it accurate. They’re reporting F1 scores above eighty percent across multiple users and environments, which is a huge jump from their previous version.

Lu: The previous version, Helios one point zero, was already a proof of concept. But this one is a full product-grade system. They’ve got a simulator that generates synthetic training data, they’ve got a quantization-aware training pipeline, and they’ve tested it on real people in real environments. It’s the difference between a lab demo and something you could actually ship.

Meng: And that’s what I want to dig into — how they actually got the accuracy up while keeping the power down. Because that’s the hard part. But maybe we should save that for the next segment.

Jane: Good call. We’ll get into the methodology in a moment. For now, just remember the name: “Helios two point zero” — gesture control for smart glasses, running at six milliwatts. That’s the headline.

Summary: Tom: So we’ve established that “Helios two point zero” is a big deal for smart glasses. But let’s talk about what the paper actually does, because the summary is pretty dense. Jane, can you break down the core idea for someone who hasn’t read it?

Jane: Sure. The paper’s main contribution is a complete pipeline for gesture recognition on wearable devices. It starts with a simulator that generates synthetic training data — because collecting real-world data from hundreds of users in different lighting conditions is expensive and slow. Then they train a neural network on that synthetic data, and finally they deploy it on a low-power chip.

Lu: And the simulator is actually one of the most interesting parts. They’re using a hand-tracking system to record real human hand motions, then they render those motions in a three dee environment with different lighting and backgrounds. They even model the event camera’s behavior — how it responds to contrast and motion — so the synthetic events look like real events.

Meng: But here’s the thing that impressed me: they don’t just generate single gestures. They generate long sequences — two seconds long — with multiple gestures in a row, using a Markov chain to decide what comes next. So the model learns not just what a swipe looks like, but what a swipe followed by a return to rest looks like. That’s how you avoid false positives.

Jane: Exactly. And that’s a huge deal for real-world usability. Imagine you swipe right to go to the next menu, and then when you bring your hand back to rest, the system thinks you swiped left and goes back. That would be infuriating. So they added “return” classes — like swipe-left-return — so the model learns to distinguish between an intentional gesture and just moving your hand back.

Tom: And they also expanded the gesture set from seven classes to ten. They’ve got pinch, double pinch, swipe left, swipe right, and the return versions of those, plus unknown, untracked, and rest. That gives the model a much richer understanding of what’s happening.

Lu: And the training methodology is clever too. They use quantization-aware training, which means the model learns to work with eight-bit integers from the start, rather than being trained in floating point and then converted. That’s what lets them run on the DSP without losing accuracy.

Meng: Right, and that’s the key to the power savings. The DSP can do eight-bit integer math much faster and more efficiently than a CPU doing thirty-two-bit floating point. They’re getting the same accuracy with a fraction of the energy.

Jane: And then they tested it on real people. Twenty users in a controlled office environment, then an expert user in four different home environments, and then a group of users outdoors. And the model held up — F1 scores above eighty percent across the board, even in the outdoor setting with three thousand lux of sunlight.

Tom: Which is wild because event cameras are supposed to struggle in bright light. But the paper shows the model actually performed better outdoors. Jane, why do you think that is?

Jane: The paper suggests it’s because the gesture signals are stronger against natural backgrounds — there’s more contrast between the hand and the environment. And in the controlled office, the background was static, so the model might have been relying on background cues rather than the hand itself. Outdoors, it had to focus on the actual motion.

Lu: That’s a really interesting finding. It means the model is learning the gesture, not the background. That’s exactly what you want for a wearable device that’s going to be used in all sorts of places.

Meng: And the latency is just two point three five milliseconds. That’s imperceptible. You swipe, and the response is instant. That’s what makes it feel natural.

Tom: So we’ve got the summary: synthetic data, smart training, low-power deployment, and real-world validation. But I want to know more about the specific improvements over the previous version. That’s our next topic.

Improvements: Tom: We’ve covered what “Helios two point zero” does. Now let’s talk about what’s actually new compared to the original Helios. Jane, what are the big upgrades?

Jane: The biggest one is the power consumption. Helios one point zero ran at three hundred fifty milliwatts on a CPU. Helios two point zero runs at six milliwatts on a DSP. That’s a twenty-five times reduction. And the accuracy went up by twenty percent — from around sixty-five percent F1 to over eighty percent. So they made it both faster and more accurate at the same time.

Lu: And the architecture is completely redesigned. It’s now a five-stage model. The first stage downsamples the input, the second stage does feature extraction on the low-resolution image, the third stage predicts a bounding box and crops the hand, the fourth stage does detailed feature extraction on that crop, and the fifth stage combines everything to make the final gesture prediction.

Meng: And the clever part is that stages two and four — the heavy convolutional layers — are quantized to eight-bit integers and run on the DSP. Stages one, three, and five are lightweight and run in floating point on the CPU. That split is what lets them get the power down without sacrificing accuracy.

Jane: And they also improved the training data. Instead of single gestures in short windows, they’re using two-second sequences with up to six gestures. That’s a much harder task for the model, but it learns to handle transitions between gestures, which is what happens in real life.

Tom: And they added that fine-tuning step with rotation augmentation. They realized the synthetic data didn’t have enough variety in hand orientations, so they rotated the sequences by twenty-five to forty degrees and fine-tuned the model. That made a big difference in real-world performance.

Lu: The ablation studies are really informative too. They tried replacing the standard convolutions with depthwise separable convolutions to reduce the model size, but the accuracy dropped. They tried adding more layers, but it didn’t help. So they found a sweet spot — a model that’s small enough to be efficient but big enough to actually learn the gestures.

Meng: And they also tested different quantization training strategies. Training from scratch with quantization-aware training worked just as well as training a full-precision model and then quantizing it. That’s a nice result because it simplifies the pipeline.

Jane: And the simulator itself got an upgrade. They’re using a Markov chain to decide which gestures come next, which means the training data has realistic sequences rather than random gestures. They also added kinematic blending between gestures, so the transitions look natural — not like the hand is teleporting from one position to another.

Tom: And that matters because if the training data looks artificial, the model will learn artificial patterns. By making the synthetic data more realistic, they’re closing the sim-to-real gap.

Lu: And the results on real-world data confirm it. They tested on a human variability dataset with twenty users, a scene variability dataset with different home environments, and an outdoor dataset. The model performed consistently across all of them, which is exactly what you need for a product.

Meng: The outdoor result is particularly impressive. Most event-based systems struggle in bright sunlight, but they got F1 scores above zero point eight eight for right swipe and zero point eight four for left swipe. That’s better than the indoor results in some cases.

Jane: And the latency is just two point three five milliseconds. That means the system can respond to a gesture before you even finish making it. It feels instant, which is crucial for something like navigating a menu on your glasses.

Tom: So the improvements are clear: lower power, higher accuracy, better training data, and a smarter architecture. But what does this mean for the future? That’s where I want to go next.

Conclusion: Tom: We’ve spent this whole episode on “Helios two point zero: A Robust, Ultra-Low Power Gesture Recognition System for Event-Sensor based Wearables.” And I think we can all agree this is a major step forward for wearable computing. Jane, how would you sum it up?

Jane: I’d say the paper proves that event-based vision can be practical for always-on wearable devices. They’ve shown you can get accurate gesture recognition at just six milliwatts of power, with two point three five milliseconds of latency, and it works across different users, environments, and lighting conditions. That’s a complete package.

Lu: And the implications go beyond smart glasses. If you can run gesture recognition this efficiently, you can run other vision tasks too — object detection, hand tracking, maybe even scene understanding. This architecture could be a template for a whole class of low-power vision systems.

Meng: From an engineering standpoint, the most exciting part is the quantization-aware training pipeline. They’ve shown you can train a model that’s optimized for a specific chip from the start, rather than training a general model and trying to squeeze it onto the hardware later. That’s a much more efficient workflow.

Jane: And the simulator is a huge contribution too. Being able to generate realistic training data without collecting thousands of hours of real-world video is a game-changer. It means you can iterate on the model design quickly, test new gestures, and adapt to new environments without expensive data collection campaigns.

Tom: And for the user experience, this could finally make smart glasses feel natural. No more tapping your temple or wearing a separate ring. You just gesture with your hand, and the glasses respond. It’s the kind of interaction that feels like magic but is actually just really good engineering.

Lu: And I think the cultural impact is worth noting. As these devices become more accessible, the way we interact with technology changes. Gesture control is more intuitive than touch — it’s how we communicate with each other. So this could make wearable tech feel less like a gadget and more like an extension of yourself.

Meng: The main limitation right now is the gesture vocabulary — three microgestures is pretty limited. But the framework they’ve built can scale. Add more gestures to the simulator, train the model, and you’ve got a richer interaction language.

Jane: And that’s what I’m excited about for the future. This paper is a foundation, not a final product. The team at Ultraleap has shown what’s possible, and now it’s up to the rest of the field to build on it.

Tom: Well said. “Helios two point zero” is a paper we’ll be referencing for a while. It’s a clear demonstration that low-power, real-time gesture recognition is not just possible — it’s ready for prime time. Thanks for joining us, and we’ll see you next time with another exciting paper.

More episodes

← Home