Real-Time sEMG-Based Telecontrol of an Assistive Robotic Arm Using a 1D Convolutional Neural Network

arXiv:2607.16310 · cs.RO · Submitted 2026-07-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Real-Time sEMG-Based Telecontrol of an Assistive Robotic Arm Using a 1D Convolutional Neural Network".

Dev: Real-time sEMG-based telecontrol of an assistive robotic arm using a 1D Convolutional Neural Network addresses the challenge of providing intuitive, reliable,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: Well, according to what we've read from "Real-Time sEMG-Based Telecontrol of an Assistive Robotic Arm Using a 1D Convolutional Neural Network," the thesis is that a system using multichannel sEMG signals and a classifier can enable reliable and sufficiently responsive control of an assistive robotic arm through several discrete commands <ref:2607.16310#pg0,Real-Time sEMG-Based Telecontrol of an Assistive Robotic Arm Using a>. Dev That seems to be the big claim, Rosa; they are proposing this pipeline covers everything from signal acquisition all the way to robotic action, which is pretty comprehensive. Taro It matters because it addresses the need for more intuitive interfaces than just joysticks or position sensors when upper limb function is limited in daily living activities.

Rosa: Exactly, Taro; these simple actions like reaching or grasping become much harder without assistance, and this paper suggests sEMG provides a viable way forward for compensating for that loss of function. Dev The authors claim that by combining coherent preprocessing, appropriate segmentation, a robust muscle activation detection strategy, and a stabilized decision logic, they can achieve interpretable and usable robotic behavior under semi-real-time conditions.

Taro: I'm thinking about the broader implications here; if this kind of control becomes feasible for more people with motor impairments in daily life scenarios outside the lab, it could significantly reduce dependence on caregivers for a wider range of activities. Rosa That’s a big picture thought, Taro; it moves the technology closer to actual assistive care rather than just academic simulation.

Dev: From an engineering standpoint, the paper emphasizes that this system is designed to function under semi-real-time conditions, which speaks to the practical constraints of deploying such a system in a functional environment. Taro Does that mean we're talking about latency levels that are acceptable for someone actively trying to use the arm?

Rosa: The authors specifically test this pipeline on a real robot, showing that it achieves control through several discrete commands based on muscle activation detection. Dev So, the core message is demonstrating a complete end-to-end system capable of translating muscle activity into specific actions for an assistive robotic arm.

Conclusion: Rosa: Looking at the full title, "Real-Time sEMG-Based Telecontrol of an Assistive Robotic Arm Using a 1D Convolutional Neural Network," it really highlights how they combine signal processing, machine learning, and robotics to achieve this control <ref:2607.16310#pg0,Real-Time sEMG-Based Telecontrol of an Assistive Robotic Arm Using a>. Dev And the authors are clearly trying to show that this specific combination—the 1D CNN applied to filtered sEMG signals—is effective for achieving reliable control through several discrete commands <ref:2607.16310#pg0>. Taro What I find interesting is how they've managed the transition from simulation to a real robot, which suggests a level of robustness they aimed for in the design.

Rosa: It really shows the feasibility of using non-invasive surface electromyography to create an intuitive interface for people with upper limb motor impairments by translating muscle activity into robot movements. Dev The implications are that this approach could lead to assistive devices that are much more responsive and usable than current visual or joystick controls, provided the system can handle real-world noise and variations in user condition effectively.

Taro: If we look ahead, the challenge they flag is the difficulty in distinguishing certain similar gestures, like ulnar and radial deviations, which points toward where future work needs to go for greater practical utility. Rosa That limitation suggests that while they've shown a functional pipeline now, improving gesture differentiation will be key to making this technology truly versatile for a wider set of human movements.

Dev: My focus is on the real-time aspect; the latency measurement they provided, which was approximately zero point three two four seconds with a standard deviation of zero point zero one four six seconds, suggests that while it's stable enough for simple tasks, we have room to optimize that loop rate if we were targeting faster manipulation. Rosa So, to sum up the impact: this paper confirms that sEMG telecontrol is possible and provides a concrete architecture for achieving responsive control in this domain.

Taro: And from an autonomy research view, the next step is making sure this system can handle unexpected situations when the user misbehaves or when external factors interfere with their movement.

Dev: That’s where we need to look at those hybrid control approaches they mentioned in the introduction to improve robustness.

Department of Mechanical Engineering, Polytechnique Montréal

cs.RO

Submitted: 2026-07-14

Updated: 2026-10-01

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 83/100

The gist: Real-time sEMG-based telecontrol of an assistive robotic arm using a 1D Convolutional Neural Network addresses the challenge of providing intuitive, reliable, and responsive control for individuals

Key concepts

sEMG Acquisition
This involves using four electrodes on the skin to measure electrical activity generated by muscles. The system used a Delsys Bagnoli-4 amplifier to capture these signals at 2000 Hz. This raw muscle signal is the input data that the AI model analyzes to understand user intent.
1D CNN
A one-dimensional Convolutional Neural Network is a type of deep learning model designed to process sequences of data, like time-series signals from sEMG. It scans the signal windows to automatically learn patterns associated with specific hand gestures, making it effective for classifying complex muscle movements.
Onset Detection Algorithm
This method identifies the exact moment a muscle movement begins in the signal. The system used an adaptive threshold based on the median and standard deviation of the signal envelope to determine when a gesture starts, ensuring that only meaningful movement segments are processed by the classifier.

Terminology

Summary

Real-time sEMG-based telecontrol of an assistive robotic arm using a 1D Convolutional Neural Network addresses the challenge of providing intuitive, reliable, and responsive control for individuals with upper limb motor impairments by developing a complete real-time pipeline that translates muscle activity into robotic commands. The gist: A complete real-time pipeline covering acquisition, preprocessing, muscle activation detection, window segmentation, normalization, convolutional neural network inference, and robotic control enables reliable and sufficiently responsive control of an assistive robotic arm through several discrete commands.

System Architecture and Data Acquisition

The system relies on four-channel sEMG acquisition using a Delsys Bagnoli-4 EMG amplifier with a bandwidth of 20 Hz to 2000 Hz, sampled at 2000 Hz via a National Instruments cDAQ-9171 CompactDAQ chassis. The robot utilized for testing is the UFACTORY xArm7, a 7-DoF robotic arm. Numerical simulations were created using Robotics Toolbox for Python as a replica of the xArm7. The sEMG signals were preprocessed in real time by applying a 60 Hz notch filter to attenuate power-line interference, followed by a 6th-order Butterworth band-pass filter between 5 and 500 Hz. This preprocessing maintains methodological consistency with the training data.

Gesture Classification Pipeline

The core of the classification stage involves developing a model capable of classifying hand gestures from multichannel sEMG signals. The dataset included six classes: rest, extension, flexion, ulnar deviation, radial deviation, and grip. A crucial step in preparing the data was an onset-detection algorithm based on an adaptive threshold applied to the signal envelope, which was selected over a local-maxima percentile method because it proved better. The threshold was defined as threshold = median + k · σ, where k=2.5, and an onset required exceeding this threshold for a minimum duration of 20 ms plus a 50 ms margin.

Data Segmentation and Model Selection

The continuous signal is segmented into sliding windows to create independent training instances. The study compared three windowing configurations: a 250 ms window (50% overlap), a 300 ms window (75% overlap), and a 350 ms window (75% overlap). The configuration retained was the 300 ms/75% because it achieved the best test performance while maintaining low latency. For classification, a 1D CNN taking filtered signal windows as input was selected over RMS envelope variants or LSTM networks, as it preserves the raw nature of the signal while benefiting from prior filtering and onset-based segmentation. The final model utilized a lightweight yet sufficiently deep 1D CNN with three successive convolutional blocks.

Real-Time Control Strategy

The system employs a threshold-based detection strategy, where movement is detected when activity consistently exceeds the threshold, triggering window processing by the classifier. The classification decision logic uses a sliding 'majority vote' approach over the last three windows to determine the command sent to the robot. This approach is used for Cartesian velocity control based on discrete gestures:


Rest: no movement

Ulnar deviation: motion along +Y

Extension: motion along +Z

Radial deviation: motion along −Y

Grip: dedicated gripper

Performance Evaluation and Results

The final 1D-CNN model achieved an overall test accuracy of 0.9055 on the test set, with class-wise F1-scores ranging from 0.8596 to 0.9388 (Table 9). When tested on experimentally acquired sEMG signals, the average recognition rates were high: Extension and flexion exhibit the best performance, while some confusions appear for lateral gestures (ulnar and radial). The system latency was measured to be approximately 0.324 s with a standard deviation of 0.0146 s, reflecting stable behavior compatible with real-time operation for simple tasks. The majority vote mechanism was found to be effective in limiting occasional classification errors.

Conclusion and Limitations

The work demonstrates the feasibility of sEMG-based telecontrol, confirming that the 1D CNN applied to the raw filtered signal is the most effective architecture. However, limitations persist: the difficulty in distinguishing certain similar gestures, particularly lateral movements such as ulnar and radial deviations, remains a challenge. Furthermore, performance is highly sensitive to experimental conditions like electrode placement, contact quality, noise levels, and the user’s muscle condition. Future work suggests exploring hybrid control approaches combining sEMG with other sources of information e.g., integrating vision sensors or environmental perception. The overall system enables intuitive and responsive robot control in a real-time setting.

Improvements for AI systems

Here are specific improvements to an AI system based on this research, detailing what the improved system can achieve:


  1. The current system relies on a fixed 300ms window size with 75% overlap and a majority-vote over the last three windows for control. This introduces a minimum latency of approximately 0.324s and relies on discrete classification, which can lead to jerky movements or missed transitions (e.g., during rapid changes).

  2. The improved AI system should implement an end-to-end, continuous decision pipeline that moves beyond fixed windowing and discrete commands to achieve fluid, natural motion.

  3. The improved AI system can perform:

4.1 Continuous, low-latency trajectory generation: Instead of waiting for a full window classification (0.324s latency), the system should utilize a recurrent architecture (like an LSTM or a hybrid CNN-RNN) trained on sequential windows to predict the next few control steps continuously. This allows for velocity/position commands to be updated at a much higher frequency (e.g., 50-100Hz), mimicking natural human motor control, significantly reducing perceived latency and improving smoothness.

4.2 Adaptive Control Strategy: Integrate a dynamic switching mechanism between the three real-time strategies (Threshold-based detection, Two-stage classification, Direct six-class classification). This switch should be triggered based on real-time signal quality metrics (e.g., signal variance or confidence scores from the CNN) to ensure robustness. For example, if the system detects high inter-subject variability or noise spikes (as seen in Table 10 for radial gesture), it automatically reverts to a more conservative, robust strategy like the threshold-based method until conditions stabilize.

4.3 Proactive Error Correction via Uncertainty Estimation: Enhance the CNN model to output not just a class prediction but also an explicit uncertainty score (e.g., using Bayesian Neural Networks or Monte Carlo Dropout). The improved system can then use this uncertainty score to modulate control authority—if the system is highly uncertain about the current gesture (high uncertainty), it can dampen robotic velocity or trigger a brief hold state, preventing erroneous movements caused by transient noise or classification errors, thereby improving control stability over the majority-vote mechanism.

4.4 Context-Aware Control Integration (Hybrid Sensing): The system should be upgraded to incorporate an auxiliary input modality, such as a simple vision sensor tracking the end-effector position. The AI can then use this visual data to perform contextual correction, allowing the control system to adjust Cartesian velocity commands in real-time based on where the arm is currently located in space, making movements more goal-directed and reducing errors caused by purely muscle-driven trajectory generation.

These improvements transform the system from a discrete, latency-bound teleoperation tool into a high-bandwidth, robust, and contextually aware assistive control interface.

Abstract

Motor impairments affecting the upper limb significantly reduce autonomy in daily activities, particularly for tasks involving object manipulation. Assistive robotic arms offer a promising solution, provided they can be controlled in an intuitive, reliable, and responsive manner. Among human--machine interface approaches, surface electromyography (sEMG) enables non-invasive access to muscle activity and thus to the user's motor intentions. This work proposes a real-time sEMG-based interface for the teleoperation of an assistive robotic arm. The system relies on four-channel sEMG acquisition, signal preprocessing, segmentation into sliding windows, and classification using a one-dimensional convolutional neural network (CNN). Several real-time strategies are investigated, including threshold-based onset detection, a two-stage classification approach (rest vs movement followed by gesture recognition), and a single classifier handling both rest and five gestures. The complete pipeline is implemented and evaluated both in simulation and on a real robotic platform. The CNN-based approach achieves high classification performance, with a test accuracy above 90% and strong generalization on experimentally acquired signals. The system exhibits stable real-time behavior, with an average latency of approximately 0.32 s consistent with the chosen windowing strategy, and the robot can be controlled reliably using discrete gestures, producing coherent and smooth movements in both simulated and real environments. These findings demonstrate the feasibility of sEMG-based telecontrol for assistive robotics and highlight the importance of integrating signal processing, deep learning, and control strategies within a unified real-time framework. Future work may explore hybrid control approaches combining sEMG with additional sensing modalities to further improve robustness and usability.

Related papers