A 2-Block Architecture for Real-Time EEG Gait Decoding: A Pilot Study

arXiv:2608.02083 · cs.LG, cs.HC, eess.SP · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "A 2-Block Architecture for Real-Time EEG Gait Decoding: A Pilot Study".

Jane: The paper was written by Shantanu Sarkar, Saurabh Prasad and Jose L. Contreras-Vidal from University of Houston.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone! Today we're digging into a brand new paper that just hit the arXiv servers, and it's called "A two-BLOCK ARCHITECTURE FOR REAL-TIME EEG GAIT DECODING: A PILOT STUDY." Jane, I have to say, the title alone got me excited — real-time EEG gait decoding is one of those things that sounds like science fiction but is actually happening right now.

Jane: It really is, Tom! And when I first saw the title, I thought, okay, two blocks — that's a nice, clean way to think about a brain-computer interface. The first block is all about cleaning up the messy brain signals, and the second block is the decoder that figures out what the person wants to do. It's like having a really good noise-canceling microphone and then a translator that actually understands what's being said.

Tom: Exactly! And the team behind this is from the University of Houston — Shantanu Sarkar, Saurabh Prasad, and Jose Luis Contreras-Vidal. Contreras-Vidal is a big name in the exoskeleton and BCI world, so when I saw his name on this, I knew we were in for something serious.

Jane: And the pilot study part is important too. This isn't a huge clinical trial with hundreds of participants — it's one healthy participant, twenty-two years old, doing ten sessions. But for a pilot, what they're testing is whether the whole pipeline can actually work in real time, not just in offline analysis where you can take your time and clean up the data later.

Tom: Right, and that's the huge leap. Most EEG studies are done offline — you record the brain signals, then you go back to your computer and spend hours processing them. But this paper is about closed-loop control, which means the exoskeleton is actually responding to the person's brain in real time. That's the difference between a lab experiment and a device that could help someone walk.

Jane: And the title hints at the four gait states they're decoding — Stand, Initiate, Execute, and Terminate. That's already more sophisticated than the usual walk/stop binary that most studies use. Walking isn't just on or off; there's a whole sequence of intentions, and this paper tries to capture that.

Tom: So when we say "two-block architecture," we're really talking about a modular system where you can swap out parts and retrain them separately. That's a big deal for practical deployment, because the brain signals change from session to session, and you don't want to retrain the whole thing every time.

Jane: And that's exactly what we're going to dig into in the next segment — how that architecture actually works and what the results look like. Stick around, because this is where it gets really interesting.

Summary: Tom: So we're back, still talking about "A two-BLOCK ARCHITECTURE FOR REAL-TIME EEG GAIT DECODING: A PILOT STUDY." Jane, let's break down what this paper actually did, because the summary is pretty dense.

Jane: Okay, so the big picture is this: they built a system that takes raw EEG from twenty-eight channels placed on the head, cleans it up in real time, extracts features from it, and then classifies what the person is trying to do — stand, start walking, keep walking, or stop. And they did it with a participant wearing the Rex exoskeleton, which is this big, fully powered lower-limb device.

Tom: And the cleaning part is crucial, because EEG is notoriously noisy. You've got eye blinks, muscle activity, movement artifacts — all of that gets mixed in with the brain signals. The paper uses a combination of methods: H-infinity filtering for ocular artifacts, a band-pass filter, and then something called nASR, which is a neural Artifact Subspace Reconstruction layer. That last one is trainable, which is clever — it learns which channels are contaminated and reconstructs them from the clean ones.

Jane: Right, and then they extract features in multiple domains. They group the twenty-eight channels into nine regions of interest based on anatomy — frontal, central, parietal, occipital areas — and they use a depth-wise convolution to combine channels within each region. Then they do a redundant discrete wavelet transform to split the signal into frequency bands: delta, theta, alpha, beta, and gamma. They throw out gamma because that's where muscle artifacts live.

Tom: And that's the Feature Extraction Block. Then the Decoder Block takes those features and runs them through three parallel branches, one for each frequency band. Each branch has this new layer they invented called PolyTVL — Polynomial Time-Varying Layer — followed by average pooling and an LSTM. The outputs get concatenated and passed through a dense network that outputs probabilities for the four gait states.

Jane: And the key innovation here is the PolyTVL. Most state-space models, like the S4 or Mamba architectures that are popular right now, are linear time-invariant — meaning the system doesn't change over time. But brain signals are non-stationary; they change constantly. PolyTVL introduces learnable polynomial transforms and a time-varying state transition matrix, so it can capture those nonlinear dynamics.

Tom: And they compared four versions of the decoder: PolyTVL with a dense layer, PolyTVL with LSTM — that's the proposed one — S4D-Lin with dense, and LSTM with dense. The PolyTVL+LSTM version won. Validation Matthews Correlation Coefficient of zero point four three five, and the gap between training and validation was the smallest at zero point one eight seven, meaning it generalized the best without overfitting.

Jane: And the closed-loop results — this is where it gets real. In the closed-loop sessions, the exoskeleton was actually triggered by the brain signals. They got a fifty-five point three percent success rate for Rex-assisted gait initiation and fifty-two point seven percent for volitional walking without the exoskeleton. And the prediction time was seventy point five milliseconds on average — that's fast enough for real-time control.

Tom: Now, those success rates might sound low, but you have to remember — chance level is around eighteen percent. So they're well above chance, and for a pilot study with one participant, that's actually promising. It proves the pipeline works, and now the question is how to make it more accurate.

Jane: And that's exactly what we'll talk about next — what improvements they're suggesting and where this research is heading.

Improvements: Tom: We're still on "A two-BLOCK ARCHITECTURE FOR REAL-TIME EEG GAIT DECODING: A PILOT STUDY," and Jane, I want to get into what the authors say could be improved, because they're pretty honest about the limitations.

Jane: They are, and that's refreshing. The most obvious limitation is that it's a single participant. One healthy twenty-two-year-old male. That means we don't know how well this generalizes to other people, let alone to patients with spinal cord injuries or stroke survivors, who are the ultimate target population. The brain signals could be very different after injury.

Tom: And they also mention that some sessions performed worse than others. Session three and Session ten were notably lower, and they attribute that to the participant rushing — trying to finish early. That's a real-world problem. Motivation and attention affect EEG quality, and in a clinical setting, you can't always control that.

Jane: Right, and the hyperparameters were set heuristically. They say systematic optimization is future work. So things like the learning rate, the batch size, the number of hidden states in PolyTVL — those were chosen based on experience and intuition, not a formal search. There's probably room to squeeze out better performance with a proper hyperparameter sweep.

Tom: And there's also the penalty matrix they used in the loss function. They penalize certain misclassifications more than others — like confusing Initiate with Terminate is heavily penalized because that could cause the exoskeleton to stop when the person wants to start. But the specific values in that matrix were also set heuristically.

Jane: Another improvement they hint at is the wavelet choice. They used Symlet-two which is a reasonable choice for motor imagery, but they've done other work on selecting optimal mother wavelets, so there might be a better option for gait specifically.

Tom: And I think the biggest improvement they're pointing toward is multi-subject validation. They need to show this works across a diverse group of people — different ages, different genders, different neurological conditions. That's the difference between a pilot study and a real product.

Jane: And the closed-loop success rate of around fifty-five percent — they want to push that higher. One way might be to make the Initiate window detection smarter. Right now, they require at least two Initiate predictions within ten decoding windows. Maybe a more sophisticated decision rule could reduce false positives while catching more true intentions.

Tom: Lu, you've been quiet — what do you think about the improvements from a research perspective?

Lu: I think the most exciting direction is making PolyTVL even more expressive. The polynomial orders they learned — O1 and O2 — were close to one for the low-frequency bands but went up to about one point two for beta. That suggests the nonlinearity matters more at higher frequencies. If they can make the polynomial order adaptive per timestep, not just per band, that could capture even more of the brain's dynamics.

Meng: And from an engineering standpoint, I'd want to see the model run on embedded hardware, not just a laptop with an Intel i7. The seventy point five millisecond prediction time is great, but if you're putting this in a wearable exoskeleton, you need it to run on a small, low-power processor. That's a whole different optimization problem.

Jane: Great points from both of you. So the improvements are clear — more participants, better hyperparameters, smarter decision rules, and maybe a more flexible PolyTVL. But the foundation is solid, and that's what matters for a pilot.

Conclusion: Tom: And that brings us to the end of our discussion on "A two-BLOCK ARCHITECTURE FOR REAL-TIME EEG GAIT DECODING: A PILOT STUDY." Jane, give us the final takeaway.

Jane: The takeaway is that this paper proves real-time closed-loop EEG gait decoding is feasible. They built a two-block system — one for cleaning and feature extraction, one for decoding — and they showed it can drive an exoskeleton in real time with a single participant. The PolyTVL layer is a novel contribution that addresses the non-stationary nature of brain signals, and it outperformed the alternatives they tested.

Tom: And while the success rates of around fifty-five percent are modest, they're well above chance, and this is just the beginning. The architecture is modular, so you can improve each block independently. That's a smart design for iterative development.

Jane: The implications for the world are significant. If this technology matures, it could give people with paralysis or spinal cord injuries a way to control exoskeletons with their thoughts — not in a lab, but in daily life. That would be life-changing for so many people.

Tom: And the authors are already looking at multi-subject studies, which is the natural next step. We'll be watching for that follow-up paper.

Jane: Absolutely. So we're saying goodbye to this paper, but we're not saying goodbye to the field. Next up, we've got another exciting paper to discuss, so stay tuned.

Tom: Thanks for listening, everyone. We'll be right back.

Shantanu Sarkar, Saurabh Prasad, Jose L. Contreras-Vidal

University of Houston

cs.LG, cs.HC, eess.SP

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: Accepted for publication in the 2026 IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2026), September 28-October 1, 2026, Atlanta, GA, USA. Camera-ready version

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 58/100

Key concepts

EEG Gait Decoding
This process uses electroencephalography (EEG) signals recorded from the head to classify a person's intended movement states—such as standing, initiating, or executing gait. It aims to translate brain activity into actionable commands for devices like exoskeletons.
Two-Block Architecture
A modular system designed for processing EEG data. The first block cleans the noisy raw brain signals and extracts relevant features. The second block (the decoder) uses those cleaned features to classify the user's intended action or gait state.
PolyTVL Layer
A novel component used in the decoding block that addresses non-stationary brain signals. It introduces learnable polynomial transforms and a time-varying state transition matrix, allowing it to capture complex, changing dynamics in the brain data.
Closed-Loop Control
A system where an external device (like an exoskeleton) responds to the user's brain signals in real time. This is a significant leap from offline analysis, moving toward functional devices that assist with movement.

Terminology

Summary

Summary

This paper proposes a 2-block Brain-Computer Interface (BCI) architecture for real-time EEG-based gait decoding to enable closed-loop lower-limb exoskeleton control. The authors identify that existing EEG-based exoskeleton control is limited by motion artifacts, low signal-to-noise ratio, and binary gait formulations (walk/stop) that fail to capture the full cortical complexity of gait. The proposed architecture consists of two blocks: a trainable session-specific Feature Extraction Block and a Decoder Block built on a novel Polynomial Time-Varying Layer (PolyTVL) coupled with an LSTM for four-state gait classification (Stand, Initiate, Execute, Terminate).

The Feature Extraction Block performs real-time artifact suppression using H∞ filtering (leveraging EOG references), band-pass filtering (0.5–30 Hz), and a neural Artifact Subspace Reconstruction layer (nASR). It then extracts multi-domain features via ROI-based depth-wise convolution across nine anatomically defined ROIs spanning 28 channels, followed by an adaptive Redundant Discrete Wavelet Transform (RDWT) filtering layer implemented as a trainable neural network layer. The RDWT decomposes each signal (sampled at 100 Hz) into four sub-bands: δ, θ (0–6.25 Hz), α (6.25–12.5 Hz), β (12.5–25 Hz), and γ (>25 Hz). The γ sub-band is discarded prior to decoding due to EMG contamination above 25 Hz, leaving three physiologically relevant spectral representations.

The Decoder Block receives the extracted features through three parallel branches (one per frequency band: δ/θ, α, β), each comprising a PolyTVL, average pooling, and an LSTM-based sequence summarizer. The PolyTVL introduces two key extensions over standard state-space models: (i) dual learnable polynomial transforms with positive-constrained orders O1, O2 ≥ 1 for nonlinear dynamics, and (ii) a time-varying state transition matrix A ∈ R(W×N×N), enabling position-specific state dynamics at each timestep. The band-specific representations are concatenated and passed through a two-stage dense layer for four-class gait state decoding.

The experimental paradigm (approved by University of Houston IRB STUDY00003848) consisted of 10 sessions divided into two phases: an open-loop (OL) phase (Ses. 1–5) and a mixed open- and closed-loop (CL) phase (Ses. 6–10). In the OL phase, each session comprised five runs of 20 gait cycles, each preceded by a 2-sec rest, with a 200 ms audio cue at 12-sec intervals triggering each cycle. Four gait states were defined: Initiate (2-sec post-cue), Execute (motion period), Terminate (2-sec pre-halt), and Stand (rest). Steps 1–10 and 16–20 were used for training, and steps 11–15 for validation (75:25 split). In the mixed phase, OL Runs 1-3 (20 steps) retrained the Feature Extraction Block while the Decoder Block remained frozen, and CL Runs 4-6 (Rex-assisted) and Runs 7-9 (manual/volitional) each had 10 gait cycles with real-time exoskeleton triggering.

A single healthy right-handed male participant (S1: 22 years) was recruited. EEG from 9 ROIs (28 channels) and 4 EOG channels were recorded using Brain Products actiCAP and MOVE system at 100 Hz. The study used the Rex lower-limb exoskeleton (Rex Bionics Ltd.). EMG and IMU data were also acquired using three Trigno Avanti sensors for future analysis.

Four decoder variants were evaluated: v00 (PolyTVL+Dense), v01 (PolyTVL+LSTM, proposed), v02 (S4D-Lin+Dense), and v03 (LSTM+Dense). Training used Adam optimizer (lr = 10−3, batch size = 128) for up to 200 epochs with early stopping, and the loss combined categorical cross-entropy with a confusion-aware penalty matrix reflecting the gait sequence (Stand → Initiate → Execute → Terminate), with maximum penalties assigned to Initiate/Terminate misclassifications.

Results showed that v01 (PolyTVL+LSTM) outperformed all variants with the highest validation MCC of 0.435 and the lowest overfitting gap of 0.187, while v00 had the largest overfitting gap (0.362). v01 also achieved the strongest Execute (0.594) and Initiate (0.359) recall scores. Kruskal-Wallis analysis revealed significant discriminability (p < 0.05, Bonferroni corrected, α = 0.05/27) across gait classes in 128/135 ROI-session-band combinations, with δ and θ yielding the highest −log10(p) overall. The learned PolyTVL polynomial orders increased monotonically across frequency bands (δ, θ: 1.068, α: 1.115, β: 1.205), confirming progressively stronger nonlinear reshaping at higher frequencies.

Closed-loop deployment with v01 achieved 55.3% gait initiation success for Rex-assisted runs and 52.7% for volitional (manual) runs across Ses. 6–10, with a false-positive rate of 18.2% for Rex-assisted runs. The mean prediction time (including preprocessing) was 70.5 ms (±41.5) on an Intel i7-14650HX (32 GB RAM), validating real-time feasibility. The end-to-end latency for Rex-assisted CL valid initiations (n=83) was 1035.7 ± 409.9 ms, within the 2-sec initiation window. The complete architecture comprises 70,755 total weights (227 non-trainable) with a 0.270 MB memory footprint (Decoder: 98%).

The authors conclude that the proposed 2-block architecture addresses three core requirements for a deployable closed-loop lower-limb BCI: (1) real-time artifact suppression via H∞, BPF, nASR, and trainable RDWT filtering; (2) multi-domain feature extraction through ROI-based depth-wise spatial filtering and adaptive RDWT coefficient filtering; and (3) temporal sequence modeling of time-varying nonlinear dynamics via PolyTVL, coupled with average pooling and an LSTM-based sequence summarizer. Future work will extend to multi-subject cohorts.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:


  • Current limitation: SSMs (S4, S4D, Mamba) are linear time-invariant, failing to model non-stationary EEG dynamics.

  • Improvement: Implement PolyTVL with dual learnable polynomial transforms (orders O1, O2 ≥ 1) and a time-varying state transition matrix A ∈ R(W×N×N). This captures position-specific nonlinear dynamics.

  • Result: Achieved validation MCC 0.435 vs. 0.362 for the best LTI baseline (v02), with a smaller overfitting gap (0.187 vs. 0.362).

  • Current limitation: ICA and ASR are computationally heavy or heuristic, precluding real-time use.

  • Improvement: Add nASR as a differentiable layer that detects contaminated channels and reconstructs them from clean neighbors (based on pairwise L2 distances), with weighted reconstruction and 20% spatial dropout.

  • Result: Enables real-time artifact suppression without offline ICA, reducing preprocessing latency to 70.5 ms (±41.5) per window.

  • Current limitation: Single-channel or global spatial filtering ignores anatomical structure.

  • Improvement: Group 28 channels into 9 anatomically defined ROIs; apply independent depthwise convolution per ROI (kernel: Cr×1, tanh, max-norm 1.0) to collapse channels within each ROI.

  • Result: Consistent discriminability (p < 0.05, Bonferroni corrected) across 128/135 ROI-session-band combinations, with MFC, MCP, LCP, RCP, MPO remaining discriminative across all bands.

  • Current limitation: Fixed wavelet filters or offline decomposition.

  • Improvement: Implement RDWT with Symlet-2 mother wavelet, decomposed into 5 scales, with trainable zero-phase filters in the wavelet domain and sparsity regularization on the γ band (25–50 Hz) to suppress EMG.

  • Result: Produces clean δ, θ, α, β sub-bands; γ discarded. The learned polynomial orders increase monotonically with frequency (δ,θ: 1.068; α: 1.115; β: 1.205), confirming frequency-adaptive nonlinear reshaping.

  • Current limitation: Standard cross-entropy treats all misclassifications equally, risking unsafe transitions (e.g., Stand→Execute).

  • Improvement: Add a 4×4 penalty matrix (Table 1) that heavily penalizes non-adjacent confusions (e.g., Initiate↔Terminate: 30) and lightly penalizes adjacent ones (e.g., Stand↔Initiate: 5).

  • Result: Reduces false positives in CL runs (FPR 18.2%) and improves gait initiation success to 55.3% (Rex-assisted) and 52.7% (volitional).

  • Current limitation: Full model retraining per session is slow and risks catastrophic forgetting.

  • Improvement: After initial training (Ses. 1–5), freeze the Decoder Block and retrain only the Feature Extraction Block per session (Ses. 6–10) using OL runs 1–3.

  • Result: Maintains performance across sessions with minimal retraining time, enabling real-time deployment.

  • Current limitation: Single dense layer or simple pooling loses temporal context.

  • Improvement: For each band (δ,θ, α, β), pass through PolyTVL → Average Pooling (size 2, stride 2) → LSTM (8 units) → concatenate → Dense (16, ReLU) → Dense (4, softmax).

  • Result: Outperforms all ablation variants (v00, v02, v03) in validation MCC and recall for Execute (0.594) and Initiate (0.359).

  1. Real-Time Closed-Loop Gait Decoding: Process EEG at 100 Hz with a 256-sample window (2.56 s) and 200 ms stride, producing a prediction every 200 ms with mean latency of 70.5 ms (±41.5) — well within the 2-second initiation window.

  2. Four-State Gait Classification: Distinguish Stand, Initiate, Execute, and Terminate with a validation MCC of 0.435, significantly above chance (18.3–50.4% per class).

  3. Robust Artifact Suppression Without Offline Processing: Suppress ocular artifacts (via H∞), high-frequency noise (via BPF + Chebyshev), and channel-level artifacts (via nASR) entirely within the computational graph, enabling deployment on portable hardware.

  4. Safe Exoskeleton Control: The confusion-aware penalty matrix reduces dangerous misclassifications (e.g., Initiate→Terminate), achieving a false-positive rate of only 18.2% during CL runs.

  5. Session-Specific Adaptation: Rapidly retrain the Feature Extraction Block per session (using only 3 OL runs) while keeping the Decoder frozen, allowing the system to adapt to day-to-day EEG variability without full retraining.

  6. Frequency-Aware Nonlinear Modeling: The PolyTVL layer learns frequency-dependent polynomial orders (higher for β, lower for δ/θ), capturing the nonlinear, non-stationary dynamics of cortical gait signals that LTI SSMs cannot.

  7. Scalable to Multi-Subject Cohorts: The modular 2-block design (feature extraction + decoder) allows transfer learning — a frozen decoder can be paired with a new subject’s feature extraction block, reducing training time and data requirements.

  8. Deployable on Edge Devices: With only 70,755 total parameters (0.270 MB memory footprint, 98% in the decoder), the system is lightweight enough for real-time inference on embedded processors (e.g., Intel i7 or lower-power ARM boards).

Bottom line: The improved AI system is a real-time, closed-loop, four-state gait decoder for lower-limb exoskeletons that is artifact-robust, session-adaptive, and safe for clinical use — outperforming existing LTI SSM and LSTM-only approaches in both accuracy and latency.

Sources

Related papers