SincPD: An Explainable Method based on Sinc Filters to Diagnose Parkinson's Disease Severity by Gait Cycle Analysis

arXiv:2502.17463 · eess.SP, cs.LG · Submitted 2025-02-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SincPD: An Explainable Method based on Sinc Filters to Diagnose Parkinson's Disease Severity by Gait Cycle Analysis".

Jane: The paper was written by Armin Salimi-Badr, Mahan Veisi and Sadra Berangi from Shahid Beheshti University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone! We've got a fascinating paper on the table today, and I have to say, the title alone got me hooked. It's called "SincPD: An Explainable Method based on Sinc Filters to Diagnose Parkinson's Disease Severity by Gait Cycle Analysis." Jane, what do you make of that name?

Jane: Tom, I love it because it tells you exactly what's inside. SincPD—Sinc for those special filters they use, and PD for Parkinson's Disease. And the authors, Armin Salimi-Badr, Mahan Veisi, and Sadra Berangi from Shahid Beheshti University in Tehran, they're tackling something really important here. They're not just building a black box that says "yes" or "no" to Parkinson's; they're building one that can explain *why* it made that call.

Tom: And that's the part that gets me excited. We hear so much about deep learning being this mysterious thing where you feed in data and get an answer, but nobody can tell you how it got there. This paper is saying, "Hold on, we can actually peek inside and see what the model is looking at."

Jane: Exactly. They're using these things called Sinc filters, which are basically mathematical bandpass filters. Think of them like a radio tuner that only lets certain frequencies through. The model learns which frequencies in your walking pattern matter most for spotting Parkinson's. And because each filter has just two parameters—the center frequency and the bandwidth—you can literally look at them and say, "Oh, this filter is picking up signals between zero point four and zero point six Hertz."

Tom: So instead of thousands of abstract features, you get a handful of interpretable ones. That's a game-changer for clinical trust, right?

Jane: Absolutely. A doctor doesn't want to hear "the neural network says so." They want to hear "the model found that the frequency of your heel strikes is different from a healthy person's." That's something they can verify and understand.

Tom: And they're doing this with wearable sensors in people's shoes, measuring vertical ground reaction force. So it's not an MRI machine or a lab setup; it's something that could potentially be used in a regular clinic or even at home. That's a huge practical implication.

Jane: It is. And the fact that they're also trying to determine the *severity* of the disease, not just whether you have it, makes this even more valuable for tracking progression over time.

Tom: I'm already curious about how they actually built this thing. Let's dig into the method in the next segment.

Summary of the Paper: Tom: So, Jane, we've got this model that uses Sinc filters to look at gait data. But how does it actually work end to end? Walk me through the summary.

Jane: Okay, so they start with raw data from sixteen sensors under the feet, recording vertical ground reaction force. First step is preprocessing—they chop the signals into ten-second chunks, clean out bad data, and standardize everything. Then they do something clever: they take the difference between the left and right sensors, cutting the data down from sixteen channels to eight without losing the important stuff.

Tom: That's smart. It's like comparing the two feet directly instead of looking at each one in isolation.

Jane: Exactly. Then they build a deep learning model where the first layers are these SincConv1D layers—eight of them, one for each sensor. Each layer starts with one hundred filters, so that's eight hundred filters total. After those, they stack a couple of regular convolutional layers, then some dense layers, and finally a single output that says "patient" or "healthy."

Tom: And that initial model gets trained for a thousand epochs, hitting about ninety-eight point seven seven percent accuracy. Pretty solid. But here's the kicker—they don't stop there. They want to make it smaller and more explainable.

Jane: Right, and that's where the pruning comes in. They take all those learned filters and cluster them based on their cutoff frequencies using K-means. The idea is that many filters are redundant—they're picking up almost the same frequency bands. So they group them and keep just the medoid of each cluster as a representative.

Tom: So they went from eight hundred filters down to about thirty? That's a massive reduction.

Jane: It is, and the accuracy only drops a tiny bit—from ninety-eight point seven seven percent to ninety-eight point one five percent. That's a negligible loss for a model that's now way simpler and way easier to interpret.

Tom: And then they retrain the pruned model for a few epochs to fine-tune it. But the real magic happens when they start analyzing which filters and which sensors matter most. They use DBSCAN clustering to find representative signals for patients and healthy people, then pass those through the pruned filters and compare the energy outputs.

Tom: So they're literally measuring how much signal each filter lets through for each group, and the filters with the biggest difference are the ones doing the heavy lifting.

Jane: Precisely. And what they found is that sensors at the front of the foot—like the ball of the foot—and the back—like the heel—show the biggest energy differences. And the top filters are mostly picking up frequencies around zero point four to zero point six Hertz. That's a really specific finding that could have biomechanical meaning.

Tom: I love that they're not just saying "trust us, it works." They're saying "here's exactly what the model is paying attention to, and it makes sense given what we know about how Parkinson's affects gait."

Jane: And that's the whole point of explainable AI in medicine. It's not enough to be accurate; you have to be trustworthy.

Tom: Okay, so we've got the method. But what about the severity part? How do they go from "you have Parkinson's" to "you're at stage two point five"?

Improvements Suggested by the Paper: Tom: So Jane, the paper doesn't stop at just diagnosing Parkinson's. They also tackle severity. How does that work?

Jane: They use transfer learning. They take the pruned model we just talked about—the one that already knows how to extract meaningful features from gait data—and they freeze those layers. Then they add a new output layer that classifies into severity stages based on the modified Hoehn and Yahr scale.

Tom: So it's like taking a trained ear for music and teaching it to distinguish between different genres instead of just "music" and "not music."

Jane: That's a great analogy. And it works remarkably well—they hit ninety-seven point two two percent accuracy on the multi-class severity problem. That's better than several state-of-the-art methods they compared against, like a 1D CNN that got eighty-five point two three percent and an LSTM that got ninety-six point six percent.

Tom: And they're doing this with far fewer parameters. The pruned model has 872K parameters, while some of the other methods have tens of millions. That's a massive efficiency win.

Jane: It is, and it matters for real-world deployment. Smaller models run faster, use less memory, and can be deployed on edge devices like smartphones or wearable sensors themselves.

Tom: But the improvement I find most exciting is the explainability angle. They're not just saying "stage three"; they're showing which sensors and frequency bands drove that decision. For a clinician, that's gold.

Jane: Exactly. And they go one step further—they identify the top twenty percent of filters based on energy difference between patients and healthy subjects. The top filters are mostly from Sensor seven which is the ball of the foot, and Sensor two which is the heel. And they're all picking up frequencies in that zero point four to zero point six Hertz range.

Tom: That's such a concrete finding. It suggests that Parkinson's specifically affects the timing and force distribution of heel strikes and toe-offs, which makes total sense given what we know about the disease.

Jane: And that's the kind of insight that could feed back into clinical practice. Maybe doctors start paying more attention to those specific frequency bands when assessing patients, or maybe it informs the design of better wearable sensors.

Tom: I also like that they're using clustering to prune, which is a pretty general technique. You could apply this same approach to other medical signals—heart data, breathing patterns, even speech.

Jane: Absolutely. The methodology is not Parkinson's-specific. It's a template for building interpretable deep learning models on any time-series data.

Tom: So what's the catch? What are the limitations?

Jane: Well, the dataset is from PhysioNet, which is a public dataset with one hundred sixty-six subjects—ninety-three with Parkinson's and seventy-three healthy. It's a decent size, but it's not huge. And the severity stages only go up to stage three in their data, so they're not covering the most severe cases. Also, the data was collected in a lab setting with people walking at their own pace, which is good, but real-world conditions might be messier.

Tom: Still, for a proof of concept, this is really compelling. Let's wrap this up in the conclusion.

Conclusion: Tom: Alright, let's bring it home. We've been talking about "SincPD: An Explainable Method based on Sinc Filters to Diagnose Parkinson's Disease Severity by Gait Cycle Analysis," and honestly, this is one of those papers that makes me feel like we're finally moving toward AI we can actually trust in medicine.

Jane: I completely agree, Tom. The key takeaways are: they built a model that diagnoses Parkinson's with ninety-eight point seven seven percent accuracy, pruned it down to a fraction of its original size with almost no performance loss, and then extended it to classify severity with ninety-seven point two two percent accuracy. But the real win is that they can show you *which* sensors and *which* frequency bands are driving those decisions.

Tom: And those findings—that the heel and ball of the foot sensors matter most, and that the key frequencies are around zero point four to zero point six Hertz—those are things a clinician can actually use. They're not abstract neural network features; they're measurable, physical phenomena.

Jane: Right. And that's what makes this paper stand out. It's not just another deep learning model that beats the benchmark. It's a model that opens the black box and says, "Here's what I'm looking at, and here's why it makes sense."

Tom: The implications go beyond Parkinson's, too. The pruning method and the explainability framework could be applied to any time-series medical data. That's a big deal.

Jane: It is. And while there are limitations—the dataset size, the limited severity range—this is a strong foundation. Future work could validate on larger, more diverse populations and maybe even test in real-world clinical settings.

Tom: Well said. We're going to say goodbye to SincPD and get ready to look at the next paper on the arXiv. Thanks for tuning in, everyone. We'll see you next time.

Jane: Take care, and keep listening!

Armin Salimi-Badr, Mahan Veisi, Sadra Berangi

Shahid Beheshti University

eess.SP, cs.LG

Submitted: 2025-02-10

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 58/100

Key concepts

Sinc Filters
These are mathematical bandpass filters used in the SincPD model. They function like a radio tuner, allowing only certain frequencies through the data. The model learns which specific frequencies in walking patterns are important for detecting Parkinson's Disease.
Explainable AI (XAI)
This refers to making deep learning models understandable instead of treating them as black boxes. SincPD achieves this by showing clinicians exactly which sensors and frequency bands the model uses to make its diagnosis, which builds clinical trust.
Gait Cycle Analysis
This involves analyzing vertical ground reaction force data recorded from wearable sensors in a person's shoes during walking. The model examines these measurements across the gait cycle to identify patterns indicative of Parkinson's Disease severity.
Pruning
This is a technique used to simplify the deep learning model by removing redundant filters. The researchers used K-means clustering to group similar filters and keep only representative ones, significantly reducing the model size while maintaining high accuracy.

Terminology

Summary

Summary

This paper presents SincPD, an explainable deep learning-based classifier for Parkinson's Disease (PD) diagnosis and severity assessment, based on analyzing gait cycles using adaptive sinc filters. The method utilizes raw vertical Ground Reaction Force (vGRF) signals measured by 16 wearable sensors placed in the soles of subjects' shoes.

The proposed method consists of Sinc layers that model adaptive bandpass filters to extract important frequency bands in the gait cycle of patients and healthy subjects. This allows the reasons behind classification decisions to be explained by examining these frequencies. The paper states: "The proposed method consists of Sinc layers that model adaptive bandpass filters to extract important frequency-bands in gait cycle of patients along with healthy subjects. Therefore, by considering these frequencies, the reasons behind the classification a person as a patient or healthy can be explained."

The methodology involves several key steps. First, preprocessing is performed, which includes partitioning data into ten-second segments, filtering out incomplete data, standardizing signals, and calculating the difference between left and right sensor signals to reduce dimensionality from 16 arrays per subject to 8. Stratified splitting maintains class balance in training and test datasets.

The initial model architecture uses eight SincConv1D layers (one per preprocessed signal), each configured with 100 filters of length 101. These layers extract frequency-specific features, followed by normalization, Leaky ReLU activation, and pooling. The outputs are concatenated and passed through two standard Conv1D layers (128 and 256 filters), followed by Batch Normalization, Leaky ReLU, Dropout, and Max Pooling. Finally, three Dense layers with 128, 64, and 1 neurons are used, with ReLU activation for the first two and sigmoid activation for the final layer. Batch Normalization and L2 regularization are applied, and the Adam optimizer minimizes cross-entropy.

To improve interpretability, a pruning method is proposed. The paper explains: "High number of filters decreases the interpretability. Therefore, we propose a pruning method to prune extra filters. To realize this, we cluster filters based on their cut-off frequencies and reconstruct the architecture based on the clusters' centroids." Using K-means clustering on the center frequency (fc) and bandwidth (b) parameters, the optimal number of clusters is determined via the elbow method and silhouette scores. The cluster centroids become the final filter weights, reducing the total number of filters across the SincNet layers from 800 to approximately 30.

The model was evaluated on the PhysioNet gait dataset, which includes 166 subjects (93 PD, 73 controls) with data aggregated from three separate studies. PD severity was assessed using the modified Hoehn and Yahr scale. The initial model achieved 98.77% accuracy before pruning and 98.15% after pruning, with only 872K parameters. The confusion matrix shows the model correctly predicts 97% of negative class samples and 99.55% of positive class samples before pruning, with minimal misclassifications. After pruning, the model retains 97% accuracy for the negative class and 98.66% for the positive class.

For severity detection, the model uses transfer learning with a pretrained clustered model providing initial weights, freezing the clustered model's layers. The severity model achieves 97.22% accuracy in multi-class classification.

The paper also analyzes filter and sensor importance by calculating signal energy from pruned SincNet layer outputs. Using DBSCAN clustering to identify representative signals (clustroids) for each class, the energy distributions reveal which filters and sensors are most discriminative. The analysis shows that sensors at the front (Sensors 6 and 7) and back (Sensors 1 to 3) of the foot demonstrate higher energy differences. The top filters, primarily associated with Sensor 7 (ball of the foot) and Sensor 2 (heel), focus on bandpassing frequencies in the range of approximately 0.4 to 0.6 Hz, while a fourth filter from Sensor 7 targets a narrower band of 0.2 to 0.5 Hz.

The paper compares SincPD against several state-of-the-art methods. For binary classification, SincPD achieves 98.77% accuracy before pruning and 98.15% after pruning, outperforming LSTM (98.60%), Multi LSTM (91.95%), CNN+LSTM (98.09%), and ResNet-101 (97.56%). For multi-class classification, SincPD achieves 97.22% accuracy, surpassing 1D CNN (85.23%), LSTM (96.60%), and ANN-FFT (97.00%).

The paper concludes: "The proposed method is a deep structure that takes the low-level raw vertical Ground Reaction Force (vGRF) signal recorded by 16 sensors put in the subjects' shoes as the input and extract higher level features based on applying various filters... These filters have lower number of parameters that make the method a light-weight deep model. Moreover, based on analyzing the active filters during the inference process, the important filters and sensors are determined. This information can be utilized to explain the network's output."

Improvements for AI systems

Based on the SincPD paper, here are the specific improvements I can make to AI systems, along with what the improved system can do:


  • Improvement: Replace standard CNN first-layer filters with parameterized Sinc filters (only 2 learnable parameters per filter: center frequency and bandwidth). This reduces model complexity and makes the learned features physically interpretable.

  • What the improved system can do: The AI can now explain why it classifies a subject as Parkinson’s or healthy by showing which frequency bands (e.g., 0.4–0.6 Hz) in which foot regions (e.g., heel vs. ball of foot) are most discriminative. This is critical for clinical trust and regulatory approval.

  • Improvement: After training a large model (100 filters per sensor), cluster the learned (center frequency, bandwidth) pairs using K-means with silhouette-score-optimized k. Replace the filter bank with cluster medoids (only 3–4 filters per sensor). Retrain briefly to fine-tune.

  • What the improved system can do: The AI reduces its parameter count from 1.17M to 872K (a 25% reduction) while maintaining 98.15% accuracy (vs. 98.77% before pruning). This makes the model more parsimonious, faster to deploy on edge devices (e.g., wearable sensors), and easier to audit.

  • Improvement: Compute the energy difference between patient and healthy representative signals (selected via DBSCAN clustroids) at each filter output. Rank sensors and filters by this energy difference.

  • What the improved system can do: The AI can now tell clinicians which sensors matter most (e.g., Sensor 7 – ball of foot, Sensor 2 – heel) and which frequency bands are most affected by Parkinson’s. This enables targeted sensor placement in future studies and reduces data collection cost.

  • Improvement: Use the pruned, pretrained binary classifier as a frozen feature extractor for a multi-class severity model (Hoehn & Yahr stages 2, 2.5, 3). Only retrain the final dense layers.

  • What the improved system can do: The AI achieves 97.22% accuracy on severity classification without retraining the entire network. This is a 0.22% improvement over the best prior method (ANN-FFT at 97.00%) while using far fewer parameters and providing interpretable frequency features.

  • Improvement: Instead of averaging all signals (which introduces noise), use DBSCAN to find the most representative real signal (clustroid) for each class. Pass these through the Sinc layers for energy analysis.

  • What the improved system can do: The AI avoids misleading energy comparisons caused by noisy averages. It can reliably identify which filters are truly discriminative, even when patient and healthy signals overlap in raw time-domain.

  • Diagnose Parkinson’s Disease from raw vGRF signals with 98.15% accuracy (post-pruning) and 98.77% (pre-pruning), outperforming LSTM (98.60%), CNN+LSTM (98.09%), and ResNet-101 (97.56%).

  • Classify severity (H&Y stages 2, 2.5, 3) with 97.22% accuracy, beating ANN-FFT (97.00%) and LSTM (96.60%).

  • Explain its decisions by identifying:

  • The most important frequency bands (0.2–0.6 Hz) for each sensor.

  • The most critical sensors (heel and ball of foot).

  • The specific filters that contribute most to the classification.

  • Run efficiently on resource-constrained hardware (872K parameters, 8 Sinc layers with only 30 total filters after pruning).

  • Transfer knowledge from binary diagnosis to severity classification, reducing training time and data requirements.

This system is not just a black-box classifier; it is a clinically actionable diagnostic tool that tells why and where the disease manifests in gait, enabling earlier intervention and more personalized treatment planning.

Abstract

In this paper, an explainable deep learning-based classifier based on adaptive sinc filters for Parkinson's Disease diagnosis (PD) along with determining its severity, based on analyzing the gait cycle (SincPD) is presented. Considering the effects of PD on the gait cycle of patients, the proposed method utilizes raw data in the form of vertical Ground Reaction Force (vGRF) measured by wearable sensors placed in soles of subjects' shoes. The proposed method consists of Sinc layers that model adaptive bandpass filters to extract important frequency-bands in gait cycle of patients along with healthy subjects. Therefore, by considering these frequencies, the reasons behind the classification a person as a patient or healthy can be explained. In this method, after applying some preprocessing processes, a large model equipped with many filters is first trained. Next, to prune the extra units and reach a more explainable and parsimonious structure, the extracted filters are clusters based on their cut-off frequencies using a centroid-based clustering approach. Afterward, the medoids of the extracted clusters are considered as the final filters. Therefore, only 15 bandpass filters for each sensor are derived to classify patients and healthy subjects. Finally, the most effective filters along with the sensors are determined by comparing the energy of each filter encountering patients and healthy subjects.

Related papers