Deep Learning for Time Series Classification of Parkinson's Disease Eye Tracking Data

arXiv:2311.16381 · cs.LG, q-bio.QM · Submitted 2023-11-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Deep Learning for Time Series Classification of Parkinson's Disease Eye Tracking Data".

Jane: The study details a deep learning approach for time series classification applied to eye tracking data from subjects with Parkinson's disease (PD).

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Building on that idea of finding complex patterns, let's look at what the paper actually proposes in terms of classification. The paper introduces a deep learning approach specifically for classifying Parkinson's Disease eye tracking data, moving past earlier attempts that hadn't found a single clear biomarker.

Jane: Exactly, and what’s striking is that they are using sophisticated machine learning techniques to classify these data segments into different groups. The abstract mentions that eye-tracking provides information about motor and cognitive abilities, which is the foundation for this whole classification effort.

Lu: They utilize models like InceptionTime and ROCKET, which suggests they are looking at patterns across multiple time scales simultaneously, which is a powerful way to capture different aspects of eye movement. It’s not just about one snapshot but the dynamic sequence of movements.

Meng: I see them using specific architectures, like the InceptionTime model with its parallel convolutions and residual connections, which sounds computationally intensive; how do they manage the training process to keep that manageable for real-world deployment?

Lalam: I think the architecture itself is important because if we can design these layers to capture nuanced temporal dependencies, it could improve how our AI systems understand complex human behaviors in a way that's much richer than just simple pattern matching.

The paper's summary: Tom: Now let’s dig into the summary section of the paper to really get into what they found regarding the classification performance. The results they present are quite compelling when you look at how these models perform on their test sets, specifically mentioning accuracies like seventy-three percent and ninety-six percent at trial and subject levels, respectively.

Jane: Those accuracy numbers really make a difference because they show that the proposed method has improved performance significantly when compared to previous limitations where no single biomarker was conclusively identified. It’s showing tangible progress in the field.

Lu: The paper also points out something important about generalization, suggesting that the homogeneity of the fixation task minimizes subject identity fingerprints in the data, which implies that these models should be able to work well on unseen subjects too. That’s a big deal for real-world application.

Meng: It sounds like they are focused on making sure the model isn't just memorizing the training data, but actually learning the underlying movement principles, which is crucial for practical impact on diagnosis.

Lalam: If these models can generalize well across different subjects without being overly tied to specific individual recording styles, it means our AI could be much more robust when deployed in diverse clinical settings.

The paper's improvements: Tom: The authors also suggest several improvements they think can make this classification framework even better, which is where the real future work lies. They aren't just stopping at the initial model implementation; they are looking at ways to enhance its diagnostic power and interpretability.

Jane: What’s interesting is that they are demonstrating how the selected features can actually be used to interpret the model, which moves it away from being just a black box and toward something clinically useful. That’s a crucial step for any medical AI application.

Lu: I see them pointing toward using techniques like residual connections in InceptionTime to help mitigate potential vanishing gradient issues, which is a technical refinement aimed at making the training more stable and effective. That shows they are thinking about the deeper mechanics of the network itself.

Meng: From an engineering viewpoint, focusing on stability through better connection methods is smart because unstable models are hard to trust in a high-stakes environment like medical diagnosis; we need reliability before we worry about raw accuracy scores.

Lalam: I think focusing on interpretability and making the model’s reasoning transparent gives us a huge advantage because it builds trust with clinicians, which is something that will really shape how this technology is adopted.

Conclusion: Tom: So, to wrap things up on "Deep Learning for Time Series Classification of Parkinson's Disease Eye Tracking Data," the paper shows a solid application of deep learning to tackle a persistent challenge in diagnosing PD using eye tracking data. The overall implication is that advanced analytical techniques can indeed help us move toward a more robust and objective classification system.

Jane: Indeed, it suggests that we can use these models not just as tools for categorization, but as resources for understanding the pathophysiological substrates of the disease. It gives us a clearer picture of what's going on neurologically behind the eye movements.

Lu: I think the biggest impact here is showing that this approach is promising for other cognitive or motor disorders, expanding its potential beyond just Parkinson's disease. The framework itself has broad applicability in analyzing dynamic time series data.

Meng: Practically speaking, the implication is that we need to keep pushing these models to handle real-world noise and variability effectively before we can see widespread use in clinical settings.

Lalam: I think what really stands out about this paper is how it lays a foundation for future work by emphasizing interpretability, which will be key for making this kind of AI truly useful and trusted across the entire healthcare landscape.

Lexiang Ye, Eamonn Keogh

cs.LG, q-bio.QM

Submitted: 2023-11-28

Updated: 2026-08-24

Code: https://github.com/gon-uri/detach_rocket

Importance score: 79/100

The gist: The study details a deep learning approach for time series classification applied to eye tracking data from subjects with Parkinson's disease (PD).

Key concepts

InceptionTime
This is a specific model used in the study that employs parallel convolutions and residual connections. It is designed to capture patterns across multiple time scales simultaneously in eye tracking data.
ROCKET
Another model utilized, ROCKET suggests the researchers are analyzing patterns across different time scales at once. This helps capture various aspects of dynamic eye movements rather than just single snapshots.
Generalization
The study suggests that because the fixation task is homogeneous, these models should perform well on subjects they have not seen before. This implies the models can work effectively with new subjects in real-world applications.
Interpretability
The authors emphasize using selected features to interpret model results instead of treating the AI as a 'black box.' This makes the model's reasoning transparent, which is crucial for building clinical trust.

Terminology

Summary

The study details a deep learning approach for time series classification applied to eye tracking data from subjects with Parkinson's disease (PD).

Data Acquisition and Cohort:

The final dataset utilized 1584 MTS segments, each having a length of 440 (approximately 1.5 seconds). These data segments were derived from trials performed by subjects under three distinct conditions: "451 trials from HC [Healthy Controls], 561 trials from PD patients on their regular dose of dopamine replacement therapy (PD ON), and 572 trials from PD patients after a 12 hour withdrawal from their dopaminergic medication (PD OFF). The overall cohort consisted of 23 healthy controls, 28 PD patients on medication, and 30 PD patients off medication."

Data Preprocessing and Sanitization:

Rigorous data cleaning was implemented to ensure the quality of the recordings. To identify and remove corrupted or noisy recordings, sessions were excluded using a criterion based on the Median Absolute Deviation (MAD). The MAD was chosen because it is a measure of the variability in the data, which is more robust to outliers than standard deviation. Specifically, A threshold of 3 MADs was used to exclude sessions, with this exclusion performed independently on both the x and y channels. This sanitization process was extensive, as it removed all data from 21 subjects, resulting in a total of 63 subjects in the dataset.

Regarding the time series content, only data collected during a specific preparatory phase were utilized: We only used data from the period before the target moves, when the subject is focusing in preparation for the saccade. The time interval length was standardized across all trials. Furthermore, to counteract signal drift throughout each session, the MTS corresponding to trials were centered by subtracting the mean.

Deep Learning Models Implemented:

Two state-of-the-art deep learning architectures were employed for classification: InceptionTime and ROCKET.

1. Model 1: InceptionTime

The InceptionTime model was implemented using PyTorch Lightning, following the original authors’ TensorFlow implementation. The architecture is composed of 4 Inception modules, each with a bottleneck dimensionality of 32, a hidden dimensionality of 64, and three sets of kernels of sizes [40, 20, 10]. The output from each module is 4 times 64 = 256 (including the parallel max pooling). To enhance performance and mitigate potential vanishing gradient issues, we use residual connections between the Inception modules.

The training setup was highly controlled:

  • Optimizer: AdamW (beta 1 = 0.9, beta 2 = 0.999).

  • Learning Rate: 1 times 10-3.

  • Weight Decay: Ridge Parameter of 1 times 10-3.

  • Scheduling: The learning rate was scheduled with cosine annealing and 15 epochs of gradual warm up [Goyal et al. (2018)].

  • Training Duration: Training was conducted for 300 epochs with a large batch size of 256.

  • Loss Function: The binary cross entropy was used as the loss function.

2. Model 2: ROCKET

The

Improvements for AI systems

As a diligent AI researcher operating under high financial stakes, I have analyzed this protocol. The current pipeline—combining feature engineering (Shapelets/4D time series) with state-of-the-art CNN/Deep Learning models (InceptionTime, ROCKET)—is robust. However, several critical advancements can be made to enhance diagnostic specificity, data robustness, and clinical interpretability.

Here are the specific improvements I recommend for the AI system:


The current use of 3 MADs is effective for outlier detection, but it discards entire sessions, which results in significant data loss and potential class imbalance.

Improvement: Implement a Generative Adversarial Network (GAN)-based imputation module prior to feature extraction.

  • Methodology: Train a Conditional GAN (cGAN) on the clean time series segments from the HC group. The cGAN would take the corrupted/missing segment as input and generate a statistically plausible, high-fidelity imputation of the missing data points (e.g., filling gaps caused by blinks or motion artifacts).

  • Specifics: Use a Wasserstein loss (WGAN) for stable training. The imputation must be conditioned on the subject's behavioral state (e.g., attempting antisaccade).

  • Improved Capability: This allows the system to process nearly all available data, drastically reducing class imbalance and maximizing the usable sample size, making the model far more resilient to real-world recording imperfections.

The current approach treats kinematics (x, y, v) and shapelets as separate inputs to different models or sequentially combined. Eye movements are not isolated; they are contextually linked to cognitive load and task demands.

While InceptionTime is excellent for temporal feature extraction, it primarily treats the time series as a sequence of local features. The underlying physical process (the eye moving in 2D space) is inherently graph-like.

In a high-stakes medical context, black box predictions are unacceptable. We must know why the system classified a patient as having PD OFF versus PD ON.

Sources

Related papers