A Benchmark for Early-stage Parkinson's Disease Detection from Speech

arXiv:2605.14066 · eess.AS, cs.AI, cs.CL, cs.SD · Submitted 2026-05-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "A Benchmark for Early-stage Parkinson's Disease Detection from Speech".

Jane: Early-stage Parkinson’s disease (EarlyPD) detection from speech is clinically meaningful yet underexplored, and published results are hard to compare because studies differ in datasets, languages, tasks, evaluation protocols,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about the title of this paper, "A Benchmark for Early-stage Parkinson's Disease Detection from Speech," and the authors—Terry Yi Zhong, Cristian Tejedor-Garcia, Khiet P. Truong, Janna Maas, Louis ten Bosch, and Bastiaan R. Bloem—and what that title really means for us. It’s not just about a specific model; it’s about creating a common testing ground for speech analysis related to early Parkinson's disease detection across different methods and data setups.

Jane: That title really highlights the core contribution: they aren't just proposing one new technique, but establishing a benchmark so that other researchers can test their own ideas against a consistent set of rules, which is what we need when we want to move this research from theory into something usable for doctors.

Lu: The authors are clearly coming from different centers, which I think brings a lot of diverse perspectives to the work; having input from language studies alongside neurology gives them a strong grounding in both the linguistic and clinical aspects of speech data.

Meng: It’s interesting that they focused on creating this benchmark specifically because published results are hard to compare right now; that tells me they recognized a real gap in the literature where we lack a reliable way to gauge how different detection methods stack up against each other fairly.

Lalam: For us, having this paper means we get a standardized vocabulary for discussing EarlyPD detection from speech, which is crucial because it helps us align our internal metrics and ensure our development paths are moving in the right direction for better diagnostics.

The paper's summary: Tom: Moving into the summary of "A Benchmark for Early-stage Parkinson's Disease Detection from Speech," they propose the first benchmark for speech-based EarlyPD detection that features a speaker-independent split, which is a significant step toward making the results replicable. They are testing methods across three common speech tasks and using different training settings to see how robust those methods are.

Jane: That speaker independence is what makes this work so powerful; it means they're not just testing on data from one specific group of speakers, but rather showing how well a method performs when it encounters a new, unseen person speaking. It’s about generalizability in the real world.

Lu: They specifically focus on using PC-GITA and NeuroVoz as their open-source datasets because they found those were the only ones with enough clinical metadata to apply their specific EarlyPD definition criteria, which is quite a specific methodological choice that grounds the study.

Meng: The summary also mentions testing under four different training data settings, from all potential PD cases to subsets and specifically early-stage PD data, which gives us a clear way to see the impact of restricting or broadening the training data on performance.

Lalam: It’s interesting how they structured their summary to show that they are not just doing one test, but building a comprehensive framework that allows for multi-dimensional evaluation breakdowns later on, which sets up a really thorough analysis for us.

The paper's improvements: Tom: Now, let’s look at the suggested improvements within the paper; they are focusing on refining how we define EarlyPD—prioritizing the H andY scale because of its clear cutoffs and then using a five-year TAD cutoff as a practical compromise for disease duration, which is a specific suggestion to fix prior inconsistencies.

Jane: That's smart clinical grounding; by prioritizing the H andY scale, they are trying to use something that clinicians are already familiar with and can interpret easily, rather than relying solely on metrics that might be abstract.

Lu: They also noted the limitation of using the MDSUPDRS for measuring outcomes because it doesn't have widely accepted thresholds for separating stages, which is a critique that points toward the need for better clinical measurement tools in general.

Meng: The paper suggests restricting disease duration to unmedicated cases would be neither clinically representative nor desirable, which is a practical constraint they put on their own study design to keep the cohort relevant for actual patient care scenarios.

Lalam: I think these specific suggestions about how to define the criteria are really important because they show that even in benchmark research, we need careful consideration of what makes a definition clinically meaningful versus just mathematically neat.

Conclusion: Tom: To wrap things up with this paper, the main conclusion is that expanding speaker diversity, no matter where it comes from, looks like a very promising direction for improving EarlyPD detection performance across all three tasks they tested. They also highlight that RECA-PD achieved the highest F1 and AUC scores averaged over tasks, especially on DDK and sentence reading.

Jane: That finding about expanding speaker diversity is significant because it suggests that having a wider variety of speech patterns in the data helps any detection model generalize better, which is a very positive signal for us when we're trying to build tools that work universally.

Lu: I think the observation that sentence-level cues benefited from broader disease-stage variability is also insightful, as it suggests that different parts of the speech might require different types of data diversity to be effective.

Meng: From an engineering standpoint, the finding about aggregation improving performance for some models while others degrade under small-sample aggregation gives us a concrete rule we can use when deciding how much data pooling we need for our own systems.

Lalam: Overall, this benchmark sets up a shared reference point for all of us to build upon, and I think the focus on both aggregate-level and gender-stratified results is what makes this study critical for getting these tools ready for actual clinical adoption.

Terry Yi Zhong, Cristian Tejedor-Garcia, Khiet P. Truong, Janna Maas, Louis ten Bosch

Centre for Language Studies, Radboud University · Center of Expertise for Parkinson & Movement Disorders, Department of Neurology, Donders Institute, Radboud University Medical Center

eess.AS, cs.AI, cs.CL, cs.SD

Submitted: 2026-05-13

Updated: 2026-07-18

Comments: Accepted by Interspeech2026. Camera Ready version

DOI: 10.21437/Interspeech.2026-2850

Code: https://github.com/terryyizhongru/B-EarlyPD-Speech

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 77/100

The gist: Early-stage Parkinson’s disease (EarlyPD) detection from speech is clinically meaningful yet underexplored, and published results are hard to compare because studies differ in datasets, languages,

Key concepts

EarlyPD Criteria
The paper standardized the definition of EarlyPD using specific clinical markers: H&Y stage less than or equal to 2 and a Tremor Amplitude Duration (TAD) of five years or less. Participants not meeting these rules are classified as non-Early PD, providing a consistent baseline for comparison across studies.
Speaker-Independent Split
To ensure fair and replicable testing, the researchers used a fixed 5-fold split where each validation and test set contained exactly six EarlyPD speakers and six healthy control (HC) speakers. This method prevents bias introduced by speaker characteristics in the data division.
Benchmark Protocol
The protocol details a transparent procedure for evaluating speech detection models. It specifies tasks like sustained vowel, diadochokinetic (DDK), and sentence reading, along with standardized training settings and evaluation metrics like AUC and F1 score to ensure reproducible results.

Terminology

Summary

Early-stage Parkinson’s disease (EarlyPD) detection from speech is clinically meaningful yet underexplored, and published results are hard to compare because studies differ in datasets, languages, tasks, evaluation protocols, and EarlyPD definitions.

The gist

We propose the first benchmark for speech-based EarlyPD detection with a speaker-independent split designed for fair and replicable cross-method evaluation on researcher-accessible datasets.

Benchmark Setup and Criteria for Early-Stage PD

The paper establishes standardized criteria for defining EarlyPD to address prior inconsistencies. The eligibility criteria adopted are: (i) H&Y stage ≤ 2; (ii) TAD ≤ 5 years. Participants who do not meet these criteria are referred to as the non-Early PD group in this paper. The authors prioritize the H&Y scale for its well-defined cutoff points and straightforward interpretability, while adopting a 5 year TAD cutoff, a practical compromise used in [23]. For disease duration, restricting to unmedicated cases would be neither clinically representative nor desirable.

Open-source Datasets (Open Track) and Private Data (Private Track)

The benchmark utilizes two open-source datasets: PC-GITA and NeuroVoz. PCGITA contains recordings from 100 speakers (50 PD, 50 HC), and NeuroVoz includes 108 speakers (53 PD, 55 HC). Both datasets provide "rich clinical metadata (e.g., MDS-UPDRS, H&Y stage, and TAD). Additionally, a private dataset called PERSPECTIVE-Base is utilized for the Private Track. This dataset comprises pre-therapy recordings from participants with sufficient metadata who provided explicit consent for reuse in internal research."

Benchmark Protocol and Evaluation Metrics

The benchmark protocol describes a transparent procedure for binary speech-based EarlyPD vs. HC detection. The task selection is limited to three sub-tasks in the open track: sustained vowel, diadochokinetic (DDK), and sentence reading, specifically utilizing /a/ for sustained vowel and /pa-ta-ka/ for DDK. For splits, a fixed, speaker-independent 5-fold split is used, ensuring that each validation and test set contains 6 EarlyPD and 6 HC speakers. Evaluation metrics include AUC and F1. The protocol involves training with the same maximum number of epochs and the same earlystopping settings, selecting checkpoints based on the highest validation area under the receiver operating characteristic curve (AUC), and choosing a decision threshold by maximizing the positive-class F1 score on the validation set.

Experimental Setup and Model Comparison

The study benchmarks three open-source speech-based PD detection methods: BDHPD, InceptionPD, and RECA-PD. All experiments are conducted on an NVIDIA A10 GPU. Training configurations are standardized by fixing the maximum audio duration to 10s and using consistent FFT parameters. The paper tests four training settings: (i) AllPD (EarlyPD+non-EarlyPD), (ii) AllPD-subset, (iii) EarlyPD, and (iv) EarlyPD+Private. These settings allow for targeted comparisons such as comparing Settings 2 and 3 isolates the effect of restricting PD training data to early-stage speakers, and comparing Settings 1 and 4 tests whether external EarlyPD data from a private dataset is more beneficial than non-Early PD data within the benchmark datasets.

Results and Discussion

The main results show that expanding speaker diversity, regardless of whether it originates from external cohorts, is a promising direction. Performance differences vary across tasks; RECA-PD achieves the highest F1 and AUC scores averaged over tasks, with particularly strong performance on DDK and sentence reading. Furthermore, aggregation improves performance: for BDHPD and InceptionPD, gains generally increase with more aggregated samples, while RECA-PD can degrade under small-sample aggregation but becomes positive when aggregating 10 sentences. The analysis also reveals a consistent trend toward higher performance for female speakers in all models and tasks. Finally, the gap between disease stages is largest for the sentence task, consistent with observations that sentence-level cues may benefit from broader disease-stage variability.

Conclusion

This paper presents the first benchmark for speech-based EarlyPD detection, providing a transparent protocol under different training-resource settings and establishing a shared, replicable reference point for future research. The findings suggest that expanding speaker diversity is a promising path toward improving EarlyPD detection performance, and that explainability-oriented designs can remain competitive. The proposed protocol includes both aggregate-level and gender-stratified results, which are critical for clinical adoption. Future work is encouraged to adopt this benchmark as a shared reference point.

Acknowledgments

This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD)

Improvements for AI systems

Here are the specific improvements for AI systems based on this benchmark, categorized by capability:


) 1. Enhanced Clinical Diagnostic Precision for Early-Stage PD (EarlyPD):

The improved system will move beyond simple binary classification (PD vs. HC) to provide a clinically actionable risk stratification score for patients presenting with speech anomalies. By training on the AllPD setting and comparing it against the EarlyPD setting, the system can be fine-tuned to detect subtle acoustic biomarkers specific to prodromal or very early stages of Parkinson's disease (defined by H&Y stage ≤ 2).

) 2. Robust, Generalizable Cross-Method Evaluation Framework:

The benchmark itself forces a shift from model-specific validation to a standardized, multi-dimensional comparison. The resulting AI system can be deployed as an automated Model Selector or Benchmarker. It will be able to ingest any new speech classification method and immediately compare its performance across the established four training settings (AllPD, AllPD-subset, EarlyPD, EarlyPD+Private) and three tasks (Vowel, DDK, Sentence Reading).

) 3. Task-Specific Acoustic Feature Extraction for Clinical Utility:

Since performance varies by task (DDK vs. Vowel), the system can be modularized to extract features optimized for specific clinical assessments.

  • For motor function monitoring, it prioritizes the DDK task performance metrics derived from the benchmark.

  • For broader diagnostic screening, it leverages sentence reading and sustained vowel metrics.

) 4. Adaptive Training Strategy via Data Augmentation (Leveraging Private Data):

The system can be designed to dynamically adjust its training regimen based on data availability:

  • If only accessible data is available, it defaults to the AllPD setting for general robustness.

  • If access to private, early-stage cohort data (like PERSPECTIVE-Base) is obtained, the system can switch to the EarlyPD+ setting. This allows researchers or clinicians to quantitatively assess how much performance gains are attributable specifically to incorporating rare, high-value early-stage data versus general population data.

) 5. Explainable AI (XAI) for Clinician Trust:

By integrating models like RECA-PD (which showed strong performance), the resulting system will provide decision rationales for its predictions. This allows a clinician to see not just This patient is EarlyPD, but The prediction is driven primarily by the increased frequency of [specific acoustic feature] in the DDK task, significantly boosting clinical adoption and reducing reliance on black-box outputs.

) 6. Robustness Assessment via Aggregation Analysis:

The system will incorporate an internal mechanism to evaluate its own performance under various data pooling strategies (utterance-level vs. aggregate-level, e.g., 10 sentences per speaker). This allows the AI to report a confidence score based on how robust its prediction is when presented with real-world, multi-utterance recordings, mitigating the issue of high intra-speaker variability.

Related papers