Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

arXiv:2608.13425 · cs.CL, eess.AS, eess.SP · Submitted 2026-08-13 · Read on arXiv

Serli Kopar, Sam Gijsen, Abner Hernandez, Paula Andrea Perez-Toro, Kerstin Ritter

Hertie Institute for AI in Brain Health, University of Tübingen · Tübingen AI Center, University of Tübingen · Charité–Universitätsmedizin, Department of Psychiatry and Psychotherapy · Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg

cs.CL, eess.AS, eess.SP

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This study investigates whether self-supervised learning (SSL) speech representations capture disease-related characteristics of Parkinson’s disease (PD) or instead exploit dataset-specific

Terminology

Summary

This study investigates whether self-supervised learning (SSL) speech representations capture disease-related characteristics of Parkinson’s disease (PD) or instead exploit dataset-specific confounds, particularly under cross-lingual transfer. The authors perform a layer-wise analysis of nine SSL speech backbones using a low-capacity logistic regression probe across three languages (Spanish, German, Czech). The evaluation is structured as five scenarios that progressively introduce distribution shifts in participant identity, recording conditions, language, and pathology.

The paper reports two key findings. First, layer selection is highly corpus-dependent: the optimal representation layer is determined primarily by the source dataset rather than by the SSL architecture itself. For example, on the DDK task, WavLM-L selects layer 24 for German, layer 1 for Spanish, and layer 22 for Czech (σ = 0.43). Second, the transferred discriminative signal lacks pathological specificity: classifiers trained to detect PD assign similarly high probabilities to both PD and dementia speech in the target corpus.

The five scenarios are: REF (within-corpus cross-validation), S1 (+Re-Take, same speakers with second recordings), S2 (+Condition, clean vs. noisy Spanish corpora), S3 (+Language, cross-lingual transfer), S4 (+Task, transfer to German TREND cohort with different tasks), and S5 (+Task +Language, combined language and task shift). Results show progressive degradation: mean ∆BA of −1.9 ± 4.5 for S1, −12.5 ± 14.3 for S2, and −16.3 ± 10.6 for S3 relative to REF.

For S1, VOWEL exhibits the greatest sensitivity to recording re-takes, with average performance decreases of 3.8 ± 5.5 BA points on Spanish and 4.2 ± 3.3 BA points on Czech. For S2, cross-condition transfer is asymmetric: training on noisy recordings transfers better to clean recordings than vice versa, with mean ∆BA improving from −19.1 ± 19.7 to −8.1 ± 15.9 on DDK and from −16.3 ± 10.0 to −6.6 ± 4.6 on READ. The handcrafted eGeMAPS baseline is the most condition-sensitive (−32.4 ± 26.5). For S3, no consistent advantage of multilingual over monolingual SSL backbones is observed, and task and backbone choice have a stronger influence on performance than the monolingual versus multilingual distinction.

For S4 and S5, out of 540 transfer combinations, only 9 achieved at least moderate separation between PD and healthy controls (AUC > 0.6 and BA > 0.6). The highest-BA classifiers for each training corpus achieve only modest separation between PD and matched controls (DE = 0.67, ES = 0.66, CZ = 0.64). Critically, these classifiers fail to distinguish PD from dementia according to both permutation testing and the Mann–Whitney U test. Additionally, when adjusting PD probability scores for demographics (age, sex, education) using leave-one-subject-out regression, the residualised performance drops to chance level (AUC = 0.50–0.55), supporting the claim of non-specificity.

The authors conclude that speech-based PD detection is driven more by corpus-specific factors than by PD-related motor or cognitive pathology. Under cross-lingual and cross-task transfer, models largely preserve a generic patient–healthy separation rather than disease-specific structure. The paper states: what survives across transfer settings is primarily driven by the corpus, reflecting a general divergence from healthy speech, rather than a robust speech signature specific to Parkinson’s disease. The authors interpret their findings as absence of evidence for PD-specificity, not evidence of its absence, and note limitations due to modest cohort sizes.

Improvements for AI systems

Improvements to AI systems:

  1. Domain-shift-aware layer selection: Instead of using a fixed or architecture-derived SSL layer for downstream tasks, AI systems should dynamically select representation layers based on the target corpus’s distributional properties (e.g., recording conditions, language, speaker demographics). This can be implemented via a meta-learner that predicts the optimal layer from a small labeled or unlabeled sample of the target domain, reducing the observed σ=0.43 variance in layer choice.

  2. Confound-aware calibration for clinical speech classifiers: AI systems for disease detection should include a built-in confound detector that quantifies the similarity between the training and target corpora (e.g., via domain-adversarial training or a separate classifier for recording condition, language, and participant identity). If the confound score exceeds a threshold, the system should flag the prediction as low-confidence or abstain, preventing false positives in cross-lingual or cross-task settings.

  3. Disease-specificity verification module: After training a binary classifier (e.g., PD vs. healthy), the AI system should automatically run a secondary validation against a different neurological condition (e.g., dementia) using a small held-out set. If the classifier cannot distinguish the target disease from the alternative condition (as shown by permutation testing and Mann–Whitney U), the system should downgrade its output from “disease-specific” to “generic patient vs. healthy” and adjust the confidence score accordingly.

  4. Demographic-residualized prediction: AI systems should incorporate a post-hoc regression step that removes the influence of age, sex, and education from the raw probability scores (leave-one-subject-out). The improved system would report both raw and residualized AUC; if the residualized AUC drops to chance (0.50–0.55), the system should explicitly warn that the signal is non-specific and likely driven by demographic or corpus confounds.

  5. Asymmetric transfer-aware training: For cross-condition deployment (e.g., noisy to clean or vice versa), AI systems should be trained with a curriculum that explicitly models asymmetric transferability. The system can learn a transferability matrix between conditions and, when only one condition is available for training, automatically augment the training set with simulated samples from the harder-to-transfer direction (e.g., adding noise to clean speech) to improve robustness, as the paper shows clean-to-noisy transfer is worse.

  6. Task- and language-agnostic feature fusion: Instead of relying solely on SSL embeddings, the improved AI system should fuse SSL representations with handcrafted features (e.g., eGeMAPS) but only after applying a condition-invariance penalty. The system would learn to down-weight features that are highly sensitive to recording conditions (like eGeMAPS, which showed −32.4 ∆BA) and up-weight features that are more stable across languages and tasks, based on the layer-wise and feature-wise stability scores computed during training.

What the improved AI system can do:

  • Reliable cross-lingual clinical screening: It can deploy PD detection models across new languages (e.g., Spanish, German, Czech) with automatic layer re-selection and confound calibration, reducing false positives from corpus-specific artifacts.

  • Transparent clinical decision support: It can output a three-tier label: “disease-specific,” “generic patient deviation,” or “inconclusive due to confounds,” along with residualized AUC, so clinicians know whether the model’s signal is truly pathological or just reflects a general divergence from healthy speech.

  • Robust transfer under mismatched conditions: It can maintain acceptable performance (BA > 0.6) when moving from clean to noisy recordings or from one task (e.g., DDK) to another (e.g., reading), by explicitly modeling asymmetric transfer and avoiding the observed −16 to −19 ∆BA drops.

  • Avoid overclaiming in small cohorts: It can automatically flag when cohort sizes are too small to establish disease-specificity (as in this study’s limitation), and instead report the system’s performance as exploratory rather than confirmatory.

  • Adaptive feature selection: It can dynamically combine SSL layers and handcrafted features based on the target corpus’s condition profile, preventing the catastrophic degradation seen with eGeMAPS in cross-condition scenarios.

Abstract

Self-supervised learning (SSL) speech representations achieve strong performance for Parkinson's disease (PD) detection within individual corpora. However, it remains unclear whether these models capture disease-related characteristics or exploit dataset-specific confounds, particularly since most SSL backbones are pretrained exclusively on healthy speech. To investigate this question, we perform a layer-wise analysis of nine SSL speech backbones using a low-capacity logistic regression probe across three languages. We structure the evaluation as multiple scenarios that progressively introduce distribution shifts in participant identity, recording conditions, language, and pathology. Our results reveal two key findings. First, layer selection is highly corpus-dependent: the optimal representation layer is determined primarily by the source dataset rather than by the SSL architecture itself. Second, the transferred discriminative signal lacks pathological specificity: classifiers trained to detect PD assign similarly high probabilities to both PD and dementia speech in the target corpus. These results highlight critical limitations that must be addressed before speech-based pathology recognition models can be reliably deployed in clinical settings.

Sources

Related papers