What your brain activity says about you: A review of neuropsychiatric disorders identified in resting-state and sleep EEG data
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "What your brain activity says about you".
Nadia: Electroencephalogram (EEG) monitoring devices and online data repositories hold large amounts of data from individuals participating in research and medical studies without direct reference to personal identifiers.
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: So, this paper is titled "What your brain activity says about you: A review of neuropsychiatric disorders identified in resting-state and sleep EEG data," and it's clear from the authors that they are bringing together a lot of information from different areas. What we're looking at here is how we can use brain signals, specifically EEG, to figure out what kind of mental health issues people might have based on their brain activity while they are resting or sleeping.
Elias: I see the title suggests a broad look across both resting and sleep data, which means they're not focusing on just one type of recording but trying to find common patterns in different brain states for diagnosing various conditions. The authors listed are Scanlon, Pelzer, Gharleghi, Fuhrmeister, Köllmer, Aichroth, Göder, Hansen and Wolf.
Priya: From a privacy standpoint right off the bat, I wonder what kind of personal health information they’re actually talking about when they use this EEG data for disorder classification across such different brain states. Are we looking at something that’s sensitive enough to warrant such intense scrutiny?
Nadia: Exactly, Priya; because the authors mention that EEG and online repositories hold huge amounts of data without direct personal identifiers, it makes me wonder how deep this review goes into the actual risks when we start classifying these disorders. I want to know if they're just listing facts or if they're flagging specific vulnerabilities for AI systems right now.
Elias: The implication here is that as machine learning gets better at spotting patterns in task-free EEG data, the privacy risk of re-identification becomes more concrete, which ties into other work we’ve seen on things like SoK and traffic analysis attacks against onion services.
The paper's summary: Nadia: Now, diving into the actual summary of "What your brain activity says about you: A review of neuropsychiatric disorders identified in resting-state and sleep EEG data," it seems they found that task-free EEG data can classify several conditions with high accuracy, like Autism Spectrum Disorder, Parkinson’s disease, alcohol use disorder, and Major Depressive Disorder.
Elias: They also point out that the classification performance varies depending on the data type; for instance, resting state data seemed to be effective for a wider variety of disorders compared to sleep EEG data which tended to focus more on sleep-related issues like insomnia or REM sleep disorder.
Priya: What’s interesting from my perspective is that while they achieved high accuracy in some areas, they also highlighted that the required recording times and the number of channels needed are quite different between resting state and sleep studies. They noted that for resting state data, only about five minutes of recordings were sometimes enough to get over ninety percent accuracy.
Nadia: That's a key distinction; it shows that the underlying biological signal we’re looking at requires completely different input parameters depending on whether you’re analyzing a person at rest or during sleep. This suggests that a one-size-fits-all AI model for all EEG data probably won't work well.
Elias: And when we look at the specific disorders, the paper mentions that twenty-four disorders were primarily identified using resting state EEG data, with Major Depressive Disorder being the most common finding in those studies.
The paper's improvements: Nadia: Moving on to what the authors suggest as improvements for this field, they really emphasize the need for better tools focused directly on privacy and anonymization because they found that many studies lacked critical details needed for proper replication. They pointed out that without details about machine learning methods or proper data splitting, there's a real risk of unintended information leakage.
Elias: I agree with Nadia here; it's not just about the classification accuracy, but ensuring the process itself is transparent and secure enough to prevent people from being re-identified using meta-information. This links directly to the work on how different data sets can be de-anonymized by mentioned methods.
Priya: From my viewpoint, their suggestion of focusing on anonymization tools like removing personal names and ages is crucial because those seemingly small pieces of meta-information are exactly what the authors say can lead to de-anonymization if a single dataset reveals identification directly.
Nadia: So, the improvement they propose is shifting the focus from just getting a high classification score to building better protocols for keeping study participants and medical EEG users’ privacy safe before AI advances allow for disorder classification in reidentified datasets.
Elias: That points toward needing techniques like Differential Privacy or Federated Learning to train models on distributed data without centralizing the raw EEG signals, which is a much more robust way to handle the privacy concerns raised in "What your brain activity says about you: A review of neuropsychiatric disorders identified in resting-state and sleep EEG data."
Conclusion: Nadia: So, to wrap up this discussion on "What your brain activity says about you: A review of neuropsychiatric disorders identified in resting-state and sleep EEG data," the main thing is that task-free EEG data shows promise for classifying a wide range of conditions, but we have to be extremely careful about how we use it because re-identification is a known risk.
Elias: We see high classification accuracies in both resting state and sleep data, but the paper clearly states that many studies lacked the necessary details for proper replication, which means our tools need to focus on better quality checks and preventing information leakage during model training.
Priya: I think the most important point is that as we move toward real-world applications, we have to confront the limitation that most studies only look at one disorder per participant, which doesn't reflect how people actually present with multiple conditions in the real world.
Nadia: That’s a fair point about comorbidity; it means any system needs to be built not just for single disorders but for complex patterns, and we need those privacy mitigation tools they suggested to handle that complexity safely.
Elias: We should keep an eye on how research addresses the methodological shortcomings mentioned in this review as we explore these applications further, because the authors clearly laid out where the current methods fall short.
Priya: Exactly; understanding what’s missing from these studies is just as important as knowing what they found, and that's a vital part of this whole discussion about using brain activity data responsibly.
Scanlon, J.E.M., Pelzer, A., Gharleghi, M., Fuhrmeister, K.C., Köllmer, T., Aichroth, P., Göder, R., Hansen, C., Wolf, K.I.
Fraunhofer Institute for Digital Media Technology · Department of Neurology · Department of Psychiatry and Psychotherapy
cs.NE, cs.CR, cs.CY, q-bio.NC
Submitted: 2025-10-06
Updated: 2026-09-28
Comments: 47 pages, 5 figures, 3 tables
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 83/100
The gist: Electroencephalogram (EEG) monitoring devices and online data repositories hold large amounts of data from individuals participating in research and medical studies without direct reference to
Key concepts
- Task-free EEG Data
- This refers to brain activity recordings taken without specific tasks or instructions from the participant, such as resting state or sleep stages. This type of data is used by machine learning models to classify various mental health conditions without requiring the subject to perform any specific action during recording.
- Resting State EEG
- This involves analyzing brain activity when a person is at rest, often measured in minutes of recording. Studies found that resting state EEG data can accurately classify many disorders, such as Autism Spectrum Disorder and Parkinson’s disease, often requiring very little recording time.
- Sleep EEG Data
- This data captures brain patterns while a person is asleep, including different sleep stages like REM. Sleep EEG is effective for identifying disorders related to sleep, such as insomnia and sleep apnea, typically requiring longer recordings spanning multiple stages.
Terminology
Summary
Electroencephalogram (EEG) monitoring devices and online data repositories hold large amounts of data from individuals participating in research and medical studies without direct reference to personal identifiers. This review explores what types of personal and health information have been detected and classified within task-free EEG data, investigating the privacy risks involved with openly available EEG data as machine learning methods improve re-identification capabilities.
The gist
Task-free EEG data, including resting state (REO/REC) and sleep stages, can be used by machine learning to classify various neuropsychiatric disorders such as Autism Spectrum Disorder, Parkinson’s disease, alcohol use disorder, Major Depressive Disorder (MDD), and insomnia with high classification accuracy.
Data Classification Findings
The review synthesized studies classifying medical disorders using task-free EEG data. Key findings regarding the types of data analyzed are:
-
Resting state EEG data has been used to classify a variety of disorders, including
Autism Spectrum Disorder, Parkinson’s disease, alcohol use disorder,
often requiringonly 5 mins of data or less
for high classification accuracy (>90%). -
Sleep EEG data tends to hold classifiable information about sleep disorders such as
sleep apnea, insomnia, and REM sleep disorder,
but usually involveslonger recordings or data from multiple sleep stages.
-
In terms of disorder classification,
24 disorders were primarily identified with resting state EEG data,
with the most common beingMajor Depressive Disorder (MDD) or depressive disorders (5 studies).
-
Sleep data was primarily used to identify
12 disorders, with the most common being insomnia (10 studies).
Data Characteristics Comparison
The study compared the characteristics of resting state and sleep EEG data across various metrics:
- Sample Size and Channels:
Resting state studies tended to use more subjects, with an average of 229 participants per dataset,
while sleep studies used about 40 participants per dataset.
Resting state studies also tended to use more channels, with an average of 34.3 channels,
compared to the sleep studies' average of 3.4 channels per dataset.
- Signal Time:
In the resting-state EEG studies, relatively low amounts of signal time were required to detect disorders (Figure 3).
The minimum recording time used was 30 s,
with a mean of 4.7 min.
For sleep data, the average amount of recording time was 5.3 hours.
- Classification Accuracy:
Resting state had a mean of 85.7 % accuracy while sleep data had a mean of 90.3% accuracy.
Quality and Reproducibility Assessment
The quality analysis revealed that while many studies achieved high classification accuracies, many studies lacked critical information.
-
Quality assessment showed that
8 (31%) resting state studies and 6 (35%) of the sleep studies received a ‘somewhat’ or ‘no’ in question 7 about quality of the machine learning performed in the study.
-
Many papers were found to lack details required for replication, such as
details about machine learning and recording time,
or proper data splitting without subject overlap, which is important toprevent unintended information leakage.
Privacy Risks and Future Directions
The paper emphasizes that access to EEG can reveal sensitive personal health information, raising ethical concerns regarding re-identification.
-
Reidentification is possible using machine learning on both resting state and sleep EEG data; studies have shown the ability to
reidentify individuals simply using their resting state (Zhang et al., 2018a; Zhang et al., 2020).
-
The risk is heightened if a direct personal identifier is never made accessible, as
if just one data set reveals the identification directly or based on meta-information, other data sets could be de-anonymised by the mentioned methods.
-
To mitigate risks, efforts should focus on
anonymization,
such asthe removal of all personalized and personal information such as an individual’s name and age from the research data,
and developing tools for keeping privacy safe before machine learning advances allow for disorder classification in reidentified datasets. -
Future methods may involve
removing or disguising identifying information from the EEG data to anonymize it before publishing,
oronly release necessary non-identifying features of the data upon request.
Limitations
Despite high classification accuracies, limitations exist when applying these models to real-world scenarios:
-
There are differences between study datasets and a
real world sample,
and thebase rate of many disorders is much lower
than what is often observed in controlled studies. -
Real-world data frequently includes individuals with
comorbidity, or individuals with multiple disorders at the same time,
whereas most studies listed only one disorder within each participant.
Improvements for AI systems
Based on the review of current research regarding task-free EEG data, here are specific improvements for AI systems, categorized by domain:
)1. Enhanced Diagnostic Classification Systems (Focus: Accuracy & Robustness)
The paper highlights that while classification accuracies are generally high (e.g., 90% for insomnia), many studies lack critical information regarding machine learning protocols (e.g., epoch length, training/testing split, subject overlap).
-
Improvement: Develop
Meta-Quality Check
layers within the AI pipeline that automatically assess the metadata of incoming EEG datasets against established best practices (as outlined in Section 15). -
Improvement: Implement robust data splitting mechanisms that explicitly prevent subject overlap between training and testing sets to mitigate unintended information leakage, a known vulnerability identified in Section 16.
-
Improved AI Capability: The system can achieve higher reliability when deployed on real-world, unseen datasets by ensuring the model generalizes correctly and is not biased by information leaked from the training set.
)2. Privacy-Preserving Re-identification Mitigation (Focus: Security & Anonymization)
The core concern is that task-free data, even without direct identifiers, can be used for re-identification via meta-information or sensitive features (Section 4).
-
Improvement: Integrate
Privacy Augmentation
modules using techniques like Differential Privacy or Federated Learning specifically tailored for EEG signals. This would allow model training on distributed datasets without centralizing raw EEG data. -
Improvement: Utilize privacy-preserving neural network architectures, such as those employing Homomorphic Encryption (as mentioned in the references), to perform classification directly on encrypted EEG features.
-Improved AI Capability: The system can classify disorders or identify individuals from EEG data while mathematically guaranteeing that no identifiable personal traits (even subtle ones) are extracted or revealed by the model's parameters.
)3. Domain Adaptation for Real-World Deployment (Focus: Generalization & Bias Reduction)
The paper explicitly notes a limitation: real world data will often include individuals with comorbidity, or individuals with multiple disorders at the same time, as well as varying factors such as age, which were controlled in most of these studies
(Section 18).
-
Improvement: Develop Transfer Learning frameworks where models pre-trained on large, diverse datasets (like those from the Healthy Brain Network) are fine-tuned using small, specific clinical datasets.
-
Improvement: Incorporate age and comorbidity as explicit, controllable input features into the model architecture (as opposed to relying solely on spectral features), allowing the AI to better handle real-world complexity.
-Improved AI Capability: The system moves beyond simple binary classification (e.g., Insomnia vs. No Insomnia
) to a nuanced diagnostic tool that can predict complex multimorbidity patterns based on a patient's entire EEG profile, significantly improving clinical utility in diverse patient populations.
)4. Automated Feature Engineering & Selection (Focus: Efficiency & Reproducibility)
The review shows that different disorders benefit from different features (e.g., resting state vs. sleep data use more channels).
- Improvement: Implement an automated feature selection pipeline driven by domain knowledge, which dynamically selects the optimal set of EEG features (spectral power bands, connectivity measures, fractal dimensions) based on the specific disorder target and data modality (resting vs. sleep).
-Improved AI Capability: The system can automatically adapt its input requirements. For instance, when classifying Parkinson's disease from resting state data, it prioritizes motor cortex spectral features and higher channel counts (as suggested by Suuronen et al., 2023), while for insomnia classification from sleep data, it optimizes for temporal context learning (LSTM/CNN features).
Abstract
Electroencephalogram monitoring devices and online data repositories hold large amounts of data from individuals participating in research and medical studies without direct reference to personal identifiers. This paper explores what types of personal and health information have been detected and classified within task-free EEG data. Additionally, we investigate key characteristics of the collected resting-state and sleep data, in order to determine the privacy risks involved with openly available EEG data. We used Google Scholar, Web of Science and searched relevant journals to find studies which classified or detected the presence of various disorders and personal information in resting state and sleep EEG. Only English full-text peer-reviewed journal articles or conference papers about classifying the presence of medical disorders between individuals were included. A quality analysis carried out by 3 reviewers determined general paper quality based on specified evaluation criteria. In resting state EEG, various disorders including Autism Spectrum Disorder, Parkinson's disease, and alcohol use disorder have been classified with high classification accuracy, often requiring only 5 mins of data or less. Sleep EEG tends to hold classifiable information about sleep disorders such as sleep apnea, insomnia, and REM sleep disorder, but usually involve longer recordings or data from multiple sleep stages. Many classification methods are still developing but even today, access to a person's EEG can reveal sensitive personal health information. With an increasing ability of machine learning methods to re-identify individuals from their EEG data, this review demonstrates the importance of anonymization, and the development of improved tools for keeping study participants and medical EEG users' privacy safe.
Related papers
- Evolutionary Ensemble of Agents
- Encoding and Decoding Temporal Signals with Spiking Bandpass Wavelets
- Large Language Models and Evolutionary Computation: A Critical Review of Bidirectional Interaction, Automated Algorithm Design, and Co-Adaptive Systems
- Learning Alzheimer's Disease Signatures by bridging EEG with Spiking Neural Networks and Biophysical Simulations
- Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach
- S-AI-Recursive: A Bio-Inspired and Temporal Sparse AI Architecture for Iterative, Introspective, and Energy-Frugal Reasoning