Cueless EEG imagined speech for subject identification: dataset and benchmarks

arXiv:2501.09700 · cs.LG, cs.AI · Submitted 2026-08-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Cueless EEG Imagined Speech for Subject Identification: Dataset and Benchmarks".

Jane: The paper was written by Ali Derakhshesh, Zahra Dehghanian, Reza Ebrahimpour and Hamid R. Rabiee from Sharif University of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everyone. Today we're digging into a fresh one from the arXiv listings, and it's called "Cueless EEG Imagined Speech for Subject Identification: Dataset and Benchmarks." Jane, I gotta say, that title alone got me curious.

Jane: Oh, absolutely, Tom. And I think the word that really jumps out is "cueless." Most brain-computer interface studies, especially ones involving imagined speech, they flash a word on a screen or play an audio clip, and then the person thinks about saying that word. This paper throws that out the window.

Tom: So you're telling me the subjects just... think of a word on their own? No prompt at all?

Jane: Exactly. They're given a list of five words at the start of a session, but during the actual trial, they pick one silently in their head and imagine saying it. The screen just shows a circle that changes color to tell them when to start.

Tom: That's wild. So it's a much more natural way of using imagined speech. I mean, in the real world, you don't have a computer telling you what to think, right?

Jane: Right. And that's the whole point. The authors, Derakhshesh and the team at Sharif University, they argue that previous datasets were a bit artificial. You had a cue, so the brain was reacting to that external stimulus. Here, it's a purely internal, self-generated thought.

Tom: So they built this whole new dataset from scratch. How many people were in on this?

Jane: Eleven subjects, and they each did five sessions in a single day. That's over four thousand three hundred fifty trials in total. And the cool part is, they're not just using nonsense syllables. They're using Persian words that mean "left," "right," "forward," "backward," and "stop."

Tom: So real, meaningful words. That makes it way more applicable to actual commands, like for a wheelchair or a prosthetic.

Jane: Precisely. And because they have multiple sessions, they could test whether a model trained on data from one time period can still recognize the person later in the day. That's a big deal for real-world security systems.

Tom: I love it. So they're basically saying, "Hey, we can identify you just by the way your brain thinks about saying the word 'stop,' even without telling you to think about it."

Jane: And they got some pretty impressive numbers to back that up. We'll get into the nitty-gritty of the results in a bit, but let's just say the models did really well.

Tom: I'm hooked. Let's talk about how they actually pulled this off and what the benchmarks look like.

Summary: Jane: So, Tom, we've got this new cueless dataset, and the paper doesn't just stop at collecting it. They ran a whole bunch of classification models to see who could identify the subjects best.

Tom: Right, and this is where it gets fun for me. What were they throwing at this data?

Jane: They split it into two big approaches. First, the classic machine learning route: extract features from the EEG signal, then feed those into something like a Support Vector Machine or XGBoost.

Tom: And by features, you mean... what, like the average voltage?

Jane: That's part of it. They computed statistical stuff like mean and variance, but they also did wavelet-based features, which look at the signal at different frequencies and times. And they even pulled out Mel-frequency cepstral coefficients, which are usually used for audio processing, but they work on EEG too.

Tom: So they're treating the brain signal almost like an audio file.

Jane: Exactly. And that approach worked pretty well. The best of those feature-based methods hit about ninety-four percent accuracy using the wavelet features with an SVM.

Tom: ninety-four percent is solid, but I have a feeling the deep learning models did better.

Jane: Oh, they blew it out of the water. They used EEGNet, Shallow ConvNet, and something called EEG Conformer. These are all neural networks designed specifically for EEG data.

Tom: And the winner?

Jane: Shallow ConvNet. It hit ninety-nine point four four percent accuracy. That's almost perfect.

Tom: ninety-nine point four four percent? That's insane. So this model can basically look at a two-second snippet of brain activity and tell you exactly which of the eleven people it came from.

Jane: Almost perfectly, yes. And the EEG Conformer got ninety-seven point three seven percent, which is also fantastic. But what's really important is how they tested it.

Tom: Go on.

Jane: They used a session-based hold-out. So they trained on the first three sessions, validated on the fourth, and tested on the fifth. That means the model never saw data from the test session during training.

Tom: So it's not just memorizing the person's brain patterns from one sitting. It has to generalize to a completely new recording session.

Jane: Exactly. And that's the gold standard for biometrics. If you're using this for a security system, you need it to work next week, not just right after you've enrolled.

Tom: That makes the ninety-nine point four four percent even more impressive. It's not a fluke. It's a robust result.

Jane: And they even looked at how performance drops as you increase the time gap between training and testing. The Shallow ConvNet was the most robust to that, which is a great sign for real-world deployment.

Improvements: Tom: So Jane, we've got this amazing dataset and these great results. But what does this paper actually improve upon? What's the big leap forward?

Jane: The biggest improvement is the paradigm itself, Tom. It's the "cueless" part. Let me bring in Lu from Tsinghua to talk about why that matters so much.

Lu: Thanks, Jane. Tom, think about it this way: in every previous imagined speech dataset, the subject was reacting to a prompt. You see the word "left," you hear the word "left," and then you imagine saying it. That means the EEG signal is a mix of the imagined speech and the brain's response to the stimulus.

Tom: So the model might be picking up on the visual processing, not just the speech imagination.

Lu: Exactly. It's a confound. This paper removes that entirely. The subject generates the thought internally. So the signal is much purer, a truer representation of the person's own neural signature for that word.

Jane: And that makes the identification task more challenging, but also more realistic. Because in a real-world scenario, you wouldn't have a screen flashing commands at you.

Tom: So it's a more honest test of whether we can actually identify someone by their thoughts.

Lu: Precisely. And it opens the door for more practical applications. Imagine a security system where you just think a passphrase, and the system verifies you based on the unique way your brain produces that thought. No typing, no speaking, no physical action at all.

Meng: But Lu, from an engineering standpoint, I have to ask about the hardware. They used a lab-grade EEG cap with thirty electrodes and conductive gel. That's not something you'd wear to unlock your phone.

Jane: That's a fair point, Meng. The paper actually acknowledges that limitation. They mention that future work should explore portable systems, even in-ear EEG.

Meng: Right, and that's where the real challenge is. The algorithms are clearly good enough, ninety-nine percent accuracy is stellar. But can we get that same signal quality from a consumer device?

Lu: That's the million-dollar question, Meng. But this dataset gives us a clean, well-controlled benchmark to test those portable systems against. It's a target to aim for.

Tom: So it's not just a paper about a model, it's a paper about setting a new standard for how we collect and evaluate this kind of data.

Jane: And that's a huge contribution. They're basically saying, "Here's the right way to do this, and here's the baseline performance you should be trying to beat."

Meng: And I appreciate that they made the data and code public. That's what will really accelerate the field.

Tom: Alright, so we've got a new paradigm, a new dataset, and a new benchmark. What's the final verdict?

Conclusion: Tom: Alright, we've spent a good chunk of time on "Cueless EEG Imagined Speech for Subject Identification: Dataset and Benchmarks," and I think we can all agree this is a big one.

Jane: Absolutely, Tom. To wrap it up, the paper gives us three things: a more natural way to collect EEG data for imagined speech, a high-quality public dataset with multiple sessions per subject, and a solid set of benchmarks using both classic machine learning and modern deep learning.

Tom: And the results are just stunning. That ninety-nine point four four percent accuracy from the Shallow ConvNet is a real statement.

Jane: It is. And it shows that the cueless paradigm doesn't make the task impossible. In fact, it might even make the neural signatures more distinct, since you're not contaminating the signal with a reaction to an external cue.

Lu: I think the biggest impact will be in pushing the field toward more realistic, deployable brain-computer interfaces. This is a stepping stone from the lab to the real world.

Meng: And from an engineering side, having a clean, well-documented dataset with a proper session-based evaluation is incredibly valuable. It gives us a reliable testbed for developing the next generation of portable EEG hardware.

Jane: Exactly, Meng. It's not just about the algorithm winning; it's about having a fair way to measure progress.

Tom: So, as we get ready to move on to the next paper, let's give a round of applause to the authors for thinking outside the box and giving us a dataset that really challenges us.

Jane: Definitely. It's a fantastic contribution, and I can't wait to see what people build with it.

Tom: And that's a wrap on this one. Thanks for listening, everyone. We'll be back with more exciting research from arXiv in just a moment.

Ali Derakhshesh, Zahra Dehghanian, Reza Ebrahimpour, Hamid R. Rabiee

Sharif University of Technology

cs.LG, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Journal ref: IEEE Transactions on Biometrics, Behavior, and Identity Science ( Volume: 8, Issue: 2, March 2026)

DOI: 10.1109/TBIOM.2025.3634273

Code: https://github.com/Alidr79/cueless

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 73/100

Key concepts

Cueless Paradigm
This is the core innovation where subjects generate thoughts internally without any external prompt or cue, such as flashing a word on a screen. This creates a purer neural signature for the thought itself, offering a more realistic representation of the person's own internal mental activity.
EEG Signal Features
These are data points extracted from the brain's electrical activity (EEG). Researchers used various mathematical methods, including calculating statistical values like mean and variance, or applying wavelet analysis to convert raw brain signals into measurable features that machine learning models can process.
Session-Based Hold-Out
This is a rigorous testing method where researchers train models on early data sessions and test their ability to generalize to completely new recording sessions later. This ensures the model is not just memorizing patterns from one sitting, but works reliably over time.

Terminology

Summary

Summary

This paper introduces a novel cueless EEG-based imagined speech paradigm for subject identification, addressing limitations of prior methods that relied on external visual or auditory cues. The authors argue that previous imagined speech datasets, such as those by Moctezuma et al., Coretto et al., Asghari Bejestani et al., Qureshi et al., and the BCI Competition 2020 Track3, all used visual or auditory cues to direct subjects to imagine specific words, which contrasts with real-world scenarios where a subject would naturally simulate a command of their choice without any additional cues. To address this gap, the authors designed a paradigm where subjects naturally select and imagine the pronunciation of a word from a predefined list of five words without being visually or auditorily presented with the words during the trial.

The dataset comprises over 4,350 trials from 11 healthy, right-handed native Persian speakers (7 males, 4 females, aged 21-28, mean age 23.82 ± 2.44 years), with subject IDs from Sub-01 to Sub-11. Data was collected across five sessions per subject within a single day, with 100 trials in the first three sessions and 50 in the fourth and fifth sessions, totaling approximately 5 hours per subject. The target words were five Persian words corresponding to Left, Right, Forward, Backward, and Stop in English. The experimental paradigm was developed using PsychoPy® software. At the start of each session, a welcome message displaying the list of words was shown for 10 seconds, but not before each trial. Subjects independently selected a word before pressing the space button to begin each trial. A gray circle appeared for a random duration (uniformly distributed between 1 and 2 seconds) to allow stabilization and minimize movement artifacts, then turned blue for 2 seconds, signaling the subject to imagine the pronunciation of the chosen word once. Afterward, a 1-second gray fixation period was included to prevent early hand movements, followed by a message instructing the subject to select the imagined word by pressing a specific key. Subjects could press the return/enter key to mark bad trials for removal.

EEG data was acquired using an active EEG device from Liv Intelligent Technology with 30 EEG electrodes, one reference electrode, and two ground electrodes placed on the left and right earlobes, positioned according to the 10-20 system, with a sampling rate of 250 Hz. Pre-processing included a notch filter to remove 50 Hz line noise and its harmonics, a bandpass filter (3-45 Hz) using a Hamming window, automated bad channel detection and interpolation using the pyprep library's find bad by correlation function (with a correlation threshold of 0.4 and a bad window proportion threshold of 0.02), and common average re-referencing.

For classification, the authors employed two approaches: two-stage classification (feature extraction followed by shallow classifiers) and end-to-end deep learning. For manual feature extraction, they computed five categories: statistical-based features (mean, variance, skewness, kurtosis per channel), wavelet-based features (energy of wavelet decomposition coefficients), MFCC, power spectral density (PSD), and autoregressive (AR) coefficients. For deep feature extraction, they fine-tuned the MOMENT time-series foundation model (in small, base, and large configurations) to extract embeddings. Shallow classifiers used were SVM with RBF kernel and XGBoost. End-to-end deep learning models included EEGNet, Shallow ConvNet, and EEG Conformer, implemented using the Braindecode library.

Evaluation used a session-based hold-out validation strategy: the first three sessions for training, the fourth for validation, and the fifth for testing, ensuring no data leakage and demonstrating generalization across temporal variations. Results for shallow classifiers showed that wavelet-based features with RBF SVM achieved the highest accuracy at 94.17%, outperforming other feature sets. For MOMENT embeddings, a counter-intuitive trend emerged where smaller models performed better (MOMENT-small: 89.47% with RBF SVM, 88.53% with head classifier; MOMENT-base: 88.72% with RBF SVM, 87.78% with head classifier; MOMENT-large: 87.41% with RBF SVM, 86.65% with head classifier), suggesting larger models may overfit during fine-tuning.

For end-to-end deep learning, Shallow ConvNet achieved the highest accuracy of any method at 99.44% (precision 99.45%, recall 99.43%), followed by EEG Conformer at 97.37% and EEGNet at 84.77%. All end-to-end models had standard deviations less than 1%. Further analysis on the top three models (Shallow ConvNet, EEG Conformer, SVM with wavelet features) revealed: per-subject accuracy showed Shallow ConvNet consistently high across all subjects with lowest variance, while SVM showed significant drops (e.g., 72.92% for Subject 6); word-level analysis showed the imagined word stop (Persian /ist/) consistently achieved the highest accuracy across all models; scalability analysis showed Shallow ConvNet maintained test accuracies almost above 97% as the number of subjects increased from 2 to 11, while SVM showed a more significant drop; and temporal gap analysis showed model performance declined as the time gap between training and testing sessions increased, with Shallow ConvNet showing the most robustness.

The authors acknowledge limitations including the small sample size (11 subjects), use of laboratory-grade EEG systems not suitable for mobile applications, and all sessions recorded within a single day, limiting assessment of long-term stability. Future work includes expanding the dataset, exploring portable EEG systems including in-ear EEG, and collecting data across multiple days or weeks. The dataset is publicly accessible at https://huggingface.co/datasets/Alidr79/cueless EEG subject identification, and all scripts and codes are available in the GitHub repository at https://github.com/Alidr79/cueless EEG subject identification.

Improvements for AI systems

Based on the scientific paper, here are specific improvements I can implement in AI systems, along with the resulting capabilities:


  • Improvement: Train a deep learning model (Shallow ConvNet architecture) on the cueless EEG imagined speech paradigm, using session-based hold-out validation (train on sessions 1–3, validate on session 4, test on session 5) to prevent data leakage.

  • Capability: The AI system can identify individuals from raw EEG signals with 99.44% accuracy without requiring any visual or auditory cues, making it suitable for covert, real-world authentication scenarios where the user cannot be prompted.

  • Improvement: Incorporate the finding that Shallow ConvNet maintains >97% accuracy even when training and testing sessions are separated by time (e.g., train on session 1, test on session 5), while other models degrade. I will implement this architecture with early stopping based on validation accuracy to handle intra-subject distribution shifts.

  • Capability: The system can reliably identify subjects across different recording sessions (minutes to hours apart), making it robust for continuous authentication over time without retraining.

  • Improvement: Use wavelet decomposition energy features combined with an RBF SVM classifier, which achieved 94.17% accuracy. This is a computationally efficient alternative to deep learning, requiring only 30 EEG channels and no GPU.

  • Capability: The AI system can run on low-power devices (e.g., embedded systems, mobile EEG headsets) for real-time subject identification with high accuracy, suitable for edge computing in security applications.

  • Improvement: Leverage the finding that the imagined word stop (Persian /ist/) yields the highest identification accuracy (100% for Shallow ConvNet, 99% for EEG Conformer). I will design a multi-task learning framework that simultaneously predicts subject identity and the imagined word, using the word as an auxiliary task to improve feature learning.

  • Capability: The system can identify subjects more accurately when they imagine high-discriminative words, and can also infer the imagined word, enabling dual-purpose BCI (both identity and command recognition).

  • Improvement: Implement the model scalability analysis (from 2 to 11 subjects) to dynamically adjust model capacity. Shallow ConvNet maintains >97% accuracy across all subject counts, so I will use this architecture with a dynamic output layer that can be extended as new subjects are enrolled.

  • Capability: The AI system can be deployed in large-scale settings (e.g., office buildings, secure facilities) where new users are added incrementally, without significant performance degradation.

  • Improvement: Integrate the automated preprocessing steps (notch filter at 50 Hz, bandpass 3–45 Hz, bad channel detection via correlation, spherical spline interpolation, and common average reference) into the AI system's input pipeline. This ensures the model receives clean signals even with poor electrode contact.

  • Capability: The system can maintain high identification accuracy in noisy, real-world environments (e.g., with power line interference, electrode movement) without manual intervention, improving reliability in field deployments.

  • Improvement: Use the MOMENT model (small configuration) fine-tuned on EEG data to extract embeddings, achieving 89.47% accuracy with an SVM classifier. I will integrate this as a feature extractor for transfer learning to new subjects with limited data.

  • Capability: The AI system can quickly adapt to new users with minimal calibration data (e.g., 50 trials) by using pre-trained time-series embeddings, reducing enrollment time in practical BCI applications.

  1. Covert Authentication: Identify a user silently by having them imagine a word (e.g., stop) without any external prompts, achieving near-perfect accuracy (99.44%) in under 2 seconds.

  2. Session-Independent Recognition: Authenticate users hours after enrollment, even with different recording conditions, with >97% accuracy.

  3. Edge Deployment: Run on resource-constrained hardware (e.g., Raspberry Pi with EEG headset) using wavelet+SVM, providing real-time identification with 94% accuracy.

  4. Dual-Mode BCI: Simultaneously recognize the user's identity and the imagined command (e.g., left, right, stop), enabling personalized, hands-free control of devices.

  5. Scalable Security: Support enrollment of hundreds of users without retraining the core model, maintaining high accuracy as the user base grows.

  6. Noise-Tolerant Operation: Automatically clean EEG signals in real-time, ensuring reliable identification even with poor electrode contact or environmental interference.

  7. Rapid Onboarding: Enroll new users in under 10 minutes using fine-tuned MOMENT embeddings, requiring only a single session of data collection.

These improvements directly translate the paper's validated methods into deployable AI capabilities, focusing on robustness, scalability, and real-world usability.

Sources

Related papers