Neurai-VN Benchmark: Standardized Machine Learning Models for Multimodal Digital Phenotyping in Mental Health Classification

summary

Video file (mp4)

The gist

"Digital phenotyping (DP) using smartphones and wearable devices has shown considerable potential for mental health monitoring.

In short

The episode discusses the 'Neurai-VN Benchmark' paper, which uses multimodal digital phenotyping to classify mental health conditions using data from a Vietnamese cohort. The hosts conclude that while baseline performance is solid, the paper' introduces a crucial standardized protocol (subject-wise validation) for researchers to ensure reproducible results.

Key concepts

Digital Phenotyping
This method uses sensors found in devices like smartphones and smartwatches to monitor and understand a person's mental state. It involves tracking various data points, such as sleep patterns, activity levels, heart rate, and screen usage.
Multimodal Digital Phenotyping
This approach combines data from multiple sources—for example combining wearable measurements (like heart rate) with phone usage data (like app states). This is used to gather a comprehensive view of the user's condition for research purposes.
Subject-wise Validation
This is a rigorous testing method where all data from the same individual is kept entirely within one set, either the training set or the test set. This prevents data leakage and ensures that model performance is reliable and not artificially inflated.

Terminology used across episodes

This episode discusses

The paper

Neurai-VN Benchmark: Standardized Machine Learning Models for Multimodal Digital Phenotyping in Mental Health Classification · Read on arXiv

Quoc-Cuong Pham, Hoang-Thuy-Duong Vu, Thi-Thanh-Huong Ha, Huy-Hieu Pham

VinUniversity · Vietnam National University HCMC

Digital phenotyping (DP) using smartphones and wearable devices has shown considerable potential for mental health monitoring. However, progress remains difficult to evaluate due to heterogeneous datasets, inconsistent preprocessing pipelines. In this work, we present a reproducible benchmark built upon the Neurai-VN dataset, a high-resolution, multimodal dataset comprising passive sensing and active assessment from wearable and smartphone devices, collected from 100 Vietnamese adults over two weeks. We define four binary classification tasks evaluated using standardized subject-wise cross-validation. Representative linear, tree-based, and neural baseline models are evaluated systematically across predefined feature-group configurations. Mean subject-level F1 scores across five cross-validation folds reached 0.71 for Healthy Control vs. Depression and Healthy Control vs. Clinical, while Healthy Control vs. Anxiety and Depression vs. Anxiety achieved 0.69 and 0.56, respectively. These baseline results provide reproducible baselines for future research on multimodal DP for mental health classification tasks.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Neurai-VN Benchmark: Standardized Machine Learning Models for Multimodal Digital Phenotyping in Mental Health Classification".

Jane: The paper was written by Quoc-Cuong Pham, Hoang-Thuy-Duong Vu, Thi-Thanh-Huong Ha and Huy-Hieu Pham from VinUniversity and Vietnam National University HCMC.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everybody. We've got a fascinating new paper on the arXiv desk today, and I have to say, the title alone got me hooked. It's called "NEURAI-VN BENCHMARK: STANDARDIZED MACHINE LEARNING MODELS FOR MULTIMODAL DIGITAL PHENOTYPING IN MENTAL HEALTH CLASSIFICATION."

Jane: And Tom, I know that's a mouthful, but the idea behind it is actually pretty simple. Digital phenotyping just means using the sensors in your phone and your smartwatch to understand how you're doing mentally. This paper is trying to make that research more reliable.

Tom: Right, and the "VN" part is what really caught my eye. This is a Vietnamese dataset, which is a big deal because most of this research has been done on Western, English-speaking college students.

Jane: Exactly. The team at VinUniversity collected data from one hundred Vietnamese adults over two weeks. They had wearables tracking heart rate, sleep, activity, and the phones tracking movement, screen usage, even app states. Plus, the participants filled out daily mood logs and weekly mental health questionnaires.

Tom: And they got real clinical labels from psychiatrists. So we have healthy controls, people with depression, people with anxiety, and a group with other psychiatric conditions. That's the kind of ground truth you need to train a useful model.

Jane: The whole point of the benchmark is to give researchers a standard way to test their models. Before this, everyone was using different datasets, different preprocessing, different evaluation methods, so you couldn't compare results across studies at all.

Tom: And that's the reproducibility problem in a nutshell. The paper makes the case that a lot of the impressive numbers we've seen in digital phenotyping might be inflated because of how the data was split.

Jane: They're specifically calling out record-wise cross-validation, where data from the same person can end up in both the training set and the test set. That's like grading your own homework. The model looks great, but it falls apart in the real world.

Tom: So they're pushing subject-wise validation, where each person's data stays in one fold. It's a much harder test, but it's honest.

Jane: And that's the foundation they've built this whole benchmark on. It's not just a dataset, it's a standardized protocol. I'm really curious to see what the actual results look like.

Tom: Me too. Let's get into the numbers and see how these baseline models actually performed.

Summary: Tom: So, Jane, we've got the setup. Now let's talk about what this benchmark actually found. The paper defines four binary classification tasks, and the results are pretty sobering in a good way.

Jane: Sobering in a good way, I like that. They tested four models — Logistic Regression, Random Forest, XGBoost, and a neural network — across fifteen different feature combinations. And the best subject-level F1 scores were around zero point seven one for distinguishing healthy controls from depression, and also for healthy controls versus all clinical groups combined.

Tom: For people who don't speak metrics, an F1 score of zero point seven one is decent, but it's not magic. It means the model is getting a solid majority right, but there's real room for improvement.

Jane: Right. And the anxiety task was a bit lower at zero point six nine, and the hardest task was telling depression apart from anxiety, which only hit zero point five six. That one makes sense clinically because those conditions overlap a lot in symptoms.

Tom: I love that they reported the standard deviation too. Some of those numbers had huge variance across the folds, like plus or minus zero point two two on the depression versus anxiety task. That tells you the model is unstable depending on which people end up in the training set.

Jane: That's the kind of honesty we need in this field. A lot of papers would just report the best fold or the average and call it a day. This paper shows you the spread so you know how much to trust the result.

Tom: And the other interesting finding was about feature integration. They found that combining more modalities generally helped, but not always. The best configuration for each task was different.

Jane: For the depression task, the winning combo was wearable minute-level data, wearable daily summaries, and the daily self-reports. But for anxiety, you needed everything — all four groups together.

Tom: So there's no one-size-fits-all answer. You can't just throw every sensor at every problem and expect it to work. The signal you need depends on the condition you're trying to detect.

Jane: And that's a really important lesson for the field. It also explains why some previous studies didn't replicate — they were using the wrong features for the wrong task.

Tom: The paper also shows that the average performance goes up as you add more modalities, from zero point four eight with one group to zero point six zero with all four. So more data does help on average, even if the best individual configuration varies.

Jane: It's like cooking. More ingredients can make a better dish, but you still need the right recipe. Just dumping everything in doesn't guarantee success.

Tom: So what does this mean for the future of mental health monitoring? I'm thinking this could be a real turning point for how we validate these models.

Improvements: Jane: Tom, we've talked about the results, but I think the biggest contribution of this paper is actually the improvements it suggests for the whole field. It's not just about the numbers, it's about how we do the research.

Tom: Absolutely. And I want to bring in Lu and Meng for this one, because they're going to have strong opinions on the methodology. Lu, what do you think is the most important improvement here?

Lu: Thanks, Tom. For me, the most impactful thing is the standardized subject-wise cross-validation protocol. The paper makes a really clear case that record-wise validation inflates performance, and they cite earlier work showing that. By enforcing subject-level splits, they're setting a new standard for the field.

Meng: I agree, but I want to push back a little. The protocol is great, but what about the feature selection? They're using ANOVA F-statistics to pick the top K features, and K is determined by a formula that depends on the total number of features. That's a pretty simple approach.

Jane: So you're saying the feature selection could be more sophisticated?

Meng: Yeah. With deep learning, you could learn the feature representations directly from the raw time series instead of hand-crafting statistical descriptors like mean, min, max, skew. The paper uses those basic descriptors for the wearable and phone data, and I think that's leaving signal on the table.

Lu: That's a fair point, Meng. But I'd argue the simplicity is intentional. They want a reproducible baseline that anyone can run without a huge compute budget. If you start with complex deep learning models, it's harder to isolate whether the performance comes from the model or the data.

Tom: And they did use an MLP as one of the baselines, so it's not like they ignored neural networks entirely.

Lu: Right, but the MLP is still trained on the same hand-crafted features. There's a whole research direction waiting to be explored where you learn features end-to-end from the raw sensor streams.

Meng: And I'd add that the dataset itself is relatively small — one hundred participants. That's a limitation for deep learning. You'd need a lot more data to train a model that can learn directly from raw signals without overfitting.

Jane: So the benchmark is a starting point, not the finish line. It gives us a solid foundation, but there's room for the community to build on it with better models and more data.

Tom: And that's exactly what a good benchmark should do. It sets the bar, and then it challenges everyone to jump higher.

Lu: The other improvement I'd highlight is the demographic diversity. Most digital phenotyping datasets come from Western university students. This is a Vietnamese cohort, and that's a big step toward understanding whether these behavior-symptom associations generalize across cultures.

Meng: That's actually huge for practical deployment. If you're building a mental health monitoring app, you need to know it works for the population you're serving, not just for the population that happened to be in the original study.

Jane: So the improvements here are both technical and cultural. Better evaluation protocols, and better representation of who's being studied.

Conclusion: Tom: Well, we've covered a lot of ground on "NEURAI-VN BENCHMARK: STANDARDIZED MACHINE LEARNING MODELS FOR MULTIMODAL DIGITAL PHENOTYPING IN MENTAL HEALTH CLASSIFICATION." Let's wrap this up and get ready for the next paper.

Jane: I think the takeaway here is that this paper gives the field a much-needed reality check. The baseline results are solid but not spectacular, and that's actually a good thing. It means there's a clear benchmark to beat.

Tom: And the standardized protocol means that when someone does beat it, we'll actually be able to trust the comparison. No more cherry-picking folds or mixing training and test data from the same person.

Jane: The dataset itself is a valuable resource too. A Vietnamese cohort with psychiatrist-assigned labels, multiple sensing modalities, and daily self-reports — that's going to enable a lot of research that wasn't possible before.

Tom: And the code is available on GitHub, so anyone can reproduce the experiments. That's the kind of transparency that moves the field forward.

Jane: There are still open questions, of course. The depression versus anxiety task was the hardest, and that's a clinically important distinction. Future work will need to dig into what features can separate those conditions better.

Tom: And we haven't even talked about the OPC group — the other psychiatric conditions. The paper includes them in one of the tasks, but there's room to explore that group more deeply.

Jane: For now, though, this paper gives us a honest, reproducible foundation for digital phenotyping research. It's not claiming to have solved mental health detection, but it's giving us a fair way to measure progress.

Tom: And that's exactly what we need. Alright, we're going to say goodbye to this paper and move on to the next one on the arXiv stack.

Jane: Thanks for listening, everyone. We'll be right back with more research, more discussion, and more ideas that could change how we understand the human mind.

Tom: See you in a minute.

More episodes

← Home