Neurai-VN Benchmark: Standardized Machine Learning Models for Multimodal Digital Phenotyping in Mental Health Classification

arXiv:2607.25232 · cs.LG · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Neurai-VN Benchmark: Standardized Machine Learning Models for Multimodal Digital Phenotyping in Mental Health Classification".

Jane: The paper was written by Quoc-Cuong Pham, Hoang-Thuy-Duong Vu, Thi-Thanh-Huong Ha and Huy-Hieu Pham from VinUniversity and Vietnam National University HCMC.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everybody. We've got a fascinating new paper on the arXiv desk today, and I have to say, the title alone got me hooked. It's called "NEURAI-VN BENCHMARK: STANDARDIZED MACHINE LEARNING MODELS FOR MULTIMODAL DIGITAL PHENOTYPING IN MENTAL HEALTH CLASSIFICATION."

Jane: And Tom, I know that's a mouthful, but the idea behind it is actually pretty simple. Digital phenotyping just means using the sensors in your phone and your smartwatch to understand how you're doing mentally. This paper is trying to make that research more reliable.

Tom: Right, and the "VN" part is what really caught my eye. This is a Vietnamese dataset, which is a big deal because most of this research has been done on Western, English-speaking college students.

Jane: Exactly. The team at VinUniversity collected data from one hundred Vietnamese adults over two weeks. They had wearables tracking heart rate, sleep, activity, and the phones tracking movement, screen usage, even app states. Plus, the participants filled out daily mood logs and weekly mental health questionnaires.

Tom: And they got real clinical labels from psychiatrists. So we have healthy controls, people with depression, people with anxiety, and a group with other psychiatric conditions. That's the kind of ground truth you need to train a useful model.

Jane: The whole point of the benchmark is to give researchers a standard way to test their models. Before this, everyone was using different datasets, different preprocessing, different evaluation methods, so you couldn't compare results across studies at all.

Tom: And that's the reproducibility problem in a nutshell. The paper makes the case that a lot of the impressive numbers we've seen in digital phenotyping might be inflated because of how the data was split.

Jane: They're specifically calling out record-wise cross-validation, where data from the same person can end up in both the training set and the test set. That's like grading your own homework. The model looks great, but it falls apart in the real world.

Tom: So they're pushing subject-wise validation, where each person's data stays in one fold. It's a much harder test, but it's honest.

Jane: And that's the foundation they've built this whole benchmark on. It's not just a dataset, it's a standardized protocol. I'm really curious to see what the actual results look like.

Tom: Me too. Let's get into the numbers and see how these baseline models actually performed.

Summary: Tom: So, Jane, we've got the setup. Now let's talk about what this benchmark actually found. The paper defines four binary classification tasks, and the results are pretty sobering in a good way.

Jane: Sobering in a good way, I like that. They tested four models — Logistic Regression, Random Forest, XGBoost, and a neural network — across fifteen different feature combinations. And the best subject-level F1 scores were around zero point seven one for distinguishing healthy controls from depression, and also for healthy controls versus all clinical groups combined.

Tom: For people who don't speak metrics, an F1 score of zero point seven one is decent, but it's not magic. It means the model is getting a solid majority right, but there's real room for improvement.

Jane: Right. And the anxiety task was a bit lower at zero point six nine, and the hardest task was telling depression apart from anxiety, which only hit zero point five six. That one makes sense clinically because those conditions overlap a lot in symptoms.

Tom: I love that they reported the standard deviation too. Some of those numbers had huge variance across the folds, like plus or minus zero point two two on the depression versus anxiety task. That tells you the model is unstable depending on which people end up in the training set.

Jane: That's the kind of honesty we need in this field. A lot of papers would just report the best fold or the average and call it a day. This paper shows you the spread so you know how much to trust the result.

Tom: And the other interesting finding was about feature integration. They found that combining more modalities generally helped, but not always. The best configuration for each task was different.

Jane: For the depression task, the winning combo was wearable minute-level data, wearable daily summaries, and the daily self-reports. But for anxiety, you needed everything — all four groups together.

Tom: So there's no one-size-fits-all answer. You can't just throw every sensor at every problem and expect it to work. The signal you need depends on the condition you're trying to detect.

Jane: And that's a really important lesson for the field. It also explains why some previous studies didn't replicate — they were using the wrong features for the wrong task.

Tom: The paper also shows that the average performance goes up as you add more modalities, from zero point four eight with one group to zero point six zero with all four. So more data does help on average, even if the best individual configuration varies.

Jane: It's like cooking. More ingredients can make a better dish, but you still need the right recipe. Just dumping everything in doesn't guarantee success.

Tom: So what does this mean for the future of mental health monitoring? I'm thinking this could be a real turning point for how we validate these models.

Improvements: Jane: Tom, we've talked about the results, but I think the biggest contribution of this paper is actually the improvements it suggests for the whole field. It's not just about the numbers, it's about how we do the research.

Tom: Absolutely. And I want to bring in Lu and Meng for this one, because they're going to have strong opinions on the methodology. Lu, what do you think is the most important improvement here?

Lu: Thanks, Tom. For me, the most impactful thing is the standardized subject-wise cross-validation protocol. The paper makes a really clear case that record-wise validation inflates performance, and they cite earlier work showing that. By enforcing subject-level splits, they're setting a new standard for the field.

Meng: I agree, but I want to push back a little. The protocol is great, but what about the feature selection? They're using ANOVA F-statistics to pick the top K features, and K is determined by a formula that depends on the total number of features. That's a pretty simple approach.

Jane: So you're saying the feature selection could be more sophisticated?

Meng: Yeah. With deep learning, you could learn the feature representations directly from the raw time series instead of hand-crafting statistical descriptors like mean, min, max, skew. The paper uses those basic descriptors for the wearable and phone data, and I think that's leaving signal on the table.

Lu: That's a fair point, Meng. But I'd argue the simplicity is intentional. They want a reproducible baseline that anyone can run without a huge compute budget. If you start with complex deep learning models, it's harder to isolate whether the performance comes from the model or the data.

Tom: And they did use an MLP as one of the baselines, so it's not like they ignored neural networks entirely.

Lu: Right, but the MLP is still trained on the same hand-crafted features. There's a whole research direction waiting to be explored where you learn features end-to-end from the raw sensor streams.

Meng: And I'd add that the dataset itself is relatively small — one hundred participants. That's a limitation for deep learning. You'd need a lot more data to train a model that can learn directly from raw signals without overfitting.

Jane: So the benchmark is a starting point, not the finish line. It gives us a solid foundation, but there's room for the community to build on it with better models and more data.

Tom: And that's exactly what a good benchmark should do. It sets the bar, and then it challenges everyone to jump higher.

Lu: The other improvement I'd highlight is the demographic diversity. Most digital phenotyping datasets come from Western university students. This is a Vietnamese cohort, and that's a big step toward understanding whether these behavior-symptom associations generalize across cultures.

Meng: That's actually huge for practical deployment. If you're building a mental health monitoring app, you need to know it works for the population you're serving, not just for the population that happened to be in the original study.

Jane: So the improvements here are both technical and cultural. Better evaluation protocols, and better representation of who's being studied.

Conclusion: Tom: Well, we've covered a lot of ground on "NEURAI-VN BENCHMARK: STANDARDIZED MACHINE LEARNING MODELS FOR MULTIMODAL DIGITAL PHENOTYPING IN MENTAL HEALTH CLASSIFICATION." Let's wrap this up and get ready for the next paper.

Jane: I think the takeaway here is that this paper gives the field a much-needed reality check. The baseline results are solid but not spectacular, and that's actually a good thing. It means there's a clear benchmark to beat.

Tom: And the standardized protocol means that when someone does beat it, we'll actually be able to trust the comparison. No more cherry-picking folds or mixing training and test data from the same person.

Jane: The dataset itself is a valuable resource too. A Vietnamese cohort with psychiatrist-assigned labels, multiple sensing modalities, and daily self-reports — that's going to enable a lot of research that wasn't possible before.

Tom: And the code is available on GitHub, so anyone can reproduce the experiments. That's the kind of transparency that moves the field forward.

Jane: There are still open questions, of course. The depression versus anxiety task was the hardest, and that's a clinically important distinction. Future work will need to dig into what features can separate those conditions better.

Tom: And we haven't even talked about the OPC group — the other psychiatric conditions. The paper includes them in one of the tasks, but there's room to explore that group more deeply.

Jane: For now, though, this paper gives us a honest, reproducible foundation for digital phenotyping research. It's not claiming to have solved mental health detection, but it's giving us a fair way to measure progress.

Tom: And that's exactly what we need. Alright, we're going to say goodbye to this paper and move on to the next one on the arXiv stack.

Jane: Thanks for listening, everyone. We'll be right back with more research, more discussion, and more ideas that could change how we understand the human mind.

Tom: See you in a minute.

Quoc-Cuong Pham, Hoang-Thuy-Duong Vu, Thi-Thanh-Huong Ha, Huy-Hieu Pham

VinUniversity · Vietnam National University HCMC

cs.LG

Submitted: 2026-08-17

Updated: 2026-08-18

Code: https://github.com/neurai-vn/Neurai-VN-benchmark

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: "Digital phenotyping (DP) using smartphones and wearable devices has shown considerable potential for mental health monitoring.

Key concepts

Digital Phenotyping
This method uses sensors found in devices like smartphones and smartwatches to monitor and understand a person's mental state. It involves tracking various data points, such as sleep patterns, activity levels, heart rate, and screen usage.
Multimodal Digital Phenotyping
This approach combines data from multiple sources—for example combining wearable measurements (like heart rate) with phone usage data (like app states). This is used to gather a comprehensive view of the user's condition for research purposes.
Subject-wise Validation
This is a rigorous testing method where all data from the same individual is kept entirely within one set, either the training set or the test set. This prevents data leakage and ensures that model performance is reliable and not artificially inflated.

Terminology

Summary

Summary

The paper introduces a reproducible benchmark for mental health classification using digital phenotyping (DP) data, built upon the NEURAI-VN dataset. The authors state: "Digital phenotyping (DP) using smartphones and wearable devices has shown considerable potential for mental health monitoring. However, progress remains difficult to evaluate due to heterogeneous datasets, inconsistent preprocessing pipelines. They address these limitations by presenting a reproducible benchmark built upon the NEURAI-VN dataset, a high-resolution, multimodal dataset comprising passive sensing and active assessment from wearable and smartphone devices, collected from 100 Vietnamese adults over two weeks."

The dataset comprises "13 passive sensing modalities collected over a two-week monitoring period, spanning continuous, event-based, minute-level, and daily measurements, alongside longitudinal active self-report assessments and psychiatrist-assigned clinical diagnoses. Participants were classified into four groups: depressive disorders (Dep), anxiety disorders (Anx), healthy controls (HC), and other psychiatric conditions (OPC). The cohort characteristics include: Healthy controls (HC) 43, Individual with Depressive Depression (Dep) 31, Individual with Anxiety Disorder (Anx) 18, Other psychiatric conditions (OPC) 8, Total participants 100"; Female, n (%) 65; Age (years), mean ± SD 26.1 ± 4.6; Monitoring duration (days), mean ± SD 16.3 ± 2.7; Participant-day record entries 1730; Self-assessment entries 2096.

The methodology defines six modality groups: "Wm (wearable per-minute: ats, hrts, azmts), Ws (wearable daily: sleep, hrv, spo2, br, skinTemp), Sd (self-report daily: dailySurvey, moodLog), Sw (self-report weekly: phq9, gad7), Pe (smartphone event-based: appstate), and Pc (smartphone continuous: accelerometer, gyroscope, network, battery). For benchmarking, the two smartphone modalities, event-based sensing (Pe) and continuous sensing (Pc), are merged into a unified smartphone feature group (P), while the weekly self-report modality (Sw) is excluded because its temporal resolution is not aligned with the remaining modalities. This yields a configuration space M = Wm, Ws, Sd, P, with 15 benchmark configurations evaluated: Single-group (4): Wm, Ws, Sd, and P. Two-group (6): Wm+Ws, Wm+Sd, Wm+P, Ws+Sd, Ws+P, and Sd+P. Three-group (4): Wm+Ws+Sd, Wm+Ws+P, Wm+Sd+P, and Ws+Sd+P. All-group (1): Wm+Ws+Sd+P."

Four binary classification tasks are defined: Task B1: Classification between HC and Dep; Task B2: Classification between HC and Anx; Task B3: Classification between Dep and Anx; Task B4: Classification between HC and clinical groups (Dep+Anx+OPC). Feature engineering includes daily channel-wise statistical descriptors (mean, min, max, skew, std) for Wm and Pc, daily event counts for Pe, direct daily summary measurements for Ws, and normalized daily responses for Sd. Feature preprocessing uses feature-wise mean imputation and min–max scaling followed by univariate feature selection using ANOVA F-statistics, retaining the top-K features: K = min(max(20, ⌊0.5F⌋), F).

Four baseline models are benchmarked: Logistic Regression (LR), Random Forest (RF), XGBoost (XGB), and a Multi-Layer Perceptron (MLP). Hyperparameters are specified: "LR: Logistic Regression, C = 1.0, default L2 regularization, maximum iterations=1000, random state=42; RF: Random Forest, default hyperparameters, random state=42; XGB: XGBoost, learning rate=0.05, number of estimators=300, evaluation metric=logloss, random state=42; MLP: Multi-Layer Perceptron, default architecture, maximum iterations=500, random state=42. The evaluation protocol is five-fold subject-wise Group K-Fold cross-validation, where samples from the same participant are restricted to a single fold to prevent subject-level information leakage. The same fold partition is used across all models, feature configurations, and classification tasks. The primary metric is subject-level macro F1-score (mF1)."

The main quantitative results are reported in Table 5: "Task B1: HC vs. Dep — Best Configuration Wm+Ws+Sd, Best Model MLP, mF1 0.71 ±0.17; Task B2: HC vs. Anx — Best Configuration Wm+Ws+P+Sd, Best Model LR, mF1 0.69 ±0.17; Task B3: Dep vs. Anx — Best Configuration Wm, Best Model XGB, mF1 0.56 ±0.22; Task B4: HC vs. (Dep+Anx+OPC) — Best Configuration Wm+Ws+Sd, Best Model LR, mF1 0.71 ±0.06."

The paper reports on the effect of multimodal integration: "Overall, the average Subject-F1 increased progressively from single-modality to multimodal feature combinations, with the highest mean performance achieved when all four modalities were integrated. In contrast, the highest individual result was obtained using a three-modality configuration, suggesting that the optimal feature combination remained task-dependent. Grouped performance is: Single-group 0.48 ±0.08 (range 0.35–0.67); Two-group 0.52 ±0.08 (range 0.39–0.68); Three-group 0.57 ±0.08 (range 0.40–0.71); All-group 0.60 ±0.08 (range 0.43–0.69)."

The paper concludes: These baseline results provide reproducible baselines for future research on multimodal DP for mental health classification tasks. Data availability is noted: The Neurai-VN dataset is hosted on Zenodo and is available at https://zenodo.org/records/18976769 under restricted access upon request. Code availability: The code used to reproduce the benchmark experiments, including preprocessing pipelines and evaluation protocols, is available at https://github.com/neurai-vn/Neurai-VN-benchmark.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:

  1. Subject-wise cross-validation enforcement: I will implement Group K-Fold splitting at the participant level as a mandatory default in all digital phenotyping pipelines, preventing subject-level data leakage that inflates performance metrics.

  2. Standardized feature-group configuration framework: I will adopt the paper’s explicit modality grouping (Wm, Ws, Sd, P) and exhaustive 15-configuration evaluation to systematically test single- vs. multi-modal integration, rather than ad-hoc feature selection.

  3. Reproducible baseline model suite: I will use the exact hyperparameters (LR C=1.0, XGB lr=0.05/300 estimators, RF default, MLP default) and fixed seed 42 as a reference point for any new model comparison.

  4. Feature preprocessing pipeline: I will apply feature-wise mean imputation, min-max scaling, and ANOVA F-statistic-based top-K selection (K = min(max(20, 0.5F), F)) as a standardized preprocessing step for daily aggregated features.

  5. Task-specific optimal configuration selection: I will use the paper’s finding that the best configuration is task-dependent (e.g., Wm+Ws+Sd for HC vs. Dep; Wm for Dep vs. Anx) to guide automated configuration search rather than assuming one universal feature set.

  6. Evaluation metric standardization: I will report subject-level macro F1 with mean ± std across 5 folds as the primary metric, replacing record-level accuracy or AUC that may be misleading under class imbalance.

  • Detect depression vs. healthy controls with a subject-level F1 of 0.71 ± 0.17 using a three-modality configuration (wearable minute-level + wearable daily + self-report daily), matching the best baseline in the paper.

  • Classify healthy vs. clinical (depression+anxiety+other psychiatric) with F1 of 0.71 ± 0.06 using the same three-modality configuration, enabling reliable screening in primary care settings.

  • Distinguish depression from anxiety with F1 of 0.56 ± 0.22 using only wearable minute-level data, providing a lightweight, passive-only screening tool without requiring self-report input.

  • Benchmark any new model against four established baselines (LR, RF, XGB, MLP) across 15 feature configurations and 4 tasks, ensuring fair comparison and preventing over-optimistic claims.

  • Generalize across cultural contexts by incorporating the Vietnamese cohort’s demographic diversity (age 19–42, 65% female), testing whether Western-derived behavior–symptom associations hold in non-WEIRD populations.

  • Automate configuration selection: Given a new dataset, the system will evaluate all 15 feature combinations and select the task-optimal one, avoiding manual feature engineering bias.

  • Provide uncertainty-aware predictions: By reporting mean ± std across folds, the system can flag low-confidence classifications (e.g., high variance) for human review, reducing costly misdiagnoses.

These improvements directly address the paper’s core contributions: reproducible evaluation, standardized preprocessing, and task-specific multimodal integration for mental health classification.

Abstract

Digital phenotyping (DP) using smartphones and wearable devices has shown considerable potential for mental health monitoring. However, progress remains difficult to evaluate due to heterogeneous datasets, inconsistent preprocessing pipelines. In this work, we present a reproducible benchmark built upon the Neurai-VN dataset, a high-resolution, multimodal dataset comprising passive sensing and active assessment from wearable and smartphone devices, collected from 100 Vietnamese adults over two weeks. We define four binary classification tasks evaluated using standardized subject-wise cross-validation. Representative linear, tree-based, and neural baseline models are evaluated systematically across predefined feature-group configurations. Mean subject-level F1 scores across five cross-validation folds reached 0.71 for Healthy Control vs. Depression and Healthy Control vs. Clinical, while Healthy Control vs. Anxiety and Depression vs. Anxiety achieved 0.69 and 0.56, respectively. These baseline results provide reproducible baselines for future research on multimodal DP for mental health classification tasks.

Related papers