Robust COVID-19 Detection from Cough Sounds using Deep Neural Decision Tree and Forest: A Comprehensive Cross-Datasets Evaluation

arXiv:2501.01117 · cs.SD, cs.AI, cs.LG, eess.AS · Submitted 2025-01-02 · Read on arXiv

Rofiqul Islam, Nihad Karim Chowdhury, Muhammad Ashad Kabir

Department of Computer Science and Engineering, University of Chittagong · School of Computing, Mathematics, and Engineering, Charles Sturt University

cs.SD, cs.AI, cs.LG, eess.AS

Submitted: 2025-01-02

Updated: 2026-08-10

Comments: 39 pages

Journal ref: Expert Systems with Applications, vol. 310, p. 131235, 2026

DOI: 10.1016/j.eswa.2026.131235

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 63/100

The gist: This research presents "a robust approach to classifying COVID-19 cough sounds using cutting-edge machine learning techniques," specifically "leveraging deep neural decision trees (DNDT) and deep

Terminology

Summary

This research presents a robust approach to classifying COVID-19 cough sounds using cutting-edge machine learning techniques, specifically leveraging deep neural decision trees (DNDT) and deep neural decision forests (DNDF). The methodology is designed to demonstrate consistent performance across diverse cough sound datasets.

Methodology

The proposed framework involves several stages: data curation, feature extraction, feature selection and model evaluation.

  • Feature Extraction: The researchers capture the acoustic signal used for feature extraction at a frequency of 22 kHz and extract 193 features from each audio signal, which include 40 MFCCs, 128 mel-scaled spectrogram, 6 tonal centroid, 12 chromagram, and 7 spectral contrast features.

  • Feature Selection: To perform optimal feature selection, the study employs Recursive Feature Elimination with Cross-Validation (RFECV) in combination with the Extra-Trees classifier to identify the most critical features and minimize feature dimensions.

  • Hyper-parameter Tuning: Bayesian Optimization (BO) is used to fine-tune the hyper-parameters of our proposed method, which is noted as being more effective than more brute-force methods like Grid Search (GS) and Random Search (RS).

  • Handling Class Imbalance: To address the under-representation of the positive category for COVID-19, the researchers incorporated SMOTE [synthetic minority over-sampling technique] during training to ensure a balanced representation of positive and negative data.

  • Threshold Optimization: Model performance refinement is achieved through threshold optimization, maximizing the ROC-AUC score via a threshold moving (TM) technique.

  • Classification Models: The study utilizes deep neural decision trees (DNDT), which are structured as a tree... encompassing both decision and prediction nodes, and deep neural decision forests (DNDF), which consist of multiple DNDTs trained simultaneously and produce an output by averaging the individual outputs from each of the trees within the forest.

Training Strategies

The researchers explore five distinct training strategies to evaluate the effectiveness of different components:

  • Strategy 1: exclusively relies on our trained classifiers.

  • Strategy 2: only use the threshold moving technique... with a trained classifier.

  • Strategy 3: use both the threshold moving technique and the feature dimension reduction technique... with a trained classifier.

  • Strategy 4: use the threshold moving technique, optimal feature selection... and Bayesian Optimization.

  • Strategy 5: use the threshold moving technique, optimize feature selection... use Bayesian Optimization, and apply SMOTE to balance the data of the minority class.

Datasets

The approach is evaluated across five diverse cough datasets: Cambridge (asymptomatic and symptomatic), Coswara, COUGHVID, Virufy, and the combined Virufy with the NoCoCoDa dataset. Additionally, all five datasets are consolidated into a combined dataset comprising 3,398 samples, including 1,181 cough samples from COVID-19-positive individuals and 2,217 cough samples from COVID-19-negative individuals.

Results and Discussion

The study finds that strategy 5 consistently outperforms the other four across nearly all datasets. Specifically, the DNDF classifier consistently surpasses the DNDT classifier in all datasets. For Strategy 5, the proposed approach yields notable AUC scores of 0.97, 0.98, 0.92, 0.93, 0.99, and 0.99, alongside remarkable precision scores of 1, 1, 0.72, 0.93, 1, and 1 across the respective datasets.

When merging all datasets into a combined dataset, the method using the deep neural decision forest classifier, achieves an accuracy of 0.97, AUC of 0.97, precision of 0.95, recall of 0.96, F1-score of 0.96, and specificity score of 0.97.

Cross-Datasets Analysis

The study includes a comprehensive cross-datasets analysis, revealing demographic and geographic differences in the cough sounds associated with COVID-19. The results show that classification performance tends to decrease in cross-datasets evaluations, which is primarily due to the variations in recording quality, equipment used, demographic differences, and ethnic and environmental factors. These findings "highlight the challenges in transferring learned features across diverse datasets and underscore the potential benefits of dataset integration, improving generalizability and enhancing COVID-19 detection from audio signals."

Improvements for AI systems

1. Adversarial Domain Adaptation (ADA) Integration

  • The Improvement: Implement a domain-adversarial neural network component that uses a domain discriminator to penalize the model for learning features that identify the specific dataset, recording device, or environment.

  • What the improved system can do: It will maintain high classification accuracy (AUC/Precision) even when deployed on new, unseen hardware or in different acoustic environments (e.g., moving from a controlled lab to a noisy home setting), solving the cross-dataset performance drop identified in the paper.

2. Disentangled Representation Learning

  • The Improvement: Incorporate a loss function designed to disentangle pathological features (cough signatures of COVID-19) from nuisance features (demographic, ethnic, and gender-based vocal characteristics).

  • What the improved system can do: It will provide equitable and unbiased diagnostic accuracy across diverse global populations, ensuring that a user's ethnicity or biological sex does not skew the model's ability to detect the virus.

3. Self-Supervised Pre-training (SSL) via Audio Transformers

  • The Improvement: Replace manual feature extraction (MFCCs, chromagrams, etc.) with a self-supervised backbone, such as a Wav2Vec 2.0 or an Audio Spectrogram Transformer (AST), pre-trained on massive unlabeled audio datasets.

  • What the improved system can do: It will capture complex, high-dimensional temporal and spectral nuances in the cough signal that manual 193-feature sets might miss, leading to higher sensitivity in early-stage detection.

4. Attention-Weighted Forest Aggregation

  • The Improvement: Replace the simple averaging method in the Deep Neural Decision Forest (DNDF) with a learnable attention mechanism that assigns weights to individual trees based on their confidence and relevance to the specific input signal.

  • What the improved system can do: It will become more resilient to outlier trees that may be confused by background noise, effectively filtering out unreliable predictions to further boost precision and F1-scores.

5. Knowledge Distillation for Edge Deployment

  • The Improvement: Use the high-performing DNDF (Strategy 5) as a teacher model to train a much smaller, compressed student model (e.g., a lightweight 1D-CNN or a pruned decision tree).

  • What the improved system can do: It will enable real-time, on-device COVID-19 screening via smartphone applications, allowing for immediate, privacy-preserving diagnostics without requiring an internet connection or heavy cloud computing.

Abstract

This research presents a robust approach to classifying COVID-19 cough sounds using cutting-edge machine-learning techniques. Leveraging deep neural decision trees and deep neural decision forests, our methodology demonstrates consistent performance across diverse cough sound datasets. We begin with a comprehensive extraction of features to capture a wide range of audio features from individuals, whether COVID-19 positive or negative. To determine the most important features, we use recursive feature elimination along with cross-validation. Bayesian optimization fine-tunes hyper-parameters of deep neural decision tree and deep neural decision forest models. Additionally, we integrate the SMOTE during training to ensure a balanced representation of positive and negative data. Model performance refinement is achieved through threshold optimization, maximizing the ROC-AUC score. Our approach undergoes a comprehensive evaluation in five datasets: Cambridge, Coswara, COUGHVID, Virufy, and the combined Virufy with the NoCoCoDa dataset. Consistently outperforming state-of-the-art methods, our proposed approach yields notable AUC scores of 0.97, 0.98, 0.92, 0.93, 0.99, and 0.99 across the respective datasets. Merging all datasets into a combined dataset, our method, using a deep neural decision forest classifier, achieves an AUC of 0.97. Also, our study includes a comprehensive cross-datasets analysis, revealing demographic and geographic differences in the cough sounds associated with COVID-19. These differences highlight the challenges in transferring learned features across diverse datasets and underscore the potential benefits of dataset integration, improving generalizability and enhancing COVID-19 detection from audio signals.

Sources

Related papers