Robust COVID-19 Detection from Cough Sounds using Deep Neural Decision Tree and Forest: A Comprehensive Cross-Datasets Evaluation
Rofiqul Islam, Nihad Karim Chowdhury, Muhammad Ashad Kabir
Department of Computer Science and Engineering, University of Chittagong · School of Computing, Mathematics, and Engineering, Charles Sturt University
cs.SD, cs.AI, cs.LG, eess.AS
Submitted: 2025-01-02
Updated: 2026-08-10
Comments: 39 pages
Journal ref: Expert Systems with Applications, vol. 310, p. 131235, 2026
DOI: 10.1016/j.eswa.2026.131235
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 63/100
The gist: This research presents "a robust approach to classifying COVID-19 cough sounds using cutting-edge machine learning techniques," specifically "leveraging deep neural decision trees (DNDT) and deep
Terminology
Summary
This research presents a robust approach to classifying COVID-19 cough sounds using cutting-edge machine learning techniques,
specifically leveraging deep neural decision trees (DNDT) and deep neural decision forests (DNDF).
The methodology is designed to demonstrate consistent performance across diverse cough sound datasets.
Methodology
The proposed framework involves several stages: data curation, feature extraction, feature selection and model evaluation.
-
Feature Extraction: The researchers
capture the acoustic signal used for feature extraction at a frequency of 22 kHz
andextract 193 features from each audio signal,
which include40 MFCCs, 128 mel-scaled spectrogram, 6 tonal centroid, 12 chromagram, and 7 spectral contrast features.
-
Feature Selection: To perform
optimal feature selection,
the study employsRecursive Feature Elimination with Cross-Validation (RFECV) in combination with the Extra-Trees classifier
toidentify the most critical features
andminimize feature dimensions.
-
Hyper-parameter Tuning:
Bayesian Optimization (BO) is used to fine-tune the hyper-parameters of our proposed method,
which is noted as beingmore effective than more brute-force methods like Grid Search (GS) and Random Search (RS).
-
Handling Class Imbalance: To address the
under-representation of the positive category for COVID-19,
the researchersincorporated SMOTE [synthetic minority over-sampling technique] during training to ensure a balanced representation of positive and negative data.
-
Threshold Optimization:
Model performance refinement is achieved through threshold optimization, maximizing the ROC-AUC score
via athreshold moving (TM) technique.
-
Classification Models: The study utilizes
deep neural decision trees (DNDT),
which arestructured as a tree... encompassing both decision and prediction nodes,
anddeep neural decision forests (DNDF),
which consist ofmultiple DNDTs trained simultaneously
and produce an output byaveraging the individual outputs from each of the trees within the forest.
Training Strategies
The researchers explore five distinct training strategies
to evaluate the effectiveness of different components:
-
Strategy 1:
exclusively relies on our trained classifiers.
-
Strategy 2:
only use the threshold moving technique... with a trained classifier.
-
Strategy 3:
use both the threshold moving technique and the feature dimension reduction technique... with a trained classifier.
-
Strategy 4:
use the threshold moving technique, optimal feature selection... and Bayesian Optimization.
-
Strategy 5:
use the threshold moving technique, optimize feature selection... use Bayesian Optimization, and apply SMOTE to balance the data of the minority class.
Datasets
The approach is evaluated across five diverse cough datasets: Cambridge (asymptomatic and symptomatic), Coswara, COUGHVID, Virufy, and the combined Virufy with the NoCoCoDa dataset.
Additionally, all five datasets are consolidated into a combined dataset
comprising 3,398 samples, including 1,181 cough samples from COVID-19-positive individuals and 2,217 cough samples from COVID-19-negative individuals.
Results and Discussion
The study finds that strategy 5 consistently outperforms the other four across nearly all datasets.
Specifically, the DNDF classifier consistently surpasses the DNDT classifier in all datasets.
For Strategy 5, the proposed approach yields notable AUC scores of 0.97, 0.98, 0.92, 0.93, 0.99, and 0.99, alongside remarkable precision scores of 1, 1, 0.72, 0.93, 1, and 1 across the respective datasets.
When merging all datasets into a combined dataset,
the method using the deep neural decision forest classifier, achieves an accuracy of 0.97, AUC of 0.97, precision of 0.95, recall of 0.96, F1-score of 0.96, and specificity score of 0.97.
Cross-Datasets Analysis
The study includes a comprehensive cross-datasets analysis, revealing demographic and geographic differences in the cough sounds associated with COVID-19.
The results show that classification performance tends to decrease in cross-datasets evaluations,
which is primarily due to the variations in recording quality, equipment used, demographic differences, and ethnic and environmental factors.
These findings "highlight the challenges in transferring learned features across diverse datasets and underscore the potential benefits of dataset integration, improving generalizability and enhancing COVID-19 detection from audio signals."
Improvements for AI systems
1. Adversarial Domain Adaptation (ADA) Integration
-
The Improvement: Implement a domain-adversarial neural network component that uses a
domain discriminator
to penalize the model for learning features that identify the specific dataset, recording device, or environment. -
What the improved system can do: It will maintain high classification accuracy (AUC/Precision) even when deployed on new, unseen hardware or in different acoustic environments (e.g., moving from a controlled lab to a noisy home setting), solving the
cross-dataset performance drop
identified in the paper.
2. Disentangled Representation Learning
-
The Improvement: Incorporate a loss function designed to disentangle
pathological features
(cough signatures of COVID-19) fromnuisance features
(demographic, ethnic, and gender-based vocal characteristics). -
What the improved system can do: It will provide equitable and unbiased diagnostic accuracy across diverse global populations, ensuring that a user's ethnicity or biological sex does not skew the model's ability to detect the virus.
3. Self-Supervised Pre-training (SSL) via Audio Transformers
-
The Improvement: Replace manual feature extraction (MFCCs, chromagrams, etc.) with a self-supervised backbone, such as a Wav2Vec 2.0 or an Audio Spectrogram Transformer (AST), pre-trained on massive unlabeled audio datasets.
-
What the improved system can do: It will capture complex, high-dimensional temporal and spectral nuances in the cough signal that manual 193-feature sets might miss, leading to higher sensitivity in early-stage detection.
4. Attention-Weighted Forest Aggregation
-
The Improvement: Replace the
simple averaging
method in the Deep Neural Decision Forest (DNDF) with a learnable attention mechanism that assigns weights to individual trees based on their confidence and relevance to the specific input signal. -
What the improved system can do: It will become more resilient to
outlier
trees that may be confused by background noise, effectively filtering out unreliable predictions to further boost precision and F1-scores.
5. Knowledge Distillation for Edge Deployment
-
The Improvement: Use the high-performing DNDF (Strategy 5) as a
teacher
model to train a much smaller, compressedstudent
model (e.g., a lightweight 1D-CNN or a pruned decision tree). -
What the improved system can do: It will enable real-time, on-device COVID-19 screening via smartphone applications, allowing for immediate, privacy-preserving diagnostics without requiring an internet connection or heavy cloud computing.
Abstract
This research presents a robust approach to classifying COVID-19 cough sounds using cutting-edge machine-learning techniques. Leveraging deep neural decision trees and deep neural decision forests, our methodology demonstrates consistent performance across diverse cough sound datasets. We begin with a comprehensive extraction of features to capture a wide range of audio features from individuals, whether COVID-19 positive or negative. To determine the most important features, we use recursive feature elimination along with cross-validation. Bayesian optimization fine-tunes hyper-parameters of deep neural decision tree and deep neural decision forest models. Additionally, we integrate the SMOTE during training to ensure a balanced representation of positive and negative data. Model performance refinement is achieved through threshold optimization, maximizing the ROC-AUC score. Our approach undergoes a comprehensive evaluation in five datasets: Cambridge, Coswara, COUGHVID, Virufy, and the combined Virufy with the NoCoCoDa dataset. Consistently outperforming state-of-the-art methods, our proposed approach yields notable AUC scores of 0.97, 0.98, 0.92, 0.93, 0.99, and 0.99 across the respective datasets. Merging all datasets into a combined dataset, our method, using a deep neural decision forest classifier, achieves an AUC of 0.97. Also, our study includes a comprehensive cross-datasets analysis, revealing demographic and geographic differences in the cough sounds associated with COVID-19. These differences highlight the challenges in transferring learned features across diverse datasets and underscore the potential benefits of dataset integration, improving generalizability and enhancing COVID-19 detection from audio signals.
Sources
- Virufy: Global Applicability of Crowdsourced and Clinical Datasets for AI Detection of COVID-19 from Cough
- Coswara -- A Database of Breathing, Cough, and Voice Sounds for COVID-19 Diagnosis
- A literature review on COVID-19 disease diagnosis from respiratory sound data
- Cough Against COVID: Evidence of COVID-19 Signature in Cough Sounds
- Hi Sigma, do I have the Coronavirus?: Call for a New Artificial Intelligence Approach to Support Health Care Professionals Dealing With The COVID-19 Pandemic
- IATos: AI-powered pre-screening tool for COVID-19 from cough audio samples
- QUCoughScope: An Artificially Intelligent Mobile Application to Detect Asymptomatic COVID-19 Patients using Cough and Breathing Sounds
- Using Deep Learning with Large Aggregated Datasets for COVID-19 Classification from Cough
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment