Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions
eess.AS, cs.AI, cs.SD
Submitted: 2026-05-28
Updated: 2026-09-17
Comments: 2 figures, 4 tables, and 5 pages
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: AI-driven respiratory sound classification (RSC) is promising for automated pulmonary disease detection, yet multi-site deployment is hindered by inter-stethoscope variability.
Terminology
Abstract
AI-driven respiratory sound classification (RSC) is promising for automated pulmonary disease detection, yet multi-site deployment is hindered by inter-stethoscope variability. We introduce a federated domain generalization (FedDG) formulation for RSC in which clients hold recordings from different stethoscopes and the model is evaluated on an unseen device. Our empirical analysis shows that stethoscope-induced style and disease-relevant content are partially entangled, making deterministic style removal unreliable. In response, we propose BTS-CAFE, a framework combining (i) causality-inspired device-style interventions with constraints designed to limit content distortion, (ii) counterfactual metadata augmentation to relieve device and demographic shortcuts, and (iii) gradient alignment to promote device-invariant decision boundaries across clients. Built on BTS with CLAP, a multimodal language-audio pretraining model, BTS-CAFE improves the out-of-distribution ICBHI Score by 3.69 points on average over five held-out devices relative to its backbone, and outperforms conventional data augmentation and federated learning baselines in simulated device leave-out evaluations on the ICBHI and SPRSound datasets.
Sources
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- A Comprehensive Survey on Generative Diffusion Models for Structured Data
- Federated Out-of-Distribution Generalization: A Causal Augmentation View
- mixup: Beyond Empirical Risk Minimization
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions