Removing Speech, Keeping Activities: A Privacy Firewall for Acoustic Sensing in Assisted Living
cs.SD, cs.CR, cs.HC
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: 42 pages, 8 figures, 4 tables. Submitted to Pervasive and Mobile Computing. Preprint also available on SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7369016
Code: https://github.com/snakers4/silero-vad
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Real Time Speech Enhancement in the Waveform Domain
- SPMamba: State-space model is all you need in speech separation
- MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
- 30+ Years of Source Separation Research: Achievements and Future Challenges
- WHAM!: Extending Speech Separation to Noisy Environments
- LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
- AST: Audio Spectrogram Transformer
- DCASE 2018 Challenge - Task 5: Monitoring of domestic activities based on multi-channel acoustics
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment