MAST: Label-Efficient, Robust, and Generalizable Sound Detection for Biodiversity Monitoring via Masked Audio Pretraining and Self-Training
cs.SD, cs.LG
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 29 pages, 8 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Passive acoustic monitoring can measure biodiversity at larger scales, but time--frequency annotation of animal vocalizations is expensive, site-specific, and difficult to sustain at scale.
Terminology
Abstract
Passive acoustic monitoring can measure biodiversity at larger scales, but time--frequency annotation of animal vocalizations is expensive, site-specific, and difficult to sustain at scale. We present a label-efficient sound detection framework that combines masked audio pretraining with a lightweight detector on mel spectrograms, then further improves robustness through iterative self-training on unlabeled audio. We first pretrain a ViT-based encoder on unlabeled recordings via masked reconstruction and transfer the encoder to a detection backbone. To better separate animal sounds from confounding background, we add a box-level contrastive loss that pulls matched event regions together while pushing noisy negatives apart. We then apply a two-stage pseudo-labeling curriculum to exploit large unlabeled pools without additional annotation. We evaluate the performance on two ecologically distinct domains: tropical rainforest soundscapes (Indonesia) and bird vocalizations in Mediterranean habitats (Spain). On both domains, masked audio pretraining and contrastive learning consistently improve time--frequency detection under temporal and cross-site distribution shift, and self-training yields further gains in out-of-distribution performance. On the rainforest domain, MAST with self-training achieves +0.22 mAP and +0.24 F1 over the strongest baseline under cross-site shift. On the bird domain, self-training achieves +0.12 mAP and +0.10 F1 over the strongest baseline under cross-site shift. Overall, our results show that MAST can effectively extend self-supervised audio representations from clip-level tasks to robust box-level localization across diverse bioacoustic settings, providing a practical path for biodiversity monitoring with limited labels.
Sources
- BEATs: Audio Pre-Training with Acoustic Tokenizers
- Sound Event Bounding Boxes
- R-CRNN: Region-based Convolutional Recurrent Neural Network for Audio Event Detection
- Semi-supervsied Learning-based Sound Event Detection using Freuqency Dynamic Convolution with Large Kernel Attention for DCASE Challenge 2023 Task 4
- Robust detection of overlapping bioacoustic sound events
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics
- Can Masked Autoencoders Also Listen to Birds?
- NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
- YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
- Large-Scale Weakly Labeled Semi-Supervised Sound Event Detection in Domestic Environments
- Perch 2.0: The Bittern Lesson for Bioacoustics
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled Data
- SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment