Time-frequency localization of bird calls in dense soundscapes
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Time-frequency localization of bird calls in dense soundscapes".
Jane: Passive acoustic monitoring enables large-scale observation of wildlife, but most bioacoustic classifiers only predict species presence in a time window without localizing vocalizations precisely in time or frequency,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: The paper introduces a new approach to bird vocalization detection by treating it as an object detection task on spectrograms and training YOLO11 models specifically for dense tropical soundscapes in Singapore. This directly addresses the limitation where existing global-context methods only predict species presence without time or frequency localization, which is what this study sets out to fix.
Jane: They are essentially proposing a method that aims to localize bird calls precisely in time and frequency, which is crucial because it allows researchers to see exactly how animals adapt their vocalizations to the specific acoustic environment they are in.
Lu: The methodology involves a fairly detailed pipeline, starting with converting raw audio into spectrograms using the short-time Fourier transform, and then preparing these spectrograms for YOLO by scaling them up to one thousand twenty-four times one thousand twenty-four square images after some downsampling and padding.
Meng: I'm interested in how they prepare the input data because getting those spectrograms into a format that works well with a deep learning model like YOLO requires careful parameter selection, which is where the technical challenge lies.
Lalam: It’s interesting how they adapt the spectrogram parameters to have an equal number of time and frequency bins, creating these square images that feed directly into the detection task for the AI models.
Tom: And beyond just training a model, they also developed an open-source browser-based annotation tool called BirdWatch, which lets users listen to time- and frequency-limited sections by drawing bounding boxes directly on the spectrograms.
Jane: That annotation tool is super helpful because it lets people visually interact with the data in a way that helps them understand complex soundscapes better, especially when dealing with calls that overlap in both time and frequency.
Lu: The researchers also introduced Intersection over Minimum, or IoMin, as a new evaluation metric intended to handle ambiguous acoustic boundaries much better than the standard IoU used previously for this type of problem.
Meng: So they aren't just focusing on accuracy; they’ve created a way to measure how well the model handles those fuzzy edges where bird calls blend into ambient noise, which is a practical concern.
Lalam: It shows a real focus on making the metrics match the actual difficulties encountered when trying to isolate specific acoustic events in noisy natural settings.
Conclusion: Tom: So, wrapping up this discussion on "Time-frequency localization of bird calls in dense soundscapes," we see that the authors successfully formulated bird vocalization detection as an object detection task using YOLO11 models trained on Singaporean data to localize calls in time and frequency.
Jane: The implication here is that researchers can move past just knowing *if* a bird is there and start understanding the fine-grained acoustic mechanics of its communication within a complex natural setting.
Lu: It opens up possibilities for incredibly detailed ecological studies, allowing scientists to track vocalization changes across subtle shifts in soundscape composition that were previously invisible.
Meng: Practically speaking, this means we can build more sophisticated models of animal behavior by knowing the exact acoustic context surrounding every recorded signal, which is a step toward truly contextual understanding.
Lalam: For culture and conservation science, this level of detail could mean we can better understand how species respond to human activities or environmental changes by analyzing the specific vocalizations they produce in different noise regimes.
Tom: The authors also proposed IoMin as a better metric for evaluating these detections because it handles those tricky acoustic boundaries more effectively than standard IoU, which is a key contribution to the paper's technical framework.
Jane: Ultimately, this work suggests that using advanced object detection on spectrograms with a specialized metric like IoMin can yield much more useful data for passive acoustic monitoring in dense environments.
Simen Hexeberg, Fanghui Tong, Hari Vishnu, Mandar Chitre
Tropical Marine Science Institute, National University of Singapore
cs.SD, cs.CV, q-bio.QM
Submitted: 2026-06-09
Updated: 2026-09-28
Comments: The following changes are made in this version: 1) added comparison to SAM 3 and RF-DETR models, 2) added a literature review on object-detection based acoustic segmentation, and 3) compressed the paper to fit within 5 pages
Code: https://github.com/org-arl/birdwatch-public
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 77/100
The gist: Passive acoustic monitoring enables large-scale observation of wildlife, but most bioacoustic classifiers only predict species presence in a time window without localizing vocalizations precisely in
Key concepts
- Spectrograms
- These are visual representations of audio data that show how sound intensity changes over time and across different frequencies. They allow researchers to see the 'shape' of a sound event, making it possible for AI models to treat vocalizations like pictures.
- YOLO11 Model
- YOLO (You Only Look Once) is an object detection algorithm used here. It identifies and draws bounding boxes around specific objects—in this case, individual bird calls—within the spectrogram images, providing precise localization in both time and frequency.
Terminology
Summary
Passive acoustic monitoring enables large-scale observation of wildlife, but most bioacoustic classifiers only predict species presence in a time window without localizing vocalizations precisely in time or frequency, limiting downstream analyses.
How it works
The researchers formulate bird vocalization detection as an object detection task on spectrograms and train YOLO11 models to localize bird calls in dense tropical soundscapes from Singapore. This approach addresses the limitations of existing global-context methods that only indicate species presence without time or frequency localization, which is crucial for studying how animals adapt their vocalizations to acoustic environments.
The methodology involves several key steps:
-
Raw audio recordings are converted into spectrograms using the short-time Fourier transform (STFT).
-
Spectrograms are adapted for YOLO by selecting STFT parameters such that they contain approximately equal numbers of time and frequency bins, resulting in square images of size 1024 × 1024 after downsampling and padding.
-
Amplitude scaling is applied through a sequence: converting STFT magnitudes to log-power (decibel) values, clipping the spectrogram to the [1st, 99.8th] percentile range, and applying gamma compression with γ = 0.85 to enhance contrast in low-energy regions.
-
The single-channel spectrograms are converted into three RGB channels using the magma colormap to improve YOLO performance, as this mapping helps the model perform better on natural image statistics.
Annotation and Data
To facilitate analysis, an open-source browser-based annotation tool named BirdWatch was developed. Key features of this tool include:
** Time–frequency playback:**
The tool allows users to listen to time- and frequency-limited sections of recordings simply by drawing bounding boxes on the spectrograms, which is particularly useful in complex soundscapes with time-overlapping vocalizations.
** Bounding box annotation:**
Individual vocalizations can be annotated by drawing bounding boxes, and these annotations can be exported directly in YOLO format for model training. Furthermore, these boxes can be edited through a versatile set of tools allowing for quality checking, boundary refinement, and expert review.
The study utilized two datasets:
-
The Singapore Dataset (In-Distribution), collected across two recording sites (SBG1 and SBG2) with 18,095 bounding box annotations spanning 4 hours and 25 minutes of audio.
-
The Hawaii Dataset (Out-of-Distribution), used exclusively for out-of-distribution evaluation, containing 81,691 final labels after processing to align with the Singapore dataset's frequency range restriction (0.5–12 kHz).
Evaluation Metrics and Results
The researchers introduced Intersection over Minimum (IoMin) as an evaluation metric to better handle ambiguous acoustic boundaries than standard IoU, defined as:
** IoMin = intersection / min(areapred, areagt)**
The performance comparison against the unsupervised, energy-based Time-Frequency Event (TFE) detector showed that YOLO models significantly outperform the TFE detector baseline on both datasets and under both overlap criteria. On Singapore, the best-performing model achieved an IoMin@50 F1-score of 81.8%, nearly doubling the baseline's 42.1%.
The performance on Hawaii (OOD) was 57.6% under IoMin@50, compared to 48.6% for the TFE detector baseline, suggesting that YOLO appears to better suppress non-bird sources such as insect noise and human speech, producing tighter bounding boxes around individual calls.
Model Trade-offs and Future Directions
The study evaluated five standard-size variants of YOLO11 (nano to extra-large). While larger models showed a slight advantage on Hawaii, the YOLO11n model offered the most favorable performance–compute trade-off, falling within one percentage point of the best model on Singapore while requiring 10× fewer parameters and 13× fewer FLOPs.
Limitations identified include:
** Annotation boundary ambiguity:**
The paper notes that standard IoU can penalize predictions that only capture a target part partially, whereas IoMin yields one TP and zero FPs in such cases.
Future work suggests increasing the diversity of training data or constructing semi-synthetic soundscapes by overlaying isolated bird vocalizations onto real ambient noise recordings to gain finer control over species composition, call density, and SNR. Furthermore, exploring alternative input representations like Mel-STFT or wavelet transforms could potentially improve performance beyond the current log-STFT representation.
Conclusion
The paper presents a YOLO-based method to detect and localize bird vocalizations in both time and frequency from dense soundscapes, proposing IoMin as an alternative metric to IoU that better reflects detection quality for acoustic signals with ambiguous boundaries.
Improvements for AI systems
Here are the specific improvements for AI systems based on this research, and what these improved systems can achieve:
-
Acoustic Object Detection Preprocessing Module: The system can be used as a first-stage preprocessing module for long-duration soundscape analysis.
-
Time and Frequency Localization of Vocalizations: The system can precisely localize individual bird calls in both time and frequency within complex tropical soundscapes (e.g., Singapore).
-
Noise Rejection/Source Separation: The improved system can effectively ignore or suppress non-target sounds such as insect noise, human speech, and anthropogenic noise by producing tighter bounding boxes around bird calls compared to traditional energy-based methods (TFE detector).
-
Robust Evaluation of Acoustic Boundaries: By adopting the Intersection over Minimum (IoMin) metric instead of standard Intersection over Union (IoU), the system can be evaluated using a metric that better handles ambiguous acoustic boundaries, leading to more realistic performance assessments for time-frequency localization tasks.
-
Adaptive Soundscape Analysis: The ability to localize calls in time and frequency enables downstream analyses of how animals adapt their vocalizations in response to environmental changes (e.g., quantifying shifts in call timing or frequency based on noise source introduction).
-
Efficient Edge Deployment: Utilizing the YOLO11n architecture provides a high performance-compute trade-off, making the system suitable for low-power edge devices common in continuous, long-duration Passive Acoustic Monitoring (PAM) surveys.
Abstract
Passive acoustic monitoring enables large-scale wildlife observation. Most bioacoustic classifiers predict species presence in a time window without localizing vocalizations precisely in time or frequency, limiting downstream analyses. We formulate time-frequency localization of bird calls as object detection on spectrograms and compare three computer vision model families (YOLO11, SAM 3, RF-DETR) against a non-learnable baseline. We introduce Intersection over Minimum (IoMin), an evaluation metric that better handles ambiguous acoustic boundaries than IoU. We also open-source a browser-based tool for efficient bounding-box labeling. The best RF-DETR model nearly doubles baseline performance on in-distribution, dense soundscapes from Singapore (83.4% vs. 42.1% IoMin@50 F1-score) and generalizes better to out-of-distribution recordings from Hawaii (63.2% vs. 48.6%). These results indicate that fine-tuned computer-vision models are well suited for time-frequency localization of bird vocalizations in complex soundscapes.
Sources
- Foundation Models for Bioacoustics -- a Comparative Review
- Perch 2.0: The Bittern Lesson for Bioacoustics
- BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics
- Localized time-frequency representation learning for bioacoustic classification in complex soundscapes
- Microsoft COCO: Common Objects in Context
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment