The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification
cs.SD, cs.LG
Submitted: 2026-09-15
Updated: 2026-09-15
Code: https://github.com/AndrewPBerg/UAV
Project page: https://iver56.github.io/audiomentations
License: http://creativecommons.org/licenses/by/4.0/
The gist: As unmanned aerial vehicles (UAVs) become increasingly prevalent in consumer and defense settings, classifying them reliably from limited, modality-specific data is an urgent challenge.
Terminology
Abstract
As unmanned aerial vehicles (UAVs) become increasingly prevalent in consumer and defense settings, classifying them reliably from limited, modality-specific data is an urgent challenge. The dominant approach, large pretrained networks fully fine-tuned on task data, carries a substantial computational and memory weight that is hard to bear in resource-constrained UAV deployments, where edge inference and rapid retraining for emerging platforms are both required. This paper systematically scales across both model architectures and fine-tuning methods for UAV audio classification, asking when that weight is justified and when lighter alternatives prevail. Using a custom dataset of 3,100 audio clips spanning 31 drone classes, we evaluate transformer (ViT, AST) and convolutional (custom CNN, ResNet-18/152, MobileNet-V3-S/L, EfficientNet-B0/B7) backbones under full fine-tuning, classifier-only fine-tuning, and four parameter-efficient fine-tuning (PEFT) methods: SSF, IA3, OFT, and selective batch-norm tuning. All configurations are evaluated with 5-fold cross-validation across accuracy, training time, trainable-parameter share, and inference-time memory footprint. Selective batch-norm fine-tuning of EfficientNet-B7 with three-fold augmentations achieves the highest validation accuracy (97.65% +- 0.30) while updating under 0.5% of model parameters. Across the sweep, lightweight CNNs consistently outperform transformers on both accuracy and efficiency. For UAV audio classification under data scarcity, scaling the method outperforms scaling the model.
Sources
- Advance and Refinement: The Evolution of UAV Detection and Classification Technologies
- Rethinking CNN Models for Audio Classification
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Spectral and Rhythm Features for Audio Classification with Deep Convolutional Neural Networks
- Understanding intermediate layers using linear classifier probes
- UAV (Unmanned Aerial Vehicles): Diverse Applications of UAV Datasets in Segmentation, Classification, Detection, and Tracking
- Controlling Text-to-Image Diffusion by Orthogonal Finetuning
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
- Scaling & Shifting Your Features: A New Baseline for Efficient Model Tuning
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNs
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- AST: Audio Spectrogram Transformer
- ImageNet Large Scale Visual Recognition Challenge
- Deep Residual Learning for Image Recognition
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Searching for MobileNetV3
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment