A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases

arXiv:2608.11582 · cs.CV, cs.AI · Submitted 2026-08-12 · Read on arXiv

Danial Sharifrazi, Saadat Behzadi, Nouman Javed, Roohallah Alizadehsani, Prasad N. Paradkar, Asim Bhatti

Deakin University · University of Bologna · CSIRO Health and Biosecurity

cs.CV, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/danialsharifrazi/Video-ClassificationVisionTransformer-ConvGRU

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper introduces a three-step framework for detecting dengue and Zika virus-infected mosquitoes from control mosquitoes by analyzing their locomotion behavior in video.

Terminology

Summary

This paper introduces a three-step framework for detecting dengue and Zika virus-infected mosquitoes from control mosquitoes by analyzing their locomotion behavior in video. The problem is challenging because mosquitoes are very small in video frames and the background is complex, which causes conventional AI methods to extract erroneous features. The proposed framework first identifies mosquitoes and removes the background using the YOLO 11M model, then extracts visual features using a frozen, ImageNet-pretrained Vision Transformer (ViT), and finally classifies the videos using a convolutional Gated Recurrent Unit (ConvGRU) classifier. The data was collected by recording 15 mosquitoes in a cage over 1 to 13 days, resulting in three classes: dengue-infected, Zika-infected, and uninfected control. A comparative analysis was conducted against RNN, LSTM, GRU, and their convolutional versions (ConvRNN, ConvLSTM, ConvGRU). The results showed that the ConvGRU model achieved the best performance, with 88.88% accuracy, 84.45% precision, 82.82% recall, and 82.81% F1 score. Ablation studies confirmed that YOLO 11M was the best preprocessing detector (mAP50 of 97.8%), ViT was the best feature extractor compared to ResNet, VGG19, and EfficientNet, and hyperparameters were tuned using the Optuna platform. The authors conclude that combining convolutional models with sequence-based networks, especially ConvGRU, allows simultaneous extraction of precise spatial features and long-term temporal dependencies from mosquito movements, providing a reliable solution for analyzing mosquito behavior in complex environments. The paper also notes limitations, including evaluation on only a specific mosquito dataset, and suggests future work on larger, more diverse datasets.

Improvements for AI systems

Improvements to AI Systems:

  1. Hybrid Spatial-Temporal Feature Extraction for Small-Object Video Analysis: Integrate a frozen, pretrained Vision Transformer (ViT) as a robust spatial feature extractor, combined with a convolutional Gated Recurrent Unit (ConvGRU) for temporal modeling. This architecture reduces background noise and motion artifacts in low-resolution, high-clutter videos, improving classification of subtle behavioral differences.

  2. Two-Stage Preprocessing with Object Detection and Background Removal: Use YOLO11M (or similar lightweight detector) as a mandatory first stage to isolate the target object (e.g., mosquito) before feature extraction. This prevents the ViT from attending to irrelevant background pixels, increasing feature fidelity and reducing false positives in complex environments.

  3. Ablation-Guided Model Selection and Hyperparameter Optimization: Automate the selection of the optimal detector, feature extractor, and sequence model via systematic ablation studies (e.g., comparing YOLO variants, ViT vs. ResNet/VGG/EfficientNet, and ConvGRU vs. RNN/LSTM/GRU). Use Bayesian optimization (e.g., Optuna) to tune sequence length, hidden units, and learning rate, ensuring the best trade-off between accuracy and computational cost.

  4. Temporal Dependency Modeling for Long-Term Behavior Analysis: Replace standard recurrent units with convolutional recurrent cells (ConvGRU/ConvLSTM) to capture both spatial patterns and long-range temporal dependencies (e.g., 1–13 days of behavioral changes). This improves recall for slow-evolving infection states that are missed by frame-level classifiers.

  5. Robustness to Domain Shift via Pretrained Feature Extractors: Leverage ImageNet-pretrained ViT features that are frozen, reducing overfitting to small, domain-specific datasets. This enables the system to generalize to new insect species or camera setups without retraining the entire backbone.

What the Improved AI System Can Do:

  • Automatically detect and classify disease-infected insects (e.g., dengue vs. Zika vs. healthy) from raw video in real-time, even when targets are tiny and backgrounds are cluttered.

  • Operate with high accuracy (≈89% F1) on small datasets, while being adaptable to larger, more diverse datasets with minimal fine-tuning.

  • Provide explainable temporal behavior analysis, identifying movement patterns that correlate with infection progression over days.

  • Serve as a reusable pipeline for other small-object behavioral classification tasks (e.g., pest monitoring, laboratory animal tracking) by swapping the detector and classifier heads.

Sources

Related papers