The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis

arXiv:2608.12677 · cs.AI, cs.CV · Submitted 2026-08-13 · Read on arXiv

Danial Sharifrazi, Saadat Behzadi, Julakha Jahan Jui, Mojtaba Mohammadi, Nouman Javed, Roohallah Alizadehsani, Prasad N. Paradkar, Asim Bhatti

Institute for Intelligent Systems Research and Innovations (IISRI), Deakin University · Department of Electronic Engineering, University of Bologna · CSIRO Health and Biosecurity, Australian Animal Health Laboratory

cs.AI, cs.CV

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 72/100

The gist: The paper proposes a YOLO- and Contrastive Language-Image Pre-training (CLIP)-based vision-language framework to classify mosquito flight frames of uninfected and Dengue virus serotype 2

Terminology

Summary

The paper proposes a YOLO- and Contrastive Language-Image Pre-training (CLIP)-based vision-language framework to classify mosquito flight frames of uninfected and Dengue virus serotype 2 (DENV2)-infected mosquitoes. The study addresses the challenge that "Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make reliable feature extraction difficult." The proposed method first uses YOLO to isolate mosquito regions from the background, then aligns visual features extracted from video frames with biologically meaningful textual prompts in a shared embedding space. The multimodal model was fine-tuned using supervised bidirectional contrastive learning and evaluated through frame-level image–text similarity-based classification. The results show that the proposed method achieved 98.54% accuracy and 99.91% sensitivity at the frame level. After temporal aggregation of frame-level information, the model achieved complete video-level performance. The ablation results showed that fine-tuning and CLIP-based representations were essential for this domain, while the textual branch provided semantic image-text alignment rather than an accuracy advantage over the vision-only model. These findings suggest that vision-language models can provide a useful framework for analyzing infection-related biological behaviors from video data.

The dataset used in this study was obtained by continuously monitoring 15 mosquitoes in a laboratory chamber, with a food source containing a water and sugar solution placed inside the chamber. A camera was placed in front of the cage and data recording continued for 1-13 days under different lighting conditions. The dataset contains a total of 60 videos recorded from the flight paths of mosquitoes, with each video containing the entire group of 15 mosquitoes rather than an individually recorded mosquito. The videos are classified into two main classes: uninfected mosquitoes (control group) and mosquitoes infected with dengue fever virus, with 30 videos for each group. Environmental challenges include light fluctuations, uneven light distribution in the left and right parts of the cage, slight changes in camera angles, subjects being too small, and irregular and chaotic movements in the frame.

The proposed methodology consists of several main steps. First, in the main image processing scenario, the YOLOv11 model was used to detect and track mosquitoes in video frames, followed by a masking method to remove the background, leaving only the mosquito region in the image. In the second scenario, a motion-based method was used that indirectly extracts the mosquito's movement path based on the difference between consecutive frames. A fixed number of T = 32 frames was uniformly sampled from each video, and frames were resized to 224 × 224, converted to RGB color space, and normalized to the interval [0, 1]. The CLIP model with the clip vit base patch32 architecture was used, consisting of a Vision Transformer-based image encoder and a text encoder. Four different strategies were investigated: full fine-tuning of both image and text encoders along with projection layers; keeping both encoders fixed and only training projection layers; keeping the text encoder fixed but training the image encoder and projections; and adding an LSTM layer to the output of the image encoder. Based on experimental results, the first strategy showed the best performance and representation stability.

For semantic prompt generation, descriptive prompts were used in the training phase, such as Healthy mosquitoes remaining in the central area of the cage and DENV2-infected mosquitoes exploring cage corners, which were biologically motivated by previously reported arbovirus-associated behavioral changes in Aedes aegypti. In the inference phase, simpler prompts were used: a non-infected mosquito for the control class and a DENV2-infected mosquito for the DENV2 class. All text sequences were processed using the CLIP model tokenizer and Byte Pair Encoding method, with tokens padded or truncated to a maximum of 20 tokens. A supervised bidirectional contrastive loss function was used to solve the problem that conventional contrastive learning may mistakenly consider examples of the same class as negative examples. A normalized mask matrix M was constructed indicating whether image i and text j belong to the same class, and cross-entropy was calculated in two directions: from image to text and from text to image. The final objective function was defined as the arithmetic mean of the two two-way losses.

For evaluation, a 5-fold cross-validation was used with data splitting performed at the video level so that frames of the same video were not included in the training and test sets at the same time. The model was trained with the Adam optimizer with batch size set to 8 and learning rate set to 1 × 10-5, with early stopping with a patience of 5 epochs. Performance was evaluated at the frame level, with individual frames projected into the latent space and classified based on the highest cosine similarity to the inference text prompts.

The results show that the proposed model achieved an accuracy of 98.54%, a sensitivity of 99.91%, an F1-score of 98.28%, a specificity of 97.55%, and a precision of 96.75% at the frame level. The fold-wise results show high and stable performance across most folds, with sensitivity remaining at a very high level across all folds. The training dynamics indicate rapid and relatively stable convergence, with the closeness of training and validation trends indicating that the model does not show severe overfitting.

In comparison with alternative methods, the motion-driven method achieved 97.19% accuracy, 94.97% sensitivity, 99.56% specificity, 99.74% precision, and 97.21% F1-score. The LSTM-based model showed very low sensitivity (0.07%) and F1-score (0.10%) under the current training setting. The proposed method achieved the highest sensitivity and F1-score while maintaining high specificity and precision, resulting in the most balanced performance profile among the evaluated configurations.

The ablation study revealed several important findings. When both CLIP encoders were completely frozen, the model completely failed to identify DENV2-infected mosquitoes, with sensitivity and F1-score equal to zero, indicating that general CLIP features are not sufficient for this biological problem and full fine-tuning is necessary. The image-only model achieved frame-level performance comparable to the proposed model, with 97.60% accuracy, 100% sensitivity, 96.22% specificity, 94.38% precision, and 96.94% F1-score. The vision-only configuration achieved 99.01% accuracy, 99.80% sensitivity, 98.56% specificity, 97.80% precision, and 98.76% F1-score. These results show that the pre-trained CLIP visual encoder provides strong representations for detecting subtle mosquito flight patterns, and the textual component should not be interpreted as providing a clear accuracy advantage; its primary role is to align visual representations with semantically meaningful textual descriptions.

The main limitations of this study are the limited number of recordings and the limited biological diversity of the dataset. Each video contained the entire group of mosquitoes and was not associated with a single individual mosquito. To prevent direct information leakage, all frames extracted from the same video were assigned to the same cross-validation fold, but frames within a video remain temporally and visually correlated and should not be considered completely independent biological observations. The authors note that future evaluation using a larger number of recordings, larger and more diverse mosquito cohorts, different imaging conditions, and alternative prompts could provide a more reliable assessment of the method's generalizability.

Improvements for AI systems

Improvements to AI Systems:

  1. Domain-Adaptive Vision-Language Fine-Tuning: Implement full fine-tuning of both CLIP encoders (image and text) with supervised bidirectional contrastive loss, rather than freezing encoders or using unidirectional contrastive learning. This enables the system to learn subtle, domain-specific visual patterns (e.g., mosquito flight behavior) that generic pre-trained features miss, while preventing false negatives from same-class pairs.

  2. Temporal Aggregation for Video-Level Classification: Add a temporal pooling or attention mechanism (e.g., mean pooling, self-attention over frame embeddings, or a lightweight transformer) after frame-level image–text similarity scoring. This converts noisy, high-frequency frame predictions into robust video-level decisions, achieving perfect accuracy on full recordings despite frame-level imperfections.

  3. Biologically-Informed Prompt Engineering: Dynamically generate textual prompts that encode known behavioral phenotypes (e.g., infected mosquitoes exploring cage corners vs. healthy mosquitoes remaining central) during training, and use simpler, class-specific prompts (e.g., a DENV2-infected mosquito) at inference. This allows the system to leverage semantic priors for better feature alignment without requiring complex prompts at deployment.

  4. Motion-Based Feature Fusion: Integrate a motion-path extraction branch (via consecutive-frame differencing) as an auxiliary input to the vision encoder. This provides complementary spatiotemporal cues, improving sensitivity (from 94.97% to 99.91%) and precision, especially when visual appearance is ambiguous due to lighting or background noise.

  5. Leakage-Aware Cross-Validation and Early Stopping: Enforce video-level data splitting (all frames from one video stay in the same fold) and use early stopping with patience to prevent overfitting to temporally correlated frames. This yields stable, generalizable performance across folds and avoids inflated accuracy from data leakage.

  6. Adaptive Frame Sampling and Preprocessing: Uniformly sample a fixed number of frames (e.g., T=32) per video, resize to 224×224, convert to RGB, and normalize to [0,1]. Additionally, apply YOLO-based mosquito detection and background masking to isolate the subject, reducing environmental interference (lighting, shadows, background clutter) and improving feature extraction reliability.


What the Improved AI System Can Do:

  • Automatically classify infection status from raw video data with >98% frame-level accuracy and >99% sensitivity, and 100% video-level accuracy, even for small, fast-moving, and irregularly behaving subjects in noisy environments.

  • Generalize across lighting conditions, camera angles, and background variations by leveraging fine-tuned CLIP representations and motion-based features, reducing the need for manual feature engineering.

  • Provide interpretable classifications by aligning visual features with semantic text prompts, enabling users to understand why a video is classified (e.g., infected mosquito exploring corners) rather than just receiving a binary label.

  • Handle group-level recordings (multiple subjects per video) without individual tracking, making it applicable to real-world surveillance or ecological monitoring where individual tagging is impractical.

  • Adapt to new biological or behavioral classification tasks by swapping textual prompts and fine-tuning on new datasets, offering a reusable framework for detecting other pathogen-induced behaviors or subtle movement anomalies in video.

Sources

Related papers