A Multi-Modal Generative Model for Tomato Disease Leaves Understanding
cs.CV, cs.AI
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: In submission to Computers and Electronics in Agriculture Journal
Code: https://github.com/EnalisUs/SOLAR
License: http://creativecommons.org/licenses/by/4.0/
The gist: Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information.
Terminology
Abstract
Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical deployment in precision agriculture remains limited because most existing approaches treat disease understanding as isolated prediction tasks, failing to capture the complementary relationships among symptom recognition, severity assessment, and question-driven diagnostic reasoning. In tomato pathology, accurate interpretation of diseased leaves requires more than label prediction; it demands integrating visual symptoms with semantic context to support a comprehensive and explainable understanding. Here, we present SOLAR, a multimodal generative model that understands tomato disease spanning six question-answering tasks. SOLAR learns to align visual features with task-aware language representations by Fusion Expert module based on mixture-of-expert, enabling it to generate contextually relevant answers across diverse diagnostic tasks. By formulating tomato disease analysis as a generative Visual Question Answering (VQA) task, SOLAR provides a flexible framework that supports multi-task inference within a single model while improving performance and cross-task knowledge sharing. We evaluate SOLAR on 41,677 images, including 216,209 Question-Answering (QA) pairs to understand tomato leaf disease under both closed and open-ended QA settings. Experimental results show that SOLAR consistently outperforms state-of-the-art vision-only, vision-language, and task-specific models across all tasks, demonstrating superior accuracy, robustness, and multimodal reasoning. These findings highlight the potential of generative multimodal modeling as an effective direction for understanding of plant disease. The code for this study is available at https://github.com/EnalisUs/SOLAR.
Sources
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- Leveraging Vision Language Models for Specialized Agricultural Tasks
- LeafNet: A Large-Scale Dataset and Comprehensive Benchmark for Foundational Vision-Language Understanding of Plant Diseases
- PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
- LLaVA-OneVision: Easy Visual Task Transfer
- LFM2 Technical Report
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Geometric Regularization in Mixture-of-Experts: The Disconnect Between Weights and Activations
- TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases
- LoRA: Low-Rank Adaptation of Large Language Models
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Qwen2.5-VL Technical Report
- Qwen3-VL Technical Report
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- DenseNet: Implementing Efficient ConvNet Descriptor Pyramids
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models