TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases
Gia-Han Truong, Khang Nguyen Quoc, Luyl-Da Quach
cs.CV, cs.AI
Submitted: 2026-08-09
Updated: 2026-08-11
Comments: Accepted at ECCV CVPPA Workshop
License: http://creativecommons.org/licenses/by/4.0/
The gist: To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding.
Terminology
Abstract
To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding. TomaMMU comprises 28,808 high-quality images spanning 15 categories and 213,119 human-annotated visual question-answer pairs, generated through a three-stage pipeline comprising Data Collection, Human Annotation, and Question-Answer Generation. Building on this foundation, TomaBench organizes seven agricultural tasks into a hierarchical three-level taxonomy spanning Basic Perception, Pathology Understanding, and Expert Diagnosis, which together enable systematic evaluation from low-level visual recognition to high-level diagnostic reasoning. The tasks assess visual symptom recognition, taxonomic relationships, and diagnostic reasoning, offering a comprehensive view of how well models grasp plant pathology. Our results pronounced gaps in fine-grained recognition and factually grounded reasoning with 14 state-of-the-art VLMs, consistently underperforming on both challenging MCQs and open-ended questions. These results suggest that current VLMs struggle to translate visual perception into reliable diagnostic knowledge, motivating the need for targeted domain adaptation. Simple fine-tuning on TomaMMU substantially narrows this gap, boosting accuracy on challenging MCQs to 96.09%, outperforming recent VLMs, and pointing toward promising directions for future work. All data and code is available in https://huggingface.co/datasets/enalis/TomaMMU.
Sources
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- LFM2 Technical Report
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- LoRA: Low-Rank Adaptation of Large Language Models
- GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
- Visual Instruction Tuning
- Unified Deployment-Aware Evaluation of Open Reasoning Language Models
- AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
- LeafNet: A Large-Scale Dataset and Comprehensive Benchmark for Foundational Vision-Language Understanding of Plant Diseases
- A Vision-Language Foundation Model for Leaf Disease Identification
- Learning Transferable Visual Models From Natural Language Supervision
- PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
- BioCLIP: A Vision Foundation Model for the Tree of Life
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Qwen3 Technical Report
- Qwen2.5 Technical Report
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models