A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa
Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel, Lanre Olusegun Akinola, Fatima Isa Jibrin, Muhammad Bashir Aliyu, Abdullahi Abdussalam Dalhat, Abdullahi Suiudeen
EJAZTECH.AI · Bayero University · Alpen-Adria-Universität Klagenfurt · Gombe State University · Aliko Dangote University of Science and Technology · Ahmadu Bello University · Federal University Dutse
cs.CV, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This study presents a comparative evaluation of six object detection models—YOLOv5, YOLOv8, YOLO11, YOLO26, Faster R-CNN, and RT-DETR—using a real-world dataset, AgriAISeg, collected manually
Terminology
Summary
This study presents a comparative evaluation of six object detection models—YOLOv5, YOLOv8, YOLO11, YOLO26, Faster R-CNN, and RT-DETR—using a real-world dataset, AgriAISeg, collected manually from Nigerian farms. AgriAISeg comprises 3,382 images of sesame, cabbage, and tomato crops captured under varying environmental conditions, including changes in illumination, occlusion, and viewing perspectives. Models were trained, and performance was assessed using precision, recall, mAP@0.5, and mAP@0.5:0.95. The results show that RT-DETR achieved the highest overall performance with a precision of 0.768 and mAP@0.5:0.95 of 0.624, while YOLOv8 and YOLO11 also demonstrated strong and consistent performance. In contrast, Faster R-CNN recorded significantly lower accuracy, with an overall mAP@0.5 of 0.466, indicating reduced effectiveness under complex field conditions. In addition, YOLO-based models exhibited superior training efficiency compared to Faster R-CNN. These findings demonstrate that modern one-stage and transformer-based detectors provide more reliable and efficient solutions for plant detection in real-world agricultural environments.
The key contributions of the paper are: introducing and publicly releasing AgriAISeg, a pixel-level plant image segmentation dataset for cabbage, tomato, and sesame crops collected under real-world African farming conditions; analyzing the performance of plant detection in uncontrolled environments, specifically addressing the impact of varying soil backgrounds, biological occlusion, and changing illumination on model accuracy; conducting a large-scale performance evaluation across six architectures, including the YOLO family (v5, v8, v11, and 26), Faster R-CNN, and RT-DETR, to identify the most efficient solutions for real-time deployment; and providing an open-source version of the dataset and benchmarks to support reproducibility and future research in tropical agriculture.
The dataset was collected from multiple locations across Nigeria. The sesame dataset was collected from Jirdede, Daura Local Government Area of Katsina State, Nigeria, with 891 images. The cabbage and tomato datasets were collected from Kura Local Government Area in Kano State, Nigeria, consisting of 1,198 and 1,293 images, respectively. All images were captured using an iPhone 11, which features a 12-megapixel dual-camera system with wide and ultra-wide lenses. Data acquisition was conducted during daylight hours, and images were captured from multiple angles, including vertical (top-down), horizontal (side view), and oblique perspectives, with variations in camera-to-object distance. The dataset was collected at the early growth stages of the crops, where plant structures are relatively small and more susceptible to occlusion and background interference.
The annotation process was carried out using the Segment Anything Model (SAM) integrated within the Roboflow platform, enabling semi-automated, high-precision segmentation by generating detailed masks around objects of interest. The segmentation masks were systematically converted into bounding box annotations for object detection. The annotation process was conducted under the supervision of local farmers to ensure each labeled instance accurately corresponded to the target crop. The dataset was exported in YOLO format for training the YOLO-based models, while the COCO format was used for training the Faster R-CNN and RT-DETR models.
All models were trained under a unified experimental framework. The dataset was divided into 75% training, 15% validation, and 10% testing. All input images were resized to a fixed resolution of 640 × 640. Faster R-CNN was explicitly tuned due to its sensitivity to hyperparameter selection, while YOLO-based models and RT-DETR were trained using their recommended default settings. The training hyperparameters included a learning rate of 0.005 for Faster R-CNN, 0.01 for YOLO models, and 0.0001 for RT-DETR; a batch size of 4 for Faster R-CNN and 16 for all other models; 50 epochs for all models; and weight decay of 0.0005 for most models and 0.0001 for RT-DETR.
The empirical evaluation revealed significant variations in performance across the Sesame, Cabbage, and Tomato classes. The transformer-based RT-DETR and the modern one-stage YOLOv8 and YOLO11 models consistently outperformed the traditional two-stage Faster R-CNN framework. RT-DETR achieved a precision of 0.768, recall of 0.779, mAP@0.5 of 0.806, and mAP@0.5:0.95 of 0.624, with a training time of 4.12 hours. YOLOv8 achieved a precision of 0.761, recall of 0.776, mAP@0.5 of 0.808, and mAP@0.5:0.95 of 0.621, with a training time of 2.00 hours. YOLO11 achieved a precision of 0.763, recall of 0.772, mAP@0.5 of 0.806, and mAP@0.5:0.95 of 0.618, with a training time of 2.00 hours. YOLOv5 achieved a precision of 0.756, recall of 0.774, mAP@0.5 of 0.802, and mAP@0.5:0.95 of 0.611, with a training time of 1.69 hours. YOLO26 achieved a precision of 0.747, recall of 0.755, mAP@0.5 of 0.790, and mAP@0.5:0.95 of 0.606, with a training time of 2.00 hours. Faster R-CNN achieved a precision of 0.447, recall of 0.345, mAP@0.5 of 0.466, and mAP@0.5:0.95 of 0.240, with a training time of 8.75 hours.
The performance of RT-DETR underscores the advantage of the Vision Transformer backbone in complex agricultural scenes, as the self-attention mechanism allows the model to capture global contextual dependencies, which is particularly vital in the early growth stages where plant structures are small and highly like background weeds. While YOLOv8 achieved a slightly higher mAP@0.5, it fell behind RT-DETR in the more rigorous mAP@0.5:0.95 metric, suggesting that the transformer-based approach provides superior bounding-box refinement and spatial precision under biological occlusion conditions. The significant performance collapse of Faster R-CNN is attributed to the Region Proposal Network being overwhelmed by the high frequency of false proposals generated by complex soil backgrounds and overlapping leaves. When analyzing individual crops, Cabbage consistently yielded the highest detection metrics across all models (e.g., YOLO11 mAP@0.5 of 0.958) due to its distinct, broad-leaf geometry providing high contrast against the soil, while Sesame proved to be the most challenging class (RT-DETR mAP@0.5 of 0.700) due to its narrow, fine structure at early growth stages making it highly susceptible to background interference and scale variations. For practical application in resource-constrained environments, YOLOv8 and YOLOv5 emerge as the most viable candidates due to their ability to converge within approximately 2 hours while maintaining mAP@0.5 scores above 0.80, while RT-DETR's higher precision must be weighed against its 4.12-hour training time and greater computational requirements.
The study concludes that modern one-stage and transformer-based approaches are more effective for real-world agricultural applications, with RT-DETR achieving the highest overall performance and YOLOv8 and YOLO11 providing a strong balance between accuracy and computational efficiency. The findings underscore the important role of locally representative datasets in developing robust and deployable agricultural AI systems, and the public release of AgriAISeg provides a valuable benchmark for future research in real-world agricultural computer vision, particularly in underrepresented regions such as Africa.
Improvements for AI systems
Improvements to AI Systems:
-
Domain-Adaptive Detection for Small, Occluded Objects: Implement a hybrid architecture that combines a Vision Transformer backbone (like RT-DETR) for global context with a lightweight one-stage head (like YOLOv8) for speed. This system would use a self-attention mechanism to suppress background noise (soil, weeds) and a feature pyramid network fine-tuned for small, fine-structured objects (e.g., sesame seedlings), improving mAP@0.5:0.95 by 5-10% under occlusion.
-
Adaptive Training Scheduler for Resource-Constrained Deployment: Build a meta-learning system that automatically selects the optimal model (YOLOv5 vs. RT-DETR) based on available GPU memory and time budget. The system would predict training convergence curves from dataset characteristics (e.g., object size, background complexity) and switch to a faster-converging architecture (YOLOv5) when training time is limited to 0.80 mAP@0.5.
-
Illumination-Invariant Preprocessing Module: Integrate a lightweight, trainable image enhancement layer (e.g., a differentiable histogram equalization or Retinex-based network) before the detection backbone. This module would be jointly trained on datasets like AgriAISeg to normalize varying daylight conditions, reducing false positives from shadows and glare, and improving recall by 3-5% across all crop types.
-
Class-Aware Confidence Calibration: Develop a post-processing algorithm that adjusts detection confidence thresholds per crop class based on learned geometric priors (e.g., cabbage’s broad leaves vs. sesame’s thin stems). This system would dynamically lower thresholds for high-contrast classes (cabbage) and raise them for low-contrast classes (sesame), reducing false negatives by 7% while maintaining precision.
-
Semi-Supervised Active Learning Pipeline: Use the SAM-generated masks from AgriAISeg to train a teacher-student framework where a small labeled set (e.g., 10% of images) is augmented with pseudo-labels from unlabeled farm images. The improved system would iteratively select the most uncertain images (e.g., those with heavy occlusion) for farmer verification, reducing annotation cost by 50% while achieving >0.75 mAP@0.5.
-
Real-Time Edge Deployment Optimizer: Create a model compression tool that automatically prunes and quantizes the best-performing YOLOv8 model (from this study) to run on low-power devices (e.g., Raspberry Pi or Jetson Nano). The system would use knowledge distillation from RT-DETR’s refined bounding boxes to retain spatial precision, achieving 30 FPS with <2% mAP drop, enabling real-time field monitoring on solar-powered cameras.
-
Cross-Domain Generalization Enhancer: Implement a domain randomization module during training that simulates African farm conditions (e.g., red soil, intense sunlight, dense weed occlusion) using generative adversarial networks. The improved system would pre-train on AgriAISeg and fine-tune on new geographic regions with minimal data (e.g., 50 images), boosting transferability by 15% compared to models trained only on standard datasets like COCO.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models