UltraLight Luma: A Novel Edge-Deployable Perception Network for Crop-Row Segmentation in Agricultural Robotics

arXiv:2610.11764 · cs.RO · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "UltraLight Luma: A Novel Edge-Deployable Perception Network for Crop-Row Segmentation in Agricultural Robotics".

Dev: The gist The proposed approach uses an ultra-lightweight segmentation model to predict crop-row regions from RGB images in real time on constrained hardware, improving parameter efficiency over U-Net, YOLOv8,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Now, let's talk about the title and who wrote this. The paper is titled "UltraLight Luma: A Novel Edge-Deployable Perception Network for Crop-Row Segmentation in Agricultural Robotics."

Dev: It’s written by Balasingham, Ginigaddarage, Rajahrajasingh, Kodagoda, and de Silva. These are the folks who developed this approach.

Rosa: The authors are focused on solving the core challenge: reliable crop-row perception for agricultural robots while keeping deployment costs low. It’s all about that edge deployment aspect of their work.

Dev: They’re presenting a solution that combines a specific segmentation network with a method for extracting navigation lines from those detected rows, which is key for the next steps in robot movement.

Rosa: It suggests this new model should be able to handle the visual information needed for both identifying the rows and figuring out exactly where to go.

Dev: So, they're not just giving us a segmentation model; they’re proposing a complete pipeline from image input all the way to generating movement commands.

The paper's summary: Rosa: So what’s the actual gist of UltraLight Luma? It frames crop-row detection as a binary semantic segmentation problem, meaning it just tries to draw a mask over the image separating the rows from everything else.

Dev: The goal is that pixels labeled one are crop-row regions, and zero is just background. This mask then feeds into a post-processing step for navigation line extraction.

Rosa: And the core of their innovation is this UltraLight Luma architecture which uses a U-Net inspired structure but replaces standard convolutions with depthwise separable blocks to cut down on parameters and memory usage.

Dev: They use this depthwise separable design across four stages: encoder, downsample, decoder, and output head. The encoder starts by reducing the image resolution using a stride of two to get that initial low-level spatial feature map <ref:2610.11764#pg3>.

Rosa: The decoder then restores that spatial detail through three successive bilinear upsampling steps, and it uses skip connections to keep those important row structures intact during the process.

Dev: They also used a two-phase progressive learning strategy, starting with a lower resolution input of sixty-four by sixty-four and slowly training up to higher resolutions like one hundred fifty epochs.

The paper's improvements: Rosa: The authors claim this method improves parameter efficiency compared to models like U-Net or YOLOv8, and they put a number on that: the proposed model uses only 13 point 1k trainable parameters for crop-row segmentation <ref:2610.11764#pg1,only 13.1k trainable parameters>.

Dev: That’s a big reduction when you compare it to something like YOLOv8n, which requires three point one five seven million parameters in that context. It’s about making things much smaller for deployment on low-cost platforms.

Rosa: Plus, they achieved an average crop-row detection score of seventy-four point eight one percent after training for fifty to one hundred fifty epochs, which shows it’s still performing reliably despite being so small.

Dev: They also found its parameter efficiency score is at ninety-one point zero percent, which they show in Figure five as being superior to other comparisons they ran. It’s a solid balance between performance and size, I guess.

Rosa: The post-processing step using the Triangle Scan Method takes the predicted mask and estimates the angular error, which feeds directly into a visual servoing controller to generate velocity commands for movement.

Conclusion: Dev: So, to wrap up this UltraLight Luma paper, it seems they’ve put together a practical system that balances segmentation accuracy with the strict computational constraints of edge devices.

Rosa: They’ve managed to achieve that balance by using a compact encoder-decoder network with depthwise separable blocks and a progressive training strategy.

Dev: The key finding is that you can get a detection score of seventy-four point eight one percent while keeping the parameter count very low at only 13 point 1k parameters, which is what makes it suitable for those resource-constrained agricultural robots we're interested in.

Rosa: It’s an effective way to get the necessary crop-row data needed for downstream tasks like navigation line extraction, which helps generate velocity commands for the robot.

Dev: This UltraLight Luma work shows a very practical application of lightweight networks when dealing with specific problems in robotics, and it sets a benchmark for what we can expect from models designed purely for edge deployment.

Taro: From my side, I’m curious about how robust this works when the world misbehaves—like when there are shadows or weird lighting conditions. The paper shows results, but does it hold up when the scene appearance changes drastically?

Rosa: That's a good point, Taro. The authors did show some results under varied conditions like front shadow and horizontal shadow versus sunny versus cloudy scenarios in Figure five <ref:2610.11764#pg3>.

Dev: So it suggests the model has some inherent robustness to illumination changes, but we’d need more testing on truly challenging visual inputs to know how reliable that detection stays outside the controlled lab environment.

Taro: Right. Because when you’re actually out there on a field, those conditions aren't just simple variations; they can be completely unpredictable. We need to see if this architecture handles those misbehaviors in a real-world operational sense, not just a benchmark sense.

Rosa: Yeah, so the focus now shifts to how we can integrate this detection with more complex navigation logic that accounts for those kinds of environmental surprises.

Dev: Exactly. It’s one piece of the puzzle that gets us closer to having truly autonomous systems operating reliably in varied outdoor settings.

Dhanushka Balasingham, Shamod Ginigaddarage, Hanojhan Rajahrajasingh, Nuwan Kodagoda, Rajitha de Silva

Faculty of Computing, Sri Lanka Institute of Information Technology · Lincoln Centre for Autonomous Systems (L-CAS), University of Lincoln

cs.RO

Submitted: 2026-10-08

Updated: 2026-10-08

The gist: The gist The proposed approach uses an ultra-lightweight segmentation model to predict crop-row regions from RGB images in real time on constrained hardware, improving parameter efficiency over

Key concepts

Crop-Row Detection Formulation
This defines the problem as a binary segmentation task where the goal is to create a mask identifying crop rows (label 1) versus background (label 0). The pipeline involves four steps: data preparation, model architecture design, defining training objectives, and finally processing the output for navigation.
UltraLight Luma Architecture
This model uses a U-Net-like structure enhanced with depthwise separable convolutions. These specialized blocks significantly reduce the number of parameters and memory needed for deployment on edge devices compared to standard segmentation networks.
Progressive Resolution Training Strategy
The training starts with low-resolution images (64x64) before gradually increasing the resolution. This two-phase approach helps the network learn efficiently, utilizes parameters effectively, and speeds up convergence by reducing total training time by 28%.

Terminology

Summary

The gist The proposed approach uses an ultra-lightweight segmentation model to predict crop-row regions from RGB images in real time on constrained hardware, improving parameter efficiency over U-Net, YOLOv8, and YOLOv26 while maintaining reliable crop-row detection performance

Crop-Row Detection Formulation

Crop-row detection is formulated as a binary semantic segmentation problem where the objective is to predict a binary mask M where pixels with label 1 correspond to crop-row regions and pixels with label 0 correspond to background The complete crop-row perception pipeline consists of four phases: A. Dataset Preparation, B. UltraLight Luma Architecture, C. Training Objective, and D. Inference and Post-Processing

UltraLight Luma Architecture

UltraLight Luma uses a U-Net-inspired encoder–decoder structure [12] and extends the lightweight depthwise-separable design of UltraLight+ [13] to segmentation Standard convolutions are replaced with depthwise-separable (DW-Sep) blocks to reduce parameters and activation memory for edge deployment Each DW-Sep block applies a 3 × 3 depthwise convolution followed by a 1 × 1 pointwise convolution, with Batch Normalization (BN) and ReLU after each step The network processes inputs through four sequential stages namely encoder, downsample, decoder, and output head The encoder stage begins with a stem convolution that applies a standard 3 × 3 convolution with stride 2, reducing the 64 × 64 × 3 input to a 32 × 32 × 16 feature map while extracting initial low-level spatial features The decoder stage restores spatial resolution through three successive bilinear upsampling steps, following the Feature Pyramid Network paradigm [14] The final decoder output d0 in R 64×64×16 is produced by a third upsampling step without a skip connection

Training Strategy and Progressive Learning

The model is trained using the Adam optimizer with an initial learning rate of 10−3 and a cosine decay learning-rate schedule The total loss is defined as L = Lfocal + Ldice to address class imbalance in crop-row segmentation, where Focal loss emphasizes hard foreground pixels and Dice loss encourages overlap between the predicted and ground-truth masks To further optimize this process, the network incorporates the two-phase progressive resolution training strategy established in paper [15], which initiates training with low-resolution inputs (64x64) before progressing to higher resolutions This progressive training approach enables highly efficient parameter utilization, accelerates convergence, and reduces total training time by 28%

Performance and Efficiency Comparison

The model achieved IoU values of 15.63%, 17.77%, 18.72%, 19.66%, 19.88%, and 20.68% after training for 50, 100, and 150 epochs, respectively The proposed Luma model achieved an average crop-row detection score of 74.81% while using only 13.1k trainable parameters and reaching a parameter-efficiency score of 91.0% The comparative analysis of ηi,j values is presented in Fig. 5, where Luma achieved 91.0%, demonstrating its superior performance In terms of computational efficiency, UltraLight Luma achieved lower resource consumption across all metrics compared to YOLOv8n, with only 0.013M parameters compared to 3.157M for YOLOv8n

Post-Processing and Navigation Line Extraction

The predicted crop-row mask is processed by the TSM to estimate the angular error, which is then fed to the visual-servoing controller to generate velocity commands The TSM exploits the crop-row vanishing point caused by perspective distortion to define the top point of the centerline, and it performs a sweeping scan to candidate points along the bottom edge For each image, two parameters are extracted from the detected croprow: the angular deviation from the vertical axis (∆θ) and the lateral offset of the lower endpoint from the image center (∆Lx2) The proposed Luma model achieved an average ϵ score of 74.81% using this metric This demonstrates that UltraLight Luma offers a practical balance between segmentation performance, navigation-line extraction, and edge-deployable efficiency for resource-constrained agricultural robots

Conclusion

UltraLight Luma presents an ultra-lightweight encoder–decoder segmentation network for crop-row perception on low-cost agricultural robotic platforms The proposed approach formulates crop-row detection as a binary segmentation problem and combines the predicted mask with a Triangle Scan Method-based navigation-line extraction module to support downstream visual servoing Despite its modest maximum IoU of 20.68%, UltraLight Luma preserved sufficient crop-row structure to achieve an average crop-row detection score of 74.81% while using only 13.1k trainable parameters and reaching a parameter-efficiency score of 91.0% These results show that UltraLight Luma offers a practical balance between segmentation performance, navigation-line extraction, and edge-deployable efficiency for resource-constrained agricultural robots>

Acknowledgements

The authors acknowledge the use of OpenAI ChatGPT for language editing, grammar improvement, and clarity enhancement during manuscript preparation The AI tool was used only to refine the presentation of author-provided technical content and final manuscript content were reviewed and verified by the authors

References

[1] V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient processing of deep neural networks: A tutorial and survey,” Proceedings of the IEEE, vol 105, no 12, pp 2295–2329, 2017

[4] Y. Hu and H. Huang, “Extraction method for centerlines of crop row based on improved lightweight yolov4,” in 2021 6th International Symposium on Computer and Information Processing Technology (ISCIPT), pp 127–132, 2021

[7] R. de Silva, G. Cielniak, G. Wang, and J. Gao, “Deep learning-based crop row detection for infield navigation of agri-robots,” Journal of Field Robotics, vol 41, no 7, pp 2299–2321, 2024

[8] R. Doha, M. A. Hasan, and S. Anwar, “Semantic segmentation approaches in crop row detection,” in 2022 IEEE International Conference on Big Data (Big Data), pp 2224–230, 2022

[9] A. Chaurasia and E. Culurciello, “Linknet: Exploiting encoder representations for efficient semantic segmentation,” CoRR, vol abs/1707.03718, 2017

[12] O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 (N. Navab, J. Hornegger, W. M. Wells, and A. F. Frangi, eds.), (Cham), pp 234–241, Springer International Publishing, 2015

[13] S. Ginigaddarage and H. Rajahrajasingh, “Ultralight+: An optimized progressive cnn architecture for efficient image classification,” in 2025 IEEE 19th International Conference on Industrial and Information Systems (ICIIS), vol 19, pp 401–406, 2026

[15] S. Ginigaddarage and H. Rajahrajasingh, “Ultralight: Progressive resolution training for ultra-efficient cnns on edge devices,” 01 2026

[16] R. de Silva, G. Cielniak, and J.

Improvements for AI systems

  1. To improve navigation under challenging conditions, integrate UltraLight Luma's predicted binary mask with a refined Triangle Scan Method (TSM) analysis to extract a more accurate crop-row centerline, as described by the TSM is a post-processing algorithm that exploits the crop-row vanishing point caused by perspective distortion to define the top point of the centerline. This allows for better estimation of the angular deviation from the vertical axis (∆θ) and the lateral offset of the lower endpoint from the image center (∆Lx2) which are used to compute the overall detection performance ϵ as defined in 10.

  2. To enhance robustness against illumination and scene appearance variations, leverage UltraLight Luma's demonstrated capability under diverse conditions, such as Front shadow, Horizontal shadow, and Sunny vs. Cloudy scenarios shown in Figure 5, to create an adaptive perception system for autonomous navigation. This ensures that the segmentation performance remains reliable even when facing challenging visual inputs.

  3. To drastically reduce computational overhead on embedded devices, deploy UltraLight Luma's architecture, which uses depthwise-separable (DW-Sep) blocks to reduce parameters and activation memory for edge deployment, specifically for real-time crop-row detection on low-cost platforms. This results in a model with only 13.1k trainable parameters and an inference requirement of only 25.23 mJ per inference, making it suitable for hardware with limited memory, power, and inference time.

  4. To optimize training convergence and parameter utilization during deployment, implement the two-phase progressive resolution training strategy established in paper [15], which initiates training with low-resolution inputs (64x64) before progressing to higher resolutions, to accelerate convergence and reduce total training time by 28%. This strategy enables highly efficient parameter utilization for a more stable final model.

Related papers