Containerized Vertical Farming Using Cobots

arXiv:2310.15385 · cs.RO · Submitted 2023-10-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Containerized Vertical Farming Using Cobots".

Rosa: Containerized vertical farming (CVF) presents challenges due to space limitations and labor intensity, necessitating automation for key operations like sapling transplantation and harvesting.

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: Well, the title itself tells us they are focusing on using collaborative robots in containerized vertical farming because space is super tight there. Dev It’s interesting that they bring in cobots since traditional mobile manipulators just don't fit into those shipping containers, right? Taro The authors are Mahalingam, Patankar, Phi, Chakraborty, McGann, and Ramakrishnan—they seem to be a solid mix of robotics and autonomy expertise.

Rosa: They're aiming to automate two specific operations: transplanting saplings and harvesting plants. It sounds like they’re trying to solve the labor issue in that environment by automating the physical manipulation steps. Dev And what's exciting about their approach is that they’re not programming a new motion plan for every single plant; they are trying to learn constraints from just one human demonstration.

Taro: That idea of extracting motion constraints from a single human demo to generalize it is where my focus lies, because it suggests a level of adaptability that goes beyond task-specific programming. It implies the system can handle variations in the physical setup without needing entirely new code for every single growing tube configuration.

Rosa: Right, so the core idea is using that demonstration to derive rules for movement, rather than explicitly coding every single insertion or extraction path they need to perform. Dev That moves us away from rigid programming and toward a more learned behavior based on geometric understanding.

The paper's summary: Dev: The summary highlights how they combine a deep learning model, specifically the Segment Anything Model, with geometric knowledge of the tubes and screw-geometric representations of motion into their planning system. Rosa That combination is key because it’s not just relying on vision; it’s using that visual data to define mathematical constraints in SE(three) space <ref:2310.15385#pg0>. Taro So, when the robot needs to transplant something, it uses SAM to figure out where the slot is in three dee, and then those visual features are combined with the demonstration data <ref:2310.15385#pg0>.

Dev: And they represent the demonstration as a sequence of constant screw motions or one-parameter subgroups of SE(three), which they argue is a coordinate-invariant way to describe movement <ref:2310.15385#pg0>. It’s interesting because that representation should theoretically be robust to how you define your starting point in space. Rosa That sounds like a strong theoretical foundation for transferring those constraints from the training demonstration to the actual task instance, which is exactly what they set out to do with different slots.

Taro: I wonder how that mathematical transfer works when the environment changes significantly; if we move outside the exact geometry of the demo, can this constraint transfer still be accurate? It sounds like a major area where you need high autonomy to handle those discrepancies.

Dev: That's a valid concern about robustness. The paper claims this method allows them to define a new sequence of motion subgroup constraints, G′, based on identifying the constant screws that fall inside the sphere around the new objects. It’s like they are extracting a localized rule set for the specific task instance and applying it to plan the next movement.

Rosa: So, in essence, they’ve built a system where one demonstration teaches you *how* to move relative to an object, and then their vision system tells you *where* that object is now so you can apply those learned rules correctly. Taro That dependency on both the visual localization via SAM and the learned motion geometry seems like a clever way to bridge perception and action for this constrained manipulation task.

The paper's improvements: Rosa: The improvements they propose are really about achieving that generalization we talked about earlier, moving past simple task programming. They suggest using the deep learning foundation model, SAM, alongside geometric knowledge of the tubes to define the slot pose estimate in R3. Dev That estimation step is crucial because if the robot doesn't accurately know where the slot is in three dee space, none of that motion constraint transfer will work properly <ref:2310.15385#pg0>.

Taro: I'm interested in how this impacts real-world scenarios where things aren't perfect; for example, what happens when the RGBD data is noisy or if the lighting changes significantly? Does this framework handle those kinds of sensing errors well?

Dev: The experimental validation suggests it's quite resilient, achieving an overall success rate of eighty-three point eight percent in their tests with a Franka Emika Panda manipulator. They showed it could successfully insert saplings into slots with different diameters, like thirty mm and thirty-five mm, while still satisfying those constraints they learned from the demonstration.

Rosa: That's a solid result for handling physical variations in tube sizes, which is exactly what a farming operation needs to do. But what about the harvesting task? Taro Harvesting involves occlusion with foliage, so I wonder if the method can handle that visual ambiguity well when trying to extract those constraints for extraction from the tube.

Dev: For harvesting, they found that even when views were occluded by leaves, the system could still perform it successfully because it used the pose estimates of the planting slots derived from the transplantation task. It seems like reusing prior information helps compensate for temporary visual obstructions.

Conclusion: Rosa: So to wrap up, this paper on "Containerized Vertical Farming Using Cobots" shows a way to use a single human demonstration and deep learning segmentation alongside screw-geometric representations in SE(three) to plan for constrained manipulation tasks <ref:2310.15385#pg0,Containerized Vertical Farming Using Cobots>. Dev The main implication is that we can move toward robots that adapt their motion plans based on learned constraints instead of needing bespoke programming for every single growing tube configuration.

Taro: I think the real impact here is demonstrating how we can create a flexible planning system where the autonomy handles task instance variation by leveraging learned motion subgroups, which opens up possibilities for deploying these systems in less controlled settings.

Rosa: Right, and that’s what makes it so compelling for applications like CVF where labor is scarce and space is limited. It shows a path toward building more versatile robotic systems that can operate without constant manual reprogramming.

Dev: From my end, the success rate of eighty-three point eight percent across different tube specifications gives us a concrete baseline for how reliable this constraint-based transfer method is in practice, provided the initial pose estimation doesn't fail catastrophically.

Taro: I just want to emphasize that while it works well under the tested conditions, future work needs to focus on improving gripper geometry and making the system even more robust against those environmental uncertainties we discussed earlier.

Rosa: Well, it’s clear this research lays a solid foundation for automating repetitive physical tasks in vertical farming using collaborative robots. We’ll keep an eye on how they refine this approach in future iterations of "Containerized Vertical Farming Using Cobots."

Department of Mechanical Engineering, Stony Brook University, USA · Department of Computer Science, Stony Brook University, USA · CubicAcres LLC

cs.RO

Submitted: 2023-10-23

Updated: 2023-10-23

Journal ref: IEEE International Conference on Robotics and Automation (ICRA), 2024

DOI: 10.1109/ICRA57147.2024.10609985

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 71/100

The gist: Containerized vertical farming (CVF) presents challenges due to space limitations and labor intensity, necessitating automation for key operations like sapling transplantation and harvesting.

Key concepts

Collaborative Robots (Cobots)
These are robots designed to work safely alongside humans in shared workspaces. In this context, they are used in vertical farming containers to perform delicate manipulation tasks like inserting saplings or picking greens, requiring them to follow learned motion constraints.
Segment Anything Model (SAM)
SAM is a deep learning foundation model that excels at identifying and segmenting objects in images. It helps the system quickly recognize and isolate specific parts of the scene, such as the growing tubes or saplings, providing geometric knowledge for planning.
Screw-Geometric Representation
This method describes robot movements not as traditional paths but as sequences of constant screw motions within a mathematical framework called SE(3). This representation is coordinate-invariant and allows the system to capture the underlying physical constraints of a movement, making it easier to generalize across different task setups.
Motion Subgroups in SE(3)
These are specific sets of allowed movements or constraints that define how a robot can move relative to its environment. By extracting these subgroups from a demonstration, the system can create flexible motion plans that satisfy the required physical rules for transplanting or harvesting, even when the exact objects change.

Terminology

Summary

Containerized vertical farming (CVF) presents challenges due to space limitations and labor intensity, necessitating automation for key operations like sapling transplantation and harvesting. This paper explores using collaborative robots (cobots) to automate these tasks in CVF by leveraging single human demonstrations to extract motion constraints, enabling the generalization of motion planning for new task instances without task-specific programming.

The Gist

Using RGBD camera images and a single demonstration for each task, it is feasible to perform transplantation of saplings and harvesting of leafy greens using a cobot, without task-specific programming.

Problem Context and Motivation

Vertical hydroponic farming in mobile shipping containers (CVF) offers hyperlocal food production but requires manual labor for steps like transplantation of saplings and harvesting. The tight space constraints preclude the use of traditional automated devices or mobile manipulators. The core challenge lies in reliably performing manipulation tasks, such as insertion of the sapling within the growing tube for transplantation and extraction from the growing tube for harvesting, where motion constraints must be estimated from image data. Programming these constraints typically requires specialized robotics knowledge, which farmers lack.

Novel Method Overview

The proposed method combines three key components: (a) a deep learning-based foundation model for image segmentation, namely, the Segment Anything Model (SAM) from Meta AI, (b) geometric knowledge of the slots in the growing tubes, and (c) a screw-geometric representation of the demonstration as a sequence of constant screw motions or one-parameter motion subgroups of SE(3). This combination allows for generalizing constraints from one demonstration to different task instances.

Key Steps in the Solution Approach

The solution follows three main steps:

  1. Screw Extraction: The process involves extracting task constraints embedded in the demonstrated motion of the end-effector as a sequence of constant screws. This representation is chosen because any path in SE(3) can be approximated by constant screw motions, and it is a coordinate-invariant representation. The demonstration is expressed as a sequence of constant screw segments which are sub-sequences of the original demonstration.

  2. Goal Estimation: To define the task instance (the objects whose poses affect planning), the robot needs the pose of the planting slot and other relevant objects like the sapling pod or collection tray. The pose of a slot is estimated by segmenting pixels in an RGB image using SAM, de-projecting them to 3D points, and then fitting a bounding box to obtain a rough estimate of the slot position in R3. This process allows defining the task instance, such as Ot = [gp, gs] for transplantation.

  3. Transfer of Demonstration: The extracted constraints are transferred by identifying the constant screws that lie inside the region-of-interest surrounding the task-related objects (sapling and planting slot) within a sphere. These local sequences of constant screws are then transformed relative to the new task instance's frame of reference, resulting in a new sequence of motion subgroup constraints, denoted as G′ = [G′p, G′s], which can be used for planning.

Experimental Validation and Results

The approach was tested using a Franka Emika Panda manipulator with an eye-in-hand RGBD camera. The experiments showed an overall success rate of 83.8% when using our approach. The method proved robust, as the robot performed tasks successfully even when tested against different specifications of growing tubes and slots than those used in the demonstration. For the transplantation task, the system successfully inserted a pod into slots of different diameters (30 mm and 35 mm) while satisfying constraints. Failures were mainly attributed to gripper geometry and error in the pose estimate of the planting slots. For harvesting, where foliage occludes views, the robot utilized the pose estimates of the planting slots from the transplanting task to successfully perform harvesting. Failures in harvesting were due to leaves catching between gripper tips.

Conclusion

The paper demonstrates that representing task constraints as motion subgroups in SE(3) is a viable way to plan for constrained manipulation tasks in CVF. The combination of SAM for geometric knowledge and screw-geometry representation with ScLERP-based planners ensures that the generated motion plans satisfy the required task constraints, achieving a high success rate. Future work will focus on optimizing gripper dimensions to improve robustness against observed failures.

Improvements for AI systems

As a fastidious and diligent AI researcher, I have analyzed the core contributions of this paper: leveraging kinesthetic demonstrations, deep learning (SAM) for geometric feature extraction from RGBD data, screw-geometric representation of motion constraints (SE(3)), and Screw Linear Interpolation (ScLERP) for motion planning.

Here are specific improvements to existing AI systems based on this research, and what the improved system can achieve:


The proposed method fundamentally improves the capability of robotic manipulation in highly constrained, unstructured environments like Containerized Vertical Farming (CVF). The core improvement lies in shifting from task-specific programming to a generalized, constraint-based planning paradigm derived from human demonstration.

Here are specific improvements and capabilities:

  1. Improved Generalization of Manipulation Planning using Constraint Subgroups:

  2. Enhanced Pose Estimation for Slot Localization via Geometric Segmentation:

  3. Robust Motion Generation using Screw Geometry for Task Transfer:

  4. Feasibility in Real-World, Unprecedented Scenarios (CVF):

The improved AI system can perform the following specific actions:

  1. A robot equipped with this system can perform the task of transplanting saplings into growing tubes and harvesting plants from them using only a single demonstration for each task, even when presented with an entirely new set of planting slot locations (a different task instance) and different physical specifications for the growing tubes or slots.

  2. The robot can reliably determine the precise 6D pose (position and orientation) of a target planting slot within a shipping container's interior by utilizing RGBD camera data, specifically by employing the Segment Anything Model (SAM) to identify slots in 3D point clouds and projecting them back into pixel space for accurate localization.

  3. The system can generate a precise sequence of joint movements for the robot arm that guarantees successful insertion or extraction of an object (sapling or plant) into a target slot, by mathematically transferring motion constraints extracted from a human demonstration (the screw constraints) to the new task instance using Screw Linear Interpolation (ScLERP).

  4. The system can handle inherent uncertainties in the environment and sensing, such as noisy RGBD point clouds or occlusions caused by foliage during harvesting. The method is robust enough to plan paths that respect the learned motion constraints even when environmental factors (like slight variations in tube dimensions) cause deviations from the initial demonstration geometry, leading to an overall success rate of approximately 83.8% in experimental trials.

Abstract

Containerized vertical farming is a type of vertical farming practice using hydroponics in which plants are grown in vertical layers within a mobile shipping container. Space limitations within shipping containers make the automation of different farming operations challenging. In this paper, we explore the use of cobots (i.e., collaborative robots) to automate two key farming operations, namely, the transplantation of saplings and the harvesting of grown plants. Our method uses a single demonstration from a farmer to extract the motion constraints associated with the tasks, namely, transplanting and harvesting, and can then generalize to different instances of the same task. For transplantation, the motion constraint arises during insertion of the sapling within the growing tube, whereas for harvesting, it arises during extraction from the growing tube. We present experimental results to show that using RGBD camera images (obtained from an eye-in-hand configuration) and one demonstration for each task, it is feasible to perform transplantation of saplings and harvesting of leafy greens using a cobot, without task-specific programming.

Sources

Related papers