Clothes Grasping and Unfolding Based on RGB-D Semantic Segmentation

arXiv:2305.03259 · cs.CV, cs.AI · Submitted 2023-05-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Clothes Grasping and Unfolding Based on RGB-D Semantic Segmentation".

Jane: The paper was written by Xingyu Zhu, Xin Wang, Jonathan Freer, Hyung Jin Chang and Yixing Gao from Jilin University and University of Birmingham.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Hey everyone, welcome back to the show. I’m Tom, and as always I’m joined by my co-host Jane. Today we’re looking at a fresh arXiv paper that’s got us both pretty excited. It’s called “Clothes Grasping and Unfolding Based on RGB-D Semantic Segmentation.”

Jane: And I have to say, Tom, the title alone tells you a lot. We’re talking about robots that can actually grab a piece of clothing and unfold it. That sounds simple, but if you’ve ever tried to teach a robot to handle something floppy, you know it’s a nightmare.

Tom: Right, a rigid block is easy to pick up. A shirt that’s crumpled and hanging on a hanger? That’s a whole different beast. The paper comes from a team at Jilin University in China, with collaborators at the University of Birmingham. They’re using a Baxter robot, which is a pretty common research platform.

Jane: And what they’re doing is using both a regular color camera and a depth camera together. So RGB plus depth, hence the RGB-D in the title. The idea is to segment the clothing into different regions, like the collar, the outer edge, and the rest of the garment, so the robot knows where it can safely grab.

Tom: That’s the key move here. Instead of trying to find a single point on the clothing, they find a whole region that’s graspable. That gives the robot more options, especially when parts of the clothing are folded over or occluded.

Jane: Exactly. And that’s a big deal because a lot of earlier work just looks at depth images alone. But depth sensors can be noisy, and if the clothing is all bunched up, the depth image can look like a mess. Adding the color image gives the network a lot more to work with.

Tom: So they’re not just throwing both images into a neural network and hoping for the best. They actually designed a new fusion module that combines the color and depth features in a clever way. We’ll get into that in a minute, but first, why should anyone care about robots unfolding clothes?

Jane: Well, the big motivation here is robot-assisted dressing. Think about elderly people or people with limited mobility who need help getting dressed. A robot that can reliably grab a shirt and unfold it is a necessary first step before it can help someone put it on. So this isn’t just a lab curiosity, it’s a practical step toward assistive robotics.

Tom: And the results are pretty solid. They report an eighty-four percent success rate on real robot trials. That’s not perfect, but for a task this hard, that’s a strong number. And they did a hundred trials, so it’s not just a lucky run.

Jane: Yeah, and we should mention they also tested their segmentation model on a public dataset called NYUDv2, which is a standard benchmark for indoor scene understanding. They got results comparable to the state of the art, which shows the model isn’t just tuned for clothing.

Tom: So we’ve got a new network, a new way to fuse RGB and depth, and a real robot doing real tasks. I’m curious about how they actually train this thing, because labeling real images of clothes is time-consuming. That’s going to be our next topic.

Jane: Good hook. Let’s talk about the data problem and how they tackled it.

Summary: Tom: So Jane, we just mentioned the data problem. Labeling hundreds of real images of clothes by hand is slow and expensive. The authors of “Clothes Grasping and Unfolding Based on RGB-D Semantic Segmentation” came up with a clever workaround.

Jane: Right, and it’s not just about collecting more data. They built a data augmentation method that’s based on an adversarial strategy. That sounds fancy, but let me break it down. Normally, you might flip an image, change its brightness, or rotate it a bit to create more training examples. That works for a single image, but here they have both RGB and depth images, and they need to transform them together while keeping the labels aligned.

Tom: So if you rotate the RGB image, you also have to rotate the depth image and the segmentation label the same way. Otherwise, the network gets confused. But the clever part is that they don’t just use random transformations. They train a small network to learn which transformations are hardest for the main segmentation network.

Jane: Exactly. So you have two networks playing against each other. The augmentation network tries to make the segmentation network’s job harder by applying tricky color and geometric changes. The segmentation network tries to get better despite those changes. They take turns improving, and the result is a model that generalizes much better.

Tom: And that’s a big deal because it means they can train on just one hundred fifty labeled image pairs and still get good performance. They actually ran an ablation study where they compared training with and without this augmentation. Without it, the success rate on the robot dropped to fifty-seven percent. With it, they got eighty-four percent.

Jane: That’s a huge jump. And they even compared against just collecting more real data. They collected four hundred fifty labeled images and trained without augmentation, and that only got them to seventy-four percent success. So the augmentation method beat having three times as much real data.

Tom: That’s a really strong result. It suggests that the quality and diversity of the training signal matters more than raw quantity. And it also means that if you want to apply this to a different type of clothing, you don’t need to start from scratch with thousands of labeled images.

Jane: Right, and that’s important because the type of clothing matters. A t-shirt folds differently than a jacket or a pair of pants. The authors used a white medical suit in their experiments, but the method should transfer to other garments if you collect a modest amount of new data.

Tom: So the data augmentation is one piece. The other piece is the network architecture itself, which we mentioned earlier. They call it BiFCNet, and it has this Fractal Cross Fusion module. That’s where things get interesting from a technical standpoint.

Jane: And I’m glad we’re getting to that, because the fusion of RGB and depth is really the heart of this paper. Let’s talk about why just stacking the two images together isn’t enough.

Tom: Good point. We’ll get into the details of the network next.

Improvements: Tom: So we’re back, and we’re still talking about “Clothes Grasping and Unfolding Based on RGB-D Semantic Segmentation.” Jane, you just said that simply stacking RGB and depth images isn’t enough. Why not?

Jane: Because the two modalities are so different. RGB gives you color and texture, depth gives you shape and distance. If you just concatenate them and feed them into a network, the network has to figure out on its own how to combine them. Often, it ends up relying too much on one modality and ignoring the other.

Tom: And that’s a problem, especially when the depth image is noisy, which happens a lot with real sensors. So the authors designed a module called the Fractal Cross Fusion, or FCF, that explicitly fuses the features from both branches.

Jane: Right. The idea is to look at global features, not just local ones. Most fusion methods focus on local patterns, like edges or small textures. But the authors use something called fractal geometry to capture complex, self-similar patterns across the whole image. That helps the network understand the overall structure of the clothing, not just little patches.

Tom: And they do this by taking the feature maps from each layer of the network and passing them through a fractal process. That process applies convolutions of different kernel sizes, from one to six, which captures information at multiple scales. Then they combine those scales in a way that emphasizes the most informative parts.

Jane: It’s a bit like looking at a fern leaf. Each little leaflet looks like a smaller version of the whole leaf. Fractal geometry is about finding those repeating patterns. The network uses that idea to find structure in the feature maps that a simple convolution might miss.

Tom: And they also have a channel attention mechanism, which basically tells the network which features are most important for the task. So they’re not just fusing features blindly; they’re weighting them based on what matters for segmentation.

Jane: Exactly. And the results show it works. On their clothing dataset, they got a mean IoU of eighty-five point one three percent and a pixel accuracy of ninety-seven point nine five percent. When they compared against just stacking the features, the FCF module gave a clear improvement in the robot’s success rate, from seventy-nine percent to eighty-four percent.

Tom: So the architecture is doing real work. But there’s one more piece I want to get to, which is how they actually pick the grasping point from the segmented regions. That’s where the robot’s behavior comes in.

Jane: Right, because segmentation alone doesn’t tell the robot where to grab. They need a strategy to pick a point that’s flat and stable, so the gripper can actually close on it without slipping.

Tom: And they do that by measuring the flatness of the area around each candidate point on the outer edge of the collar. They pick the point where the surrounding area is most flat, and then they compute a grasping direction based on the geometry of the collar.

Jane: And that’s a nice touch, because grabbing at the wrong angle can cause the clothing to slip or bunch up. They also compared their point selection against a random selection method, and the difference was stark. Random selection only got a fifty-seven percent success rate, while their method got eighty-four percent.

Tom: So it’s not just about having a good segmentation model. The way you use the segmentation output matters just as much. That’s a lesson that applies beyond clothing, I think.

Jane: Definitely. Any task where you need to manipulate a deformable object could benefit from this kind of region-based reasoning. We’ll wrap up with some thoughts on the bigger picture next.

Conclusion: Tom: Alright, we’re in the final stretch. Let’s take a step back and look at the whole picture of “Clothes Grasping and Unfolding Based on RGB-D Semantic Segmentation.”

Jane: So to recap, the paper gives us three main contributions. First, a new network architecture, BiFCNet, that fuses RGB and depth features using fractal geometry. Second, an adversarial data augmentation method that reduces the need for large labeled datasets. And third, a complete pipeline for grasping and unfolding clothing on a real robot, with a smart point selection strategy.

Tom: And the results speak for themselves. An eighty-four percent success rate on a hundred real robot trials, plus competitive performance on the NYUDv2 benchmark. That’s a solid package.

Jane: I think the most exciting implication is that this could make robot-assisted dressing more practical. If you can reliably get a shirt off a hanger and unfold it, you’re one step closer to helping someone put it on. That has real potential for assistive care.

Tom: And the data augmentation piece is important beyond just clothing. Any robotics task that involves RGB-D data and limited real-world examples could benefit from that adversarial training approach.

Jane: The authors also mention that in the future they want to explore unsupervised methods for semantic segmentation, which would further reduce the need for labeled data. That would be a big step forward.

Tom: It would. And honestly, the fact that they’re thinking about that shows they’re aware of the practical constraints of deploying these systems in the real world.

Jane: So we’re saying goodbye to this paper, but I think it’s one we’ll be referencing for a while. It’s a nice example of combining clever architecture design with practical engineering.

Tom: Couldn’t agree more. Thanks to everyone listening, and we’ll be back soon with another paper to break down. Until then, keep your eyes on the robots.

Jane: And maybe keep your clothes folded, just in case. See you next time.

Xingyu Zhu, Xin Wang, Jonathan Freer, Hyung Jin Chang, Yixing Gao

Jilin University · University of Birmingham

cs.CV, cs.AI

Submitted: 2023-05-08

Updated: 2026-08-17

Comments: This paper is accepted to ICRA 2023

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 63/100

The gist: This paper proposes a novel Bi-directional Fractal Cross Fusion Network (BiFCNet) for RGB-D semantic segmentation, applied to the task of clothes grasping and unfolding in robot-assisted dressing.

Key concepts

RGB-D Semantic Segmentation
This technique uses both color (RGB) and depth information (D) simultaneously to classify different parts of clothing, such as the collar or outer edge. This helps a robot understand where it can safely grab a piece of fabric.
Adversarial Data Augmentation
This is a method where two networks compete: one tries to make the other's segmentation task harder by applying tricky color and geometric changes to images. This process helps the main network generalize better without needing massive amounts of labeled clothing data.
Fractal Cross Fusion (FCF)
This is a module in the BiFCNet architecture that fuses RGB and depth features. It uses fractal geometry to capture complex, self-similar patterns across the whole image, allowing the network to understand overall clothing structure rather than just small local patches.
Grasping Strategy
After segmentation, the robot selects a grasping point by measuring the flatness of the surrounding area on the outer edge of a garment. This ensures it picks a stable spot for gripping and avoids slippage.

Terminology

Summary

This paper proposes a novel Bi-directional Fractal Cross Fusion Network (BiFCNet) for RGB-D semantic segmentation, applied to the task of clothes grasping and unfolding in robot-assisted dressing. The authors state: We propose a novel Bi-directional Fractal Cross Fusion Network (BiFCNet) for semantic segmentation, enabling recognition of graspable regions in order to provide more possibilities for grasping. Instead of relying solely on depth images, the network uses both RGB and depth data, where the Fractal Cross Fusion (FCF) module fuses RGB and depth data by considering global complex features based on fractal geometry. The FCF module is designed to address limitations of existing RGB-D fusion methods, which mainly consider the fusion of local features and ignore the fusion of global features, while fractal geometry has been shown to be an effective method for extracting globally complex features.

To reduce the cost of real labeled data collection, the authors propose a data augmentation method based on an adversarial strategy, in which the color and geometric transformations simultaneously process RGB and depth data while maintaining the label correspondence. This method trains the data augmentation networks and BiFCNet alternately, maximizing and minimizing the loss respectively, as described by the equation: max ϕ min θ Loss [fθ (gϕ (cϕ)), gϕ (l)].

The paper also presents "a pipeline for clothes grasping and unfolding from the perspective of semantic segmentation, through the addition of a strategy for grasp point selection from segmentation regions based on clothing flatness measures, while taking into account the grasping direction." The pipeline segments the clothing into three regions: outer edge, inner edge, and other parts of the garment. Grasp points are selected by computing a flatness measure Fn for each point on the outer edge, based on the variance of direction vectors to the inner edge, and the point with the smallest Fn is chosen as the grasp point. The grasping direction is then calculated by transforming the vector of the selected point into robot coordinates, adjusting the angle with the positive x-axis to 45 degrees.

The authors evaluate BiFCNet on the public NYUDv2 dataset, achieving mIoU and PA of 51.8% and 77.9% respectively, reaching a level comparable to the state-of-the-art methods. On their own clothing dataset, consisting of 350 RGB-D image pairs of a white medical suit (150 for training, 200 for testing), the model achieves mIoU and PA on the test set reached 85.13% and 97.95%, respectively. The model is deployed on a Baxter robot, and after 100 experiments with randomly changed hanging points, the success rate reached 84%.

Ablation studies are conducted to evaluate each component. First, regarding data augmentation, the authors compare training with and without augmentation on 150 images, and also with 450 images without augmentation. Results show that Our data augmentation method results in a higher percentage of successful grasps and unfolds outperforming both cases where data augmentation is not used, even when additional training data is provided. Specifically, with augmentation on 150 images, the success rate is 84%, while without augmentation on 150 and 450 images, the rates drop to 57% and 74%, respectively. Second, regarding the fusion method, the authors compare RGB-only input, RGB-D with stacking fusion, and RGB-D with FCF fusion. The success rates are 73%, 79%, and 84% respectively, demonstrating that our FCF module is able to outperform stacking to fuse RGB and depth data. Third, regarding grasp point selection, the authors compare their proposed selection method against random selection from the outer edge and against a ResNet-101 model that directly regresses grasp points from depth data. The success rates are 84%, 57%, and 61% respectively, showing that first identifying graspable regions before identifying graspable points, results in higher quality grasp point prediction.

The main contributions are summarized as: (1) proposing BiFCNet with the FCF module for RGB-D fusion based on fractal geometry; (2) proposing an adversarial strategy-based data augmentation method for RGB-D semantic segmentation; (3) proposing a complete pipeline for clothes grasping and unfolding combining the above with a flatness-based grasp point selection strategy; and (4) evaluating the method on NYUDv2 and through extensive real-world robotic experiments on Baxter.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Implementation: I will modify the feature fusion layer in my segmentation networks to include a Fractal Process (FP) module that extracts global complex features using fractal geometry, rather than relying solely on local feature fusion.

What the improved system can do:

  • Achieve 51.8% mIoU and 77.9% PA on NYUDv2, matching state-of-the-art performance

  • Better handle self-occluded objects by leveraging both RGB color features and depth spatial information simultaneously

  • Maintain performance even when depth data contains significant noise, as the FCF module cross-propagates features between modalities

Sources

Related papers