DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter

arXiv:2602.05513 · cs.RO, cs.AI · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter".

Jane: The paper was written by Xukun Li, Yu Sun, Lei Zhang, Bosheng Huang, Yibo Peng et al. from XYZ Embodied AI and Beijing Academy of Artificial Intelligence and Technical University of Munich and University of Chinese Academy of Sciences and Tsinghua University and Nanjing University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a new paper that just hit arXiv, and it's got a title that's a mouthful: "DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter." Jane, I'm going to need you to translate that for our listeners.

Jane: Happy to, Tom. Let's break it down. "Bimanual dexterous manipulation" means a robot with two hands, like a human, doing fine motor tasks. "Multimodal" means it uses different senses—sight, touch, and its own body position. "Diffusion Transformer" is the brain architecture that generates actions. And "decoupled" is the clever part—it means the brain treats each sense differently instead of mashing them all together.

Tom: And that's the part I find fascinating. Most robot policies just dump all the sensor data into one big pile and let the model figure it out. DECO says, no, vision gets one pathway, the robot's joint positions get another, and touch gets a third. Each one has its own dedicated way of influencing the action.

Jane: Exactly. The paper argues that these senses play different roles. Vision is the main driver—it tells the robot where things are. The joint positions are like the robot's internal sense of its own body. And touch is the fine detail—the pressure, the contact—that you need when you're really manipulating something delicate.

Lu: And I'd add, Tom, that the tactile part is especially interesting. They built a "plugin adapter" for it. That means you can take a robot policy that was already trained on vision alone, and then add touch later without retraining the whole thing. You just freeze the main brain and bolt on this small adapter.

Tom: So it's like upgrading a car with a new sensor without rebuilding the engine.

Lu: Precisely. And they show it works—they only train about ten percent of the model's parameters to get the tactile benefit. That's a huge efficiency win.

Meng: I'm curious about the practical side, though. They released a dataset called DECO-fifty with fifty hours of real robot data. That's not trivial to collect. How many tasks are we talking about?

Jane: Four main scenarios, twenty-eight sub-tasks, and over eight thousand successful trajectories. And they ran more than two thousand real-world rollouts to evaluate. This isn't a simulation paper—this is real hardware, real hands, real contact.

Tom: And that's what gets me excited. We're seeing robots with actual dexterous hands—five fingers, tactile pads—doing tasks like assembly and sorting. That's the kind of thing that could eventually translate to real-world applications. But let's not get ahead of ourselves. We'll dig into the actual results and what this means in the next segment.

Summary: Tom: We're back, still talking about "DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter." Jane, we've covered the title—now let's get into what the paper actually found.

Jane: The headline result is a seventy-two point two five percent average success rate across all four tasks, which is a twenty-one percent improvement over the best baseline they compared against. And when they added the tactile adapter on top of that, they got another ten point two five percent boost on average. But the really striking number is on the contact-rich tasks—waste disposal and assembly—where tactile brought a twenty percent gain.

Lu: And I think the more interesting finding is that not all tasks need touch. For pick-and-place and material sorting, vision and the robot's own joint positions were basically enough. The tactile adapter only gave marginal gains there. But for assembly, where you're inserting a plug into a socket and you can't see the contact point, touch became essential.

Meng: That matches what I'd expect from an engineering standpoint. When you're doing precision insertion, vision gets occluded—your own hand blocks the view. That's when you need force feedback. The paper actually shows this: without tactile, the robot would sometimes think it had closed a lid when it hadn't, because it couldn't feel the resistance.

Jane: Right, and they have a great example in the waste disposal task. Closing the lid is a compound action—you need to align it, press it down, and then press a button to secure it. Vision can guide you to the right spot, but it can't tell you whether the lid is actually seated. The tactile signal tells you when you've made solid contact.

Tom: So the decoupled design isn't just a theoretical choice—it's practically motivated. Different tasks, different senses matter more or less.

Lu: Exactly. And their ablation study confirms this. They tried injecting tactile the "coupled" way—just adding it to the same conditioning pathway as everything else—and it performed no better than the vision-only baseline. But when they gave tactile its own dedicated cross-attention pathway, performance jumped significantly.

Meng: That's a strong result. It says the way you fuse the modalities matters as much as whether you fuse them at all.

Jane: And it's not just about the tactile pathway. They also showed that the specific assignment matters—proprioception through one mechanism, tactile through another. When they swapped them, performance collapsed, especially on the hardest assembly stage, going from twenty-nine out of forty successes down to zero.

Tom: Zero out of forty. That's a dramatic demonstration that the architecture design is doing real work. So we've got the results—now let's talk about what this means for the future of robotics. That's coming up next.

Improvements: Tom: Welcome back. We're still on "DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter." We've covered the results—now let's talk about what this paper actually improves and why it matters.

Jane: The biggest improvement, I think, is the plugin tactile adapter itself. It's a lightweight module—a small encoder plus cross-attention layers—that can be added to an already-trained vision policy. You don't need to retrain the whole model. You just freeze the main brain and train the adapter.

Lu: And that's a big deal for the field. Think about it: there are already large vision-language-action models trained on massive datasets. Retraining those from scratch with a new sensor modality would be prohibitively expensive. But with this adapter approach, you could take an existing pretrained policy and add tactile awareness with less than ten percent of the parameters being trainable.

Meng: From a deployment perspective, that's huge. It means you can upgrade a robot's capabilities without redoing the entire training pipeline. The paper shows the adapter achieves performance comparable to training from scratch with tactile, but with far less compute and time.

Tom: And the dataset itself is an improvement. DECO-fifty—fifty hours, over five million frames, eight thousand successful trajectories—with both tactile and active vision. The active camera part is interesting because the robot can move its head to look at what it's doing, which adds another layer of complexity.

Jane: That's right. And they were careful to include both tasks that benefit from tactile and tasks where vision is sufficient. That lets the community study when touch actually matters, rather than assuming it's always necessary.

Lu: I think the deeper improvement here is conceptual. The paper challenges the assumption that more modalities fused together is always better. Instead, it shows that structured, decoupled integration—where each sense has its own role—can outperform naive fusion. That's a design principle that could influence how we build all kinds of multimodal systems, not just robot policies.

Meng: And the practical implication is that tactile sensing becomes more accessible. If you can add it to an existing policy with a small adapter, you don't need to build a whole new system from scratch. That lowers the barrier for labs that want to experiment with tactile feedback.

Tom: So we've got a new architecture, a new dataset, and a new training paradigm. What does this mean for the world beyond the lab? Let's bring in Lalam for that perspective in our final segment.

Conclusion: Tom: We're wrapping up our discussion of "DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter." Lalam, you've been listening—what's the big picture here?

Lalam: The big picture is that we're moving toward robots that can handle the physical world the way humans do—with two hands and a sense of touch. This paper shows a practical path to get there: don't just throw all the sensor data together, but give each sense its own role, and make it easy to add new senses later.

Jane: And that has real cultural implications. Think about what robots could do if they could reliably assemble products, sort materials, or handle delicate objects. Manufacturing could become more flexible. Elder care could benefit from robots that can handle objects with appropriate force. Even household tasks—like taking out the trash, which is literally one of the tasks in this paper—could become automatable.

Meng: I'd add that the efficiency angle matters too. The fact that you can add tactile to an existing policy with minimal retraining means this could be adopted more quickly in industry, where retraining costs are a real concern.

Lu: And scientifically, the decoupling principle is a contribution that goes beyond this specific robot setup. It's a way of thinking about multimodal learning that could apply to other domains—anywhere you have multiple sensors that play different roles.

Tom: So to summarize: DECO introduces a decoupled architecture for bimanual dexterous manipulation, a plugin tactile adapter that works with minimal training, and a substantial new dataset with fifty hours of real-world data. The results show a seventy-two percent average success rate, with tactile adding significant gains on contact-rich tasks.

Jane: And the key insight is that how you combine modalities matters as much as whether you combine them. That's a lesson that could shape the next generation of robot learning.

Tom: We've covered the title, the summary, the improvements, and the implications. We're saying goodbye to DECO and getting ready to move on to the next paper. Thanks for listening, everyone—see you next time.

Xukun Li, Yu Sun, Lei Zhang, Bosheng Huang, Yibo Peng, Yuan Meng, Haojun Jiang, Shaoxuan Xie, Guocai Yao, Alois Knoll, Zhenshan Bing, Xinlong Wang, Zhenguo Sun

XYZ Embodied AI · Beijing Academy of Artificial Intelligence · Technical University of Munich · University of Chinese Academy of Sciences · Tsinghua University · Nanjing University

cs.RO, cs.AI

Submitted: 2026-08-17

Updated: 2026-08-18

Comments: 17 pages, 8 figures. Project Page: https://baai-humanoid.github.io/DECO-webpage/

Code: https://github.com/BAAI-Humanoid/DECO

Project page: https://baai-humanoid.github.io/DECO-webpage

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

The gist: DECO is a decoupled multimodal diffusion transformer paradigm for bimanual dexterous manipulation that disentangles vision, proprioception, and tactile signals through specialized conditioning

Key concepts

Bimanual dexterous manipulation
This refers to a robot with two hands capable of performing fine motor tasks similar to humans, such as assembling objects or sorting materials.
Decoupled architecture
This design treats different sensory inputs separately, giving each sense its own dedicated pathway for influencing the action instead of combining all sensor data into one single input.

Terminology

Summary

DECO is a decoupled multimodal diffusion transformer paradigm for bimanual dexterous manipulation that disentangles vision, proprioception, and tactile signals through specialized conditioning pathways, enabling structured and controllable integration of multimodal inputs, with a lightweight adapter for parameter-efficient injection of additional signals. Alongside DECO, the authors release the DECO-50 dataset for bimanual dexterous manipulation with tactile sensing, consisting of 50 hours of data and over 5M frames, collected via teleoperation on real dual-arm robots. The authors train DECO on DECO-50 and conduct extensive real-world evaluation with over 2,000 robot rollouts. Experimental results show that DECO achieves the best performance across all tasks, with a 72.25% average success rate and a 21% improvement over the baseline. Moreover, the tactile adapter brings an additional 10.25% average success rate across all tasks and a 20% gain on complex contact-rich tasks while tuning less than 10% of the model parameters.

The contributions are threefold: DECO, a decoupled multimodal DiT that conditions on visual, proprioceptive states, and tactile modalities via separate, decoupled injections, which improves policy performance over coupled fusion; a plugin tactile adapter that significantly improves vision-based DECO on contact-rich tasks by training only a small fraction of the parameters, demonstrating that tactile can be effectively added to pretrained visuomotor policies; and DECO-50, a bimanual dexterous manipulation dataset with tactile and active vision, comprising 4 scenarios, 28 sub-tasks, over 8K successful trajectories and 5M frames, including both contact-rich tasks that benefit from tactile and tasks where visual and proprioceptive information suffice.

The method follows a two-stage training paradigm: first learning a vision-based policy, and then extending it with tactile sensing via a lightweight adapter while keeping the pretrained policy frozen. The core block is the Multimodal Diffusion Transformer (MMDiT), which fuses visual, proprioceptive, tactile, and action-related information through modality-specific conditioning mechanisms. Vision plays a dominant role, so MMDiT adopts joint self-attention between visual tokens and action tokens. Binocular images are encoded using a shared ResNet-34 backbone, and the resulting feature maps are flattened into token sequences with rotary positional embeddings applied independently. The noisy action sequence is embedded into action tokens with learnable position embeddings. Visual tokens and action tokens are separately projected through linear layers to obtain modality-specific query, key, and value representations, which are concatenated to construct a unified attention space. After root mean square normalization, self-attention is computed over the combined representation, and the outputs are split back into their respective modalities. Proprioceptive states, diffusion timesteps, and optional task-level conditions are injected via adaptive layer normalization (AdaLN), where the conditioning parameters are generated by 2-layer MLPs with SiLU activation. Tactile signals are incorporated through a dedicated cross-attention module, enabling lightweight and plug-and-play integration without modifying the self-attention structure of the pretrained vision-based policy.

The plugin tactile adapter contains a tactile encoder and cross-attention modules for injecting tactile information, while Low Rank Adaptation (LoRA) selectively fine-tunes the attention layers of the pretrained vision-action backbone. During the second training stage, the pretrained policy is frozen, and only the adapter parameters are optimized. The tactile encoder produces region-level features using two complementary approaches: averaging raw values within each tactile pad and projecting raw values through a learnable linear layer. Features from both approaches are concatenated, and a gating mechanism controls their relative influence to produce the final tactile embeddings. These tactile embeddings are injected into the pretrained vision-based DECO via cross-attention, where the LoRA-based adapter modulates the attention projections in a low-rank manner, enabling the model to selectively attend to tactile cues relevant to contact-rich interactions while maintaining the original visual manipulation capabilities.

The DECO-50 dataset comprises four tasks: Pick and Place, Material Sorting, Waste Disposal, and Assembly. Pick and Place involves moving a plate with the left hand while picking up objects from the table and placing them on the plate with the right hand, evaluating basic bimanual coordination. Material Sorting involves grasping moving objects on a conveyor belt and placing them into corresponding containers, introducing dynamic robot-object interaction. Waste Disposal involves picking up and throwing trash into a bin while using the other hand to open and close the bin by pressing, evaluating the benefit of tactile sensing in contact-rich phases. Assembly involves simultaneously controlling a socket and a plug with both hands to complete assembly, requiring precise bimanual coordination and force control.

Experiments were conducted on four tasks with baselines including ACT, DP, and DP.t (which concatenates tactile information into the input and trains from scratch). DECO.p integrates tactile modality with the tactile adapter. Results show that DECO achieves 72.25% average success rate across all tasks, a 21% improvement over the baseline, while DECO.p achieves 82.50% average success rate, a 31.25% improvement. On contact-rich tasks (Waste Disposal and Assembly), DECO achieves 53.13% and DECO.p achieves 73.13%, a 53.75% improvement. The experiments demonstrate that not all tasks need tactile to achieve high performance: pick-and-place and material sorting are weakly tactile-relevant where vision and proprioception usually suffice, while waste disposal and assembly are strongly tactile-relevant where tactile signal is essential for key subgoals such as detecting contact, monitoring applied forces, and handling visual occlusion.

Ablation studies on assembly tasks compared DECO.cs (tactile injected in a coupled way with proprioception and trained from scratch), DECO.ds (tactile injected in a decoupled way and trained from scratch), and DECO.p. Results show that DECO.ds and DECO.p consistently outperform the baseline DECO, while DECO.cs behaves similarly to DECO, indicating that simply appending tactile embeddings is insufficient. The plugin adapter DECO.p achieves performance comparable to DECO.ds despite updating far fewer parameters (7.97M trainable parameters versus 89.41M total for DECO.ds). Further ablation on conditioning pathway assignment shows that swapping the injection mechanisms of proprioceptive states and tactile signals causes a dramatic drop, especially on Stage 3 (0/40 vs. 29/40), confirming that the specific modality-to-pathway assignment in DECO is crucial.

The paper concludes that tactile information benefits primarily contact-rich tasks, where it enables contact detection, force monitoring, and handling of visual occlusion, while simpler tasks achieve strong performance with vision and proprioception alone. The plugin tactile adapter brings significant improvement on complex contact-rich tasks while requiring fewer than 10% of the model parameters and minimal training time compared to training from scratch. Limitations include evaluation restricted to a single hardware platform with one type of dexterous hand and tactile sensor. Future work will validate the tactile adapter on larger-scale pretrained VLAs and across diverse datasets, hardware platforms, and dexterous hands, and incorporate temporal modeling and memory mechanisms for long-horizon contact-rich manipulation.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities:


Improvement:

Replace the standard single-pathway fusion of vision, proprioception, and tactile signals with a decoupled architecture where:

  • Vision tokens and action tokens interact via joint self-attention.

  • Proprioceptive states and task conditions are injected via adaptive layer normalization (AdaLN).

  • Tactile signals are injected via cross-attention.

Resulting Capability:

The AI system can generate more accurate and stable action sequences for bimanual dexterous manipulation, especially in tasks where modalities contribute unequally (e.g., vision-dominant tasks vs. contact-rich tasks). This yields a 21% average success rate improvement over coupled baselines (72.25% vs. 51.25% on the DECO-50 benchmark).

The improved AI system can:

  • Perform bimanual dexterous manipulation with a 72.25% average success rate across diverse real-world tasks.

  • Achieve 82.50% success on pick-and-place and 73.13% on contact-rich tasks when tactile is available.

  • Add tactile sensing to existing policies with <10% parameter overhead and minimal training time.

  • Distinguish when tactile is needed vs. when vision/proprioception suffice, enabling efficient resource allocation.

  • Handle visual occlusion, detect contact, monitor force, and maintain stable grasps on small or slippery objects.

  • Generalize across 28 sub-tasks with 4 distinct scenarios, including dynamic conveyor-belt sorting and precision assembly.

Abstract

Bimanual dexterous manipulation relies on integrating multimodal inputs to perform complex real-world tasks. To address the challenges of effectively combining these modalities, we propose DECO, a decoupled multimodal diffusion transformer that disentangles vision, proprioception, and tactile signals through specialized conditioning pathways, enabling structured and controllable integration of multimodal inputs, with a lightweight adapter for parameter-efficient injection of additional signals. Alongside DECO, we release DECO-50 dataset for bimanual dexterous manipulation with tactile sensing, consisting of 50 hours of data and over 5M frames, collected via teleoperation on real dual-arm robots. We train DECO on DECO-50 and conduct extensive real-world evaluation with over 2,000 robot rollouts. Experimental results show that DECO achieves the best performance across all tasks, with a 72.25% average success rate and a 21% improvement over the baseline. Moreover, the tactile adapter brings an additional 10.25% average success rate across all tasks and a 20% gain on complex contact-rich tasks while tuning less than 10% of the model parameters.

Sources

Related papers