InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation
summary
The gist
InCoM is an intent-driven perception and structured coordination framework for mobile manipulation that jointly considers stage-adaptive perception and coordinated action generation, addressing two
In short
The episode discusses InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation. Hosts explore how this framework uses inferred intent to dynamically adjust perception focus across multiple scales and decouples base and arm actions to solve strong coupling issues. They conclude that the framework shows promise for robust, adaptive mobile manipulation in unstructured environments.
Key concepts
- Intent-Driven Perception
- This involves a system inferring the robot's goal or intent before looking at the environment. This allows the system to dynamically reweight different levels of perceptual features based on what it is currently doing, shifting focus from broad navigation cues to fine details needed for specific tasks.
- IDPPM
- The Intent-Driven Pyramid Perception Module extracts features using three abstraction levels: shallow, mid, and deep. This multi-scale approach allows the robot to view the world through different lenses simultaneously to adapt its perception based on the task stage.
- Decoupled Decoder
- This strategy factors the high-dimensional action space into separate base actions and arm actions. Training separate decoders for each part, linked by a learnable Trend Token, helps prevent strong coupling between the base and arm movements during coordination.
- Auxiliary Kinematic Supervision
- This component ties dynamic perceptual weighting to the robot's physical state. It uses action increments to define a target distribution, ensuring deep features are emphasized during rapid motion and shallow features are prioritized for precision work.
Terminology used across episodes
This episode discusses
- InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation · Paper Radio
- Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
- AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
- DSPv2: Improved Dense Policy for Effective and Generalizable Whole-body Mobile Manipulation
- Harmonic Mobile Manipulation
- SPIN: Simultaneous Perception, Interaction and Navigation
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- Integrated Task and Motion Planning
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World
- WildLMa: Long Horizon Loco-Manipulation in the Wild
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- OpenVLA: An Open-Source Vision-Language-Action Model
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Learning Transferable Visual Models From Natural Language Supervision
- DINOv2: Learning Robust Visual Features without Supervision
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks
- Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
- From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
- Feature Pyramid Networks for Object Detection
The paper
InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation · Read on arXiv
Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · Anyverse Dynamics
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation".
Dev: InCoM is an intent-driven perception and structured coordination framework for mobile manipulation that jointly considers stage-adaptive perception and coordinated action generation, addressing two key challenges:
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So we're diving into this paper called "InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation," which sounds like it's tackling some really tough problems in mobile robotics. What are your initial thoughts on the title and who wrote this work, Dev?
Dev: I think the title immediately tells you what it's about, Rosa; "Intent-Driven Perception and Structured Coordination for Mobile Manipulation" points directly at those two major sticking points in mobile robots—the perception side and how the base and arm coordinate their actions. The authors are Jiahao Liu, Wenbo Cui, Zhongpu Xia, Yongliang Wang, Haoran Li (corresponding author), and Dongbin Zhao from various institutes.
Taro: From an autonomy standpoint, I'm interested in what they mean by "intent-driven perception"; does this mean the system is actually predicting *why* it needs to look at something before it even looks?
Rosa: That's a really good question, Taro; and that’s exactly what InCoM seems to aim for. The authors are proposing a framework that infers latent motion intent so they can dynamically reweight different levels of perceptual features, which means the attention shifts based on what the robot is doing at any given moment.
Dev: Exactly, Rosa; and this addresses that dynamic perceptual attention problem mentioned in their introduction where existing policies often struggle to allocate perceptual focus correctly as viewpoints change during movement. It tackles the strong coupling between base and arm actions that complicates control optimization too.
Taro: If we look at what they're doing, is this just about better feature fusion, or are they fundamentally changing how the system decides what information matters for a specific stage of the task?
Rosa: It's more than just feature fusion; they’re building this structure where perception adapts to the manipulation stage. Think about it: during navigation toward a bed, attention should be global to find free space, but when grasping a bin, that focus needs to narrow down to local details on the rim.
Dev: And they achieve that adaptation through their Intent-Driven Pyramid Perception Module, or IDPPM; it extracts features using both a sparse three dee encoder and a pretrained 2D visual encoder like DINOv2 across three abstraction levels: shallow, mid, and deep.
Taro: That multi-scale approach sounds smart because it gives the system different lenses through which to view the world simultaneously, but how do they actually decide how much weight to put on each of those three scales?
Rosa: They have this Intent Modulation component that takes the historical action sequence and global visual features and maps them into an implicit intent vector. Then, an MLP-based ScaleGater network uses that intent to produce normalized hierarchical weights, which are what you see in equation (one) as wt =
wS, wM, wD: = Softmax(ϕ(ht)).
Dev: And they tie this dynamic weighting to the robot's physical state with Auxiliary Kinematic Supervision. They calculate L2 norms of action increments for both the base and the arm to define a target distribution where deep features get emphasized during rapid motion, and shallow features are prioritized when doing fine-grained manipulation, as shown in equation (three).
Title and authors: Taro: So it's not just guessing; they are trying to supervise that perceptual focus against what the robot is physically doing in real time. That’s a significant step beyond just learning a static policy.
Rosa: Precisely, Taro; and this is where the structure really shines because they minimize the KL divergence between their predicted weights and this target distribution using equation (four), which ensures deep features are emphasized during fast movements while shallow features handle precision work.
Dev: Beyond perception, InCoM also addresses coordination through a decoupled decoder. They factorize the high-dimensional action space into base actions and arm actions, building a linear path between noise and expert actions in equation (six).
Taro: That decoupling sounds like it directly combats that strong coupling issue we talked about earlier where base and arm actions are tangled up, especially when the robot is moving around.
Rosa: It does; they train separate Transformer decoders for the mobile base and the manipulator to minimize a flow matching loss, but they make it explicit by giving each branch a learnable Trend Token that exchanges information via cross-attention with stop-gradient to ensure bidirectional coordination.
Dev: That stop-gradient exchange is crucial because it prevents gradient interference when coordinating between the two parts of the policy; it’s how they achieve that stable whole-body behavior without getting stuck in local minima due to coupling issues.
Taro: So, if we consider the broader implications for real-world deployment, Rosa, does this framework suggest that mobile manipulation can finally handle truly unstructured environments where perception demands change constantly?
Rosa: Yes, Taro; because they show it performs well in challenging scenarios like SetTable without any privileged information. Their experimental results are pretty compelling too: they showed success rate gains of twenty-eight point two percent, twenty-six point one percent, and twenty-three point six percent across those ManiSkillHAB scenarios, and they maintained a mean success rate of fifty-one point two five percent on the Cobot-Magic robot platform in real-world tasks like "Close Drawer."
Dev: From an engineering standpoint, I have to ask about the loop rate here; how does the complexity of calculating that intent vector and then reweighting features impact the latency we're dealing with? We need to know if this runs fast enough for reliable control.
Taro: That’s a practical concern, Dev; but given that they are using components like DINOv2 and structured attention mechanisms, the focus seems to be on efficiency in how those computations are structured rather than just brute force speed.
Rosa: And that efficiency is part of the whole point; because they also introduced the Dual-stream Affinity Refinement Module, or DARM, which explicitly models geometric consistency and semantic correspondence between three dee point clouds and 2D images.
Dev: DARM sounds like it adds another layer of complexity to the perception pipeline, but if it’s effectively modeling that cross-modal alignment using scaled dot-product attention for both geometric affinity Ageo and semantic affinity Asem, does that add significant computational overhead compared to just a standard end-to-end VLA model?
Title and authors: Taro: It adds structure; it ensures the system isn't just guessing the relationship between what the camera sees and what’s physically there. The geometric affinity being used to construct a transport cost matrix C for Sinkhorn–Knopp regularization is an interesting way to enforce spatial plausibility.
Rosa: And they connect that geometric prior back into semantic alignment by minimizing KL divergence, which is equation (five), making sure the semantic attention distribution Psem aligns with that geometrically derived prior.
Dev: So, to summarize the improvements, we have dynamic perceptual attention driven by intent inference and a decoupled coordination strategy for base and arm actions. But what about the limitations? What doesn't this framework manage effectively?
Taro: The paper notes they still treat the mobile base and manipulator as a unified action vector without fully modeling their mutual dependencies in all cases, although they introduce that Trend Token mechanism to help with bidirectional coordination.
Rosa: They also explicitly mention that removing components hurts performance; for instance, removing IDPPM causes the largest drop in performance, showing its necessity. Also, they point out that without the stop-gradient or the Trend Token consistently degrades success rates during training.
Dev: I think those are clear limitations regarding stability and necessary components for achieving those high results; it shows that this architecture is highly tuned and sensitive to how you set up the coordination mechanism.
Taro: Considering everything, what do you see as the bigger impact of InCoM on the future of mobile robotics outside of just incremental improvements in success rates?
Rosa: I think its implication is moving us toward systems that can operate robustly in complex real-world tasks where perception and control must adapt simultaneously to changing environments. It moves beyond pretraining for tabletop scenarios into true general mobile manipulation.
Dev: From an engineering view, if we can reliably get this kind of coordination running at a decent loop rate, it means we could see robots performing much more intricate tasks autonomously in dynamic settings, which is a huge leap for deployment feasibility.
Taro: For the world at large, having mobile agents that can reason about intent and adapt their sensory focus on the fly feels like it’s bringing us closer to truly versatile robotic assistants capable of handling messy, unstructured environments.
Rosa: Well said, Taro; so that's a lot of exciting stuff we've covered about InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation. Thanks to Dev and Taro for weighing in on the technical details and the autonomy side.
Dev: I agree, Rosa; it’s a solid piece of work showing how structured coordination can actually solve those intractable coupling problems we've seen in mobile robots.
Taro: Indeed; the way they structure the coordination via decoupled decoders really shows a path forward for handling these coupled systems reliably.
The paper's summary: Rosa: So, to recap, InCoM is this framework that uses inferred intent to guide how the AI perceives things and then structures how the robot's base and arm work together during manipulation tasks.
Dev: Exactly; it moves away from having a single policy that tries to do everything at once and instead separates the high-level goal from the low-level execution, which seems key for stability.
Taro: I'm really curious about how this intent inference actually helps when things go wrong in a messy real-world situation, Rosa; does it have mechanisms for reacting when the world doesn't behave as expected?
Rosa: That’s where the IDPPM comes in; because it tracks historical actions and global context, it can dynamically shift its perceptual focus from broad navigation cues to fine details needed for grasping.
Dev: I see how that dynamic weighting helps with the coupling issue we talked about; if the intent signals a fast movement, it emphasizes those deep features that help predict momentum, which keeps the control loop responsive.
Taro: And what about when things get weird? If a sensor fails or an unexpected object appears mid-task, does this intent-driven system have any built-in logic to reconfigure its coordination strategy on the fly?
Rosa: The DCFM decoder is designed for bidirectional coordination, and by having that Trend Token exchange information between the base and arm, it suggests a level of mutual compensation that might allow it to handle minor disturbances better than traditional coupled systems.
Dev: From my side as the controls engineer, the structure of the flow matching loss seems robust because it factorizes those action spaces; if one part drifts, you can still rely on the other branch's trend token to guide its motion without completely breaking the whole system.
Taro: So it’s not just about pre-planned coordination but a more fluid, intent-based coordination that adapts to immediate environmental cues while maintaining structural integrity? That sounds like something we need to think about when we look at systems that need to be safe in unstructured settings, like those drone mapping papers.
Rosa: It really is; the implication here is moving toward robots that aren't just following a script but are actively inferring what's needed for success at every moment during complex tasks.
Dev: If this works reliably outside of a controlled lab setting, Rosa, the loop rate and latency are definitely where we need to keep an eye on to make sure it’s viable for actual deployment.
Taro: I'm looking forward to seeing how they handle those failure modes in more detailed papers; if the intent inference can predict a necessary correction before the error manifests physically, that would be significant progress for autonomy.
The paper's improvements: Rosa: So, to wrap up this discussion on InCoM, we've looked at how it uses intent to steer perception and coordination, and now we're going over what they actually improved in terms of performance and functionality.
Dev: The paper highlights a few key structural improvements; primarily the decoupling of base and arm actions within the DCFM decoder, which tackles that strong coupling we discussed earlier.
Taro: I noticed they also added these auxiliary supervision terms during training, like KL divergence on the scale weights, which sounds like they're stabilizing the learning process against those conflicting perceptual demands.
Rosa: Right; those regularization terms help ensure that the system actually learns to prioritize deep features when needed for fast movement and shallow ones for precision work.
Dev: And I’m paying attention to how they use the Stop-Gradient mechanism in cross-attention between the base and arm branches; that should significantly reduce gradient interference during backpropagation, which is a big win for training stability.
Taro: Beyond training, they show that this framework maintains superior success rates even when tested on novel scenarios without any privileged information, which suggests the intent modeling is robust enough to generalize beyond the specific training data.
Rosa: That generalization capability is huge; it means we can deploy these robots in real-world settings where we don't have perfect maps or ideal initial conditions.
Dev: However, I still have my concerns about the latency; if calculating that intent vector and performing the multi-scale attention reweighting adds too much computation, it might not be suitable for high-speed control loops on embedded hardware.
Taro: That’s a valid point, Dev; but if the structure is efficient enough, maybe we can leverage parallel processing to keep that latency manageable while still getting the benefits of intent-driven adaptation.
Rosa: Exactly; the real-world test will tell us if this level of sophisticated coordination and perception is practical for long-term field work.
Dev: And I want to know more about their limitations concerning long-term stability; they mention that removing key components like IDPPM causes the biggest performance drop, which tells us how sensitive the whole architecture is to losing one of those core pieces.
Taro: So while the wins are impressive, we need to see if this structure can handle unexpected physical interactions or sensor noise in a sustained operation without degrading that coordination.
Rosa: That’s what we need to keep an eye on moving forward; it’s about seeing if this intent-driven approach translates into reliable, long-term operational capability.
Conclusion: Tom: To wrap up this discussion on InCoM, we've looked at how it uses intent to steer perception and coordination, and now we're going over what they actually improved in terms of performance and functionality.
Rosa: So, to recap, InCoM is this framework that uses inferred intent to guide how the AI perceives things and then structures how the robot's base and arm work together during manipulation tasks.
Dev: Exactly; it moves away from having a single policy that tries to do everything at once and instead separates the high-level goal from the low-level execution, which seems key for stability.
Taro: I'm really curious about how this intent inference actually helps when things go wrong in a messy real-world situation, Rosa; does it have mechanisms for reacting when the world doesn't behave as expected?
Rosa: That’s where the IDPPM comes in; because it tracks historical actions and global context, it can dynamically shift its perceptual focus from broad navigation cues to fine details needed for grasping.
Dev: I see how that dynamic weighting helps with the coupling issue we talked about; if the intent signals a fast movement, it emphasizes those deep features that help predict momentum, which keeps the control loop responsive.
Taro: And what about when things get weird? If a sensor fails or an unexpected object appears mid-task, does this intent-driven system have any built-in logic to reconfigure its coordination strategy on the fly?
Rosa: The DCFM decoder is designed for bidirectional coordination, and by having that Trend Token exchange information between the base and arm, it suggests a level of mutual compensation that might allow it to handle minor disturbances better than traditional coupled systems.
Dev: From my side as the controls engineer, the structure of the flow matching loss seems robust because it factorizes those action spaces; if one part drifts, you can still rely on the other branch's trend token to guide its motion without completely breaking the whole system.
Taro: So it’s not just about pre-planned coordination but a more fluid, intent-based coordination that adapts to immediate environmental cues while maintaining structural integrity? That sounds like something we need to think about when we look at systems that need to be safe in unstructured settings, like those drone mapping papers.
Rosa: It really is; the implication here is moving toward robots that aren't just following a script but are actively inferring what's needed for success at every moment during complex tasks.
Dev: If this works reliably outside of a controlled lab setting, Rosa, the loop rate and latency are definitely where we need to keep an eye on to make sure it’s viable for actual deployment.
Taro: I'm looking forward to seeing how they handle those failure modes in more detailed papers; if the intent inference can predict a necessary correction before the error manifests physically, that would be significant progress for autonomy.
Rosa: That’s what we need to keep an eye on moving forward; it’s about seeing if this intent-driven approach translates into reliable, long-term operational capability.
Dev: I agree; the way they structure the coordination via decoupled decoders really shows a path forward for handling these coupled systems reliably.
Taro: Indeed; the way they structure the coordination via decoupled decoders really shows a path forward for handling these coupled systems reliably, and I think that's where we should focus our attention next.
More episodes
- 2610.12154-Stochastic Distribution Network Reconfiguration under Load Uncertainty
- 2607.00148-3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
- 2607.02403-ACID: Action Consistency via Inverse Dynamics for Planning with World Models
- 2510.26623-A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation
- 2406.13267-The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots
- 2511.02147-Census-Based Population Autonomy For Distributed Robotic Teaming
- 2603.08260-Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning
- 2602.14032-RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
- 2602.15397-ActionCodec: What Makes for Good Action Tokenizers
- 2607.01819-Koopman operator theory: fundamentals, control, and applications