InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation

arXiv:2602.23024 · cs.RO · Submitted 2026-02-26 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation".

Dev: InCoM is an intent-driven perception and structured coordination framework for mobile manipulation that jointly considers stage-adaptive perception and coordinated action generation, addressing two key challenges:

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we're diving into this paper called "InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation," which sounds like it's tackling some really tough problems in mobile robotics. What are your initial thoughts on the title and who wrote this work, Dev?

Dev: I think the title immediately tells you what it's about, Rosa; "Intent-Driven Perception and Structured Coordination for Mobile Manipulation" points directly at those two major sticking points in mobile robots—the perception side and how the base and arm coordinate their actions. The authors are Jiahao Liu, Wenbo Cui, Zhongpu Xia, Yongliang Wang, Haoran Li (corresponding author), and Dongbin Zhao from various institutes.

Taro: From an autonomy standpoint, I'm interested in what they mean by "intent-driven perception"; does this mean the system is actually predicting *why* it needs to look at something before it even looks?

Rosa: That's a really good question, Taro; and that’s exactly what InCoM seems to aim for. The authors are proposing a framework that infers latent motion intent so they can dynamically reweight different levels of perceptual features, which means the attention shifts based on what the robot is doing at any given moment.

Dev: Exactly, Rosa; and this addresses that dynamic perceptual attention problem mentioned in their introduction where existing policies often struggle to allocate perceptual focus correctly as viewpoints change during movement. It tackles the strong coupling between base and arm actions that complicates control optimization too.

Taro: If we look at what they're doing, is this just about better feature fusion, or are they fundamentally changing how the system decides what information matters for a specific stage of the task?

Rosa: It's more than just feature fusion; they’re building this structure where perception adapts to the manipulation stage. Think about it: during navigation toward a bed, attention should be global to find free space, but when grasping a bin, that focus needs to narrow down to local details on the rim.

Dev: And they achieve that adaptation through their Intent-Driven Pyramid Perception Module, or IDPPM; it extracts features using both a sparse three dee encoder and a pretrained 2D visual encoder like DINOv2 across three abstraction levels: shallow, mid, and deep.

Taro: That multi-scale approach sounds smart because it gives the system different lenses through which to view the world simultaneously, but how do they actually decide how much weight to put on each of those three scales?

Rosa: They have this Intent Modulation component that takes the historical action sequence and global visual features and maps them into an implicit intent vector. Then, an MLP-based ScaleGater network uses that intent to produce normalized hierarchical weights, which are what you see in equation (one) as wt =

wS, wM, wD: = Softmax(ϕ(ht)).

Dev: And they tie this dynamic weighting to the robot's physical state with Auxiliary Kinematic Supervision. They calculate L2 norms of action increments for both the base and the arm to define a target distribution where deep features get emphasized during rapid motion, and shallow features are prioritized when doing fine-grained manipulation, as shown in equation (three).

Title and authors: Taro: So it's not just guessing; they are trying to supervise that perceptual focus against what the robot is physically doing in real time. That’s a significant step beyond just learning a static policy.

Rosa: Precisely, Taro; and this is where the structure really shines because they minimize the KL divergence between their predicted weights and this target distribution using equation (four), which ensures deep features are emphasized during fast movements while shallow features handle precision work.

Dev: Beyond perception, InCoM also addresses coordination through a decoupled decoder. They factorize the high-dimensional action space into base actions and arm actions, building a linear path between noise and expert actions in equation (six).

Taro: That decoupling sounds like it directly combats that strong coupling issue we talked about earlier where base and arm actions are tangled up, especially when the robot is moving around.

Rosa: It does; they train separate Transformer decoders for the mobile base and the manipulator to minimize a flow matching loss, but they make it explicit by giving each branch a learnable Trend Token that exchanges information via cross-attention with stop-gradient to ensure bidirectional coordination.

Dev: That stop-gradient exchange is crucial because it prevents gradient interference when coordinating between the two parts of the policy; it’s how they achieve that stable whole-body behavior without getting stuck in local minima due to coupling issues.

Taro: So, if we consider the broader implications for real-world deployment, Rosa, does this framework suggest that mobile manipulation can finally handle truly unstructured environments where perception demands change constantly?

Rosa: Yes, Taro; because they show it performs well in challenging scenarios like SetTable without any privileged information. Their experimental results are pretty compelling too: they showed success rate gains of twenty-eight point two percent, twenty-six point one percent, and twenty-three point six percent across those ManiSkillHAB scenarios, and they maintained a mean success rate of fifty-one point two five percent on the Cobot-Magic robot platform in real-world tasks like "Close Drawer."

Dev: From an engineering standpoint, I have to ask about the loop rate here; how does the complexity of calculating that intent vector and then reweighting features impact the latency we're dealing with? We need to know if this runs fast enough for reliable control.

Taro: That’s a practical concern, Dev; but given that they are using components like DINOv2 and structured attention mechanisms, the focus seems to be on efficiency in how those computations are structured rather than just brute force speed.

Rosa: And that efficiency is part of the whole point; because they also introduced the Dual-stream Affinity Refinement Module, or DARM, which explicitly models geometric consistency and semantic correspondence between three dee point clouds and 2D images.

Dev: DARM sounds like it adds another layer of complexity to the perception pipeline, but if it’s effectively modeling that cross-modal alignment using scaled dot-product attention for both geometric affinity Ageo and semantic affinity Asem, does that add significant computational overhead compared to just a standard end-to-end VLA model?

Title and authors: Taro: It adds structure; it ensures the system isn't just guessing the relationship between what the camera sees and what’s physically there. The geometric affinity being used to construct a transport cost matrix C for Sinkhorn–Knopp regularization is an interesting way to enforce spatial plausibility.

Rosa: And they connect that geometric prior back into semantic alignment by minimizing KL divergence, which is equation (five), making sure the semantic attention distribution Psem aligns with that geometrically derived prior.

Dev: So, to summarize the improvements, we have dynamic perceptual attention driven by intent inference and a decoupled coordination strategy for base and arm actions. But what about the limitations? What doesn't this framework manage effectively?

Taro: The paper notes they still treat the mobile base and manipulator as a unified action vector without fully modeling their mutual dependencies in all cases, although they introduce that Trend Token mechanism to help with bidirectional coordination.

Rosa: They also explicitly mention that removing components hurts performance; for instance, removing IDPPM causes the largest drop in performance, showing its necessity. Also, they point out that without the stop-gradient or the Trend Token consistently degrades success rates during training.

Dev: I think those are clear limitations regarding stability and necessary components for achieving those high results; it shows that this architecture is highly tuned and sensitive to how you set up the coordination mechanism.

Taro: Considering everything, what do you see as the bigger impact of InCoM on the future of mobile robotics outside of just incremental improvements in success rates?

Rosa: I think its implication is moving us toward systems that can operate robustly in complex real-world tasks where perception and control must adapt simultaneously to changing environments. It moves beyond pretraining for tabletop scenarios into true general mobile manipulation.

Dev: From an engineering view, if we can reliably get this kind of coordination running at a decent loop rate, it means we could see robots performing much more intricate tasks autonomously in dynamic settings, which is a huge leap for deployment feasibility.

Taro: For the world at large, having mobile agents that can reason about intent and adapt their sensory focus on the fly feels like it’s bringing us closer to truly versatile robotic assistants capable of handling messy, unstructured environments.

Rosa: Well said, Taro; so that's a lot of exciting stuff we've covered about InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation. Thanks to Dev and Taro for weighing in on the technical details and the autonomy side.

Dev: I agree, Rosa; it’s a solid piece of work showing how structured coordination can actually solve those intractable coupling problems we've seen in mobile robots.

Taro: Indeed; the way they structure the coordination via decoupled decoders really shows a path forward for handling these coupled systems reliably.

The paper's summary: Rosa: So, to recap, InCoM is this framework that uses inferred intent to guide how the AI perceives things and then structures how the robot's base and arm work together during manipulation tasks.

Dev: Exactly; it moves away from having a single policy that tries to do everything at once and instead separates the high-level goal from the low-level execution, which seems key for stability.

Taro: I'm really curious about how this intent inference actually helps when things go wrong in a messy real-world situation, Rosa; does it have mechanisms for reacting when the world doesn't behave as expected?

Rosa: That’s where the IDPPM comes in; because it tracks historical actions and global context, it can dynamically shift its perceptual focus from broad navigation cues to fine details needed for grasping.

Dev: I see how that dynamic weighting helps with the coupling issue we talked about; if the intent signals a fast movement, it emphasizes those deep features that help predict momentum, which keeps the control loop responsive.

Taro: And what about when things get weird? If a sensor fails or an unexpected object appears mid-task, does this intent-driven system have any built-in logic to reconfigure its coordination strategy on the fly?

Rosa: The DCFM decoder is designed for bidirectional coordination, and by having that Trend Token exchange information between the base and arm, it suggests a level of mutual compensation that might allow it to handle minor disturbances better than traditional coupled systems.

Dev: From my side as the controls engineer, the structure of the flow matching loss seems robust because it factorizes those action spaces; if one part drifts, you can still rely on the other branch's trend token to guide its motion without completely breaking the whole system.

Taro: So it’s not just about pre-planned coordination but a more fluid, intent-based coordination that adapts to immediate environmental cues while maintaining structural integrity? That sounds like something we need to think about when we look at systems that need to be safe in unstructured settings, like those drone mapping papers.

Rosa: It really is; the implication here is moving toward robots that aren't just following a script but are actively inferring what's needed for success at every moment during complex tasks.

Dev: If this works reliably outside of a controlled lab setting, Rosa, the loop rate and latency are definitely where we need to keep an eye on to make sure it’s viable for actual deployment.

Taro: I'm looking forward to seeing how they handle those failure modes in more detailed papers; if the intent inference can predict a necessary correction before the error manifests physically, that would be significant progress for autonomy.

The paper's improvements: Rosa: So, to wrap up this discussion on InCoM, we've looked at how it uses intent to steer perception and coordination, and now we're going over what they actually improved in terms of performance and functionality.

Dev: The paper highlights a few key structural improvements; primarily the decoupling of base and arm actions within the DCFM decoder, which tackles that strong coupling we discussed earlier.

Taro: I noticed they also added these auxiliary supervision terms during training, like KL divergence on the scale weights, which sounds like they're stabilizing the learning process against those conflicting perceptual demands.

Rosa: Right; those regularization terms help ensure that the system actually learns to prioritize deep features when needed for fast movement and shallow ones for precision work.

Dev: And I’m paying attention to how they use the Stop-Gradient mechanism in cross-attention between the base and arm branches; that should significantly reduce gradient interference during backpropagation, which is a big win for training stability.

Taro: Beyond training, they show that this framework maintains superior success rates even when tested on novel scenarios without any privileged information, which suggests the intent modeling is robust enough to generalize beyond the specific training data.

Rosa: That generalization capability is huge; it means we can deploy these robots in real-world settings where we don't have perfect maps or ideal initial conditions.

Dev: However, I still have my concerns about the latency; if calculating that intent vector and performing the multi-scale attention reweighting adds too much computation, it might not be suitable for high-speed control loops on embedded hardware.

Taro: That’s a valid point, Dev; but if the structure is efficient enough, maybe we can leverage parallel processing to keep that latency manageable while still getting the benefits of intent-driven adaptation.

Rosa: Exactly; the real-world test will tell us if this level of sophisticated coordination and perception is practical for long-term field work.

Dev: And I want to know more about their limitations concerning long-term stability; they mention that removing key components like IDPPM causes the biggest performance drop, which tells us how sensitive the whole architecture is to losing one of those core pieces.

Taro: So while the wins are impressive, we need to see if this structure can handle unexpected physical interactions or sensor noise in a sustained operation without degrading that coordination.

Rosa: That’s what we need to keep an eye on moving forward; it’s about seeing if this intent-driven approach translates into reliable, long-term operational capability.

Conclusion: Tom: To wrap up this discussion on InCoM, we've looked at how it uses intent to steer perception and coordination, and now we're going over what they actually improved in terms of performance and functionality.

Rosa: So, to recap, InCoM is this framework that uses inferred intent to guide how the AI perceives things and then structures how the robot's base and arm work together during manipulation tasks.

Dev: Exactly; it moves away from having a single policy that tries to do everything at once and instead separates the high-level goal from the low-level execution, which seems key for stability.

Taro: I'm really curious about how this intent inference actually helps when things go wrong in a messy real-world situation, Rosa; does it have mechanisms for reacting when the world doesn't behave as expected?

Rosa: That’s where the IDPPM comes in; because it tracks historical actions and global context, it can dynamically shift its perceptual focus from broad navigation cues to fine details needed for grasping.

Dev: I see how that dynamic weighting helps with the coupling issue we talked about; if the intent signals a fast movement, it emphasizes those deep features that help predict momentum, which keeps the control loop responsive.

Taro: And what about when things get weird? If a sensor fails or an unexpected object appears mid-task, does this intent-driven system have any built-in logic to reconfigure its coordination strategy on the fly?

Rosa: The DCFM decoder is designed for bidirectional coordination, and by having that Trend Token exchange information between the base and arm, it suggests a level of mutual compensation that might allow it to handle minor disturbances better than traditional coupled systems.

Dev: From my side as the controls engineer, the structure of the flow matching loss seems robust because it factorizes those action spaces; if one part drifts, you can still rely on the other branch's trend token to guide its motion without completely breaking the whole system.

Taro: So it’s not just about pre-planned coordination but a more fluid, intent-based coordination that adapts to immediate environmental cues while maintaining structural integrity? That sounds like something we need to think about when we look at systems that need to be safe in unstructured settings, like those drone mapping papers.

Rosa: It really is; the implication here is moving toward robots that aren't just following a script but are actively inferring what's needed for success at every moment during complex tasks.

Dev: If this works reliably outside of a controlled lab setting, Rosa, the loop rate and latency are definitely where we need to keep an eye on to make sure it’s viable for actual deployment.

Taro: I'm looking forward to seeing how they handle those failure modes in more detailed papers; if the intent inference can predict a necessary correction before the error manifests physically, that would be significant progress for autonomy.

Rosa: That’s what we need to keep an eye on moving forward; it’s about seeing if this intent-driven approach translates into reliable, long-term operational capability.

Dev: I agree; the way they structure the coordination via decoupled decoders really shows a path forward for handling these coupled systems reliably.

Taro: Indeed; the way they structure the coordination via decoupled decoders really shows a path forward for handling these coupled systems reliably, and I think that's where we should focus our attention next.

Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · Anyverse Dynamics

cs.RO

Submitted: 2026-02-26

Updated: 2026-09-28

Comments: The project website is available at https://liujiahao2077.github.io/InCoM.github.io

Project page: https://incom-anonymous.github.io/InCoM.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: InCoM is an intent-driven perception and structured coordination framework for mobile manipulation that jointly considers stage-adaptive perception and coordinated action generation, addressing two

Key concepts

Intent-Driven Perception
This involves a system inferring the robot's goal or intent before looking at the environment. This allows the system to dynamically reweight different levels of perceptual features based on what it is currently doing, shifting focus from broad navigation cues to fine details needed for specific tasks.
IDPPM
The Intent-Driven Pyramid Perception Module extracts features using three abstraction levels: shallow, mid, and deep. This multi-scale approach allows the robot to view the world through different lenses simultaneously to adapt its perception based on the task stage.
Decoupled Decoder
This strategy factors the high-dimensional action space into separate base actions and arm actions. Training separate decoders for each part, linked by a learnable Trend Token, helps prevent strong coupling between the base and arm movements during coordination.
Auxiliary Kinematic Supervision
This component ties dynamic perceptual weighting to the robot's physical state. It uses action increments to define a target distribution, ensuring deep features are emphasized during rapid motion and shallow features are prioritized for precision work.

Terminology

Summary

InCoM is an intent-driven perception and structured coordination framework for mobile manipulation that jointly considers stage-adaptive perception and coordinated action generation, addressing two key challenges: strong coupling between base and arm actions complicating control optimization, and poor allocation of perceptual attention as viewpoints shift during mobile manipulation.

The framework consists of three main components:

  1. A multi-scale perception architecture called the Intent-Driven Pyramid Perception Module (IDPPM) that infers latent motion intent to dynamically reweight multi-scale perceptual features, enabling stage-adaptive allocation of perceptual attention. This module extracts features using a sparse 3D encoder and a pretrained 2D visual encoder (DINOv2), aligning them across three abstraction levels: shallow, mid, and deep. The Intent Modulation component maps the historical action sequence and global visual feature into an implicit intent vector, which is then mapped by an MLP-based ScaleGater network to produce normalized hierarchical weights: wt = [wS, wM, wD] = Softmax(ϕ(ht)) (1). Auxiliary Kinematic Supervision is introduced to align perceptual focus with the robot’s kinematic state by computing L2 norms of action increments for the base and arm, defining a target weight distribution where w∗ D ∝ Ibase, w∗ S ∝ Iarm, w∗ M ∝ p w∗ Sw∗ D (3). Minimizing the KL divergence between the predicted weights and this target distribution (Lscale = DKL(wtw∗ t) + λent · X k∈[S,M,D] wk log wk (4)) ensures that deep features are emphasized during rapid motion and shallow features are prioritized during fine-grained manipulation.

  2. A Dual-stream Affinity Refinement Module (DARM) that enhances cross-modal alignment by explicitly modeling geometric consistency and semantic correspondence between 3D point cloud and 2D image features. It computes a geometric affinity Ageo from positional encodings and a semantic affinity Asem from feature representations, both implemented via scaled dot-product attention in distinct representation spaces. These affinities are concatenated and fused through a lightweight convolutional module frefine to produce the final cross-modal attention distribution. Furthermore, geometric affinity is used for regularization: Ageo is normalized to construct a transport cost matrix C = 1 − Softmax(Ageo), based on which the Sinkhorn–Knopp algorithm produces a soft correspondence distribution T∗. The semantic attention distribution Psem is then encouraged to align with this geometric prior by minimizing the KL divergence: Lalign = DKL(Psem ∥ sg(T∗)) (5).

  3. A Decoupled Coordinated Flow Matching (DCFM) decoder designed to explicitly capture bidirectional coordination between the mobile base and the manipulator. The high-dimensional action space is factorized into base actions a base and arm actions a arm. A direct linear path is constructed between noise distribution a0 ∼ N (0, I) and expert action distribution a1: at = (1 − t)a0 + ta1, t ∈ [0, 1] (6). Separate Transformer decoders are trained for the mobile base and the manipulator to minimize the flow matching loss: Lf low = X k∈[base,arm]∥v k θ(a k t, t, c) − (a k 1 − a k 0)∥2 (7), where c is the multimodal conditioning embedding. To ensure explicit bidirectional coordination, each decoder branch maintains a learnable Trend Token that aggregates predicted motion trends and exchanges this information via cross-attention with stop-gradient: At each Transformer layer, the mobile base and the manipulator exchange this information via cross-attention with stop-gradient to ensure bidirectional coordination.

The overall optimization objective is a weighted combination of these losses: Ltotal = Lf low + λscaleLscale + λalignLalign (8).

Experimental results demonstrate that InCoM significantly outperforms state-of-the-art methods, achieving success rate gains of 28.2%, 26.1%, and 23.6% across three ManiSkillHAB scenarios without privileged information. The framework is validated in real-world mobile manipulation tasks, where it maintains a superior success rate over existing baselines, such as ACT [7] and π0.5 [38], achieving a mean success rate of 51.25% on Cobot-Magic robot platform across four representative tasks. Ablation studies confirm the necessity of each component: removing IDPPM leads to the largest performance drop (19.5% in SetTable), replacing DARM with Q-Former reduces performance, and removing stop-gradient or the Trend Token consistently degrades success rates. Qualitative analysis shows that IDPPM dynamically adjusts perceptual focus from deep features during navigation to shallow features during grasping, and DCFM successfully enables stable base–arm coordination through bidirectional information exchange.

Improvements for AI systems

Here are the specific improvements and capabilities enabled by implementing InCoM, based on the provided scientific paper:


) Improved AI System Capabilities Enabled by InCoM:

  1. 】Robust Coordinated Whole-Body Control in Dynamic Environments:

A major limitation in existing mobile manipulation policies is the strong coupling between base and arm actions, leading to control instability and execution errors when viewpoints shift.

InCoM introduces the Decoupled Coordinated Flow Matching (DCFM) decoder, which explicitly models bidirectional coordination between the mobile base and the manipulator. This allows for:

  • Predicting smooth, expressive actions by factoring the high-dimensional action space into independent base and arm components.

  • Enabling mutual compensation between base motion (for stability/navigation) and arm movement (for manipulation), resulting in significantly more stable whole-body behaviors during complex tasks like Close Drawer or Move Block.

  1. 】Stage-Adaptive, Intent-Driven Perceptual Attention:

Existing policies suffer from perceptual dispersion because they use fixed-scale perception encoders, failing to adapt focus between global navigation and local grasping.

InCoM's Intent-Driven Pyramid Perception Module (IDPPM) infers latent motion intent from historical actions and global context to dynamically reweight multi-scale perceptual features. This enables the system to:

  • Prioritize deep, global features during base motion (e.g., navigation toward a bed).

  • Shift attention to shallow, local details during fine-grained manipulation (e.g., grasping the rim of a trash bin).

  1. 】Enhanced Cross-Modal Perception via Decoupled Alignment:

To improve the fusion of 2D image features and 3D point cloud data, InCoM utilizes the Dual-stream Affinity Refinement Module (DARM). This module separates geometric consistency modeling from semantic correspondence modeling:

  • It computes geometric affinity (using positional encodings) to generate a geometrically plausible prior for cross-modal alignment.

  • It models semantic affinity separately, allowing the system to adaptively balance spatial consistency and semantic relevance without relying on large-scale, domain-specific pretraining for every new task.

  1. 】Improved Generalization and Robustness in Real-World Deployment:

The framework is designed to be effective even under challenging real-world sensing constraints (e.g., excluding simulator privileged information like exact target poses).

InCoM achieves superior success rates (up to 28% gains) on novel scenarios without privileged information, demonstrating its ability to infer necessary cues from raw visual and proprioceptive data alone.

  1. 】Improved Training Stability through Gradient Decoupling:

The training objective incorporates auxiliary kinematic supervision and regularization terms (e.g., entropy regularization on scale weights). Furthermore, the DCFM decoder uses a Stop-Gradient mechanism in cross-attention between base and arm branches, preventing gradient interference during backpropagation.

This results in:

  • Reduced optimization difficulties caused by control coupling.

  • Smoother and more stable learning trajectories for the entire policy.

) Summary of System Improvement:

The improved AI system, InCoM, is a unified end-to-end framework that moves beyond simple action prediction by integrating three critical components: intent inference (IDPPM), decoupled coordination (DCFM), and structured cross-modal alignment (DARM). This allows the robot to exhibit highly coordinated, stable whole-body behaviors in complex, dynamic real-world mobile manipulation tasks that require simultaneous navigation and fine motor skills.

Sources

Related papers