CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning".
Dev: The gist The proposed framework synthesizes diverse, stable grasps strictly conditioned on specific contact topologies by projecting local object features into a feature-based canonical workspace,
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So, this paper is called CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning. Basically, they're trying to create a way to generate diverse and stable grasps that are strictly conditioned on specific contact topologies.
Dev: Right, the main idea is to get around the problem of needing tons of object-annotated data when you want different kinds of functional grasps. The authors claim this framework synthesizes these grasps without needing that expensive training data.
Taro: So, what's the core mechanism for decoupling that functional intent from the actual shape of the object, if they're not using object geometry directly?
Rosa: Well, they introduce a feature-based canonical workspace anchored to the gripper frame. This workspace acts like a bridge so you can project local object features into one unified gripper-centric domain.
Dev: And how do they get those features into that space? They extract local geometric features using a modified DGCNN encoder from either the training gripper or the inference object, and then aggregate those points into fixed locations in the workspace using k-Nearest Neighbors.
Taro: So, it's not looking at every single point on the object geometry directly for learning contact maps, but rather summarizing it into these fixed spatial tokens within that workspace?
Rosa: Exactly. This unified spatial representation is what effectively decouples the functional intent from the specific object identity. They learn a latent manifold within this workspace that models what the gripper's intrinsic contact capabilities are, allowing for zero-shot generalization to different target geometries.
Dev: But they also condition this entire process on structured contact topologies derived from human grasp taxonomies, specifically adapting the Gonzalez taxonomy based strictly on the hand's active contact surfaces.
Taro: So, so the conditioning isn't just about learning a map; it’s about imposing a specific required contact pattern onto that learned gripper capability space?
Rosa: Precisely. They define a semantic mask where each template acts as a semantic mask, assigning a Zone ID to points needed for that specific contact topology. Then they use a Transformer encoder to model those non-local dependencies between potential contact regions, explicitly concatenating their learnable topology embedding onto every workspace point feature.
Paper summary: Dev: That sounds like they're forcing the AI to pay attention to where the required contacts need to be spatially located before it even tries to optimize the actual physical grasp.
Taro: And what about testing if that synthesized contact pattern is actually good? How do they ensure physical viability after the synthesis step?
Rosa: They have a strict cascading validation pipeline in their inference phase. First, a Label-Consistency Check which prevents generating ill-posed grasps by checking for a maximum deviation of one missing or hallucinated contact zone against the ground truth.
Dev: Then they do a Contact Points Force-Closure Validation to assess stability by computing the grasp wrench space based on the active workspace points, and they discard anything that doesn't meet that force-closure condition.
Taro: That’s important because it shows they aren't just guessing contacts; they are checking if those predicted contacts actually hold the object stably under physics constraints.
Rosa: And finally, there’s a Joint Configuration Optimization where they minimize an energy function that aligns the gripper’s active surfaces with the predicted spatial workspace points while respecting kinematic constraints and adding a repulsive term to enforce a minimum safety margin of five millimeters.
Dev: It sounds like they’ve built this whole loop from feature extraction, through topological conditioning, into a validation check, and finally into an energy minimization for the actual joint movement. That’s a lot of moving parts for the loop rate.
Taro: So, it solves the problem of generating diverse grasps without massive datasets by learning gripper capabilities and imposing contact requirements in a canonical space?
Rosa: It does that, and the experimental results show they achieve state-of-the-art performance on DexGraspNet compared to existing taxonomy-guided planners. They also managed to keep fifty-eight point seven two percent of their physical stability score when dealing with non-convex objects, compared to thirty-seven point four two percent for Dexonomy.
Dev: That gap in stability is significant when you're talking about real-world application on things that aren't perfectly smooth or simple shapes. They also managed to cover an average of eighty point one four percent of the objects per requested contact topology across all their generation attempts.
Taro: So, from my side as someone interested in autonomy, this means the system is robust enough to handle unseen geometries while still sticking to the functional requirements defined by a human grasp taxonomy.
Paper summary: Rosa: It's definitely showing that structural properties of these generated contact topologies are physically executable on real robot platforms when they test them with the Allegro Hand.
Dev: The authors did mention a limitation in their discussion, and it’s about Topology Compliance scores being modest because of how strict their metric is. They said this can lead to incidental contacts where intended precision pinches get reclassified due to kinematic constraints.
Taro: So, the system is great at diversity but maybe not perfectly precise when you push those kinematic limits?
Rosa: That’s right. But they argue that it still maintains a substantially more balanced and faithful distribution of functional grasps compared to the Dexonomy baseline regarding topological distribution.
Dev: It really highlights that the execution of these synthesized precision grasps on physical hardware is where the fundamental mismatch between deterministic kinematic planning and stochastic physics comes into play, suggesting future work needs to look at contact-aware interaction paradigms.
Taro: So for someone listening who just wants to know what this means for real use, it means we can generate a huge variety of functional grasp ideas that are physically possible on a robot without needing specific training examples for every single object type.
Rosa: That's the big picture here with CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning. The authors essentially decoupled the idea of *what* you want to grasp functionally from *what* the object actually looks like geometrically.
Dev: They did this by building that gripper-oriented framework and using kNN aggregation to build that unified spatial representation in the canonical workspace, which is what lets it generalize across different objects.
Taro: It’s about learning the intrinsic capabilities of your hand within a fixed space and then using those learned capabilities to satisfy a semantic requirement for contact topology.
Rosa: So, it’s a generative framework for contact-topology-conditioned dexterous grasp synthesis that uses object-agnostic training to bypass the data collection bottleneck.
Dev: The authors argue that this approach provides superior semantic diversity and physical stability compared to prior methods on unseen objects without needing object-annotated datasets.
Taro: It’s a way to get high-quality, diverse grasps from scratch relying purely on raw object geometry and a discrete semantic label.
Conclusion: Rosa: So, we’ve been looking at CoToGrasp, this paper by Rosa and Dev about this new grasp synthesis framework.
Dev: Yeah, it’s all about using a canonical workspace to separate what you *want* to grasp functionally from the actual shape of the object.
Taro: Basically, they’re decoupling semantic intent—like "I need a power grip"—from arbitrary object geometry by learning gripper-centric contact manifolds.
Rosa: Right, and they condition this whole generation process on structured contact topologies based on human grasp taxonomies, like the Gonzalez taxonomy.
Dev: That means the AI isn't just guessing random contacts; it’s being told exactly which contact zones are needed based on a learned pattern of how hands actually grab things.
Taro: So what does this mean for autonomy when the world throws a weird object at you? If you can synthesize grasps conditioned on those topologies, that’s a step toward handling unseen geometries much more reliably.
Rosa: Exactly. They show it works well on large datasets like DexGraspNet and even shows good robustness on non-convex objects.
Dev: The results are promising, but the authors did flag something important about the Topology Compliance scores being modest because of how strict their measurement is.
Taro: So, if you're using this for a real robot, you have to be careful because those strict metrics can sometimes force contacts that aren't actually what you intended just because of kinematic limits.
Rosa: That’s the caveat; the system is very good at diversity but it might over-enforce precision in ways that aren't always optimal physically.
Dev: It really points to the gap between deterministic planning and the physics of actual contact, which is a big thing for real-world deployment.
Taro: So, it’s not a perfect solution yet; we still need to figure out how to make those synthesized precision grasps behave more naturally under physics.
Université Paris-Saclay · Ecole Centrale Lyon, CNRS, LIRIS, Institut Universitaire de France
cs.RO, cs.AI
Submitted: 2026-08-20
Updated: 2026-10-08
Comments: Project website at https://cea-list.github.io/cotograspweb/
Journal ref: 19th European Conference on Computer Vision (ECCV), 2026
DOI: 10.1007/978-3-032-37574-2_7
Project page: https://cea-list.github.io/cotograspweb
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: The gist The proposed framework synthesizes diverse, stable grasps strictly conditioned on specific contact topologies by projecting local object features into a feature-based canonical workspace,
Key concepts
- Feature-Based Canonical Workspace
- This is a standardized, fixed spatial area anchored to the gripper frame. The model learns features from both the gripper and objects and projects them into this common space using kNN aggregation. This acts as a universal bridge, allowing the system to understand local geometry regardless of what object it is grasping.
- Contact Topology Conditioning
- Instead of depending on an object's shape, CoToGrasp uses human grasp taxonomies (like Gonzalez) to define required contact patterns. A semantic mask assigns specific Zone IDs to points needed for a target topology, which the network uses as a hard constraint during synthesis.
- Cascading Validation Pipeline
- Before finalizing a grasp, the framework enforces strict checks. This includes a Label-Consistency Check to ensure predicted contacts match the required topology closely and Force-Closure Validation to confirm physical stability by calculating the grasp wrench space. This prevents generating physically impossible or unstable grasps.
Terminology
Summary
The gist The proposed framework synthesizes diverse, stable grasps strictly conditioned on specific contact topologies by projecting local object features into a feature-based canonical workspace, effectively decoupling semantic functional intent from arbitrary object geometry
CoToGrasp Framework Overview
CoToGrasp is a novel generative framework designed to synthesize diverse, stable grasps strictly conditioned on specific contact topologies. It operates in two distinct phases: Object-Agnostic Training and Grasp Synthesis. During the training phase, the model learns an intrinsic, gripper-centric contact manifold within a canonical feature-based workspace independent of object geometry. The network takes as input only the gripper’s local surface geometry and a semantic contact topology, learning to reconstruct the corresponding physical contact template mask.
Feature-Based Canonical Workspace Learning
The framework introduces a feature-based canonical workspace anchored to the gripper frame, which acts as a domain-agnostic bridge. This workspace is constructed by extracting local geometric features via a modified DGCNN encoder from either the gripper (training) or the object (inference), and then aggregating these features into fixed points within the canonical workspace using k-Nearest Neighbors (kNN). To capture the complex spatial structure of functional grasps, each populated workspace point is treated as a discrete token and processed using a Transformer encoder.
Contact Topology Conditioning
The framework conditions grasp synthesis on structured contact topologies derived from human grasp taxonomies, specifically adapting the Gonzalez taxonomy [12] which categorizes grasps based strictly on the hand’s active contact surfaces rather than the object’s shape. This is achieved by defining a semantic mask where each template Am acts as a semantic mask, assigning a Zone ID to points required for contact topology m. The network models non-local dependencies between potential contact regions using a Transformer encoder, explicitly concatenating the learnable topology embedding FT to every workspace point feature FW.
Grasp Synthesis and Validation Pipeline
In the inference phase, CoToGrasp predicts a contact template Λˆm over the workspace by sampling from a prior N (0, I). To ensure physical viability before costly optimization, a strict cascading validation pipeline is enforced. This includes a Label-Consistency Check to prevent generating ill-posed grasps by applying a binary check on the symmetric difference between ground truth and predicted templates, allowing a maximum deviation of one missing or hallucinated contact zone. A Contact Points Force-Closure Validation assesses stability by computing the grasp wrench space based on active workspace points, discarding predictions if the force-closure condition is not met. Finally, a Joint Configuration Optimization minimizes an energy function that aligns the gripper’s active surfaces with predicted spatial workspace points while respecting kinematic constraints and introducing a repulsive term Erep to enforce a minimum safety margin of 5 mm.
Experimental Evaluation and Results
Extensive evaluations on the large-scale DexGraspNet dataset demonstrate that CoToGrasp achieves state-of-the-art performance, outperforming existing taxonomy-guided planners. The framework achieves the highest semantic entropy (HT C) and generation speed among taxonomically unconditioned baselines. Furthermore, on challenging non-convex objects, CoToGrasp demonstrates remarkable robustness, retaining 58.72% of its physical SR compared to Dexonomy’s 37.42%. The framework successfully covers an average of 80.14% of the objects per requested contact topology across all generation attempts. Successful real-world deployments on the Allegro Hand confirm that the structural properties of our generated contact topologies are physically executable on a real robot platform.
Discussion and Limitations
The analysis shows that while CoToGrasp achieves superior semantic diversity, its Topology Compliance (TC) scores can be modest, reflecting the extreme stringency of the asymmetric metric. The rigid metric enforces strict mathematical boundaries that can lead to incidental contacts where intended precision pinches are reclassified due to kinematic constraints. However, CoToGrasp maintains a substantially more balanced and faithful distribution of functional grasps compared to the Dexonomy baseline regarding topological distribution. The execution of synthesized precision grasps on physical hardware exposes the fundamental mismatch between deterministic kinematic planning and the stochastic physics of mechanical contact, suggesting future work must transition toward contact-aware interaction paradigms.
The paper is a novel generative framework for contacttopology-conditioned dexterous grasp synthesis that fundamentally decouples functional intent from object identity. It addresses the limitations of current planners by learning intrinsic gripper capabilities within a canonical workspace and conditioning synthesis on human grasp taxonomies, leading to state-of-the-art performance on unseen objects without requiring costly object-annotated datasets.
How it works
-
Object-Agnostic Training: The model is trained exclusively on gripper point clouds and operates entirely within the canonical gripper frame, rendering it independent of the global grasp pose.
-
Geometric Transfer: Local geometric features are extracted via a modified DGCNN encoder and projected onto a fixed set of spatial basis points W using weighted kNN aggregation to construct consistent spatial tokens.
-
Semantic Conditioning: The workspace embeddings are processed by a Transformer encoder, where the learnable topology embedding FT is concatenated to every workspace point feature FW to enforce persistent, hard conditioning across the entire spatial domain.
-
Synthesis and Optimization: The framework synthesizes grasps by predicting a contact template Λˆm, followed by a cascading validation pipeline (Label-Consistency Check and Force-Closure Validation) before retrieving the optimal joint configuration Q∗ via an energy-based optimization that balances distance, penetration, self-collision, joint limits, and a repulsive energy term.
Evaluation Metrics
The paper introduces novel metrics to quantify functional accuracy and distributional fairness: Topology Compliance (TC), which evaluates how accurately the generated grasp adheres to the functional constraints of the requested contact topology, Entropy (H), which assesses topological diversity, and Semantic Entropy (HT C), which evaluates if the model preserves true functional intent without defaulting to simpler power grasps.
Key Findings
CoToGrasp achieves superior performance over taxonomy-unaware baselines by synthesizing diverse topologies regardless of object affordances. It outperforms state-of-the-art taxonomy-aware methods like Dexonomy [5] in both physical stability (SR) and semantic accuracy (TC), particularly on highly constrained precision grasps. The framework's validation pipeline is critical, as the ablation of the Label-Consistency check demonstrates a sharp performance drop, emphasizing that verifying topological viability against the object’s local geometry is essential before optimization.
Conclusion
CoToGrasp introduces a novel object-agnostic grasp planner that synthesizes grasps conditioned on structured contact topologies derived from human taxonomies, successfully mitigating severe mode collapse and achieving state-of-the-art semantic diversity and physical stability. Successful real-world deployments on the Allegro Hand confirm that the structural properties of our generated contact topologies are physically executable on a real robot platform. The framework provides a robust method for generating functionally compliant grasps from scratch relying purely on raw object geometry and a discrete semantic label.
The paper is CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning. It is published as arXiv:2608.19776v2 [cs.RO] 21 Aug 2026. This paper is a novel generative framework for contacttopology-conditioned dexterous grasp synthesis that fundamentally decouples functional intent from object identity, allowing the synthesis of highly constrained, topology-compliant precision and power grasps on unseen geometries without relying on costly object-annotated datasets. It is published as arXiv:2608.19776v2 [cs.RO] 21 Aug 2026. This paper is a novel generative framework for contacttopology-conditioned dexterous grasp synthesis that fundamentally decouples functional intent from object identity, allowing the synthesis of highly constrained, topology-compliant precision and power grasps on unseen geometries without relying on costly object-annotated datasets. It is published as arXiv:2608.19776v2 [cs.RO] 21 Aug 2026. This paper is a novel generative framework for contacttopology-conditioned dexterous grasp synthesis that fundamentally decouples functional intent from object identity, allowing the synthesis of highly constrained, topology-compliant precision and power grasps on unseen geometries without relying on costly object-annotated datasets. It is published as arXiv:2608.19776v2 [cs.RO] 21 Aug 2026. This paper is a novel generative framework for contacttopology-conditioned dexterous grasp synthesis that fundamentally decouples functional intent from object identity, allowing the synthesis of highly constrained, topology-compliant precision and power grasps on unseen geometries without relying on costly object-annotated datasets. It is published as arXiv:2608.19776v2 [cs.RO] 21 Aug 2026.
Improvements for AI systems
- Bold header: Object-Agnostic Training Paradigm for Generalization
The improved system can synthesize functionally diverse and physically stable grasps
conditioned on specific contact topologies without needing object-annotated datasets.
This directly addresses the limitation where existing planners are constrained by explicitly mapping grippers to specific objects biases these models, hindering shape generalization.
- Bold header: Feature-Based Canonical Workspace Learning
The system can project local object features into a unified gripper-centric domain,
which effectively decouples the semantic functional intent from the arbitrary object geometry.
This allows for zero-shot generalization to unseen objects at inference
by learning a latent manifold that models the intrinsic contact capabilities of the gripper within this workspace.
- Bold header: Taxonomy-Conditioned Grasp Synthesis
The improved system can generate grasps strictly conditioned on human grasp taxonomies, moving beyond symmetric pinch and enveloping power grasps. This is achieved by conditioning the generative pipeline on structured contact topologies derived from human grasp taxonomies
like Gonzalez [12], allowing it to synthesize precision grasps required for a downstream task.
- Bold header: Robust Validation Pipeline
The system can rigorously filter out topologically invalid configurations through a sequential validation strategy involving a Label-Consistency Check
and Contact Points Force-Closure Validation.
This prevents generating ill-posed grasps
by applying a check where the tolerance threshold is set to allow for a maximum deviation of one missing or hallucinated contact zone.
- Bold header: Semantic Compliance Metric (HT C)
The system can be evaluated on semantic accuracy using the novel metric Semantic Entropy (HT C),
which measures if it preserves true functional intent without defaulting to simpler power grasps.
This metric is critical for assessing performance against unconditioned baselines, where CoToGrasp achieved an entropy of 0.83
compared to lower values in unconditioned planners.
- Bold header: Kinematic Feasibility and Real-World Deployment
The improved system can synthesize contact topologies that are not only stable in simulation but also physically executable on a physical robot platform,
as demonstrated by successful execution on the Allegro Hand using a UR10 manipulator. This confirms the physical viability and kinematic feasibility
of the synthesized contact topologies.
Abstract
Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it should be grasped to support downstream functional tasks. However, conditioning grasp synthesis on specific human grasp taxonomies typically requires prohibitively expensive, object-annotated datasets. To address these limitations, we propose CoToGrasp, a novel generative framework that synthesizes diverse, stable grasps strictly conditioned on specific contact topologies. To bypass the data collection bottleneck, CoToGrasp is trained entirely in an object-agnostic manner. We introduce a feature-based canonical workspace that projects local object features into a unified gripper-centric domain, effectively decoupling the semantic functional intent from the arbitrary object geometry. By learning the intrinsic contact manifold of the gripper within this workspace, our model achieves zero-shot generalization to unseen objects at inference. Extensive evaluations on the large-scale DexGraspNet dataset demonstrate that CoToGrasp achieves state-of-the-art performance, outperforming existing taxonomy-guided planners. Finally, we demonstrate the physical viability and kinematic feasibility of our synthesized contact topologies on a physical robot platform. Code is available on our project website at https://cea-list.github.io/cotograspweb/.
Sources
- AnyDexGrasp: General Dexterous Grasping for Different Hands with Human-level Learning Efficiency
- Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation
- Deep Differentiable Grasp Planner for High-DOF Grippers
- OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance
- DeepViT: Towards Deeper Vision Transformer
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving