ContactExplorer: Contact Coverage-Guided Exploration for General-Purpose Dexterous Manipulation

arXiv:2603.10971 · cs.RO, cs.AI · Submitted 2026-03-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "ContactExplorer: Contact Coverage-Guided Exploration for General-Purpose Dexterous Manipulation".

Rosa: ContactExplorer is a contact-centric exploration framework for general-purpose dexterous manipulation that explicitly models and incentivizes hand–object interaction on novel contact patterns, namely which fingers contact which object regions.

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So we're looking at the paper "ContactExplorer: Contact Coverage-Guided Exploration for General-Purpose Dexterous Manipulation," and the title itself really sets the stage by focusing on using contact coverage to guide exploration in dexterous manipulation tasks. It suggests they're moving away from just exploring arbitrary states and focusing specifically on how fingers interact with different parts of an object.

Dev: I agree, Rosa, that focus on contact coverage is key because traditional methods often struggle with the sheer variety of possible ways a hand can grasp or touch something. It sounds like they are trying to create a systematic way for the AI to discover novel physical interaction patterns instead of just random movements.

Taro: From an autonomy perspective, that makes sense; if the system understands *where* it's touching and what it's touching, it gains a much richer understanding of the object's surface and its own capabilities in interacting with that surface. It moves exploration from vague state novelty to concrete physical interaction data one.

Rosa: Exactly, Taro, and I wonder how this structured approach will translate when we take these models out of the lab and into a messy, real-world environment where the object shapes are constantly changing.

Dev: That's my main concern for the engineering side; if we rely too heavily on contact coverage metrics conditioned on specific learned object states, what happens when those states drift or become inaccurate in practice?

Taro: The paper addresses that by conditioning the contact counters on discretized object states obtained through learned hash codes, which should help manage that state representation issue one.

Rosa: That sounds like a clever way to handle the complexity of object configurations, and I'm curious about how robust those learned states are when we move to different types of objects.

The paper's summary: Dev: What the ContactExplorer paper really boils down to is that it tackles the difficulty of exploration in manipulation by combining two distinct signals: a count-based reward that pushes the agent toward novel contact patterns, and an energy-based reaching reward that pulls it towards areas of the object it hasn't explored much.

Rosa: That combination sounds very strategic; one signal is about finding new things, and the other is about making sure we don't get stuck in familiar spots. They condition these counters on both the current and goal object states, which I think adds a layer of context that should be helpful for planning.

Taro: Conditioning on both current and goal states means the exploration isn't just about finding *any* new contact; it’s about finding novel contacts relevant to reaching the final configuration one.

Dev: And the reward mechanisms themselves are quite specific, using a count-based contact novelty score based on a weighting function g(c) = one/√c + one and an energy term where they measure finger keypoint positions against object surface points.

Rosa: That weighting function sounds important because it makes the exploration density inversely related to how often something has been seen, which should give us a good sense of novelty when combined with the goal-directed energy reaching reward.

Taro: The authors point out that prior work sometimes uses hand-object distance as a proxy for novelty, but this paper argues that measuring which fingers touch which object regions is a much more reliable and interaction-centric exploration signal one.

The paper's improvements: Rosa: Looking at the specific improvements they detail, one major point is how they use progress-based shaping for both rewards, ensuring the agent only gets rewarded for actual increases in novelty rather than just repeated contacts.

Dev: That prevents reward saturation, which is a common problem in exploration setups; if we keep rewarding the same thing repeatedly, the signal becomes weak quickly. The paper uses Rcontact(t) = α

Scontact(t) − S max contact: + for that purpose one.

Taro: I think that progress-based shaping is crucial because it prevents the agent from getting stuck in local optima where it keeps finding slightly better versions of a known contact pattern. It forces genuine discovery of new patterns.

Rosa: And on the reaching side, they apply a similar episodic progress-based shaping to the energy-based reaching reward, Renergy(t) = β

Senergy(t) − S max energy: +, which keeps guiding it toward under-explored regions effectively.

Dev: The paper claims significant improvements in sample efficiency and success rates across various tasks, specifically noting that ContactExplorer achieves the highest average success rate with the lowest variance and is the only method to solve Constrained Object Retrieval one.

Taro: That claim about solving Constrained Object Retrieval is quite strong because that task demands a very specific, constrained interaction pattern that this method seems adept at discovering one.

Rosa: It sounds like they've managed to create a principled exploration mechanism that balances the need to find new things with the need to make progress toward a specific goal configuration.

Conclusion: Dev: So, wrapping up, the core idea of ContactExplorer is using contact coverage guided exploration rewards alongside energy-based reaching rewards, both shaped by progress metrics, to effectively discover diverse and meaningful hand-object interaction patterns across various object states one.

Rosa: It seems they've established a framework that is quite effective at balancing finding new contact types with directing the agent toward less explored areas of interaction space. I think this approach has serious implications for general-purpose manipulation systems.

Taro: If this method proves robust across the diverse set of tasks mentioned, it means we can build systems that don't just learn one specific way to grasp an object but can adapt their entire interaction strategy based on the context and the goal.

Dev: From an engineering standpoint, it suggests a path toward much faster training, with they reporting reaching seventy percent success with two to three times fewer steps than intrinsic-reward baselines on some challenging tasks one.

Rosa: And that efficiency gain is huge for physical robots; if we can train these systems in significantly less time and steps, it makes deployment a lot more feasible for real-world applications.

Taro: I just hope that this ability to discover diverse interaction strategies actually translates into reliable behavior when the system encounters unexpected physical interactions or novel environments outside of the training set one.

Zixuan Liu, Ruoyi Qiao

School of Computing, National University of Singapore

cs.RO, cs.AI

Submitted: 2026-03-11

Updated: 2026-09-29

Comments: 12 pages

Project page: https://contact-explorer-anonymous.github.io

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 88/100

The gist: ContactExplorer is a contact-centric exploration framework for general-purpose dexterous manipulation that explicitly models and incentivizes hand–object interaction on novel contact patterns,

Key concepts

ContactExplorer
A contact-centric exploration framework designed for general-purpose dexterous manipulation. It explicitly models and incentivizes novel contact patterns by focusing on which fingers touch which object regions.
Contact Coverage Guided Exploration
A method where the agent is rewarded based on the coverage of contacts it has made. This guides the system to discover new ways fingers can interact with an object, moving beyond random movements.
Progress-Based Shaping
A technique used to shape rewards by rewarding agents only for actual increases in novelty or progress. This prevents reward saturation and keeps the agent from getting stuck in local optima where it repeats known contact patterns.

Terminology

Summary

ContactExplorer is a contact-centric exploration framework for general-purpose dexterous manipulation that explicitly models and incentivizes hand–object interaction on novel contact patterns, namely which fingers contact which object regions. It abstracts objects into surface regions and tracks contact coverage between fingers and object regions. To address the sparsity of contact, ContactExplorer combines two complementary signals: a post-contact count-based reward that promotes exploration of novel contact patterns and an energy-based reaching reward that guides the agent toward under-explored contact regions.

The framework maintains a contact counter conditioned on discretized object states obtained via learned hash codes, capturing how frequently each finger interacts with different object regions. This counter is leveraged in two complementary ways: (1) to assign a count-based contact coverage reward that promotes exploration of novel contact patterns, and (2) an energy-based reaching reward that guides the agent toward under-explored contact regions. To make the exploration of contact pattern depend on the task phase and object configuration, ContactExplorer conditions its contact counters on both the current and goal object states.

The design consists of three main components: (1) a learned state hashing module that clusters continuous object states (Sec. 3.3); (2) a contact coverage counter that tracks state-conditioned finger-region interactions (Sec. 3.4); and (3) a structured exploration reward mechanism, decomposed into contact coverage (Sec. 4.1) and energy-based reaching terms (Sec. 4.2).

The Contact Coverage Reward is provided only upon physical contact: "At timestep t, for each finger f that contacts the object (I contact t(f) = 1), we map the contacted point to its corresponding surface region k under the current object-state cluster s, and compute a contact novelty score: Scontact(t) = 1/F X F f=1 I contact t(f) · g(Cs,f,kf), where g(c) = 1/√c + 1 is a monotonically decreasing count-based weighting function. To avoid repeatedly exploiting previously discovered interaction trajectories, we further adopt a progress-based shaping scheme that rewards only improvements over the best previously achieved score within the current episode: Rcontact(t) = α [Scontact(t) − S max contact]+, where S max contact denotes the maximum contact novelty score previously achieved in the current episode, α is a scaling coefficient, and [x]+ = max(x, 0) denotes the positive-part operator."

The Energy-Based Reaching Reward guides pre-contact exploration: "For each finger f, we define a contact energy: Φf = X m g(Cs,f,ξ(m)) exp −∥plf − pm∥2/δ, where δ controls the spatial decay, plf denotes the finger keypoint position, and pm denotes a point on the object surface. The energy-based exploration score is then defined as: Senergy(t) = 1/F X F f=1 Φf. Similar to the contact reward, we apply episodic progress-based shaping: Renergy(t) = β [Senergy(t) − S max energy]+, where S max energy denotes the maximum energy score achieved within the current episode and β is a scaling coefficient."

The method is evaluated on a diverse set of dexterous manipulation tasks, including Cluttered Object Singulation, Constrained Object Retrieval, In-Hand Reorientation, Bimanual Object Opening, Bimanual Board Lifting, Grasping, and Tiled Object Retrieval. Experimental results show that ContactExplorer substantially improves sample efficiency and success rates over existing exploration methods. Specifically: ContactExplorer achieves the highest average success rate with the lowest variance and is the only method to solve Constrained Object Retrieval. It also improves sample efficiency, reaching 70% success with 2×–3× fewer steps than intrinsic-reward baselines and exceeding 80% success within 3M–9M steps on challenging tasks, while baselines plateau lower or require more interactions. Furthermore, the contact patterns learned with ContactExplorer transfer robustly to the real world.

The paper makes two key contributions: (1) introducing ContactExplorer, a contact coverage-guided exploration reward that explicitly models and encourages diverse hand–object contact patterns across task regions; and (2) demonstrating that ContactExplorer significantly improves training efficiency and final success rates across a wide range of dexterous manipulation tasks, serving as a principled reward exploration for general-purpose dexterous manipulation. The method is also robust to changes in hand keypoint selection, maintaining stable performance under both low- and high-level perturbations. In cross-embodiment experiments using the Allegro Hand, ContactExplorer consistently improves performance over all baselines on both Cluttered Object Singulation and Constrained Object Retrieval when transferring from the LEAP Hand.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be implemented in AI systems, along with what those improved systems will be able to do:


  1. Improve Sample Efficiency and Training Speed for Dexterous Manipulation:

  2. Improve Robustness to Real-World Contact Dynamics (Sim-to-Real Transfer):

  3. Enable Generalization Across Diverse Object Configurations and Hand Types (Cross-Embodiment Robustness):

  4. Achieve Autonomous Discovery of Novel Interaction Strategies Without Task Priors:

  5. The improved AI system will be able to perform complex, general-purpose dexterous manipulation tasks (e.g., Cluttered Object Singulation, Constrained Object Retrieval, In-Hand Reorientation) with significantly higher success rates and much faster convergence compared to existing methods.

  6. This system will achieve superior learning efficiency by requiring substantially fewer training steps (e.g., reaching 70% success within 3M–9M steps on challenging tasks), effectively reducing the massive computational cost associated with RL exploration in physical systems.

  7. The system will be able to learn and transfer contact patterns robustly from simulation to the real world, ensuring that interaction strategies discovered in a digital environment are applicable to physical robots like the uFactory xArm with LEAP Hand.

  8. The system will be capable of performing complex manipulation tasks (like Bimanual Board Lifting or Object Opening) where coordinated, constrained contact is necessary, as it explicitly models and rewards these crucial interaction patterns.

  9. The system can autonomously discover novel ways to interact with objects simply by seeking diverse and new contact patterns (finger-region contacts), rather than relying on manually designed heuristics or task-specific priors that limit generalization.

  10. The system will exhibit robustness to variations in object state and configuration by maintaining separate contact counters conditioned on learned object states, allowing it to effectively reuse meaningful interactions across different spatiotemporal contexts of the same object (e.g., retrieving a cube from a top-opening box from different initial positions).

  11. The system can maintain stable performance even when the underlying hand keypoints are slightly perturbed or shift away from their predefined palmar face location, indicating robustness to moderate spatial variations in keypoint definition during deployment.

Abstract

Reinforcement learning explores effectively in domains such as Atari games, navigation, and locomotion, where novelty over states or dynamics is a sufficient signal. In contrast, dexterous manipulation requires rich physical hand--object interactions, but existing methods often suffer from unstable contact-based novelty signals, inefficient distance novelty signals, or reliance on task-specific priors. We propose ContactExplorer, a general exploration method for dexterous manipulation tasks. ContactExplorer represents contact as the intersection between object surface points and hand keypoints, encouraging dexterous hands to discover diverse and novel contact patterns, namely which fingers contact which object regions. It maintains a contact counter conditioned on discretized object states obtained via learned hash codes. This counter is leveraged in two complementary ways: (1) a count-based contact coverage reward that promotes exploration of novel contact patterns, and (2) an energy-based reaching reward that guides the agent toward under-explored contact regions. We evaluate ContactExplorer on seven contact-rich manipulation tasks and five dexterous hand embodiments. Experimental results show that ContactExplorer substantially improves sample efficiency and success rates over existing exploration methods, that it reduces the need for task-specific priors, and that it remains effective across hand embodiments and transfers to the real world. Project page is https://contact-explorer.github.io.

Sources

Related papers