FingerEye: Learning Dexterous Manipulation with Continuous Vision-Tactile Sensing

arXiv:2604.20689 · cs.RO · Submitted 2026-04-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "FingerEye: Learning Dexterous Manipulation with Continuous Vision-Tactile Sensing".

Dev: Dexterous robotic manipulation requires perception that remains informative from pre-contact approach to contact initiation and post-contact control, which is addressed by introducing FingerEye,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: Basically, this paper is proposing FingerEye as a way to strengthen robotic dexterity by providing continuous vision-tactile feedback throughout an entire interaction. Dev The core claim is that this approach improves performance across seven different contact-rich tasks by integrating both pre-contact vision and post-contact wrench sensing. Taro What matters here is how it handles the shift from just looking at something to actually interacting with it and then stabilizing that interaction afterward. Rosa The system uses binocular RGB cameras for close-range visual cues before contact, and then a marker-tracked deformation of a compliant ring after contact to sense forces. Dev It also addresses limitations in existing systems, specifically mentioning that relying solely on vision or tactile sensing can be unreliable when there are changes in lighting, occlusions, or subtle motion induced by contact eighteen nineteen <ref:2604.20689#pg1>.

Taro: That’s crucial because if the visual cues get noisy or lost right at the moment of contact initiation, the robot's plan falls apart quickly. Rosa They also designed a specific learning interface using group-structured modality fusion to avoid what they call "modality shortcuts," which happens when policies rely too much on easier global cues instead of local fingertip feedback. Dev That sounds like they are trying to make sure the policy actually uses the rich, continuous feedback FingerEye provides and doesn't just ignore it for simplicity.

Rosa: It seems like the main thrust is combining these sensing capabilities with a smarter way of learning so the robot can handle complex manipulation tasks more robustly across different object properties. Dev The authors built a real-and-sim infrastructure specifically to collect data and evaluate this integrated approach systematically, which is important for proving its reliability. Taro So the paper is really about creating a system where perception isn't just an input, but an active part of the manipulation loop itself during every phase.

Conclusion: Rosa: Thinking about the title, "FingerEye: Learning Dexterous Manipulation with Continuous Vision-Tactile Sensing," it really captures the essence of what they did—it’s about making perception a constant stream of feedback during manipulation rather than just snapshot data points. Dev The authors, Xu et al., have put forward a framework that specifically addresses the gap where dexterity requires information from pre-contact approach all the way through post-contact control. Taro I wonder what this means for future autonomous systems when they are dealing with highly delicate or unpredictable objects in unstructured environments. Rosa If this works well outside of a lab, it suggests that robots could become much more capable of performing intricate tasks where they have to constantly adjust based on real-time tactile and visual information during the entire process. Dev From an engineering perspective, the implication is that if we can manage the loop rate and latency effectively with this continuous feedback, we might see a significant improvement in success rates for complex manipulation tasks compared to systems relying on less integrated sensing.

Taro: The impact could be seen in applications like fine assembly or delicate handling where traditional methods struggle because they lack that continuous, multi-modal understanding of the interaction. Rosa It really points toward a future where robotic dexterity is built not just on fast movements, but on having a much richer, more persistent understanding of what’s happening at every single point of contact. Dev We need to keep checking the latency figures; if this continuous sensing introduces significant delay, the benefits might be lost in high-speed maneuvers. Taro I agree that sustained performance across diverse tasks is what really matters for real-world autonomy, not just passing a single benchmark in simulation.

Rosa: So, in simple terms, FingerEye proposes a way to give robots better eyes and better sense of touch simultaneously throughout the whole process of picking up or manipulating something. Dev The authors showed that this continuous feedback significantly helps the robot perform tasks that require precision and adjustment after contact. Taro The real-world implication is that we might see robots handle much messier, more varied objects in less controlled settings if they can maintain this level of perception during the interaction.

National University of Singapore

cs.RO

Submitted: 2026-04-22

Updated: 2026-10-06

Comments: Project website: https://nus-lins-lab.github.io/FingerEyeWeb/

Journal ref: Conference on Robot Learning (CoRL) 2026, Spotlight Presentation

Project page: https://nus-lins-lab.github.io/FingerEyeWeb

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 89/100

The gist: Dexterous robotic manipulation requires perception that remains informative from pre-contact approach to contact initiation and post-contact control, which is addressed by introducing FingerEye, a

Key concepts

Complementary Binocular RGB Cameras
This setup uses two cameras: one focused on the transparent object surface at a very close distance (10mm) for fine detail, and another further away (80mm) tilted to capture a wider view during the approach. This provides crucial implicit stereo depth cues necessary for accurately localizing objects and aligning the gripper before physical contact is made.
Peripheral Compliant Ring
Instead of placing tactile sensors directly in the visual path, a flexible silicone ring surrounds a clear acrylic window. When external forces or torques are applied, this ring deforms, causing measurable changes in the object's pose. This allows the system to sense contact and force indirectly without obstructing the primary RGB vision.
Group-Structured Modality Fusion
This learning interface method organizes different sensor inputs (like wrist vision and fingertip feedback) into structured groups. Instead of letting easy global cues override local, detailed feedback, this fusion ensures that local fingertip views integrate contact data first before being combined with broader sensory information, preventing shortcuts in policy design.
Group-Conditioned Decoding
This decoding technique ensures that every sensor group contributes separately to the final action decision. Each group generates its own update residual, which are then averaged together. This prevents one modality from dominating the policy and competition between local fingertip observations and global wrist information is reduced.

Terminology

Summary

Dexterous robotic manipulation requires perception that remains informative from pre-contact approach to contact initiation and post-contact control, which is addressed by introducing FingerEye, a sensing and learning framework that strengthens robotic dexterity through continuous vision-tactile feedback throughout interaction.

The gist: FingerEye provides continuous fingertip feedback for contact-rich dexterous manipulation by combining binocular RGB cameras with a transparent contact surface for natural visual perception before contact and marker-tracked deformation of a compliant ring for proxy contact wrench sensing after contact, leading to improvements in success rates across seven task settings.

Sensing Interface Design

The design of the FingerEye sensing interface focuses on providing continuous feedback by integrating complementary vision and tactile sensing modalities. This is achieved through five key highlights:

  1. Complementary binocular RGB cameras for wide view and implicit depth: The tip-facing camera observes the transparent contact surface at a short working distance (approximately 10mm), while the root-side camera operates at a longer working distance (roughly 80mm) and is tilted to preserve a broader view during approach. This provides implicit stereo depth cues for object localization and contact alignment.

  2. Peripheral compliant ring for contact sensing without blocking vision: Instead of placing a deformable gel directly in the visual path, FingerEye keeps a clear central acrylic window surrounded by a compliant silicone ring. External forces and torques deform this ring, inducing measurable plate-pose changes while leaving the central RGB view unobstructed. This peripheral compliant interface lets both frontal and peripheral contacts affect the estimated plate pose.

  3. Multi-tag pose tracking for robust deformation sensing: A custom AprilTag array on the acrylic plate converts contact-driven ring deformation into a compact six-dimensional plate-pose change. The system uses a joint multi-tag solve, where all visible corners are optimized under one shared rigid-plate transform, which is found to be more stable than averaging individual tag estimates when some tags are occluded or distorted.

  4. Compact wedge-shaped form factor for dexterous fingertips: The module adopts a wedge-shaped, fingertip-scale form factor with an overall size of approximately 41.7×30.0×25.0 mm, ensuring compatibility with compact dexterous hands and confined spaces.

  5. Low-cost and reproducible implementation: The system uses only low-cost off-the-shelf and 3D-printed components, with a bill of materials totaling roughly 60 USD per module, supporting a low-cost and reproducible implementation.

Learning Interface Design

The learning interface is designed to effectively leverage distributed FingerEye feedback by addressing issues in policy design. The core innovation is the use of group-structured modality fusion to reduce modality shortcuts that occur when relying on easier global cues. This involves two main components:

  1. Modality-group encoding for structured representations: Each observation group (wrist RGB, robot proprioception, per-fingertip FingerEye RGB, and optional plate poses) is tokenized into structured representations. The GroupEnc interface allows fingertip-local views to integrate local geometry and contact-state cues before being fused with wrist vision and proprioception.

  2. Group-conditioned decoding for balanced modality fusion: To prevent policies from bypassing local feedback, the decoder uses group-conditioned decoding. This method ensures that each modality group produces a separate residual update, and we average these residuals before adding them to the action-query states, which reduces competition between fingertip-local observations and global wrist or proprioceptive cues.

Policy Interface Ablations

The paper systematically evaluates different policy interfaces to determine which best exploits the distributed feedback. The comparison focuses on how observation tokens are encoded (NoEnc, FlatEnc, GroupEnc) and how action queries fuse them (FlatDec vs. GroupDec). Diagnostic analysis of flat decoders showed that a single Proprio token receives 35.6%–59.0% of the attention mass, indicating a proprioceptive shortcut. The authors conclude that the GroupEnc+GDec interface provides the best result, reaching a mean success rate of 65.9% across all tasks, demonstrating that group-structured modality fusion better preserves and exploits FingerEye’s feedback.

Experimental Results

FingerEye's effectiveness is validated across seven contact-sensitive task settings (Coin Standing, Peg-in-Hole, Nut Picking, Letter Retrieving, Chip Picking, Syringe Manipulation), showing significant improvements over wrist-only policies.

  1. Sensing Modality Comparison: Adding binocular FingerEye to wrist observations improves mean success rates by 39.1% in simulation and 33.8% in real-world tasks (Q1). Binocular sensing reduces monocular ambiguity near contact, improving performance over monocular FingerEye (Q2).

  2. Continuous Sensing vs.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the FingerEye: Learning Dexterous Manipulation with Continuous Vision-Tactile Sensing paper. The core innovation lies in fusing continuous vision (binocular RGB) with robust contact sensing (marker-tracked deformation) into a novel learning framework called FingerEye Policy.

Here are specific, actionable improvements to AI systems based on this research, categorized by the capability they enable:


),

  1. The integration of the Group-structured Modality Fusion (GEnc+GDec) policy interface. This allows the AI system to explicitly maintain and exploit local fingertip feedback (fingertip RGB and deformation proxies) while balancing it against global cues (wrist vision, proprioception).

  2. The use of Cached Visual Summaries via a frozen RADIO backbone for efficient multiview training. The AI system can perform rapid policy updates by using pre-computed visual summaries instead of repeatedly running computationally expensive image inference on every frame.

  3. The design of the sensing interface that combines complementary binocular RGB views with a peripheral compliant ring for contact sensing. This allows the AI system to perceive both close-range geometry/depth (before contact) and measurable force/torque proxies from frontal and lateral contacts (during/after contact).

  4. The derivation of a compact, analytically derived wrench-deformation mapping. The AI can use this mathematical model to predict the applied wrench with high accuracy, even when the sensors are noisy or occluded.

  5. The implementation of Contact-aware Stopping based on continuous fingertip normal displacement detection during approach (using the calibrated force/torque proxy). The AI system can execute damage-free lifting or delicate grasping behaviors by stopping motion precisely when a contact threshold is met, rather than relying on post-contact reactive control.

The improved AI system—a FingerEye-Enhanced Dexterity Policy—can achieve the following specific capabilities:

  1. It can perform complex, high-precision tasks like coin standing or chip picking with significantly higher success rates (up to 30 percentage points improvement in mean success rate) compared to wrist-only vision policies.

  2. It can execute delicate grasping and manipulation of fragile objects (e.g., eggshells, wafers) by detecting the precise moment of contact onset via fingertip deformation, preventing crushing or dropping before full contact is established.

  3. It can effectively handle contact-rich tasks that require sequential perception—such as edge engagement and post-contact stabilization—by receiving continuous feedback throughout the entire interaction lifecycle (pre-contact, contact initiation, refinement, and release).

  4. It can operate robustly in cluttered or occluded environments because the system uses complementary binocular views to resolve ambiguities in object localization and contact alignment that would otherwise be missed by a single camera.

  5. It can achieve efficient learning on large datasets by utilizing cached visual summaries, allowing for faster policy training and better generalization across diverse simulated and real-world tasks.

Sources

Related papers